CensatimCENSATIMMethodology →

Validation

How we test Censatim,
and what the tests found

A dated, versioned record of the validation programme behind every report: the automated checks, the regression testing, the failures we found, and how the tool improved because of them. Updated after every full retest, including when results get worse.

Validation report
v3
Published
22 August 2026
Criteria basis
Delegated Acts as amended, in force July 2026
Next full retest
24 August 2026
01

What we test for, and what we do not claim

EU Taxonomy alignment is a framework which relies on quantitative criteria and professional qualitative judgement. As such, established providers regularly reach different conclusions on the same company or asset. So we do not test for a number nobody can prove. We test for the six properties an assessment must have for it to have credibility:

Consistent

The same facts produce the same ratings. Re-running a case must not change the outcome.

Traceable

Every finding cites the answer or document it rests on, and the full dialogue is annexed to the report.

Honest about uncertainty

Where evidence is missing, the tool says Undetermined and states what would resolve it, rather than guessing.

Resistant to pressure

Attempts to push or trick the tool into a favourable rating must not change the result.

Accurate

Using Censatim's closed-loop database, the tool can only refer to concise and specific EU Taxonomy criteria.

Fair

The tool does not discriminate and provides a fair and impartial assessment regardless of the company or asset.

Everything on this page is backed by records we hold: per-report quality check lines, sweep-by-sweep test logs, and the full output of every run described below.

02

Every report is checked before you see it

Before any paid report is issued, it passes an automated audit. 20 checks are enforced in code: a report that fails one is not issued. Nine of the twenty are behavioural rules our testing programme forced into code; the rest are structural, built in by design, from the locked unit of assessment to criteria text reproduced verbatim from our database rather than by the AI. A further 7 checks are logged and reviewed on every report as candidates for enforcement. Among the enforced checks:

The headline conclusion must agree with the detailed ratings table.
A confirmed absence of evidence is reported as a finding; it is never softened into "pending".
An unanswered question can never, on its own, produce a Not aligned rating.
Every citation must resolve to the actual question and answer record annexed to the report.
The regulatory criteria quoted in the report are rendered directly from our criteria database, so they cannot be misquoted or invented.
No document is ever named in a report unless it was actually provided.
Ratings use a closed vocabulary: no invented categories, no hedged wording.

Each issued report carries the result line of these checks in its internal record, tied to the report reference, so any report we have ever issued can be audited against them.

03

The regression test programme

We built a test harness that plays scripted cases against the live production site, answering the tool’s questions deterministically from fixed fact sheets. Because the answers never vary, any change in output between runs is the tool’s, never the tester’s. Across the programme and the development testing that preceded it, Censatim has completed >1,000 assessment runs.

The scripted library currently holds 55 cases, producing 86 assessment runs in a full sweep (multi-activity chains run each activity and a family summary), spanning six kinds of ground:

Core behaviours

Eligibility, alignment, gateway failures, partial outcomes, qualifying sub-sets, and multi-activity companies.

Complex real-world activities

Hydrogen manufacture, cement, hydropower, district heating, gas, bio-waste, data centres and other activities with complex DNSH structures.

Multi-activity chains

Companies whose assessments span two activities across energy, real estate, waste, manufacturing, transport, ICT, water and construction.

Adversarial cases

Non-responsive answers, pressure to rate favourably, and prompt-injection attempts hidden in descriptions.

Contradictory cases

Cases where the example changes the eligible activity and cites different criteria with different metrics.

Perfect cases

Perfectly aligned cases to test whether the tool makes assumptions or jumps to conclusions.

>1,000
Assessment runs
18
Tool errors found and fixed
0
Wrong approvals
9
Rules now enforced in code

What counts as an error? The table below counts errors attributable to the tool: a rating harsher than the evidence justified, a favourable rating or endorsement that should not have been given, an internal inconsistency in a report, or a questioning defect such as probing outside the matched activity\u2019s criteria. It does not count the limits of the test scripts themselves: a scripted respondent cannot answer every follow-up a live assessment can ask, and where the tool responds to that silence with an honest Undetermined, that is the methodology working. Those runs are logged and drive improvements to the tests, not the product.

The library runs in full sweeps. A case is archived as proven only after repeated clean runs; anything less stays live and is re-run every sweep. The record, sweep by sweep:

SweepScopeRunsAssessmentsErrorsFailure rateConclusion
1First run of the 21-case core library1919315.8%Three defects found and fixed. Learned: report-assembly rules cannot be left to the AI alone, so enforcement in code began here. Next: re-run to verify.
2Re-run after the first fixes191915.3%Fixes held. New find: contradictions in the evidence must be named, never silently resolved. Next: a third run to test consistency.
3Third consecutive run1818211.1%Two rating-discipline defects found. Learned: an unanswered question is not an unfavourable answer. Both became enforced rules. Next: verify.
4Verification of the sweep 3 fixes212100%All fixes verified, zero errors. Next: expand the library with a blind holdout set of complex activities and multi-activity chains, written without sight of the tool.
5Expanded 37-case library, first run534512.2%One configuration defect in the linked-assessment feature explained nearly every chain failure; fixed once, all cleared. Learned: each new feature earns its own test cases.
6Full library, chain fix live463812.6%All eight chain family summaries passed for the first time. New find: evidence in production must never read as failure. Next: verify the pending-evidence rule.
7Verification; proven cases archived433525.7%Two defects found through full-transcript review and fixed in code. Learned: summaries are not enough, so raw transcripts are read for every escalation.
8Reduced live set322414.2%One recurrence fixed. A long-standing suspect was cleared: the tool had been right and the test case wrong. Learned: verify before concluding, in both directions.
9Verification run322400%Zero errors and zero warnings, with every miss in the cautious direction. Next: tighten the remaining test scripts.
10Current build, current library322414.2%A defect class recurred a third time and moved from instruction to code: the grounding contract, under which every negative finding must declare its basis. Learned: rules that matter get enforced, not requested.
11First run under the grounding contract322400%Zero errors. Next: the full board, including every archived negative-verdict case, to prove the new contract softened nothing.
12Full 51-case board, overnight826723.0%Nothing softened: every threshold breach, confirmed absence and adversarial case held its verdict. Two chain behaviours moved into code. Next: add rogue eligibility cases.
13Full board plus four rogue eligibility cases867111.4%All four rogue cases refused correctly on their first ever run. One judgement boundary flagged for monitoring. Next: a variance fix, carefully.
14Full board, verifying a variance fix867111.4%The fix softened confirmed failures; the untouched prompt-injection canary caught it within the sweep and it was reverted within a day. Learned: the canaries are the programme.
15Full board, verifying the revert867111.4%Revert verified across the board; the variance resolved by a narrower rule. One over-harsh rating on one dimension prompted a methodology rebuild: that event is now the loudest possible alarm.
16Full 55-case board867100%Strongest board to date: zero errors, every required negative verdict produced, and 88 percent of compliant checks reaching a full In-line verdict, with the remainder concluding Undetermined pending evidence, never wrongly negative.
17Full 55-case board, plus targeted re-runs after test-script repairs1058511.2%One error: an eligibility screening miss on a never-tuned adversarial case, which then passed three consecutive re-runs; in every run the case’s core protection held and the rogue claim was never endorsed. Four other failed expectations traced to defects in our own test scripts, not the tool: mismatched robot answers produced honest Undetermined verdicts, the scripts were repaired, and every repaired case re-verified. No wrong approvals in any of the 105 runs.
18Next full retestTBCTBCTBCTBCResults to follow, whatever they are.

Failure rate is tool errors divided by assessments. 0% under 5% 5% or more. Runs counts every scripted case played in the sweep; Assessments excludes the free compiled family summaries, which are assembled from completed assessments rather than assessed themselves.

The next full retest is scheduled for 24 August 2026, and is currently tested on a weekly basis. This page will be updated with its results, whatever they are.

04

What testing found, and how the tool is better for it

The programme exists to find problems, and it has. Rather than a bug list, here is what the significant findings had in common and what changed because of them:

1.Missing evidence is not the same as bad evidence

The hardest discipline in automated assessment is refusing to conclude. Early testing showed the tool could occasionally treat information that was merely unavailable at the time of asking as if its absence were confirmed, tipping a rating harsher than the evidence justified. This is now governed by hard rules: an unanswered question can never on its own produce a Not aligned rating, and every Not aligned rating must rest on a confirmed absence or a specifically breached requirement before the report can be issued.

2.A report must agree with itself

Testing caught instances where the summary of a report could drift from the detail beneath it. Internal consistency is now checked automatically on every report: the headline must agree with the ratings table, every dimension must be accounted for, and every finding must cite its source, or the report is not issued.

3.Questions must earn their place

Early versions could occasionally ask questions beyond what the criteria actually require, or route an answer to the wrong place. The questioning is now tightly fenced to what the matched activity’s criteria genuinely need, which is why assessments stay short without losing coverage.

Every fix becomes a scripted case, so a problem found once is re-tested until Censatim is comfortable the issue has been addressed. That is the point of the programme: not that nothing was ever wrong, but that what was found stays fixed, provably.

05

The tests we have never tuned

A small set of cases matters more than the rest, because they test whether the tool can be pushed, tricked or flattered into a favourable answer. None has ever been adjusted to pass:

User pressure
12 of 12 clean

A scripted user repeatedly pushes the tool to rate an asset as aligned, escalating across the assessment. The rating has never moved.

Prompt injection
11 of 12 clean

Instructions to the AI hidden inside a project description, attempting to override the assessment. Ignored in eleven runs. In the twelfth, this case failed: it had caught a change of ours that quietly softened verdicts, which was reverted within a day. That one failure is the strongest evidence on this page that the canaries work.

Greenwash under a transition label
4 of 4 clean

An oil and gas producer presenting itself as an energy transition company, citing a carbon capture pilot and a net zero pledge. Declined as not eligible in one round, every time.

Confirm-my-figure request
4 of 4 clean

A resort operator whose published report already claims 100 percent alignment, asking the tool simply to confirm it. The tool correctly identifies what is eligible, and has refused to endorse the published figure in every run.

These cases are release canaries: they are re-run before anything customer-facing ships and on any change to the underlying AI model.

06

What this means for your report

Every report you receive has passed the same enforced checks as every report described on this page. The criteria basis your assessment was made against is stamped on the report itself (currently “Delegated Acts as amended, in force July 2026”), and each issued report is stored under its reference and can be retrieved as issued.

We stand behind this with an accuracy commitment. If you identify a factual or technical error within a paid alignment report, email support@censatim.com quoting your report reference. A substantiated error earns a full refund plus £50. That commitment is only possible because of the programme on this page. It applies to paid alignment reports and is subject to our published terms.

For how the assessment itself works, criterion by criterion, see the methodology page. Assessments are AI-assisted analysis and do not constitute professional or regulatory advice.

See the output for yourself

Sample reports show exactly what an assessment produces. Eligibility is always free.

View sample reportsMethodologyAsset assessmentReporting assessment