Method

Evaluative AI: assessing what has been made.

Nearly every AI application is built to produce: generate the text, write the report, fill in the answer. Evaluative AI does the reverse. It assesses what has been made: a file is tested against a normative framework, and the result is a findings report. We assess; we produce nothing. The human decides.

This page explains why that requires a different kind of system, what kinds of findings such a system delivers, and why the human remains the decision-maker in our architecture.

The difference

Why a generative model is poor at finding faults.

A generative model is trained to produce plausible text. It wants to complete, round off and smooth over. That is exactly the wrong reflex for finding faults: whoever has to find a gap must not fill it in, and whoever has to see a contradiction must not move along with the text.

An evaluative system is built differently. It starts from the normative framework, looks for the supporting evidence for each obligation in the file, and stays silent where it finds no anchor. Better to miss a remark than to make one claim that does not hold up: every finding points to its source and can be verified.

Findings

Four types of findings.

What an evaluative review delivers falls into four types. They apply to any file, from a sustainability statement to an acquisition dossier.

Missing

An obligation or expectation for which no document or passage exists. The file assumes something that is not there.

Contradictory

Two documents or passages that contradict each other: a figure that differs between two tables, a policy that the process description contradicts.

Unsupported

A claim without a source that carries it. The statement is there, the evidence is not.

Outdated

A reference to a superseded version, deadline or threshold. Regulation moves; files rarely move with it.

Every finding is formulated as an observation with a source reference: what we found, where it sits and why it matters. Never as a verdict, and never as a score.

Architecture

Human-in-the-loop as architecture, not as a disclaimer.

Our systems can assess, not decide. That is not a footnote at the bottom of a report but the way they are built: the system delivers findings with source references, the expert weighs them and decides. Sharpness for whoever judges, no authority that takes the judgement over.

Just as fundamental: we assess the document, not the person. Our systems make no decisions about people and assign no scores to whoever submits a file. They read what is on paper and test whether it is supported.

How we relate to the EU AI Act, and what we do and do not do with your documents, is set out in black and white:

In production

Not a concept, a running system.

Our engine has been in production since March 2026 at Esmond, our evaluation platform for higher education, active in three markets.

We apply the same evaluative approach to corporate files: document sets where a mistake costs money and where, until today, the only answer has been expensive human hours.

Applications

Where we apply this today.

Pre-audit →

An evaluative pre-check of your CSRD and ESRS document set: the dress rehearsal before the formal review.

MacLeod, real estate →

A site's document set read the way an objector would read it, before you sign.

Advisory →

Founder-led expert engagements on evaluative AI, AI-assisted document review and EU AI Act readiness.

Contact

Want to know what an evaluative review would surface in your files? That is a conversation, not a form funnel.