Skip to main content

Announcing Field Financials, Fieldguide MCP and Our Refreshed Brand

Learn More

Announcing Field Financials, Fieldguide MCP and Our Refreshed Brand

Learn More

Auditing an audit agent

Dhruv Dhingra
Kunal Sinha
Dhruv Dhingra and Kunal Sinha
5 min read
  • AI
  • Blog

The measurement problem

If you want to know whether an audit agent is any good, the obvious test is to compare what it produced against what the auditor concluded. We do exactly that. But the comparison is harder than it looks, because the thing we are measuring against is itself a judgment call, one that depends on firm methodology, engagement history, and evidence the agent may never have seen. Sometimes the auditor is right for reasons the agent had no way to know. Occasionally the auditor is wrong.

That makes agreement with the auditor a useful signal, but also a poor scoreboard. Most of what we have learned about improving our agents has come from the cases where the two disagreed, and from the work of figuring out why.

This post covers how we evaluate control testing agents: what we measure offline, what we measure in live engagements, and what the disagreements told us.

What a testing agent is actually doing

A quick grounding for readers who don't spend their days in engagement files. In a control-based audit, the auditor tests whether a control operated effectively over a period. That means reading the control description, identifying what evidence would demonstrate it, examining that evidence, and recording a finding, typically whether an exception exists, with support for the conclusion.

The finding then moves through auditor sign-off and reviewer sign-off before it reaches a report. Our testing agents perform that work for the auditor to review: retrieve the relevant methodology and evidence, evaluate it, and produce a draft finding with its reasoning. The auditor accepts, revises, or rejects it.

What a testing agent is actually doing

What we mean by quality

We’re not trying to replace auditor judgment. We’re trying to make it faster to exercise and better supported. So we measure along two lines:

  • Time saved in delivering sound audit work.
  • Job completion quality: completing a specific task at least as well as an experienced auditor familiar with both the firm's methodology and the client's particulars.

Job completion quality decomposes further: correctness, completeness, grounding in evidence, instruction adherence, and coherence. There's also a final dimension we care about: knowing when not to answer (abstaining).

Offline: building evaluations alongside the agent

Evaluating an agent is not the same as evaluating a model. As Anthropic puts it, when you evaluate an agent you are evaluating the harness and the model together. Safeguards, instructions, retrieved context, and deterministic tooling all move the result, often more than a model swap does.

Four things we hold to:

  • Define success at the level of the job. Translate the work a practitioner actually performs into explicit inputs, outputs, and success criteria. "Produces a good assessment" is not a criterion. "Identifies the exception present in the evidence sample and cites the supporting document" is.
  • Build representative, expert-validated sets. Ordinary cases, hard cases, and cases where the correct behavior is to decline to conclude. That last category is small in most eval sets and shouldn't be.
  • Fit the evaluator to the question. Deterministic scoring where the answer is checkable. Structured model graders where it isn't. Both calibrated against expert review.
  • Build evals with the product, not after it. The eval set and pipeline are first-class components, not instrumentation bolted on once the agent ships.

Online: measuring against signed-off findings

Offline evaluations tell us what an agent can do under defined conditions, and they're how we validate architecture changes before shipping. What they can't tell us is how the agent holds up against the real distribution of work. Our eval sets are curated from real engagement traces, so the evidence in them is real, but the cases are chosen, and the evidence arrives pre-assembled.

Live engagements give us whatever controls, firms, and client types show up that week, in their real proportions, with evidence arriving incrementally.

So we track agent findings against the auditor's at every stage: when the agent is re-run with clarifications, at auditor sign-off, at reviewer sign-off, and in the final report. Tracking all four stages matters because a finding can change at any of them, and where it changes tells you something different about what went wrong.

What the mismatches told us

A match means the agent reached the answer the auditor ultimately recorded. A mismatch is a starting point, not a verdict. Every one gets reviewed with an audit expert, and they resolve into a few buckets:

  • The auditor applied firm-specific methodology the agent didn't have access to.
  • The auditor incorporated evidence outside the agent's context.
  • The agent made an error: bad retrieval, bad reasoning, or an unsupported conclusion.
  • The auditor made an error, caught later in review.

Across a set of 100 mismatches , roughly 66% fell into the first two buckets, 25% were agent errors, and 9% were corrections to the auditor's original finding.

The dominant pattern was agents were missing context that lived in the firm's own prior work. The reasoning was sound, but the inputs were incomplete.

Agent Knowledge

That finding drove the first generation of Agent Knowledge, our context retrieval layer that grounds agents in a specific firm's institutional knowledge. Ingesting a firm's prior-year workpapers into a firm-specific knowledge base lets a testing agent retrieve the precise methodology that firm applies to a given control, rather than reasoning from a generic standard.

Testing agent accuracy over time

Methodology: Accuracy measured as binary alignment of agent finding with signed off auditor findings

Where this leaves us

  • Offline evaluations show what an agent can do under defined conditions. Online scoring shows whether that carries into real audit work. Neither substitutes for the other.
  • Agreement with the auditor is a signal, not a ground truth. The review process behind each disagreement is where the actual learning happens.
  • Firm-level variation is not noise to be averaged away. It's the thing our agents most needed access to, and measuring it per firm is how we found that out.

So: can you audit an audit agent? Only in the sense that auditors audit anything, by defining what good looks like, gathering evidence, and being honest about the limits of your evidence. The uncomfortable part is that our benchmark is a human judgment that is usually right and occasionally isn't, and no amount of measurement makes that go away. It just makes it visible.

Dhruv Dhingra

Dhruv Dhingra

VP, Product at Fieldguide

Kunal Sinha

Kunal Sinha

Machine Learning Engineer at Fieldguide