At Fieldguide we evaluate models on tasks based on the real work audit and advisory practitioners perform: finding evidence across a client's document set, performing tests, and citing and annotating documents.
We evaluated OpenAI’s GPT-6 Astra against our audit procedure benchmark, a set of 938 of those tasks. GPT-6 Astra set our best recorded numbers on evidence grounding and citation precision while using about a quarter of our baseline's tokens.
Where GPT-6 Astra beat our baseline
We expected a familiar capability story: a new frontier model would be able to complete more complex work with higher accuracy. GPT-6 Astra did that, but what surprised us was what it declined to answer. When evidence didn't fully support a conclusion, it said so. We’ve seen elements of this behavior in previous models, but GPT-6 Astra did this far more often.
Our documents sub-agent beat our production baseline on citation precision by up to 10 percentage points in some use cases. This is because GPT-6 Astra made narrower claims, cited the specific evidence behind each, and didn't fill gaps with assumptions.
A higher bar for evidence
Many audit procedures end in one of three states: the evidence supports the conclusion, the evidence reveals an exception, or there is not enough evidence to conclude either way (abstention).
GPT-6 Astra elected to abstain more often than previous models. This resembles behavior we've already engineered into our long-horizon Field Agents, so we’re excited to see GPT-6 Astra adopting this natively.
Audit rewards caution up to a point. If the bar is too low, a model may elect to confirm a result without sufficient evidence. If too high, the model may flag reasonable judgment cases as exceptions and create more review than a team can absorb.
The goal is the right conclusion at the level of scrutiny that the procedure calls for. This is the kind of decision an experienced auditor can make almost instinctively, but one that off-the-shelf AI systems tend to struggle with.
Determining which response is appropriate and how much inference to permit is a crucial part of the methodology we build into Field Agents and calibrate with practitioners.
Where this behavior has room for improvement
On two evaluation categories scored by agreement with a stored reference judgment, GPT-6 Astra dropped about 5 percentage points.
As expected, it returned complete answers, used the appropriate tools, and explained its reasoning, but it also applied a different standard of evidence than the stored reference. It wouldn't state a conclusion that seems obvious to an experienced auditor unless the document fully established it. Some of those refusals were correct. In others GPT-6 Astra read the document too literally and rejected an inference an auditor would make without thinking twice about it.
How Fieldguide customers benefit
Our Field Orchestrator powers Fieldguide’s long-horizon agents. Field Orchestrator drives dozens of downstream tools and sub-agents that make up our Field Agents, most of them running models tuned for specific use cases.
Because the Orchestrator picks the model for each task, we have the ability to deploy a release like GPT-6 Astra specifically to the best use cases for customers, and only roll it out as customers are ready for these improvements. Interested Fieldguide customers can request access to GPT-6 Astra in the coming weeks.
For more on Field Agents, talk to our team. Interested in helping us build this? See our open positions.

Jaeyoon Kim
Senior Machine Learning Engineer