AI in Finance

What 64% on a finance agent benchmark actually means for your team (and what 52% on v2 changes)

Published 18 May 2026

The Anthropic finance launch on 5 May came with a benchmark figure. Claude Opus 4.7 scored 64.37% on the Vals AI Finance Agent benchmark, which Anthropic described as state-of-the-art on financial tasks. That figure was quoted across the announcement coverage for a fortnight, including by me.

On Saturday 16 May, Vals AI shipped version 2 of the benchmark. It is now a 927-question test set, the questions are harder, and the leaderboard looks different. GPT-5.5 sits at 51.76%. Claude Opus 4.7 sits at 51.51%. Claude Sonnet 4.6 sits at 51.03%. No model clears 52% under partial-credit scoring, and no model clears 40% under the stricter “All-Pass” grading. (Sources: vals.ai/benchmarks/fab for v1.1, vals.ai/benchmarks/fabv2 for v2.)

If you are running a finance function and trying to make a decision about whether to put an agent in front of real work, the move from 64.37% to 51.51% in ten days tells you more than either number does on its own. This post is about what the benchmark actually measures, what changed in v2, and how to read a figure like this when you are the one deciding whether to deploy.


What the benchmark actually measures

The Vals AI Finance Agent benchmark is a structured test set of financial tasks scored against reference answers. Version 1 ran on 537 expert-authored questions; v2 expanded to 927. The scoring is severity-weighted partial credit with dealbreaker gating, which means a load-bearing fact missed in an otherwise correct answer scores zero. The reference standard is set by a three-model LLM jury (GPT-5.4, Gemini-3.1-Pro, Claude Sonnet 4.6), not a human expert panel. (Source: vals.ai/methodology; benchmark paper at arxiv.org/abs/2508.00828.)

The task mix is sell-side and equity-research weighted. Nine categories: general qualitative analysis, general quantitative analysis, market analysis, comparables, precedents, GAAP-to-non-GAAP adjustments, earnings analysis, MD&A disclosure tracking, and financial modelling (DCF, LBO, accretion/dilution). There is no close, no consolidation, no statutory reporting, no audit-facing work, no FP&A budgeting, no board commentary.

That is a real benchmark of a real slice of finance work. It is also a specific slice. The first thing the figure tells you is how the model does on the work the benchmark covers. The second thing it does not tell you is how the model does on the work your team actually spends its time on.

Your team’s work includes a lot of structured analysis. It also includes a lot of work where the right answer is not knowable from the data, the inputs are not standardised, and the output is judged on whether the audit committee accepts it rather than whether it matches a reference. The benchmark does not measure that.


What 52% on v2 means in practice

A 52% score is not the same thing as a 52% success rate on your work. Two adjustments matter, and they pull in opposite directions.

The first adjustment is upward. On the tasks where your work resembles the benchmark, the model will probably do better than 52% in the deployment, because you can fine-tune the prompt, add context, provide the connectors, and review the output. The benchmark is a cold-start measurement. Real deployments are warm. Fintool, which has built a finance-specific tooling stack on top of the models, reports its agents at roughly 90% on the same benchmark set (fintool.com/benchmark/finance-agent-benchmark-fintool). That figure is the vendor’s own and worth treating with the appropriate skepticism. The principle is real: a well-tooled deployment outperforms a generic model.

The second adjustment is downward. On the tasks where your work does not resemble the benchmark, the score tells you almost nothing. The model’s performance on judging whether a variance is “concerning” in the context of your business is not on the benchmark. The model’s performance on writing a board commentary in a way the chair will accept is not on the benchmark. The model’s performance on a reconciliation against a chart of accounts that does not match the standard is not on the benchmark.

The hardest categories on v2, financial modelling and precedents, top out at around 23%. That is not the headline most readers came away with. It is also the part of the result that is most informative if your team is being told that agents can handle modelling work.

The net is task-by-task. For high-volume, well-structured work with good tooling, the deployment performance is likely above the benchmark. For work that requires context the benchmark cannot capture, the score is not your number.


What to read alongside the headline

Three things to look at before treating a benchmark figure as a deployment signal.

The task mix. What kind of finance work does the benchmark test? Vals AI publishes the breakdown. Read it. If the benchmark is heavy on equity research and your team is heavy on group consolidation, the figure is less informative than the headline suggests. The Vals AI mix is heavy on equity research, M&A, and modelling. If you are running a corporate finance function, half of the benchmark’s task categories are not what your team does.

The error distribution. A model that scores 52% is also a model that gets the other 48% either wrong or incomplete. Where? The categories of error matter more than the headline pass rate. As noted above, on v2 the modelling and precedents categories are at around 23%. A model that fails predictably on edge cases is more useful in your deployment than a model that fails unpredictably across the population. Look at the failure modes, not the average.

The reference standard. What does “correct” mean in the benchmark? Vals uses an LLM jury of three models, not a human expert panel. That is a defensible methodology and it is not the same thing as expert grading. Many readers assume the latter when they see the figure. Knowing it is the former is the difference between treating the score as a strong external signal and treating it as a useful but partial one. (Source: vals.ai/methodology.)

The vendor evaluation post covers the broader decision framework. The benchmark is one signal in that decision. The danger is treating it as the signal.


What this implies for deployment

A 64% benchmark figure is not a clearance to put the model in front of work that needs to be right.

In practice, the gap between benchmark performance and acceptable deployment performance depends on what “acceptable” means in your function. The audit-facing work has a much higher acceptable bar than the meeting prep work. The journal entry has a much higher bar than the market research. The same model is deployable in one of those use cases and not yet in the other.

That is why the first ninety days post argues for a shadow-mode rollout on the high-stakes use cases. The shadow tells you what the agent gets wrong on your work, on your data, with your people reviewing. That is a more informative signal than any external benchmark.

The benchmark tells you the model has crossed a threshold that makes it worth your attention. Your shadow deployment tells you whether it is worth your trust.


What changes when the benchmark moves

The v1.1 to v2 transition this month is a case study in why the headline figure is a poor anchor for deployment decisions.

Two weeks ago the top score was 64.37% and Anthropic led. Today the top score is 51.76% and OpenAI leads by a quarter of a point. The model has not become materially worse. The benchmark has become harder, and the part of the test set that the previous generation was scoring well on has been replaced by harder questions where no model is yet competent. That is exactly how benchmarks should evolve. It is also exactly how vendor marketing gets out of date in ten days.

What this means for a finance function. The benchmark figure is a track-over-time measurement, not a fixed standard. The decision you make this quarter about whether to deploy an agent is not the decision you would make in twelve months. The conversation needs to be revisited as the figure moves and as the benchmark mix changes.

The discipline is to keep the conversation honest about what the figure is, and to avoid the trap of letting a single headline number set the deployment policy. The CFO who quoted 64.37% to the board on 6 May now has to update the number for the audit committee on 20 May. The CFO who quoted “the model is approaching a deployable level on structured finance work, but is below 25% on modelling-heavy tasks” is making a more durable point.


What I would do with this

Read the benchmark methodology before quoting the figure to your board. Match the task mix in the benchmark to the task mix in your function. Pilot the model on the highest-overlap use cases first. Run shadow mode on the high-stakes use cases. Treat the benchmark as an input to the deployment timeline, not a substitute for the work of testing the model on your own work.

Most importantly, do not let “industry-leading 64% benchmark score” land in the boardroom without the qualifier. The qualifier is “on a structured benchmark of tasks Vals AI selected, which is informative but is not a forecast of how the model behaves on our actual work.” That sentence is worth saying out loud, and worth saying before someone makes a deployment decision on the basis of the headline.


Where this lands

A 64% finance agent benchmark is a meaningful technical achievement and a partial signal for finance leaders. It tells you the model is competent on structured finance work to a degree that warrants a serious pilot. It does not tell you the model is ready for your deployment.

The answer to that question lives in your data, your governance, and your team’s review capacity. The benchmark gets you to the table. The work after the table is still yours.


Maebh Collins is a Fellow Chartered Accountant (FCA, ICAEW) with Big 4 training and twenty years of operational experience as a founder and senior finance leader.

Back to Blog | AI in Finance →