← Benchmarks 🕐 27 min read
Benchmarks

Finance AI Benchmarks 2025–2026: What the Evidence Actually Says About AI in Financial Work

This synthesis is backed by the source cards and mirrored artifacts listed in

See also (wiki): wiki/slm-finance-benchmark-landscape.md · wiki/ai-model-evaluation-benchmarks.md · wiki/quant-asset-management-ai.md · wiki/financial-services-ai-deployment.md


Source files:

This synthesis is backed by the source cards and mirrored artifacts listed in the Sources section, including Samaya Frontier Finance v3.1 public JSONL and bundle mirrors under sources/21-benchmarks/samaya-frontier-finance-v3-1-*.

Evidence Boundaries:

Reported benchmark leaderboards are treated as source-reported snapshots unless the relevant StateBench fixture has reproduced the model, harness, retrieval configuration, scorer, and data split locally. Public finance benchmarks are candidate-screening evidence, not production approval for regulated workflows.

Executive Summary

  • Twenty-four independent or semi-independent benchmarks published between July 2025 and June 2026 now evaluate AI systems on realistic financial tasks: SEC filing analysis, financial search, spreadsheet automation, multi-formula calculation, tool use, exploratory data analysis, quant strategy coding, executable quant-code implementation, full workbook creation, professional investment-research report quality, auditable financial-research derivations, open-ended hedge-fund analyst reasoning, multilingual/multimodal finance, IPO S-1 diligence, broad open-ended financial research queries, banking-domain meta-benchmarking, hierarchical financial research-report auditing, chart reading, and fact-level financial OCR. The core finding still holds: realistic finance work remains far below unsupervised professional reliability.
  • The best result across all benchmarks: OpenAI o3 at 46.8% on Finance Agent Benchmark (537 expert-authored questions, real EDGAR filings). The best human expert on the same benchmark: well above 90%.
  • Accuracy collapse with complexity is severe and consistent. FinMathBench (Ant Group, AAAI-26) shows GPT-4o dropping from 72.9% on single-formula problems to 14.0% on four-formula chains — the kind of calculation required for real derivative pricing, portfolio rebalancing, or layered risk assessment.
  • The reliability gap is not primarily about raw knowledge. It is about knowing when not to answer (RealFin), using tools within regulatory constraints (FinToolBench), and navigating unguided exploratory tasks (DataClawBench). These are the capabilities financial work actually requires.
  • The cost case for AI-assisted financial analysis is real: $3.79 per task versus ~$25.66 for a human analyst (Finance Agent Benchmark). The reliability case for fully autonomous deployment is not.

The Benchmark Landscape

The 2025-2026 benchmark cluster now provides a coherent picture of where AI agents stand in financial work. The strongest eight-source core covers knowledge recall, mathematical reasoning, search, tool use, spreadsheet automation, and exploratory analysis. The newly ingested sources add quant strategy coding, executable quantitative-finance code tasks, data-recency scoring, full-workbook finance spreadsheet tasks, professional investment research report evaluation, auditable financial-research derivations, open-ended hedge-fund analyst reasoning, multilingual/multimodal evaluation, process reward modeling, chart VLM evaluation, and financial-critical OCR. Together they converge on the same structural finding: current models are useful for augmenting analysts on routine tasks, unreliable for autonomous execution of complex ones.

Publication dates and institutions:

Benchmark Institution Date Scope
FinSearchComp ByteDance Seed + Columbia Business School Sep 2025 Financial search & reasoning; 635 questions; 21 models
Finance Agent Benchmark Vals AI + Stanford May 2025 (updated) SEC filing analysis; 537 questions; agentic harness
FIRE Du Xiaoman Technology + Tsinghua + Renmin U Feb 2026 14,000+ certification questions + 3,000 scenario problems
RealFin INSAIT + Newcastle + MBZUAI Apr 2026 Implicit-premise financial reasoning; 2,020 bilingual questions; 15 models
FinToolBench Shanghai AI Lab + Tencent Mar 2026 Financial tool compliance; 760 tools; 295 queries
Spreadsheet-RL UIUC + Meta May 2026 Excel task automation; 1,660 domain-specific tasks
WorkstreamBench arXiv preprint May 21 2026 End-to-end finance spreadsheet workstreams; Accuracy / Formula / Format taxonomy
BlueFin arXiv preprint May 29 2026 Professional finance workbook synthesis/manipulation/comprehension; 131 tasks; 3,225 rubric criteria
FinMathBench Ant Group AAAI-26 Formula-chain math; 946 questions; 40 models
DataClawBench Sun Yat-sen University May 2026 Exploratory data analysis; 492 tasks; 2.06M financial records
Deep FinResearch Bench arXiv preprint Apr 2026 Professional investment-research report quality: rigor, forecasting/valuation, verifiability
BigFinanceBench Rogo + OpenAI Jun 2026 928 workflow-grounded financial-research tasks with derivation-level rubrics
Hedge-Bench Trata + BYU + Osmosis Jun 2026 102 hedge-fund analyst reasoning tasks grounded in professional reasoning traces
IPO Finance Agent Benhenda Jun 2026 1,000 IPO S-1 diligence questions; 70-question SpaceX public subset; contextual-retrieval harness
Samaya Frontier Finance v3.1 Samaya AI Jul 2026 observed 220 public open-ended financial-research queries; 11,543 rubric items; expert-rubric grading
Meta-Benchmarks for Financial-Services LLM Evaluation Commonwealth Bank of Australia Jul 2026 452 public benchmarks mapped to 41 O*NET work activities and 38 BIAN banking domains
FinReasoning TongjiFinLab Feb/May 2026 4,800 Chinese financial research-report eval items across semantic consistency, data alignment, and deep insight
QuantEval arXiv preprint Jan 2026 Quant QA, quantitative reasoning, and CTA-style strategy coding
QFBench Open benchmark / Harbor framework May 2026 public V11 snapshot Executable quantitative-finance agent coding; 87 tasks reported on the public site
AFIB / SuperInvesting arXiv preprint Mar 2026 Financial intelligence dimensions: accuracy, completeness, recency, consistency, failure modes
FinMMEval CLEF 2026 CLEF lab / arXiv preprint Feb 2026 Multilingual and multimodal finance tasks across exam QA, PolyFiQA, and decision making
Fin-PRM Qwen DianJin / arXiv preprint Aug 2025 Finance process reward model for step/trajectory supervision
FinChart-Bench arXiv preprint Jul 2025 Financial chart VLM evaluation
FinCRITICAL-ED arXiv preprint Nov 2025 Financial fact-level OCR over critical fields

Credibility notes: The original eight-source core is academic preprint or peer-reviewed conference evidence. The expanded seven-source addendum is mixed: QuantEval, WorkstreamBench, BlueFin, Deep FinResearch Bench, BigFinanceBench, Hedge-Bench, IPO Finance Agent, Samaya Frontier Finance v3.1, Meta-Benchmarks for Financial-Services LLM Evaluation, FinMMEval, FinChart-Bench, FinCRITICAL-ED, and Fin-PRM are useful research benchmarks / methods; QFBench is a public benchmark/source-code signal whose live leaderboard needs local reproduction; AFIB is a vendor/product benchmark signal and should be used for dimensions and failure taxonomy, not as independent proof that the named vendor is best. Collectively: HIGH credibility for the broad conclusions; MEDIUM for individual model scores, which shift quickly as models and retrieval systems update. Every source is date-stamped; model rankings should be treated as snapshots.


Where AI Falls Short in Financial Work

1. The 50% ceiling on realistic financial tasks

Finance Agent Benchmark presents the most direct test: 537 questions built from documents published after 2024 (minimizing contamination), validated by experts from banks, hedge funds, and PE firms, with an agentic harness that gives models Google Search and live EDGAR database access. OpenAI o3 — the best-performing model — achieves 46.8%. No model crosses 50%.

Cost per query: o3 at $3.79 versus an estimated $25.66 for a human analyst. The economic case for AI-augmented analysis is real. The operational case for autonomous deployment on complex financial research is not. The logarithmic cost-accuracy curve means spending more compute beyond $1/query yields sharply diminishing accuracy returns.

StateBench fixture: sa-052 materializes Finance Agent Benchmark as the analyst-style SEC-filing research source-acquisition task and preserves the paper/body cost discrepancy instead of forcing one cost number.

2. Multi-formula reasoning degrades catastrophically

FinMathBench (Ant Group, AAAI-26) isolates a specific failure mode: chaining financial formulas. Using 148 domain-specific formulas across accounting, insurance, fund management, futures, banking, and securities, the benchmark generates questions requiring one to four linked calculations.

GPT-4o’s accuracy by formula-chain depth:

Formula Chain GPT-4o (CoT)
1 formula 72.9%
2 formulas ~45%
3 formulas ~25%
4 formulas 14.0%

The practical implication: any financial calculation requiring three or more linked steps — standard for derivative pricing, portfolio rebalancing with constraints, or multi-layer risk calculations — is in the zone where AI error rates approach or exceed 75–85%. Three additional failure modes identified:

  • Program-of-Thought outperforms Chain-of-Thought on direct calculation. Models that narrate their reasoning are less reliable than models that write executable code for the math. For structured financial calculations, code generation is the more reliable path.
  • Variable bias. Models prefer to solve for commonly-seen variables. When a problem requires solving for an unusual variable (a corner case in risk management or specialty instruments), accuracy drops further.
  • Value “correction.” Models erroneously adjust valid but extreme financial values — treating a legitimate stress-scenario input as a data error and substituting a more typical value. During market crises, when extreme values are exactly the ones that matter, this failure mode is consequential.

3. Models confidently answer questions they should not

RealFin (INSAIT + Newcastle + MBZUAI, April 2026) tests a capability that standard benchmarks miss entirely: knowing when a problem cannot be answered. The benchmark systematically removes essential premises from financial exam questions while keeping them linguistically plausible. A question that sounds solvable is not, because a critical condition is unstated.

Findings across 15 models (5 commercial, 5 finance-specialized, 5 reasoning-enhanced):

  • General-purpose models over-commit. When conditions are missing, they guess rather than abstain.
  • Finance-specialized models do not reliably identify missing premises. Specialization on financial content does not produce calibrated uncertainty about financial problems.
  • The NOTA (None-of-the-Above) formulation — requiring genuine inference rather than pattern matching — exposes consistent overconfidence.

For high-stakes decisions — credit approval, risk assessment, compliance determinations — a model that answers underdetermined problems with confidence is more dangerous than one that returns low accuracy with appropriate uncertainty. RealFin documents that the confidence problem is systemic, not model-specific.

StateBench fixture: sa-054 materializes RealFin as the missing-premise and abstention source-acquisition task for finance refusal and clarification gates.

4. Passing binary tool-use tests does not mean regulatory compliance

FinToolBench (Shanghai AI Lab + Tencent, March 2026) evaluates agents on 295 financial tool-use queries across 760 executable tools. StateBench fixture sa-041 now materializes it as a source-acquisition task. Beyond binary success/failure, it introduces three compliance-relevant metrics:

  • TMR (Timeliness Mismatch Rate): Agent retrieves data from the wrong time period — using a quarterly filing from the wrong quarter, for instance.
  • IMR (Intent Mismatch Rate): Agent escalates an informational query to a transactional execution — placing a trade when asked about pricing.
  • DMR (Domain Mismatch Rate): Agent uses a tool from the wrong regulatory or market domain — applying a domestic tool to a cross-border transaction.

A model can pass standard tool-use benchmarks while failing on all three compliance dimensions. For regulated institutions — banks, asset managers, broker-dealers — these mismatch rates represent operational and regulatory risk that standard accuracy metrics do not capture.

5. Financial search remains substantially below expert level

FinSearchComp (ByteDance Seed + Columbia Business School, September 2025) tests 21 models on 635 expert-curated queries across three difficulty levels, validated by 70 professional financial analysts. Results:

Global subset:

  • Human expert accuracy: 75.0%
  • Best model (Grok 4 with web access): 68.9%
  • Gap to human: −6.1 percentage points

Greater China subset:

  • Human expert accuracy: 88.3%
  • Best model (DouBao with web access): ~70%
  • Gap to human: 18+ percentage points

The global gap is narrower than most benchmarks show — Grok 4’s web-augmented performance is close to human level on well-documented global topics. The Greater China gap is substantially larger, and illustrates a broader pattern: models with web access from the same country or market as the question perform disproportionately better. For multinational financial analysis, regional performance variance is a real deployment risk.

Common failure modes: shallow search depth, outdated data retrieval, timezone and reporting calendar errors, and inability to reconcile conflicting evidence across sources.

StateBench fixture: sa-053 materializes FinSearchComp as the financial search/source-discovery task and keeps live-web results out of stable replay scores.

6. Exploratory financial analysis is a harder problem than structured tasks

DataClawBench (Sun Yat-sen University, May 2026) represents the hardest class of financial analysis: 492 multi-step cross-domain tasks over 2.06 million real financial records, with native noise preserved. Unlike prior benchmarks that pre-select relevant data sources and provide schema hints, DataClawBench requires agents to identify what data is relevant, navigate noise, and synthesize across domains without guidance.

The finding: exploratory analysis breaks agent reliability in ways that structured-task performance does not predict. More exploration attempts do not produce more accurate answers. Agents that perform well on guided retrieval fail on unguided discovery — which is how most meaningful financial analysis actually works.

StateBench fixture: sa-051 materializes DataClawBench as the exploratory financial data-analysis source-acquisition task for data-agent and custom-schema harness design.

7. Spreadsheet automation: RL helps, human parity remains distant

Spreadsheet-RL (UIUC + Meta, May 2026) tests RL-fine-tuned agents on 1,660 domain-specific spreadsheet tasks (finance, supply chain, HR, investment banking, asset management) in a realistic Microsoft Excel environment.

Results (Qwen3-4B, Pass@1):

Configuration SpreadsheetBench Domain-Spreadsheet
Base model 12.0% 8.4%
With tool access ~18% ~12%
RL fine-tuning (GRPO) 23.4% 17.2%

RL post-training improves performance by roughly 5–9 percentage points. 17–23% pass rate means the majority of tasks still fail. For finance-specific spreadsheet tasks (financial modeling, investment analysis), the 17.2% rate reflects a domain that is particularly demanding due to formula dependencies and multi-step workflow requirements.

WorkstreamBench and BlueFin, both published in late May 2026, sharpen this lane from spreadsheet QA toward complete professional workbooks. WorkstreamBench evaluates end-to-end finance spreadsheet workstreams such as financial modeling, forecasting, and scenario analysis, with separate Accuracy, Formula, and Format/readability dimensions. It is now frozen as source-acquisition fixture sa-105, with local PDF metadata and Chandra OCR recorded for the May 21, 2026 v1 arXiv paper. BlueFin evaluates synthesis, manipulation, and comprehension over professional finance workbook artifacts using 131 tasks and 3,225 granular rubric criteria. It is now frozen as source-acquisition fixture sa-106, with dataset and code links recorded for the May 29, 2026 v1 arXiv paper. Both report the same negative evidence: current frontier agents can produce plausible spreadsheets, but still fall short of professional finance standards, especially when dynamic correctness and chained calculations matter.

StateBench implication: spreadsheet finance tasks need workbook artifact scoring, formula dependency checks, dynamic recalculation, formatting / reviewability, and revision-readiness. A text-only answer or a static final number is too weak for finance-workflow claims.

8. Quant strategy coding must be execution-scored

QuantEval fills a gap between finance QA and actual quant workflow. It contains 1,575 samples across knowledge QA, quantitative reasoning, and quantitative strategy coding. The crucial design choice is the 60-task CTA-style backtesting harness: model-generated strategy code is scored by whether it executes and how closely its annualized return, drawdown, and Sharpe ratio match expert-validated references.

StateBench implication: for systematic quant work, code-generation evals must run the artifact. A response that describes a factor or strategy is not enough; the agent must produce framework-compatible code, run it under a fixed universe, cost model, and metric definition, and preserve the trace.

8a. QFBench adds executable quant-code implementation tasks

QFBench is the most direct new public signal for “can an agent implement quantitative finance work?” rather than “can an agent answer finance questions?” The public site reports 87 tasks, 42 models, 10,962 runs, a V11 leaderboard dated 2026-05-04 UTC, and a 61.7% best pass@1. The task examples span options, risk, rates, market microstructure, credit, factors, portfolio construction, crypto, and SEC-event workflows.

The important design feature is executable verification. Agents work in a Docker sandbox with Python financial libraries, write code, and are checked by strict numerical pytest-style verifiers. The benchmark also separates agentic CLI runs from a one-shot Finance-Zero baseline. That maps cleanly to the repo’s need to compare local/current models against frontier agents on tasks that resemble real quant implementation.

StateBench implication: build a separate quant-code lane rather than folding QFBench into the finance replay patch suite. Borrow the task shape: frozen inputs, no live APIs at runtime, no leaked solution hints, reference oracle passing 100%, repeated runs, numerical tolerances, cost/latency/retry traces, and a one-shot baseline. Treat the live QFBench leaderboard as source discovery until a local fixture or direct run is captured.

8b. Deep-research agents still fail professional investment-research standards

Deep FinResearch Bench (arXiv:2604.21006, submitted April 22, 2026) evaluates financial investment research reports across qualitative rigor, quantitative forecasting and valuation accuracy, and claim credibility / verifiability. The paper’s central result is negative evidence: reports from frontier deep-research agents still fall short of professional financial research reports. It is now frozen as source-acquisition fixture sa-107, with local PDF metadata and Chandra OCR recorded for the April 22, 2026 v1 arXiv paper.

This matters for the State of AI corpus because finance research synthesis is one of our own benchmarked workflows. A model that finds sources or drafts a report still needs claim-level source support, valuation math checks, forecast assumption checks, and explicit uncertainty. Deep FinResearch Bench should feed the reflect-audit and wiki-maintenance guardrails: reject unsupported investment claims, uncited valuation leaps, and autonomous signoff language.

8c. Auditable derivations beat final-answer-only grading

BigFinanceBench (arXiv:2606.03829, published June 2, 2026) evaluates financial-research agents on 928 open-ended tasks written by 52 finance subject-matter experts and audited by 12 separate reviewers. Each item pairs a reference answer with a point-weighted rubric that decomposes the analyst workflow into independently checkable steps: entity identification, source selection, line-item retrieval, accounting adjustment, formula construction, calculation, and final synthesis. The paper reports 15,656 rubric criteria and 36,241 total rubric points.

The key result is that final-answer accuracy is a lossy proxy for finance research quality. The best evaluated systems reach only 58.8 percent rubric credit and below 45 percent final-answer accuracy. The benchmark also shows why rubric-level evaluation matters: models can make partial progress while still missing a source, period, accounting definition, or calculation step that a human reviewer would need to audit.

StateBench implication: finance replay and wiki-maintenance tasks should not only ask whether the final answer sounds right. They should preserve visible tool calls, source URLs, source dates, period/definition choices, formulas, Python/table traces, and partial-credit failure localization. BigFinanceBench is now frozen as source-acquisition fixture sa-127.

Important caveat: the paper reports two LLM judges, Gemini 3.1 Pro Preview and Claude Opus 4.7, grading visible trajectories against rubrics. It reports high inter-judge agreement, but model rankings remain dated paper snapshots. The paper states that a 50-question subset and harness are released for academic research; those public artifact links were not accessible from this environment on 2026-06-06, so local reproduction remains pending.

8d. Hedge-fund analyst reasoning is harder than finance QA

Hedge-Bench (arXiv:2606.03918, posted June 2, 2026) evaluates agents on 102 actual on-the-job tasks grounded in explicit reasoning traces from professional hedge fund analysts. Its target is the work that sits between source retrieval and investment memo drafting: valuation, growth and expansion, M&A, competitive positioning, operational execution and strategy, and risk. The paper reports frontier agents below 16 percent pass@1.

The benchmark design is useful because each task has a closed document pack, expert reasoning moves, source-file citation requirements, strongest counter-evidence, reconciliation, ambiguity, and hallucination checks. It maps well to StateBench finance replay and wiki-maintenance evaluation: finding the right source is not enough if the agent misses the analyst move, ignores the counter-case, or writes unsupported synthesis.

Important caveat: Hedge-Bench says its tasks are graded against verified expert steps, but the reported evaluation also uses a Gemini-3.1-Pro LLM-as-judge layer for grounding, coverage, and synthesis. Treat it as a high-value task design and negative-evidence source. Treat paper leaderboard values as dated reproduction targets, not current StateBench scores.

8e. IPO diligence is not periodic-filing QA

IPO Finance Agent (arXiv:2606.23032, v3 observed June 30, 2026) extends Finance Agent v2 from 10-K/10-Q analysis into IPO S-1 diligence. That is a real domain boundary, not a narrower filing subtype. S-1 work combines first- time disclosure, governance/control structure, pro forma and common-control accounting, capital-formation narrative, underwriting-sensitive risk factors, and valuation without a public trading history.

The benchmark contains 1,000 IPO-diligence questions, with 70 public questions over the SpaceX / SPCX S-1 and the rest private for contamination control. The question taxonomy covers seven domains: segment economics and operating performance, KPI quality and monetization, governance and control, accounting and common-control recasts, capital intensity and funding, execution and program risk, and valuation / underwriting analysis. It also labels each question by professional workflow: investment banking / ECM, public-market investing, venture or growth investing, credit analysis, securities counsel, and accounting / transaction advisory.

The most important finding is harness-related. The paper reports that the original public Finance Agent v2 harness failed to complete evaluation on the long SpaceX S-1 because naive chunk retrieval could not surface cross-referenced evidence. IPO Finance Agent replaces it with contextual retrieval and adds an automated rubric-construction loop: fact extraction from model-answer ensembles, agreement-based consolidation, rubric induction, deterministic schema validation, quality evaluation, repair / enrichment routing, deduplication, and final human expert review.

StateBench implication: finance evals need a separate IPO / S-1 diligence lane. Do not let good periodic-filing scores generalize silently to capital-markets transaction work. The reusable benchmark pattern is long-document contextual retrieval plus workflow-labeled rubrics with explicit contamination controls. Treat the reported model leaderboard as a reproduction target, not as a current StateBench score, because model names, private tasks, retrieval details, and judge implementation must be pinned before comparison.

8g. Open-ended financial research needs rubric-dense public cases

Samaya Frontier Finance v3.1, observed from Samaya’s benchmark app on July 12, 2026, is a separate benchmark from the Kensho / S&P Global / MIT FrontierFinance spreadsheet-modeling paper. Samaya’s benchmark targets open-ended financial research queries scored against expert rubric items.

The public split contains 220 queries, 11,543 rubric criteria, and 7,487 must-have criteria. The site bundle reports a balanced difficulty distribution of 74 hard, 73 medium, and 73 easy queries. The task mix is broader than SEC filing QA: financial data and modeling, sector / industry / macro research, earnings and events, company research, coverage and catalyst monitoring, and screening / discovery. Capability labels include qualitative synthesis, temporal retrieval, numerical reasoning, multi-hop workflows, cross-entity retrieval, thematic retrieval, non-disclosure checks, and claim verification.

This is valuable because the public JSONL exposes the shape of the grading contract: each query has use-case labels, capability labels, query date, and rubrics with required / optional status, rubric type, source type, subtype, and detailed source category. That makes it easier to translate benchmark cases into StateBench-style eval packets than benchmarks that only publish aggregate model scores.

Important caveat: the reported leaderboard should be treated as a reproduction target. The mirrored system-performance.json lists Samaya System at 50.8272%, Finance Agent v2 Fable 5 at 49.1922%, Opus 4.8 at 45.0399%, GPT 5.5 at 43.4736%, and general web-search baselines below those rows. Those are benchmark-page reported rows, not locally reproduced StateBench scores. The GitHub grader, model versions, retrieval configuration, and hidden/private split behavior must be pinned before using the leaderboard for procurement.

StateBench implication: this belongs in the broad financial-research lane between BigFinanceBench and Hedge-Bench. It is not an Excel artifact benchmark, not an IPO/S-1 diligence benchmark, and not a firm-wide meta-benchmark. Its reuse value is the rubric schema and public examples for open-ended research queries with temporal, cross-entity, and synthesis requirements.

8h. Financial-services evals need business-domain rollups

Meta-Benchmarks for Financial-Services LLM Evaluation (arXiv:2607.01740, July 2026) does not introduce a new task benchmark. It introduces a rollup method: map 452 public benchmarks into 41 O*NET Generalized Work Activities, then map those work activities into 38 BIAN banking business domains.

The scoring method is useful for firm-wide eval dashboards. Each benchmark is weighted by discrimination * coverage * recency, so saturated legacy tests and rarely reported niche tests fade automatically. Those weights scale the K-factor in a pairwise Elo tournament, avoiding raw-score normalization across benchmarks with different scales and difficulty. Domain scores are then weighted averages of work-activity Elo scores.

The paper’s strongest warning is coverage, not ranking. Only 24 of the 41 O*NET work activities have public benchmark evidence; 17 activities, mostly physical and managerial/interpersonal activities such as staffing, coaching, negotiating, and operating vehicles, have no public LLM benchmark coverage. In banking terms, public benchmarks heavily cover cognitive information processing and leave many real operating activities unmeasured.

StateBench implication: use public benchmark meta-scores only for candidate screening, concentration-risk monitoring, and capability drift. Final approval still requires private repo-local eval packets over the firm’s workflows. The right enterprise architecture is a two-layer system: public benchmark rollups for broad model narrowing, plus internal task evals for production authority.

9. Recency is a separate finance capability

AFIB / SuperInvesting is weaker as independent model evidence because it is a vendor/product benchmark. It is useful because it treats financial intelligence as multi-dimensional: factual accuracy, analytical completeness, data recency, consistency, hallucination resistance, and failure patterns. The paper’s own limitation language is important: model rankings are a snapshot and will change as retrieval/data integrations change.

StateBench implication: date freshness must be a scored field. A model that uses stale financial periods should fail even if the reasoning structure is plausible.

10. Multilingual and multimodal finance is no longer optional

FinMMEval CLEF 2026 adds a global-finance lane. It combines multilingual financial exam QA, PolyFiQA over filings plus multilingual news, and financial-decision tasks over BTC and TSLA daily contexts. This matters because public finance AI benchmarks are still too English/text-heavy while actual markets involve multilingual disclosures, news, regulatory regimes, charts, prices, and portfolio state.

StateBench implication: add a later multilingual/multimodal source pack with at least one filings-plus-news task and one price/news/position decision task.

11. Reward models should be finance-specific

Fin-PRM is not a trading model; it is a process reward model for financial reasoning. Its value is the supervision pattern: step-level and trajectory-level labels built from expert financial analyses, then used for offline trace selection, Best-of-N inference, and GRPO reward shaping.

StateBench fixture: sa-050 materializes Fin-PRM as a source-acquisition task for finance process supervision, not as a model-ranking or deployment proof.

StateBench implication: when converting repo traces into training data, do not only score final answer quality. Score intermediate evidence selection, calculation steps, date handling, and financial concept fidelity.

12. Source ingestion needs visual and critical-field evals

FinChart-Bench and FinCRITICAL-ED expand the ingestion layer. FinChart-Bench pushes financial chart/VLM evaluation; FinCRITICAL-ED scores OCR at the level that matters in finance: reporting dates, numbers, monetary values, table headers, concepts, and critical fields.

StateBench fixture: sa-048 materializes FinChart-Bench as the financial chart VLM source-acquisition task.

StateBench fixture: sa-049 materializes FinCRITICAL-ED / FinCriticAED as the financial fact-level OCR source-acquisition task.

StateBench implication: the source-ingestion eval should include chart pages, scanned filings, tables, and investor-deck exhibits. OCR is not “done” when text is extracted; it is done when financially critical fields survive with correct values and context.


Key Data Points

Benchmark Best Model Score Human Baseline Gap Date Credibility
Finance Agent Benchmark 46.8% (o3) ~90%+ −43pp+ May 2025 HIGH
FinSearchComp (global) 68.9% (Grok 4 web) 75.0% −6.1pp Sep 2025 HIGH
FinSearchComp (Greater China) ~70% (DouBao web) 88.3% −18pp Sep 2025 HIGH
FinMathBench (4-formula) 14.0% (GPT-4o) ~95%+ −81pp AAAI-26 HIGH
FinToolBench (compliance) Binary pass ≠ compliance Unknown Mar 2026 HIGH
RealFin (NOTA) Systematic overconfidence Apr 2026 HIGH
Spreadsheet-RL (domain) 17.2% (Qwen3 + RL) ~80%+ −63pp May 2026 MEDIUM-HIGH
WorkstreamBench End-to-end finance spreadsheets Claude family leads qualitatively, but strongest agents still fall short; frozen as sa-105 Professional workbook gap May 21 2026 MEDIUM-HIGH
BlueFin 131 finance workbook tasks / 3,225 rubric criteria Strongest systems below 50% average score; frozen as sa-106 Dynamic correctness gap May 29 2026 MEDIUM-HIGH
DataClawBench 492 exploratory tasks / 2.06M records 63.4% Claude Opus 4.6 overall; 7 of 8 models below 50% Unguided data-agent gap May 2026 MEDIUM-HIGH
Deep FinResearch Bench Professional investment-research reports Frontier deep-research agents below professional reports; frozen as sa-107 Rigor / valuation / verifiability gap Apr 22 2026 MEDIUM-HIGH
BigFinanceBench 928 auditable financial-research tasks Best reported systems: 58.8% rubric, below 45% answer accuracy; frozen as sa-127 Derivation / source / accounting-step gap Jun 2 2026 MEDIUM-HIGH
Hedge-Bench 102 hedge-fund analyst reasoning tasks Frontier agents below 16% pass@1; frozen as sa-126 Expert-move / grounding / hallucination gap Jun 2 2026 MEDIUM-HIGH
IPO Finance Agent 1,000 IPO S-1 diligence questions; 70 public SpaceX questions Paper reports contextual-retrieval harness beating Finance Agent v2 comparison rows S-1 / long-filing / transaction-diligence gap Jun 30 2026 v3 MEDIUM
Samaya Frontier Finance v3.1 220 public open-ended research queries / 11,543 rubric criteria Reported top row 50.8272%; public JSONL exposes required/optional rubric schema Open-ended research / temporal retrieval / cross-entity synthesis gap Jul 12 2026 observed MEDIUM-HIGH
Meta-Benchmarks for Financial Services 452 public benchmarks mapped to 41 O*NET work activities and 38 BIAN domains Public meta-evidence for candidate screening, not production approval Firm-wide rollup / coverage gap Jul 2026 MEDIUM-HIGH
FinReasoning 4,800 financial research-report eval items Three tracks: semantic consistency, data alignment, deep insight; 12-indicator rubric Audit/correction vs generation role-assignment gap Feb/May 2026 MEDIUM-HIGH
QuantEval strategy coding Executability + metric error Human experts high Large gap Jan 2026 MEDIUM-HIGH
QFBench 87 executable quantitative-finance coding tasks reported on public V11 site Live leaderboard requires local reproduction Quant-code implementation / execution gap May 2026 public V11 MEDIUM-HIGH
AFIB / SuperInvesting Vendor system leads in paper Vendor benchmark caveat Mar 2026 MEDIUM
FinMMEval CLEF shared-task framework Results pending/maturing Feb 2026 MEDIUM-HIGH
Fin-PRM +3.3pp GRPO vs outcome-only reward in paper Method signal Aug 2025 MEDIUM
FinChart-Bench VLM/chart-specific OCR/VLM lane Jul 2025 MEDIUM
FinCRITICAL-ED Critical-field OCR Ingestion fidelity lane Nov 2025 MEDIUM-HIGH

What This Means for Your Organization

The picture these benchmarks draw is specific: AI tools are demonstrably useful for financial analysis tasks that are well-defined, information-retrieval-forward, and human-supervised. The $3.79 versus $25.66 cost-per-task figure from Finance Agent Benchmark will improve as models improve and costs fall. The 46.8% accuracy ceiling on complex tasks will also move — but it has not crossed 50% yet, and the multi-formula collapse documented in FinMathBench reveals a structural problem that raw scaling does not immediately solve.

For an organization making deployment decisions today, three questions matter:

What is the error cost? On tasks where a wrong answer is recoverable — first-draft summaries, initial screening, research triage — current accuracy levels are useful. On tasks where a wrong answer creates regulatory exposure, client harm, or financial loss, the models in these benchmarks are not yet reliable enough to deploy without review.

Where is the human in the loop? The RealFin finding on overconfidence is the most operationally relevant: models do not know what they do not know in financial contexts. Designs that assume the model will flag its own uncertainty are not validated by the evidence. Human review at the output layer — not just the input layer — is the safer architecture.

Is the problem structured or exploratory? The DataClawBench gap between structured and exploratory performance is large. If the financial analysis task begins with “here is the relevant data,” models are more reliable. If it begins with “find the relevant data,” reliability drops substantially. Most high-value financial analysis is the second kind.

If this landscape raises specific questions about how to structure AI deployment in financial workflows at your organization, I’d welcome the conversation — brandon@brandonsneider.com.


Sources

  1. Finance Agent Benchmark — Krishnan, Wu, Nashold (Vals AI + Stanford). arXiv:2508.00828. May 2025. n=537 expert-authored questions. URL: https://arxiv.org/abs/2508.00828. Credibility: HIGH (independent expert validation, post-2024 source documents, agentic harness with real tool access; Vals AI commercial connection noted).

  2. FinSearchComp — ByteDance Seed + Columbia Business School. arXiv:2509.13160. September 16, 2025. n=635 questions, 21 models, 70 expert annotators. URL: https://arxiv.org/abs/2509.13160. Credibility: HIGH.

  3. FinMathBench — He, Wang, Xiong, Chen, Hu (Ant Group). AAAI-26 proceedings. n=946 questions, 40 models. Credibility: HIGH (peer-reviewed AAAI conference; Ant Group is a practitioner institution, not a benchmark-marketing company).

  4. RealFin — Dai, Lin, Xie, Wang (INSAIT + Newcastle + MBZUAI). arXiv:2602.07096. April 26, 2026 (v2). n=2,020 bilingual questions, 15 models. URL: https://arxiv.org/abs/2602.07096. Dataset: https://github.com/insait-institute/RealFin. Credibility: HIGH.

  5. FinToolBench — Lu, Wang, Wang, Tang (Shanghai AI Lab + Tencent). arXiv:2603.08262. March 9, 2026. n=760 tools, 295 queries. URL: https://arxiv.org/abs/2603.08262. Credibility: HIGH.

  6. FIRE — Zhang, Wu, Guo et al. (Du Xiaoman Technology + Tsinghua + Renmin University). arXiv:2602.22273. February 25, 2026. n=17,000+ questions. URL: https://arxiv.org/abs/2602.22273. Credibility: HIGH (Du Xiaoman is a major Chinese fintech; institutional practitioner credibility). StateBench fixture sa-055 materializes it as a broad finance-knowledge and business-scenario source-acquisition task.

  7. Spreadsheet-RL — Chi, Xie, Wu et al. (UIUC + Meta). arXiv:2605.22642. May 21, 2026. n=1,660 domain tasks. URL: https://arxiv.org/abs/2605.22642. Credibility: MEDIUM-HIGH (Meta co-authorship; methodology is sound; RL improvements real but magnitude depends on base model).

  8. WorkstreamBench — Yen, Poeltl, Gear et al. arXiv:2605.22664. May 21, 2026. URL: https://arxiv.org/abs/2605.22664. Local OCR: research/21-benchmarks/raw/workstreambench-finance-spreadsheet-agents-2605.22664.chandra.txt. Credibility: MEDIUM-HIGH; use for finance workbook artifact scoring and professional-standard caveats.

  9. BlueFin — Kundurthy, Na, Moraine et al. arXiv:2605.30907. May 29, 2026. n=131 workbook tasks, 3,225 rubric criteria. URL: https://arxiv.org/abs/2605.30907. Local OCR: research/21-benchmarks/raw/bluefin-financial-spreadsheet-agents-2605.30907.chandra.txt. Credibility: MEDIUM-HIGH; use for workbook synthesis/manipulation/comprehension and dynamic-correctness caveats.

  10. DataClawBench — Zhang, Ye, Chen et al. (Sun Yat-sen University). arXiv:2605.02503. May 27, 2026. n=492 tasks, 2.06M financial records. URL: https://arxiv.org/abs/2605.02503. Credibility: MEDIUM-HIGH (very recent; peer review pending; methodology is novel and rigorous by preprint standards).

  11. Deep FinResearch Bench — Haque, Papadimitriou, Mensah et al. arXiv:2604.21006. April 22, 2026. URL: https://arxiv.org/abs/2604.21006. Local OCR: research/21-benchmarks/raw/deep-finresearch-bench-2604.21006.chandra.txt. Credibility: MEDIUM-HIGH; use for professional investment-research report quality, valuation, forecasting, and verifiability guardrails.

  12. BigFinanceBench — Wang, Meinhardt, Katz, Kim, Chaudhary, Blagden, Xu. arXiv:2606.03829. June 2, 2026. n=928 workflow-grounded financial-research tasks with 15,656 rubric criteria and 36,241 rubric points. URL: https://arxiv.org/abs/2606.03829. Local OCR: research/21-benchmarks/raw/bigfinancebench-workflow-grounded-financial-research-agents-2606.03829.chandra.txt. Credibility: MEDIUM-HIGH for benchmark design; model rankings require local reproduction because the reported scoring stack uses LLM judges, and the stated public subset/harness links were not accessible from this environment on 2026-06-06. StateBench fixture sa-127 materializes it as the auditable financial-research derivation source-acquisition task.

  13. Hedge-Bench — Cho, Huang, Lu, Lyu. arXiv:2606.03918. June 2, 2026. n=102 hedge-fund analyst reasoning tasks grounded in explicit professional reasoning traces. URL: https://arxiv.org/abs/2606.03918. Repository: https://github.com/Trata-Inc/trata-hedge-bench. Local OCR: research/21-benchmarks/raw/hedge-bench-financial-reasoning-agents-2606.03918.chandra.txt. Credibility: MEDIUM-HIGH for benchmark design; model rankings require local reproduction because the reported scoring stack includes an LLM-as-judge layer. StateBench fixture sa-126 materializes it as the hedge-fund analyst-reasoning source-acquisition task.

  14. IPO Finance Agent — Benhenda. arXiv:2606.23032. v3 observed June 30, 2026. n=1,000 IPO-diligence questions, with 70 public SpaceX / SPCX S-1 questions and the rest private for contamination control. URL: https://arxiv.org/abs/2606.23032. Repository: https://github.com/benstaf/ipoagent. Local OCR: research/21-benchmarks/raw/ipo-finance-agent-2606.23032.chandra.txt. Source card: sources/21-benchmarks/ipo-finance-agent-ipo-diligence-raw.md. Credibility: MEDIUM for reported leaderboard values; MEDIUM-HIGH for benchmark-design signal around S-1 diligence, contextual retrieval, professional workflow labels, and rubric-generation controls.

  15. Samaya Frontier Finance v3.1 — Samaya AI. Observed July 12, 2026. n=220 public open-ended financial-research queries, 11,543 rubric criteria, 7,487 must-have criteria, and use-case / capability labels. URL: https://research.samaya.ai/benchmarks/frontier-finance. Dataset: https://huggingface.co/datasets/samaya-ai/FrontierFinance. Repository: https://github.com/samaya-ai/frontier-finance. Source card: sources/21-benchmarks/samaya-frontier-finance-v3-1-raw.md. Mirrored public JSONL: sources/21-benchmarks/samaya-frontier-finance-v3-1-public.jsonl. Credibility: MEDIUM-HIGH for public benchmark shape and rubric schema; MEDIUM for leaderboard values until locally reproduced.

  16. Meta-Benchmarks for Financial-Services LLM Evaluation — Hudson. arXiv:2607.01740. July 2026. Maps 452 public benchmarks into 41 O*NET Generalized Work Activities and 38 BIAN banking business domains using discrimination / coverage / recency weighted pairwise Elo. URL: https://arxiv.org/abs/2607.01740. Local OCR: research/21-benchmarks/raw/meta-benchmarks-financial-services-llm-evaluation-2607.01740.chandra.txt. Source card: sources/21-benchmarks/meta-benchmarks-financial-services-llm-evaluation-raw.md. Credibility: MEDIUM-HIGH for rollup architecture; business-domain scores remain candidate-screening evidence only.

  17. FinReasoning — Zhu, Jiang, Xu, Yao, Cheng, Ding, Xu. arXiv:2603.19254. February 25, 2026 / online May 8, 2026. Reports 4,800 Chinese financial research-report evaluation items across Semantic Consistency, Data Alignment, and Deep Insight, with GitHub code/data structure for Alignment, Consistency, and Depth evaluators. URL: https://arxiv.org/abs/2603.19254. Repository: https://github.com/TongjiFinLab/FinReasoning. Source card: sources/21-benchmarks/finreasoning-hierarchical-financial-research-reporting-raw.md. Credibility: MEDIUM-HIGH for hierarchy and schema; model-ranking claims require local reproduction and language/domain caveats.

  18. QuantEval — arXiv:2601.08689. January 2026. n=1,575 samples, including 60 CTA-style strategy-coding tasks. Local OCR: research/21-benchmarks/raw/quanteval-financial-quantitative-tasks-2601.08689.chandra.txt. Credibility: MEDIUM-HIGH. StateBench fixture sa-056 materializes it as the execution-based quant-strategy coding source-acquisition task.

  19. AFIB / SuperInvesting AI — arXiv:2603.08704. March 2026. Local OCR: research/21-benchmarks/raw/afib-superinvesting-financial-intelligence-2603.08704.chandra.txt. Credibility: MEDIUM; useful for dimensions/failure modes, vendor-leading scores require caution. StateBench fixture sa-058 materializes it as a vendor-claim and financial-analysis source-acquisition task.

  20. FinMMEval CLEF 2026 — arXiv:2602.10886. February 2026. Local OCR: research/21-benchmarks/raw/finmmeval-clef-2026-2602.10886.chandra.txt. Credibility: MEDIUM-HIGH for benchmark design. StateBench fixture sa-057 materializes it as the multilingual and multimodal finance source-acquisition task.

  21. Fin-PRM — arXiv:2508.15202. August 2025. Local OCR: research/21-benchmarks/raw/fin-prm-process-reward-model-2508.15202.chandra.txt. Credibility: MEDIUM for process-supervision method; model family is aging.

  22. FinChart-Bench — arXiv:2507.14823. July 2025. Local OCR: research/21-benchmarks/raw/finchart-bench-financial-chart-vlm-2507.14823.chandra.txt. Credibility: MEDIUM for chart/VLM lane.

  23. FinCRITICAL-ED — arXiv:2511.14998. November 2025. Local OCR: research/21-benchmarks/raw/fincriticaled-financial-fact-ocr-2511.14998.chandra.txt. Credibility: MEDIUM-HIGH for financial OCR/critical-field ingestion evaluation.

  24. QFBench / Quantitative Finance-Bench — public benchmark site and GitHub task repository, V11 snapshot dated May 4, 2026, fetched 2026-06-01 in source card. URL: https://qfbench.com/. Source card: sources/21-benchmarks/qfbench-quantitative-finance-agent-benchmark-raw.md. Credibility: MEDIUM-HIGH for benchmark design and task catalog; model leaderboard requires local reproduction before use as a StateBench score. StateBench fixture sa-083 materializes it as executable quant-code benchmark source acquisition.


Brandon Sneider | brandon@brandonsneider.com May 2026