Finance AI · 20 Independent Evaluations · June 2026

The Finance AI Evidence Base —
Every Score Sourced

Twenty independent evaluations spanning SEC filing analysis, financial search, spreadsheet automation, hedge-fund reasoning, quant coding, OCR, embeddings, tool-use compliance, and multilingual finance. No vendor benchmarks. No invented numbers.

Benchmark Field Guide →
Top Finding

No model crosses 50% on realistic financial research tasks. The best published result is o3 at 46.8% on Finance Agent Benchmark (May 2025 leaderboard — newer frontier models not yet re-evaluated on this task). Human experts on the same tasks: >90%.

Vals AI + Stanford · arXiv:2508.00828 · May 2025 · n=537 · EDGAR + web access
Finance AI Benchmark Landscape
20 benchmarks mapped by task type and best model score · size = benchmark credibility

The horizontal axis runs from retrieval (find a fact) through reasoning (combine facts) to execution (produce a verifiable artifact). Most finance AI tasks cluster bottom-left: complex reasoning, low scores. Quant coding is the lone outlier where models consistently cross 50%. OCR and embeddings are effectively solved.

SEC / Research Search / Knowledge Spreadsheets / Data Math Reasoning Tool Use Hedge / Quant OCR / Embeddings

SEC Filing & Complex Research

Vals AI + Stanford · arXiv:2508.00828 · May 2025 · n=537 · EDGAR + web access
⚠ Leaderboard last updated May 2025 — not re-run on models released since then

The most demanding finance agent evaluation: multi-hop SEC filing retrieval, web search, and quantitative reasoning in a single task. Every frontier model available in May 2025 was tested; none crossed 50%. The ceiling has not moved in subsequent smaller evaluations.

No model crossed 50% as of May 2025. Cost: o3 at $3.79/task vs ~$25.66 human analyst. The economic case for AI assistance is real; the case for autonomous deployment is not.
Rogo + OpenAI · arXiv:2606.03829 · Jun 2026 · n=928 tasks · 15,656 rubric criteria

Rubric-graded scoring exposes a critical gap: models that appear to pass binary tests fail when graded on derivation quality. The two bars below measure the same systems on the same tasks — only the grading method changes.

Best rubric score 58.8%; best answer accuracy below 45%. Derivation-level grading exposes gaps that binary pass/fail hides.
Deep FinResearch Bench (arXiv:2604.21006, Apr 2026) — evaluates professional investment-research report quality: rigor, forecasting/valuation, verifiability. Qualitative finding: frontier models consistently fail explicit uncertainty quantification and assumption checks. No clean per-model leaderboard scores published.

Spreadsheet & Workbook Automation

UIUC + Meta · May 2026 · n=1,660 tasks · finance, supply chain, HR, IB, asset management

RL post-training with GRPO boosts Excel task completion by 5–9pp versus base models. Finance tasks are the hardest domain because formula dependencies chain across cells — the same multi-step failure mode that collapses FinMathBench scores from 73% to 14%.

RL post-training (GRPO) gains ~5–9pp over base models. 17% pass rate on finance-specific tasks still means 83% failure. Formula dependencies and multi-step workflows make finance the hardest domain in the benchmark.
WorkstreamBench (arXiv, May 2026) — end-to-end finance spreadsheet workstreams (income statement build, ratio analysis, DCF). Claude family leads qualitatively; no model reliably completes a full professional workbook without errors.
BlueFin (arXiv, May 2026 · n=131 tasks · 3,225 rubric criteria) — professional finance workbook synthesis, manipulation, and comprehension. Strongest systems score below 50% average across rubric criteria.

Exploratory Data Analysis

Sun Yat-sen University · arXiv:2605.02503 · n=492 tasks · 2.06M real financial records

Tasks require a model to independently plan and execute multi-step analysis over real financial record sets — no scaffolding, no prescribed steps. The frontier/non-frontier gap is stark: this is a task class where model capability tier directly determines usability.

Claude Opus 4.6 leads at 63.4%; 7 of 8 models below 50%. Frontier models are the appropriate choice for exploratory data analysis — small models are not competitive on unguided, open-ended tasks over large real datasets.

Math Reasoning

Ant Group · AAAI-26 · n=946 questions · 40 models · 148 domain-specific formulas
⚠ Leaderboard last updated AAAI-26 (GPT-4o era) — not re-run on models released since then

As formula chains lengthen, accuracy collapses at each step. This is not a knowledge gap — GPT-4o knows the formulas. It's a chaining failure: each additional step multiplies the error probability. Real finance calculations (derivative pricing, rebalancing with constraints, layered risk) require 3–5 linked steps.

−58.9pp from one formula to four. Any finance calculation requiring 3+ linked steps is in the zone where AI error rates approach 75–85%. Single-formula accuracy is a demo metric.
Higher is better · accuracy % · Microsoft Research arXiv:2504.21233

Small open-weight reasoning models now match or beat frontier API models on structured math. The 3.8B Phi-4-mini leads — a result that changes the build-vs-buy calculus for any finance workflow involving formula chains.

3.8B Phi-4-mini-reasoning leads at 94.6% — beating o1-mini (frontier API) by 4.6pp. Small reasoning models are production-ready for single-formula structured math.
Higher is better · accuracy % · competition-level problems requiring multi-step proof chains

AIME tests competition-grade mathematical reasoning — closer to FinMathBench's multi-formula chains than to single-step recall. The same small model that leads MATH-500 nearly matches o1-mini here at a fraction of API cost.

Phi-4-mini-reasoning closes to within 6pp of o1-mini at a fraction of API cost and zero data-residency risk.
🔬
Fin-PRM (Qwen DianJin · arXiv:2508.15202 · Aug 2025) — Process Reward Model training signal for financial math. Training a verifier on intermediate reasoning steps (not just final answers) adds +3.3pp over GRPO outcome-only training on FinMathBench. This is a training methodology finding, not a task benchmark: any team fine-tuning a finance math model should use step-level rewards. Models evaluated on FinMathBench's task scores above are not Fin-PRM trained.

Tool Use

FinToolBench (Shanghai AI Lab + Tencent, Mar 2026 · n=295 queries · 760 tools) — three compliance metrics: Tool Method Reliability (TMR), Invocation Method Reliability (IMR), and Decision-Making Reliability (DMR). Key finding: models pass binary success/fail tests but fail regulatory validity checks. Binary pass ≠ compliance.
BFCL v3: Team-ACE arXiv:2409.00920 · ICLR 2025 · Finance routing: internal fixture · 10 tasks · Apple Silicon MPS

Parameter count is a hard floor for structured output reliability. On BFCL v3 (Jun 2026), the leaderboard leader is GLM 4.5 at 77.8% — the 97% figure for ToolACE-8B was from the original BFCL v1/v2 eval. Granite 4.1 8B at 68% and FunctionGemma-270M at 0% show the lower end of the spectrum.

BFCL v3 current leader: GLM 4.5 at 77.8% (Jun 2026). Granite 4.1 8B scores 68.3%. FunctionGemma-270M (4-bit MLX, internal finance fixture) scored 0/10 — all tasks produced invalid JSON. 270M parameters is below the threshold for reliable structured output on complex tool schemas.

Hedge Fund & Quant Coding

Trata + BYU + Osmosis · arXiv:2606.03918 · Jun 2026 · n=102 tasks grounded in professional reasoning traces · pass@1 · best frontier agent per category

Tasks are grounded in actual hedge fund analyst reasoning traces: the benchmark knows what a professional would do and grades deviation from it. No agent crosses 20% in any single category — showing that the failure is not domain-specific but structural.

Every frontier agent scores below 16% pass@1. Tasks cover valuation, M&A, competitive positioning, operational strategy, and risk — work sitting between source retrieval and investment memo drafting. Finding the right source is not enough; agents must execute the analyst reasoning move and avoid unsupported synthesis.
Public benchmark · Docker sandbox · V11 snapshot 2026-05-04 · n=87 tasks · 42 models · 10,962 runs · strict numerical pytest verifiers

The only finance benchmark where models consistently cross 50%: executable quantitative code that passes numerical pytest verifiers. Options pricing, factor construction, and credit modeling are tractable for frontier models when the task is code, not prose.

Best pass@1 is 61.7% on executable quant coding tasks — higher than most finance benchmarks but this is code implementation, not open-ended research. Tasks span options pricing, risk metrics, factor construction, portfolio, credit, and SEC-event workflows.
QuantEval (arXiv:2601.08689, Jan 2026 · n=1,575 · 60 CTA-style strategy-coding tasks) — models can describe factor strategies and generate plausible-looking code, but cannot close the loop to a verified, backtested, numerically accurate strategy artifact. A response that describes a strategy is not a deliverable.

OCR & Embeddings

Allen Institute for AI · arXiv:2502.18443 · 1,400 PDF pages · 7,000+ binary unit tests · deterministic scoring

The top 5 models are all small, open-weight, purpose-built OCR systems — not frontier API models. GPT-4o and Claude 3 do not appear in the top tier. For financial document ingestion (10-Ks, prospectuses, filings), dedicated OCR models outperform general-purpose LLMs by a wide margin.

Chandra 2 (5B, Datalab) leads at 85.9 — the benchmark olmOCR itself created. All top models are small, open-weight, purpose-built. Active platform pipeline: mlx-community/chandra-4bit on Apple Silicon. No API required.
Higher is better · average across 56 tasks · Alibaba arXiv:2506.05176

Open-weight embedding models now outperform both Gemini Embedding and OpenAI text-embedding-3-large on multilingual tasks. For financial RAG over SEC filings, earnings transcripts, and regulatory documents, the build-vs-API decision has shifted.

An Apache 2.0, 8B open-weight model outranks both Gemini Embedding and OpenAI text-embedding-3-large. Routing embedding through commercial APIs when open-weight models are competitive is economically indefensible.
239M parameters · Apache 2.0 · Jina AI · Jina-v5-nano vs OpenAI text-embedding-3-large

At 239M parameters, jina-v5-nano beats OpenAI's flagship on MTEB English while running entirely on-device. For financial RAG pipelines ingesting large document volumes, this is the cost-efficiency argument — API latency and per-call costs eliminated at zero quality cost.

239M parameters beats OpenAI's flagship embedding model on MTEB English. For financial RAG pipelines with large document volumes, this is the cost-efficiency argument in one number.

Multilingual Finance

FinMMEval CLEF 2026 (Feb 2026) — multilingual and multimodal finance tasks: exam QA, PolyFiQA, and financial decision making across languages. Performance degrades significantly on non-English financial questions. Models trained predominantly on English financial text miss regulatory and market terminology specific to non-English jurisdictions. For multinational deployments, language-specific fine-tuning or retrieval augmentation is required.

Finance Benchmark Inventory

BenchmarkInstitutionDatenBest ScoreCredibility
Finance Agent BenchmarkVals AI + StanfordMay 2025537 q46.8% (o3)HIGH
FinSearchCompByteDance Seed + ColumbiaSep 2025635 q · 21 models68.9% (Grok 4 web)HIGH
FinMathBenchAnt GroupAAAI-26946 q · 40 models72.9% → 14.0% collapseHIGH
RealFinINSAIT + Newcastle + MBZUAIApr 20262,020 q · 15 modelsSystemic overconfidenceHIGH
FinToolBenchShanghai AI Lab + TencentMar 2026760 tools · 295 qBinary pass ≠ complianceHIGH
FIREDu Xiaoman + Tsinghua + Renmin UFeb 202617,000+ questions~65–70% (certification)HIGH
DataClawBenchSun Yat-sen UniversityMay 2026492 tasks63.4% (Claude Opus 4.6)MED-HIGH
BigFinanceBenchRogo + OpenAIJun 2026928 tasks · 15,656 rubrics58.8% rubricMED-HIGH
Hedge-BenchTrata + BYU + OsmosisJun 2026102 tasks<16% pass@1MED-HIGH
Spreadsheet-RLUIUC + MetaMay 20261,660 tasks23.4% (RL GRPO)MED-HIGH
QFBenchPublic · Docker sandboxMay 202687 tasks · 42 models61.7% pass@1MED-HIGH
WorkstreamBencharXiv preprintMay 2026End-to-end workstreamsQualitativeMED-HIGH
BlueFinarXiv preprintMay 2026131 tasks · 3,225 rubrics<50% averageMED-HIGH
Deep FinResearcharXiv preprintApr 2026QualitativeQualitativeMED-HIGH
QuantEvalarXiv preprintJan 20261,575 samples · 60 CTA tasksQualitativeMED-HIGH
FinMMEval CLEF 2026CLEF labFeb 2026MultilingualQualitativeMED-HIGH
olmOCR-BenchAllen Institute for AI20251,400 pages · 7,000+ tests85.9 (Chandra 2)HIGH
MTEB MultilingualMTEB Leaderboard202656 tasks70.58 (Qwen3-8B)HIGH
BFCL v3Berkeley / Team-ACEICLR 2025Function callingSotA (ToolACE-8B)HIGH
Finance Tool RoutingInternal · Apple Silicon MPSJun 202610 tasks0/10 (FunctionGemma-270M)MED-HIGH
Red flags: MMLU, GSM8K, or BIG-Bench Hard as primary evidence — frontier models score 90%+ and cannot be discriminated. "Internal benchmark" without published methodology is not reproducible.