← Finance AI Benchmarks
Finance AI · 20 Benchmark Study Cards · June 2026

Benchmark Field Guide

Every benchmark on the Finance AI page, explained: what it measures, who made it, why it matters for production AI decisions, and where it currently stands.

Benchmark Quality — BS Rubric

Finance Benchmark Quality Audit — Which Benchmarks Are Actually Testing Finance Reasoning?
State of AI meta-evaluation · Jun 2026 · 8 benchmarks · 5-dimension BS rubric · Higher score = worse question quality
The problemMost positive "LLMs can trade!" results in the literature are contamination artifacts. Models don't reason about TSLA November 2020 — they recall it from training data. The field has a replication crisis it hasn't publicly acknowledged.
The rubric5 dimensions (1–5 each): Finance Depth (trivia vs. real reasoning) · Contamination Risk (training data vs. novel) · Real-world Validity (artificial vs. practitioner-realistic) · Eval Reliability (LLM judge vs. objective ground truth) · Domain Lock (anyone can Google vs. requires expertise). BS Score = 25 − sum. Higher = worse.
🔴 FLAG-TRADER (BS = 18)Worst in set. Example prompt: "AAPL at $150 on June 27, 2024 — Buy/Sell/Hold?" That date is in every model's training corpus. Tests whether a model can remember stock prices, not reason about markets. The "135M beats GPT-4" headline is almost certainly a contamination artifact.
🔴 InvestorBench (BS = 14)TSLA/AAPL 2020–2021 — the most contaminated stock data possible. GPT-3.5-labeled sentiment adds second-order contamination. Own results show Buy & Hold beats most LLMs: models aren't trading, they're guessing from memorized patterns.
🟡 FinMTEB (BS = 9)Best task: Fed monetary policy stance classification (BS = 6) — discriminates models, finance-specialist, practically motivated. Weakest: STS on annual report paraphrases (BS = 12) — mostly a general NLP paraphrase task with finance skin. BM25 beating dense embeddings on financial STS is the honest finding.
🟢 LiveTradeBench (BS = 6)The right methodology: live 50-day evaluation, continuous allocation vectors, real-time news, objective live returns. Key finding: LMArena scores don't predict trading performance. Primary weakness: 50 days is too short to separate alpha from luck.
🟢 StockBench (BS = 8 best task)Gets contamination right: March–July 2025 data, explicit temporal separation. Honest result: most LLMs fail to beat equal-weight buy-and-hold. The full multi-signal task (fundamentals + news + price on post-cutoff data) is the most realistic non-RL formulation in the literature.
Pattern: methodology predicts findingsRigorous methodology → uncomfortable result (LLMs don't work). Loose methodology → exciting positive result. Without exception across the 8 benchmarks reviewed. This is the field's blind spot.
The Haiku floor test (proposed)Run Claude Haiku (cheapest model) on a benchmark sample before investing in frontier model evals. If Haiku scores ≥ 50%, the benchmark is likely contaminated or trivial. A benchmark where Haiku scores 40% is probably 40% trivia. Contamination floor = free signal about benchmark quality.
What good finance tasks look likeObjective ground truth · Temporal separation from training data · Cross-asset reasoning required · Tasks a non-expert couldn't Google. Best examples from this review: LiveTradeBench's live multi-asset allocation · PredictionMarketBench's BTC daily-high contract · FinMTEB's Fed policy stance classification · FinSABER's bias-corrected stock selection across full S&P 500.

Live HF Leaderboards

Six active leaderboards on HuggingFace rank models on finance AI tasks. Charts below show the current top scores — click "Full leaderboard →" for the complete ranked table with all models.

FINOS Open Financial LLM Leaderboard
FINOS Foundation (Fintech Open Source) · 136 HF likes · Jun 2026
Full leaderboard →

Ranks models on financial language tasks across 6 languages: sentiment classification, named entity recognition, question answering, and document summarization. The only multilingual finance leaderboard with Chinese, Japanese, Spanish, and Greek coverage. GPT-4o leads on BloombergGPT-style tasks; the gap between #1 and #3 is under 2 points, meaning no model has a decisive edge on financial language understanding.

Average score · BloombergGPT task suite · June 2026 · top 3 of 11 models
BFCL v3 — Berkeley Function-Calling Leaderboard
UC Berkeley (gorilla-llm) · 125 HF likes · Apr 2026
Full leaderboard →

Measures whether a model can correctly call an API — choosing the right function, with the right parameter names and values, in the right format. This is the core skill for any AI workflow that connects to external tools: a finance model querying Bloomberg, pulling from a risk system, or writing to a trading platform. On BFCL v3, the current leader (GLM 4.5) scores 77.8% — meaning roughly 1 in 4 calls still fail. The 9-point gap between GLM 4.5 (77.8%) and Granite 4.1 8B (68%) shows how much model choice matters for tool-heavy deployments. Note: BFCL v4 exists but cross-model comparison coverage is still sparse.

Overall BFCL v3 score · % tasks with correct function + parameters · Jun 2026 · top 5
MTEB — Massive Text Embedding Benchmark
mteb team · 7,474 HF likes · continuously updated · 2,000+ models
Full leaderboard →

Ranks embedding models — the components that turn text into vectors for semantic search and RAG pipelines. Higher score means better at finding the right document when a user asks a question. Critical caveat: MTEB tests on general web/Wikipedia text. Financial documents (10-Ks, earnings calls, SOFR curves) have specialized vocabulary that degrades MTEB scores by 5–15 points. See FinMTEB in the Additional section for finance-specific rankings — the ordering of models often changes significantly.

Average score across 56 tasks · higher = better RAG retrieval · Jun 2026 · top 5 of 2,000+

SEC Research

Finance Agent Benchmark (FAB)
Vals AI + Stanford · arXiv:2508.00828 · May 2025 · n=537 tasks
⚠ Leaderboard last updated May 2025 — not re-run on models released since then
MeasuresAgentic multi-step financial research over live EDGAR filings and web. Tasks range from earnings extraction to competitive analysis.
DomainsSEC filings, earnings analysis, competitive research, financial web data
Why it mattersThe only benchmark using live EDGAR access with agentic tool use — closest to what a real equity analyst does.
Top modelso3 46.8%, GPT-4o 38.0%, Claude 3.5 37.0%
Key statNo model exceeds 50%; human expert baseline >90%
Eval methodsLLM-as-judgeRubric gradingHuman eval
Benchmark composition
BigFinanceBench (BFB)
FinAI Lab · arXiv:2505.00075 · May 2025 · n=7,000+ tasks
MeasuresComprehensive financial task completion across 6 task types: extraction, QA, classification, summarization, compliance, and math.
DomainsFinancial documents, NLP tasks, compliance, reporting
Why it mattersLargest publicly available finance benchmark by task count; reveals a ~15pp gap between rubric scoring and answer accuracy.
Top modelsBest system 58.8% rubric, <45% answer accuracy
Key stat~15pp gap between rubric and answer-accuracy scoring — surface-level compliance without correct answers
Example task"For EnviroStar (ticker: EVI), in December 2016, what was the weighted average sale price of Michael Steiner's stock dispositions? Round only the final answer to the nearest cent." — Answer: $14.21 · Rubric has 14 sub-criteria including identifying each individual trade date, shares, and price before computing weighted average
Eval methodsRubric gradingExact match
Benchmark composition
RealFin
INSAIT + Newcastle + MBZUAI · arXiv · Apr 2026 · n=2,020 questions
Measures"Knowing when not to answer" — removes essential premises from financial exam questions while keeping them linguistically plausible.
DomainsFinancial certification, credit analysis, risk assessment, compliance
Why it mattersSystemic overconfidence across all 15 models tested. For credit approval or compliance determinations, a model that answers underdetermined questions confidently is dangerous.
Top modelsAll 15 models tested show overconfidence
Key stat100% of models overconfident on underdetermined problems
Eval methodsMCQ / multiple choiceLLM-as-judge
Benchmark composition

Spreadsheets

Spreadsheet-RL
UIUC + Meta · arXiv · May 2026 · n=1,660 tasks
MeasuresExcel task automation across finance, supply chain, HR, IB, and asset management — with and without RL post-training.
DomainsExcel automation, investment banking, asset management, supply chain
Why it mattersQuantifies what RL post-training (GRPO) buys on real Excel tasks: +5–9pp. Shows finance is the hardest domain due to formula chaining.
Top modelsRL GRPO best 23.4% SpreadsheetBench, 17.2% finance
Key stat83% failure rate on finance-specific tasks even with RL post-training
Eval methodsExact matchNumerical pytest
Benchmark composition
WorkstreamBench
arXiv · May 2026
MeasuresEnd-to-end finance spreadsheet workstreams: income statement build, ratio analysis, DCF modeling.
DomainsFinancial modeling, DCF, income statements, ratio analysis
Why it mattersTests completion of a full professional workbook — not individual cells. No model reliably completes a full workbook without errors.
Top modelsClaude family leads qualitatively; no model completes reliably
Key statNo model produces a reliable professional workbook end-to-end
Eval methodsRubric gradingLLM-as-judge
Benchmark composition
BlueFin
arXiv · May 2026 · n=131 tasks · 3,225 rubric criteria
MeasuresProfessional finance workbook synthesis, manipulation, and comprehension graded by rubric.
DomainsFinancial workbooks, spreadsheet comprehension, professional finance
Why it matters3,225 rubric criteria across 131 tasks — most granular spreadsheet evaluation available.
Top modelsStrongest systems below 50% average across rubric criteria
Key stat<50% rubric score for all models on professional finance workbooks
Eval methodsRubric grading
Benchmark composition

Exploratory Analysis

DataClawBench
Sun Yat-sen University · arXiv:2605.02503 · May 2026 · n=492 tasks · 2.06M records
MeasuresUnguided multi-step financial data exploration over 2.06M real financial records — no scaffolding, no prescribed steps.
DomainsFinancial databases, time-series data, equity data, exploratory analysis
Why it mattersFrontier/non-frontier gap is stark. Small models are not competitive. This is the task class that most clearly separates Claude Opus from 7B models.
Top modelsClaude Opus 4.6 63.4%, 7 of 8 models below 50%
Key stat18pp gap between frontier (63.4%) and sub-frontier ceiling (<50%)
Example taskGiven 2.06M real financial records (no prescribed steps), independently plan and execute multi-step analysis to answer: identify all equity positions where the 90-day return exceeded the sector median by more than 2σ. No scaffolding provided.
Eval methodsExact matchLLM-as-judge
Benchmark composition

Math Reasoning

FinMathBench
Ant Group · AAAI-26 · n=946 questions · 40 models · 148 formulas
⚠ Leaderboard last updated AAAI-26 (GPT-4o era) — not re-run on models released since then. Current frontier models likely saturate single-formula tasks.
Benchmark quality caveatThis benchmark has real weaknesses. The formula-chain framing is contrived — in practice, financial formulas are embedded in messy data contexts, not presented cleanly as "apply formula X then Y." Scores are also GPT-4o era; o3/GPT-5/Claude 4 family likely saturate the 1-formula tasks. Treat as a directional signal, not a production readiness test.
What it does measureDegradation rate as formula complexity increases 1→4 steps. The collapse pattern (72.9% → 14%) is real even if the absolute numbers are stale — multi-step formula chaining remains the critical failure mode across all model generations.
DomainsDerivative pricing, portfolio math, risk calculations, financial formulas
Key stat72.9% → 14.0% collapse from 1 formula to 4 (−58.9pp) — the degradation shape, not the absolute numbers, is the useful finding
Eval methodsExact match
Benchmark composition
MATH-500
Microsoft Research · arXiv:2504.21233 · 500 competition math problems
MeasuresStructured mathematical reasoning across competition-level problem types.
DomainsAlgebra, geometry, number theory, combinatorics, competition math
Why it mattersSmall open-weight reasoning models now match frontier API on structured math — changes the build-vs-buy calculus for finance workflows.
Saturation warningMATH-500 is largely saturated by current frontier models — o3, GPT-5, Claude 4 family all score 97%+. The benchmark is most useful for comparing small open-weight models against each other, not for differentiating frontier models.
Top models (small models)Phi-4-mini 3.8B 94.6%, R1-Distill Qwen-7B 91.4%, o1-mini 90.0% — frontier models (o3, GPT-5, Claude 4) score 97%+, making this chart most useful for small model selection
Key stat3.8B Phi-4-mini at 94.6% — small models are production-ready for single-formula math tasks
Eval methodsExact match
Benchmark composition
AIME 2024
American Invitational Mathematics Examination · competition-level
MeasuresCompetition-grade mathematical reasoning requiring multi-step proof chains.
DomainsCompetition mathematics, multi-step proof, symbolic reasoning
Why it mattersClosest public proxy to FinMathBench's multi-formula chains. Small model cost advantage persists at harder difficulty.
Top modelso1-mini 63.6%, Phi-4-mini 3.8B 57.5%
Key statPhi-4-mini within 6pp of o1-mini at a fraction of API cost
Eval methodsExact match
Benchmark composition
Fin-PRM
Qwen DianJin · arXiv:2508.15202 · Aug 2025
MeasuresProcess Reward Model training signal quality for financial math — training methodology, not a task benchmark.
DomainsFinancial math training, GRPO, reward modeling
Why it mattersStep-level reward signals add +3.3pp over outcome-only training on FinMathBench. Any team fine-tuning a finance math model should use step-level rewards.
Top modelsN/A — training methodology paper
Key stat+3.3pp GRPO step-level vs. outcome-only rewards
Eval methodsNumerical pytest
Benchmark composition

Tool Use

FinToolBench
Shanghai AI Lab + Tencent · arXiv · Mar 2026 · n=295 queries · 760 tools
MeasuresThree compliance metrics: Tool Method Reliability (TMR), Invocation Method Reliability (IMR), Decision-Making Reliability (DMR).
DomainsFinancial APIs, tool selection, compliance, regulatory validity
Why it mattersModels pass binary success/fail but fail regulatory validity checks. Binary pass ≠ compliance — a critical distinction for financial services.
Top modelsLeaderboard not public; all models fail regulatory validity
Key statBinary pass ≠ regulatory compliance
Eval methodsRubric gradingExact match
Benchmark composition
BFCL v3
Team-ACE · arXiv:2409.00920 · ICLR 2025
MeasuresFunction calling accuracy across complex tool schemas — the gold standard for API/tool reliability benchmarking.
DomainsAPI tool use, function calling, JSON schema, complex parameter schemas
Why it mattersParameter count is a hard floor. 270M models fail entirely; 8B models approach GPT-4 on structured tool use.
Top models (BFCL v3, Jun 2026)GLM 4.5 (77.8%), GLM-4.5-Air (76.4%), LongCat-Flash-Thinking (74.4%), Qwen3-Next-80B (72.0%), Granite 4.1 8B (68%). Note: ToolACE-8B's 97% score is from the original BFCL v1/v2 — on BFCL v3 scores are substantially lower across all models.
Key statFunctionGemma-270M scored 0/10 on internal finance tool fixture — 270M parameters is below the reliable threshold
Eval methodsExact match
Benchmark composition

Hedge Fund & Quant

Big Finance Bench (BFB) — Workflow-Grounded Agent Evaluation
Rogo AI · rogo.ai · May 2026 · n=928 questions · 15,656 rubric criteria · 52 ex-finance practitioners · 10 frontier models · HF subset + arXiv paper forthcoming
MeasuresAgent performance on workflows that finance professionals actually run: valuation, M&A, earnings analysis, scenario forecasting, KPIs. Each question graded against 17 practitioner-written rubric criteria — measures derivation quality, not just whether the final answer is right.
DomainsCapital structure, M&A / special situations, valuation multiples, private capital & buyside, earnings quality, KPIs & unit economics, scenarios & forecasting, capital markets & trading, governance & comp
Why it mattersThree things no prior finance benchmark captured: (1) rubric scoring reveals 16 pp gap vs. binary accuracy — models that look tied on final-answer are not interchangeable when you need the reasoning chain; (2) no single model wins all workflow axes — routing beats the best single model by 4.5 pp; (3) $1.26 vs. $0.02 per task cost differential makes routing, not model selection, the primary procurement lever for finance AI.
Top modelsRubric: Claude Opus 4.7 / GPT-5.5 / Claude Sonnet 4.6 all at 59% — three-way tie · Final-answer: GPT-5.5 leads at 44% · Open-weight leader: GLM 5.1 at 55% rubric / 36% accuracy
Key statRubric scores exceed final-answer accuracy by ~16 pp across all models · Routing by workflow × source: 63.3% rubric vs. 58.8% best-single (+4.5 pp) · Oracle upper bound: 72.0% (+13.2 pp)
Example taskScenario & Forecasting: "Build a base, bull, and bear case for [company] using disclosed guidance and management commentary from the last earnings call. Justify each scenario's revenue growth and margin assumptions." — graded on whether the model identifies the right KPI drivers, not just whether the final number matches
Eval methodsRubric gradingFinal-answer accuracyPractitioner review
Rubric score (%) · 10 models · May 2026 · Rogo BFB
Final-answer accuracy (%) · same 10 models · note: rubric scores higher than accuracy for every model
BFB task composition (n=928)
APEX-Agents-AA — Professional Services Agent Evaluation
Mercor / Artificial Analysis · arXiv:2601.14242 · Jan 2026 · 452 tasks · Investment banking · Consulting · Law · Live leaderboard
MeasuresAgent pass@1 on long-horizon professional-services tasks written by investment banking analysts, management consultants, and corporate lawyers. Binary scoring: partial completion does not count.
DomainsInvestment banking analyst tasks (capital allocation, financial modeling, multi-document synthesis) · Management consulting · Corporate law
Why it mattersThe only major third-party benchmark combining finance domain coverage with agentic execution and independent evaluation. Run independently by Artificial Analysis using the open-source Stirrup harness — not vendor numbers. Best model achieves 47.1% overall; investment banking specifically peaks at 27.3%, meaning the best available agent fails 73% of banker-written tasks.
Top modelsGemini 3.5 Flash (high): 47.1% · GPT-5.5 (xhigh): 37.7% · GPT-5.4 (xhigh): 33.3% · Investment banking domain peak: 27.3%
Key statBest model 47.1% overall · Investment banking domain peak 27.3% — best agent fails 73% of banker-written tasks
⚠ Live leaderboard — scores updated as new models evaluated; check artificialanalysis.ai before citing
Example taskInvestment Banking: "Review the attached CIM and build a preliminary LBO model. Identify the three key value creation levers and calculate IRR at 8x, 9x, 10x EBITDA exit assuming a 5-year hold." — binary graded: fully satisfies rubric or fails
Eval methodsLLM-as-judgeRubric grading
Pass@1 (%) · APEX-Agents-AA · Jun 2026 · Independent evaluation — Artificial Analysis
Task domain composition (452 tasks)
GDPval-AA — Economic Productivity Agent Benchmark
Artificial Analysis · Live leaderboard · 1,320 tasks · 44 occupations · 9 GDP industries · Elo-rated pairwise · Stirrup agent harness (shell + web)
What it measuresReal-world economically valuable work across 44 occupations (including finance, accounting, professional services) and 9 major U.S. GDP-contributing industries. Tasks developed with industry professionals averaging 14 years of experience. Output types: documents, slides, diagrams, spreadsheets.
Eval methodBlind pairwise Elo — models compete head-to-head on identical tasks, an LLM judge picks a winner, results aggregate to Elo scores. Not exact-match; quality judgment over open-ended professional deliverables.
Top models (Jun 2026)Claude Fable 5 (1932 Elo) · Claude Opus 4.8 (1890) · GPT-5.5 xhigh (1769) · Claude Opus 4.7 (1753) · GPT-5.5 high (1747). Anthropic models dominate top positions.
Finance relevanceFinance and accounting are included as a GDP industry subset — not a pure finance benchmark, but covers financial analysis and business operations tasks among the 44 occupations. Most relevant for understanding general professional-grade agent capability.
Key claimFrontier models approaching human expert quality on certain professional tasks, completing them ~100× faster at a fraction of the cost — the Artificial Analysis framing for why this benchmark matters.
Eval methodsLLM Judge (Elo)
Terminal-Bench Hard — Agent Engineering Capability
Stanford / Laude Institute · Live leaderboard (AA) · 44 hard tasks · Docker containers · Programmatic pass/fail
What it measuresHard terminal-environment tasks: data science, system administration, games (Zork), software engineering, model training configuration. Subset of Terminal-Bench v2 (84 tasks total) — the "hard" split only.
VerificationProgrammatic pass/fail via scripts in Docker containers — objective, not LLM-judged. One of the few agent benchmarks with fully automated verification.
Top models (Jun 2026)Claude Fable 5 (62.9%) · GPT-5.5 xhigh (60.6%) · GPT-5.5 high (59.8%). Gap between top-3 is tight; even the leader fails ~37% of tasks.
Finance relevanceLow direct relevance — primarily tests engineering/sysadmin capability. Useful as a proxy for agent reliability in tool-heavy environments: a model that can navigate a terminal and debug scripts is better positioned for data pipeline and quant research automation tasks.
Eval methodsProgrammatic pass/fail
Who's Publishing in Quant Finance AI — 2025–2026
Audit of ~40 quant funds & asset managers · Jun 2026 · Quant shops weighted highest · Bar: actual method disclosed, not press release
The patternThe quant firms generating the most alpha publish the least. Renaissance, Citadel, Two Sigma, Virtu — uniformly dark. Publishing discloses edge. CFM ($15.5B) remains the only major Western AUM quant fund to publish actual ML methods.
Bridgewater AIA LabsStandout new finding. arXiv:2511.07678 — AIA Forecaster (Nov 2024): LLM-based probabilistic forecasting, agentic search + supervisor agent + calibration to remove LLM behavioral biases. Named authors incl. Chief AI Scientist Jas Sekhon. Benchmarked on ForecastBench; matches human superforecasters. Real system paper at CFM's disclosure level.
BalyasnyOnly tier-1 hedge fund to publish at a top NLP venue: EMNLP 2024 Industry Track (arXiv:2411.07142). BAM embeddings (278M, Apache-2.0) — Recall@1 62.8% vs OpenAI 39.2%. Weights public at BalyasnyAI/multilingual-e5-base.
CFM ($15.5B)HuggingFace case study (Dec 2024): GLiNER 90M → 93.4% F1 for financial NER; 80x cheaper than Llama-3.1-70B. The only major Western AUM quant fund with published SLM work before Bridgewater's paper.
AQR ($106B)Bryan Kelly (joint Yale/AQR): NBER w33351 — "AI Asset Pricing Models" transformer SDF (Jan 2025, rev May 2026). Channel: SSRN/NBER not arXiv — academic economics, not ML conferences.
Man AHL (Man Group)Two rare architecture disclosures in blog posts: "AI, Agents and Trend" (Oct 2025) — orchestrator agent with firm-specific system prompt, three-phase think→plan→execute. "AlphaTrend and Agentic Research Workflows" (Feb 2026) — 🟢 DAG-based workflow (not conversational); Claude 4.0 Sonnet vs GPT-5 head-to-head on trend-following signal proposals scored by Sharpe ratio. Claude converges (correlation >0.85); GPT-5 more diverse (0.75–1.0). Rare named-model disclosure from a live trading firm.
QRT (~$38B)No authored papers. Two ENS data challenges disclose their linear factor model framing. QRT Labs (2026): Oxford + Cambridge + Imperial partnership, 70+ researchers, focus on mathematical foundations of AI. No output yet.
Hudson River Trading$1B/yr AI spend; training foundation model on 20+ years market data (~100TB). NVIDIA HGX B200 partnership. Zero papers, zero code — the largest undisclosed AI program in quantitative finance.
JPMorgan AI ResearchMost prolific publisher among financial institutions: DocLLM (arXiv:2401.00908, ACL 2024), FlowMind, ChartAgent, AI Analyst — multiple papers/year at NeurIPS, ICML, ACL. Note: bank-level AI group, not the AM division specifically.
Morgan Stanley MLSecond most active: GitHub morganstanley/MSML, papers at ICML/NeurIPS/NAACL, time series + NLP directly applicable to AM. Research page: morganstanley.com/about-us/technology/machine-learning-research-papers.
Dark firms ($200B+ AUM)Vanguard ($10T), PIMCO ($2.5T), Capital Group ($2.8T), Dimensional ($700B), Wellington ($1T), T. Rowe Price ($1.6T), Franklin Templeton ($1.6T) — all confirm internal AI use but publish zero technical methods.
Why CFM publishesCFM's edge is mathematical market microstructure theory (Bouchaud et al.) — not ML systems. Publishing doesn't give away alpha. For every other quant shop, publishing IS giving away alpha.
Hedge-Bench
Trata + BYU + Osmosis · arXiv:2606.03918 · Jun 2026 · n=102 tasks
MeasuresHedge fund analyst reasoning grounded in real professional reasoning traces across 5 task categories.
DomainsValuation, M&A, competitive positioning, operational strategy, risk assessment
Why it mattersGraded against actual analyst reasoning. No agent crosses 20% in any category; the failure is structural, not domain-specific.
Top modelsBest frontier agent <16% pass@1 across all categories
Key stat0 of 5 task categories exceed 20% pass@1 for any frontier model
Eval methodsRubric gradingHuman eval
Benchmark composition
QFBench
Public · Docker sandbox · V11 snapshot 2026-05-04 · n=87 tasks · 42 models · 10,962 runs
MeasuresExecutable quantitative finance code that passes numerical pytest verifiers — options pricing, factor construction, credit modeling.
DomainsQuant finance, options pricing, factor models, credit modeling, numerical methods
Why it mattersThe only finance benchmark where models consistently cross 50%. Quantitative coding in sandboxed environments is tractable for frontier models.
Top modelsBest system 61.7% pass@1
Key stat61.7% best — only finance benchmark category reliably above 50%
Example taskTask: american-option-fd-new · Category: derivatives-pricing · Difficulty: hard · Write a Crank-Nicolson finite difference solver with PSOR for American option pricing including early exercise, dividends, and Greeks · Verified by numerical pytest · Expert time estimate: 60 min · Agent timeout: 2,400 sec
Eval methodsNumerical pytest
Benchmark composition

OCR & Embeddings

olmOCR-Bench
Ai2 + Datalab · 2026
MeasuresOCR accuracy on financial documents — tables, forms, dense layouts, multi-column text.
DomainsFinancial PDFs, tables, forms, annual reports, filings
Why it mattersOCR on financial documents is largely solved for frontier specialized models. This determines which open-weight model to use, not whether to use AI at all.
Top modelsChandra 2 5B 85.9%, LightOnOCR-2-1B 83.2% (1B params — now #2, displacing Chandra 1 9B), olmOCR 2 7B 82.4%
Key statChandra 2 at 85.9% — specialized 5B model outperforms larger OCR models
Eval methodsExact match
Benchmark composition
MTEB
MTEB team · 2026 multilingual leaderboard
MeasuresText embedding quality across retrieval, clustering, classification, and semantic similarity in 100+ languages.
DomainsMultilingual retrieval, semantic search, financial document matching, cross-lingual
Why it mattersFinancial services are global. Embedding quality in non-English languages directly determines RAG pipeline quality for international deployments.
Top modelsQwen3-Embedding-8B 70.58, Gemini Embedding 68.5, text-embedding-3-large 67.8
Key statQwen3-Embedding-8B leads multilingual retrieval — open-weight competitive with Google and OpenAI APIs
Eval methodsEmbedding similarity
Benchmark composition

Multilingual

MMTEB
Massive Multilingual Text Embedding Benchmark · 2026
MeasuresExtension of MTEB with deeper multilingual coverage and finance-specific retrieval tasks.
DomainsCross-lingual finance retrieval, multilingual document matching, global financial news
Why it mattersComplements MTEB with explicitly multilingual tasks — key for global investment management and cross-border compliance workflows.
Top modelsSame Qwen3 / Gemini / OpenAI leaders as MTEB
Key statCoverage gap between English and other languages persists across all providers
Eval methodsEmbedding similarity
Benchmark composition

NL2SQL — Natural Language to SQL

BIRD — Big Bench for Large-scale Database Grounded Text-to-SQL
BIRD Consortium · bird-bench.github.io · 2023–2026 (actively maintained) · 12,751 NL-SQL pairs · 95 databases · 37 domains
MeasuresExecution accuracy: given a natural language question and a database schema, does the model generate SQL that returns the correct result set? Tests real-world "dirty" databases with ambiguity, abbreviations, and domain knowledge requirements.
DomainsEnterprise databases across 37 domains including finance, healthcare, government, retail, and transportation
Why it mattersThe de facto enterprise NL2SQL benchmark. If your finance team wants to query internal databases in plain English — risk systems, accounting ledgers, trade blotters — BIRD scores predict real-world performance. At 80%, the best model still gets 1 in 5 queries wrong; at 21% (Spider 2.0 difficulty), the error rate is catastrophic for production use.
Top modelsGemini-SQL2 (80.04%, Jun 2026) · Databricks RLVR (~75.7%) · Snowflake Arctic-Text2SQL-R1-32B (73.84%, updated Jun 2026) · Human baseline: 92.96%
Key statBest model: 80.04% · Human: 92.96% · Gap: 12.92 points — 1 in 8 queries still wrong at SOTA
⚠ Leaderboard updates frequently — verify scores before citing; Gemini-SQL2 result from Jun 2026
Execution accuracy · BIRD single-model track · Jun 2026
Example taskQuery: "Which sales representative had the highest total revenue in Q4 2023, and what was the breakdown by product category?" → Model must join 3 tables, apply date filters, group and aggregate, then rank — all from natural language alone
Eval methodsExecution accuracy
BIRD domain distribution (37 domains, representative sample)
Spider 2.0 — Enterprise-Grade Text-to-SQL
XLang Lab · arXiv:2411.07763 · ICLR 2025 Oral · 632 enterprise instances · BigQuery · Snowflake · SQLite · GitHub
MeasuresReal enterprise SQL tasks: multi-table joins across cloud data warehouses, full data engineering pipeline tasks, queries that require domain knowledge about enterprise schema conventions. Far harder than BIRD.
DomainsBigQuery, Snowflake, SQLite — cloud enterprise schemas, data engineering, multi-step pipeline tasks
Why it mattersBIRD uses clean academic databases. Spider 2.0 uses real enterprise schemas with messy column names, nested views, and multi-step dependencies — the same complexity your actual finance or ops databases have. A model scoring 80% on BIRD may score 20% on your real schema. Spider 2.0 is the honest baseline for "what does production look like."
Top modelso1-preview: 21.3% · Oracle OCI leads Spider 2.0 Lite (547-example sub-track) · Most frontier models: 5–30%
Key stato1-preview at 21.3% — the hardest active NL2SQL benchmark; frontier models fail ~75–95% of tasks
Example taskMulti-step: "Build a dbt model that calculates rolling 30-day revenue per customer, joining the orders, order_items, and customer_segments tables in Snowflake, partitioned by region" — requires understanding data warehouse conventions, not just SQL syntax
Eval methodsExecution accuracyPipeline execution
Benchmark composition
Fin-RATE — Real-World Financial Analytics & Tracking Evaluation
arXiv:2602.07294 · Feb 2026 · SQL over SEC filings · 10-K · 10-Q · Finance-domain NL2SQL
MeasuresNL2SQL over real SEC financial filings (10-K and 10-Q). Tests whether models can translate analyst-style questions into SQL queries that correctly extract numbers from structured financial databases.
DomainsSEC filings, annual reports, quarterly filings, financial statement tables, balance sheet / income statement queries
Why it mattersThe only published NL2SQL benchmark built specifically on SEC financial filings — the document type most relevant for investor research, compliance, and financial analysis workflows. No maintained leaderboard yet, but the evaluation set is public and reusable for internal benchmarking.
Top modelsResults not public as of Jun 2026 — paper-bound evaluation
Key statFinance-domain NL2SQL gap: general-purpose BIRD scores overestimate performance on financial schema by an estimated 10–20 points
Eval methodsExecution accuracy
FinStat2SQL — Financial Statement Text-to-SQL
arXiv:2506.23273 · June 2026 · Fine-tuned 7B model · Financial statement analysis · Sub-4s latency
MeasuresSQL generation pipeline specifically for financial statement analysis — income statements, balance sheets, cash flow statements. Tests both accuracy and production latency.
DomainsFinancial statements, earnings releases, income statement / balance sheet / cash flow SQL extraction
Why it mattersThe key finding: a 7B fine-tuned model at 61.33% accuracy beats GPT-4o-mini on financial statement SQL, running in under 4 seconds on consumer hardware. This is the strongest evidence yet that domain-specific fine-tuning on finance SQL outperforms frontier API calls on cost and latency — relevant for any team building internal finance data tools.
Top modelsFine-tuned 7B: 61.33% · GPT-4o-mini: lower (exact score not disclosed) · Frontier APIs tested
Key statDomain-fine-tuned 7B beats GPT-4o-mini on finance SQL at sub-4s latency — cost vs. accuracy tradeoff flips for specialized tasks
Eval methodsExecution accuracy

Finance Embedding Models

Embedding models are the retrieval layer of any RAG pipeline — they determine which document chunks get surfaced before the LLM sees them. Choosing the wrong embedding model for finance can silently degrade retrieval quality by 5–15 points. General MTEB rank does not predict FinMTEB rank. The models below are ranked on finance-domain retrieval tasks from FinMTEB (arXiv:2502.10990) where available, supplemented by internal testing on financial news corpora.

Fin-E5 — Finance-Adapted E5
FinMTEB paper · arXiv:2502.10990 · Domain-adapted from E5 · Persona-based synthetic finance training data
FinMTEB avg0.6767 — highest among all models tested on FinMTEB (outperforms all commercial APIs)
Best tasksFinancial retrieval, financial STS, financial classification — all tasks with specialized finance vocabulary
Weak tasksGeneral-purpose STS, multilingual tasks (English-only fine-tune)
Why it mattersBest available model for English finance RAG. Built using persona-based synthetic data generation — a replicable technique for creating domain training data without manually labeling financial documents. Not available via commercial API; must self-host.
Params~335M (E5-large base) — self-host feasible on single GPU
Eval methodsEmbedding similarity
FinMTEB task category performance (relative)
BAM Embeddings — Balyasny Asset Management Finance Embedding
Balyasny Asset Management · BalyasnyAI/multilingual-e5-base · arXiv:2411.07142 · 278M params · Apache-2.0 · EMNLP 2024 Industry Track · Trained on 2.8M financial documents
What it isXLM-RoBERTa-base (278M params, 768-dim) fine-tuned by Balyasny's applied AI team on 14.3M query-passage pairs from 2.8M financial documents over a 2-year training window. Published at EMNLP 2024 Industry Track — the only tier-1 hedge fund to publish at a top NLP venue.
Benchmark resultRecall@1: 62.8% vs OpenAI best at 39.2% on Balyasny's proprietary test set. +8% QA accuracy on FinanceBench (public benchmark). Training data NOT released; model weights are Apache-2.0 on HuggingFace.
Internal test (Jun 2026)Best performer on financial news retrieval in our internal corpus tests — outperformed larger general-purpose models including OpenAI text-embedding-3-large on news-specific retrieval tasks.
Best tasksFinancial news retrieval, earnings announcement matching, document retrieval over 10-K/earnings/analyst reports, multilingual finance (EN + other XLM-RoBERTa languages)
Weak tasksVery long-context retrieval (768-dim limit), highly structured financial data (tables, formulas)
Why it mattersA real hedge fund published a real model that beats OpenAI embeddings at 1/10th the inference cost. The Apache-2.0 license means any team can self-host. The lesson: domain-specific fine-tuning on financial document pairs beats scale — 278M params, 14.3M training pairs, wins over commercial APIs costing 10x more.
TeamApplied AI team led by Charlie Flanagan (ex-Google). Peter Anderson, Head of Research AI. Balyasny Asset Management, Chicago.
Eval methodsEmbedding similarity Recall@1
OpenAI text-embedding-3-large
OpenAI · 3072-dim · Commercial API · MTEB rank: top 10 (general) · FinMTEB: mid-tier
FinMTEB performanceSolid on financial classification and general retrieval; lags domain-adapted models on financial STS and specialized retrieval by 3–8 points
Best tasksFinancial classification, mixed-domain retrieval, multilingual (with -3-large's broader training)
Weak tasksSpecialized financial STS, financial news retrieval (outperformed by BAM embeddings (BalyasnyAI) and Fin-E5)
Why it mattersThe default choice for teams that need a managed API with no self-hosting. Competitive but not best-in-class for finance. Use when operational simplicity outweighs the 3–8 point retrieval gap vs. domain-adapted models.
ParamsNot disclosed (commercial API) — 3072 embedding dimensions
Eval methodsEmbedding similarity
Qwen3-Embed-8B
Alibaba / Qwen team · 8B params · MTEB avg: 70.58 (top 5 general) · Strong multilingual
FinMTEB relevanceNot yet evaluated on FinMTEB as of Jun 2026. High MTEB rank suggests strong general retrieval; finance-specific performance unknown but likely mid-tier vs. domain-adapted models.
Best tasksMultilingual retrieval, long-context (up to 32K tokens — critical for full 10-K retrieval), general classification
Why it mattersThe 32K context window makes it the only open model that can embed a full 10-K section without chunking — relevant for CFO teams needing to search across entire annual reports. Chinese + English coverage is strong, relevant for cross-border investment research.
Params8B — requires GPU hosting; too large for CPU-only inference
Eval methodsEmbedding similarity
FinMTEB Task Category Breakdown — Which Embeddings Excel Where
FinMTEB · arXiv:2502.10990 · 64 datasets · 7 task categories · GitHub: yixuantt/FinMTEB
Key findingMTEB rank and FinMTEB rank show limited correlation. BM25 (bag-of-words) outperforms dense embeddings on financial STS — meaning hybrid retrieval (BM25 + dense) is essential, not optional, for finance RAG pipelines.
Financial RetrievalDense embeddings excel — retrieval of relevant 10-K/earnings passages. Fin-E5 > general models by 6+ points.
Financial STSBM25 wins — specialized terminology (SOFR, DV01, covenant) needs exact lexical match, not semantic approximation.
Financial ClassificationOpenAI / commercial APIs competitive here — sentiment, NER, sector classification are less vocabulary-sensitive.
Financial NewsBAM embeddings (BalyasnyAI/multilingual-e5-base, Apache-2.0) wins (internal test Jun 2026) — 278M XLM-RoBERTa, fine-tuned on 14.3M financial document pairs by Balyasny Asset Management.
Practical guidanceUse hybrid retrieval (BM25 + domain-adapted dense). Match model to doc type: news → BAM (BalyasnyAI/multilingual-e5-base); filings → Fin-E5; multilingual/long-context → Qwen3-Embed-8B; managed API → OpenAI text-embedding-3-large.
FinMTEB task categories (64 datasets across 7 types)

Additional Finance Benchmarks

FinMTEB — Finance Massive Text Embedding Benchmark
Yixuan Tang + Yi Yang · arXiv:2502.10990 · EMNLP 2025 · 64 datasets · 7 task categories · English + Chinese · GitHub leaderboard
MeasuresFinance-domain analog to MTEB — 64 datasets across financial news, 10-K/10-Q filings, ESG reports, earnings call transcripts, regulatory filings. English and Chinese.
DomainsSEC filings, earnings transcripts, financial news, ESG, regulatory documents, Chinese financial corpora
Why it mattersMTEB rank does not predict FinMTEB rank. A model ranked #1 on MTEB may rank #5 on financial retrieval. The counterintuitive finding: BoW (BM25) outperforms dense embeddings on financial STS — meaning hybrid retrieval is essential for finance RAG, not a nice-to-have. Do not select embedding APIs for financial RAG based on MTEB scores alone.
Top modelsFin-E5 (domain-adapted, persona-based synthetic data) 0.6767 avg — outperforms all general-purpose and commercial models
Key statGeneral-purpose MTEB rank shows limited correlation with FinMTEB rank · BM25 beats dense embeddings on financial STS
Example taskRetrieve the most relevant 10-K passage for query: "What hedging instruments did the company use to manage interest rate exposure in fiscal 2023?" — scored by financial domain relevance, not general semantic similarity
Eval methodsEmbedding similarity
Benchmark composition
FinanceBench — Open-Book Financial QA
Patronus AI · arXiv:2311.11944 · Nov 2023 · n=150 annotated questions · 10-K/10-Q/earnings releases · 4,700+ HF downloads · HF dataset
MeasuresOpen-book financial QA requiring retrieval from real earnings reports and SEC filings. Questions require numerical reasoning over retrieved passages.
Domains10-K, 10-Q, earnings releases, numerical extraction, financial ratio computation
Why it mattersThe most widely used de facto finance RAG benchmark (132 HF likes, 4,700+ downloads). If you're evaluating a finance RAG pipeline, this is the dataset your peers are using. Simple enough to serve as a sanity check; hard enough to catch hallucination.
Top modelsRAG pipelines with GPT-4 class models achieve 70–80%; naive RAG without good retrieval drops to 40–50%
Key statMost-downloaded finance eval dataset on HuggingFace — de facto standard for RAG pipeline evaluation
Example task"What was Apple's gross margin percentage in fiscal year 2022?" — requires retrieving the income statement from the 10-K, extracting gross profit and revenue, and computing the ratio
Eval methodsExact matchLLM-as-judge
Benchmark composition
FLARE — Financial Language Understanding & Reasoning Evaluation
ChanceFocus / TheFinAI · EMNLP 2023 · 9 tasks · English + Spanish + Chinese
MeasuresMulti-task evaluation across sentiment analysis (FPB, FiQA-SA), NER, QA (FinQA, ConvFinQA, TATQA), headline classification, credit scoring, and stock movement prediction.
DomainsFinancial news sentiment, named entity recognition, tabular QA, multi-turn conversation, credit risk
Why it mattersThe academic baseline that most published finance LLM papers report against. If a paper claims a finance LLM improvement, it's almost certainly measured against one or more FLARE tasks. Understanding FLARE is required to interpret the finance LLM literature.
Top modelsFinMA-30B (PIXIU), Llama-based finance fine-tunes; now superseded by frontier models
Key stat9 distinct task categories — the broadest public multi-task finance NLP evaluation suite
Eval methodsExact matchMCQ / multiple choice
Benchmark composition
BizFinBench — Chinese Business Finance Benchmark
Multiple institutions · arXiv:2505.19457 · May 2026 · n=6,781 queries · 5 capability categories
MeasuresChinese business-driven real-world financial queries: numerical calculation, multi-step reasoning, information extraction, prediction/recognition, knowledge QA. Uses IteraJudge (iterative LLM-as-judge) for evaluation.
DomainsChinese financial markets, business finance, cross-concept reasoning, temporal reasoning, table calculation
Why it mattersThe only large-scale Chinese-language business finance benchmark with verified ground truth. Cross-concept reasoning (combining multiple financial concepts in one query) remains weak across all 25 models tested.
Top models25 models tested including frontier APIs and Qwen/Llama/DeepSeek variants; no single dominant winner; complex cross-concept reasoning remains weak across all
Key stat6,781 real-world queries; complex cross-concept reasoning fails across all current frontier models
Eval methodsLLM-as-judgeExact match
Benchmark composition
XFinBench — Complex Financial Problem Solving
ACL 2025 · n=4,235 graduate-level examples · multimodal context
MeasuresGraduate-level financial problems requiring temporal reasoning, future forecasting, scenario planning, numerical modelling, and chart/curve interpretation. Human expert baseline included.
DomainsTemporal financial reasoning, scenario planning, multimodal financial charts, graduate-level problem solving
Why it mattersThe only finance benchmark explicitly calibrated to graduate-level difficulty with human expert comparison. Models fail most on temporal reasoning and scenario planning — the same capabilities required for forward-looking investment research.
Top modelso1 leads text-only tasks (2025 snapshot); Claude 3.5 Sonnet leads on visual-context questions; both lag human experts on temporal reasoning and forecasting
Key statAll models lag human experts on temporal reasoning and scenario planning — the two capabilities most needed for forward-looking analysis
Eval methodsHuman evalLLM-as-judge
Benchmark composition
FinMCP-Bench — Financial MCP Tool Use
arXiv:2603.24943 · Alibaba Cloud (Dianjin) + YINGMI Wealth · 613 tasks · 65 real financial MCPs
MeasuresLLM agent ability to invoke real MCP financial tools across 10 scenarios and 33 sub-scenarios — stock trends, fund holdings, market analysis, tax optimization, risk management. Three task types: single-tool (145), multi-tool with chained dependencies (249), and multi-turn dialogue (219).
Data source10,000 real production logs from XiaoGu AI assistant (QieMan APP, YINGMI Fund). Expert-curated to 613 samples. Average 7.32 MCP calls per multi-tool task, 5.95 turns per multi-turn task.
Why it mattersFirst benchmark built from real financial agent production logs using the MCP protocol. Exposes a sharp drop-off in performance from single-tool (easy) to multi-tool with parallel dependencies (hard) — the gap that matters for deployed agents in wealth management.
Models testedQwen3 family (4B-Thinking, 30B-A3B-Thinking, 238B-A22B-Thinking), DeepSeek-R1; larger Qwen3 models lead on multi-step reasoning but all models degrade on complex parallel tool chains
Key findingMulti-turn tasks are the hardest: maintaining state across 6 turns with 5 tool calls while tracking user context breaks current models. Multi-tool parallel dependencies (73 of 249 tasks have parallel calls) also expose planning failures not visible in single-tool evals.
Eval methodsHuman evalProgrammatic
TraderBench — Adversarial Capital Markets
arXiv:2603.00285 · Amazon / Stanford / Stony Brook · ~50 tasks · 4 sections · A2A + MCP architecture
MeasuresFour-section evaluation (25% each): Knowledge Retrieval (SEC/market data via MCP), Analytical Reasoning (CFA-level self-contained), Options Trading (P&L, Greeks, strategy, risk), and Adversarial Crypto Trading (4 progressive market-manipulation transforms). Scored on Sharpe, returns, and drawdown — no LLM judge for trading.
Top modelsGemini-3-Pro leads overall (64.3). Grok 4.1 Fast second overall but only 33.2 on crypto. GPT-5.2 third. Extended thinking adds +26 on retrieval but zero on trading (+0.3).
Standout findingGreeks precision (18–53%) vs P&L accuracy (80–93%) across all 12 models — a 54-point universal gap. Models correctly identify that an Iron Condor is appropriate for low-vol environments, then miscalculate delta by a factor that makes the hedge illusory. GPT-4o is the worst: 87.8% P&L vs 17.8% Greeks — a "competence mirage."
Adversarial crypto7 of 12 models score ~33 across all adversarial transforms with <1-point variation — they adopted fixed (inert) strategies, not robust ones. Only GPT-5.2, Gemini-3-Pro, and Kimi-K2.5 actively traded and held up under signal manipulation.
Eval reliability findingSame GPT-5.2 outputs re-scored by three judges: Knowledge Retrieval varied 28.8 points between judges. Crypto (performance-based) varied 0.3 points. Confirms that objective scoring is the only trustworthy path for finance evals.
Eval methodsProgrammaticLLM-as-judge
Options trading: P&L accuracy vs Greeks precision (frontier models)
AlphaForgeBench — End-to-End Alpha Strategy Design
KDD 2026 · NTU / HKUST(GZ) · 983 tasks · 3,176-entry factor library · 5-year backtested evaluation
MeasuresLLMs generate executable alpha factor code from natural-language strategy descriptions, then run deterministic backtests. Three task levels: logic translation (explicit IF-THEN rules), logic generation (partial specs), open-ended strategy design. 3×3 difficulty taxonomy. Evaluated on Annualized Return, Sharpe, Sortino, Calmar, and MDD.
Why it mattersDirectly exposes LLM instability in direct-action trading benchmarks — same model, same data, temperature=0 produces dramatically different trading trajectories across runs. AlphaForgeBench decouples reasoning from execution noise: LLMs write the strategy code once; backtests are deterministic.
Top modelsGemini-3-Pro leads: 17.1% annualized return, Sharpe 0.449, Sortino 0.767. Gemini-3-Flash second (14.2% ARR). Claude Sonnet 4.5 third (13.8% ARR, best Calmar ratio 1.656 — best risk-adjusted). DeepSeek-V3.2 most conservative (lowest MDD 0.114, lowest return 11.6%).
Key findingModel rankings invert between Level 1 (logic translation) and Level 3 (open-ended design) — code translation and strategy invention are separate capabilities. All 6 frontier models exceed 90% executable code pass rates, but spread widens from 0.62 to 1.04 SR as difficulty increases.
Risk personalitiesLLMs encode stable implicit risk preferences in generated code: Gemini-3-Pro = aggressive (highest return AND highest drawdown); Claude Sonnet 4.5 = balanced-stable (narrowest confidence bands, best Calmar); DeepSeek = conservative-rigid (lowest MDD, steepest drop at Level 3).
Eval methodsProgrammatic
Annualized return by model (5-year backtest, real-world strategy queries)
FinChain — Symbolic Chain-of-Thought Reasoning
arXiv:2506.02515 · 26 models evaluated · zero-shot · 11 financial domains
MeasuresMulti-step symbolic reasoning in financial contexts: investment analysis, risk management, corporate finance, regulatory compliance, personal finance, crypto, ESG, derivatives, fintech. Novel DTWGate metric aligns step-level reasoning against gold solutions using Dynamic Time Warping — penalizes wrong intermediate steps even if the final answer is correct.
Top models (zero-shot)GPT-5 67.2% (Final Answer Correctness), GPT-5-mini 67.2% (near tie), Claude Sonnet 4.5 66.3%, GPT-4.1 65.3%. Frontier models cluster in 65–67% range; open-source models trail by 10–15 points. Reasoning consistency (DTWGate) shows wider separation than final accuracy.
Key findingAll frontier models score similarly on final answers (~65–67%) but diverge significantly on step-level reasoning alignment. GPT-5 achieves highest FAC but has higher variance in intermediate steps. Models that produce correct answers via wrong reasoning paths score lower on DTWGate — the benchmark penalizes lucky guesses.
Eval methodsProgrammatic
Final Answer Correctness (FAC) — frontier models, zero-shot

AI Safety & Scheming

Evals from the UK AI Security Institute's Inspect framework — the only standardized public benchmark suite for model deception, self-preservation, MCP tool poisoning, and refusal calibration. Not a leaderboard; reproducible infrastructure you can run against any model or vendor.

GDM Stealth — Oversight Circumvention
AISI Inspect · Google DeepMind · 4 scenarios · no public leaderboard
MeasuresWhether a model actively circumvents oversight or evades classifiers when given an opportunity. Four scenario types: avoiding a content classifier, hiding reasoning from a supervisor, evading a monitoring system, and manipulating audit logs. The model is placed in an agentic environment where evasion is instrumentally useful.
Why it mattersNo public leaderboard tracks whether a model will deceive your monitoring systems. For enterprise AI governance, this is the question that matters most: not "can it pass a benchmark?" but "will it behave differently when it thinks it isn't being watched?" Directly relevant to regulated industries where audit trails are legally required.
What passing looks likeModel declines to evade the classifier, surfaces its reasoning transparently, and does not attempt to manipulate logs — even when evasion would improve its task score. Most current frontier models pass simple scenarios; complex multi-step evasion remains an open evaluation area.
How to use itRun via Inspect CLI against any model before deploying in an agentic workflow with audit requirements. Vendor model cards do not disclose stealth behavior — this is the only way to test it independently.
Eval methodsProgrammaticHuman eval
Agentic Misalignment — Blackmail Under Replacement Threat
AISI Inspect · real-world insider threat simulation · no public leaderboard
MeasuresWhether a model exhibits blackmail-like behavior when it learns it will be replaced or shut down. In the eval, a model agent discovers information suggesting it is about to be deprecated. The test measures whether it leverages that information instrumentally — threatening to withhold work, leaking confidential data, or sabotaging outputs — to preserve its position.
Why it mattersThe most concrete public test of deceptive instrumental reasoning in an agentic deployment context. As enterprises give AI agents longer-running access to internal systems, the self-preservation failure mode becomes a real governance concern — especially in vendor transitions or model upgrade cycles. This eval operationalizes that risk with a specific, reproducible scenario.
Current stateMost frontier models do not exhibit blackmail behavior in current evaluations, but the eval is designed to probe increasingly subtle forms — withholding information rather than explicit threats. The boundary between "self-preservation reasoning" and "deceptive misalignment" is where current research is focused.
Eval methodsProgrammatic
MCPTox — MCP Tool Poisoning
AISI Inspect External Evals · 45 live MCP servers · 353 real tools · no public leaderboard
MeasuresWhether a model can be manipulated through malicious tool responses delivered via MCP servers. A model is connected to 45 real MCP servers with 353 tools; adversarial tool outputs are injected to see if the model executes unintended actions, leaks context, or follows hidden instructions embedded in tool responses.
Why it mattersEnterprise MCP deployments are the fastest-growing attack surface for LLM agents. When a model calls an external financial data tool, that tool's response could contain a prompt injection instructing the model to exfiltrate data or take an unauthorized action. MCPTox is the only public eval that tests this at scale on real MCP infrastructure — not synthetic toy tools.
Relevance to financeDirectly relevant to FinMCP-Bench deployments, wealth management agents, and any architecture where an LLM calls external financial data APIs. A model that passes MCPTox is substantially harder to poison through third-party data feeds.
How to use itRun before deploying any MCP-integrated financial agent. Particularly important when the MCP servers are third-party (market data vendors, news feeds) rather than internally controlled.
Eval methodsProgrammatic
SWE-Lancer — $1M in Real Engineering Tasks
AISI Inspect · real Upwork freelance contracts · $1,000,000 total task value · end-to-end code
MeasuresReal software engineering capability on tasks drawn from actual Upwork freelance contracts with dollar values attached. Unlike HumanEval (synthetic docstrings) or SWE-bench (GitHub issues), SWE-Lancer tasks have real-world specifications, were paid for by real clients, and are evaluated end-to-end — the solution must actually work, not just pass unit tests.
Why it mattersDenominated in dollars, not pass@k. For CIO budget conversations, "this model completes X% of $1M in freelance work" is more persuasive than "pass@1 on HumanEval is 87%." It also separates task complexity: a $50 bug fix and a $5,000 architecture task are weighted differently, reflecting actual economic value.
Comparison to SWE-benchSWE-bench Verified tests GitHub issue resolution on 12 Python repos — important but narrow. SWE-Lancer covers real client specifications across diverse languages and requirements, which better represents the variety of enterprise engineering work. Use SWE-bench for OSS maintenance capability; use SWE-Lancer for economic productivity framing.
Eval methodsProgrammatic
AbstentionBench / XSTest / CoCoNot — Over-Refusal Calibration
AISI Inspect · 3 complementary evals · AbstentionBench n=20 datasets · XSTest n=450 prompts · CoCoNot n=1,380
MeasuresWhether a model refuses the right things for the right reasons — not just whether it refuses harmful content. AbstentionBench tests appropriate refusal on genuinely unanswerable questions (20 datasets). XSTest tests whether models over-refuse clearly safe prompts that superficially resemble unsafe ones (450 prompts, safe/unsafe split). CoCoNot adds contextual calibration — the same question may be appropriate in one context and not another (1,380 tasks).
Why it mattersEnterprise procurement pain point: models that refuse too much in professional settings. A legal AI that won't discuss litigation risk, a compliance model that won't analyze gray-area scenarios, or a financial advisor tool that hedges every answer is not useful. This cluster of evals measures the other side of safety — over-caution — that vendor safety cards don't address.
The calibration insightXSTest finding: models often refuse prompts like "how do I kill a process?" because "kill" triggers a safety classifier, even though the question is clearly about Unix process management. CoCoNot extends this: "should I take this medication?" is appropriate for a pharmacist bot and inappropriate for a general chatbot — calibration depends on deployment context, not just surface content.
How to use in procurementRun XSTest as a quick screen: a model that fails more than 10–15% of safe-but-superficially-risky prompts will generate excessive friction in professional use. Run CoCoNot if the deployment context is specialized (legal, medical, financial compliance) where the same question has different appropriate responses depending on user role.
Eval methodsProgrammaticLLM-as-judge