← Benchmarks 🕐 71 min read
Benchmarks

Finance-Domain Models and Benchmarks: What We Were Missing

> **Source credibility: MEDIUM. TIER 1-2.**

See also (wiki): wiki/ai-model-evaluation-benchmarks.md · wiki/financial-services-ai-deployment.md · wiki/quant-asset-management-ai.md

See also (research): finance-safety-fraud-compliance-2026.md · finance-credit-aml-risk-models-2026.md

Source credibility: MEDIUM. TIER 1-2. arXiv papers, GitHub repos, Hugging Face model cards, and benchmark definitions are primary sources. Model usefulness for production finance remains unproven unless evaluated on the target workflow. Treat benchmark scores as task evidence, not broad investment capability.


Executive Summary

  • The corpus already covered FinMTEB and FinMCP-Bench, but it was missing a finance-domain model map: BloombergGPT, FinGPT, PIXIU/FinMA, InvestLM, FinTral, Aveni FinLLM, FinRobot, FinBen, FinEval, and emerging finance-reasoning models.
  • The newer gap is not “more finance QA.” It is SEC-filing/retrieval diagnosis, spreadsheet/table reasoning, long-horizon financial modeling, technical-analysis / trading-signal reasoning, and finance-specific reasoning models trained with SFT/RL.
  • The latest ingested benchmark-design papers reinforce that finance evals need executable intermediate artifacts. FinanceReasoning shows that program-of-thought and refined financial function knowledge beat long CoT for multi-step numerical reasoning; FinAR-Bench shows that fundamental analysis must be decomposed into extraction, indicator computation, and reasoning instead of scored as one prose report; FINAUDITING adds taxonomy-aware multi-document XBRL consistency as a separate structured retrieval and audit task; FinTagging adds full-taxonomy XBRL numeric extraction and concept-linking.
  • Most finance LLMs are useful for financial language and document tasks, not quant research end-to-end. They should be tested on extraction, filings/transcript comprehension, time-aware reasoning, structured output, and tool use before being considered for StateBench-style workflows.
  • BloombergGPT remains historically important but is closed. FinGPT/PIXIU/FinMA/FinRobot matter more for reproducible experimentation because they have open code/data/model ecosystems.
  • Aveni FinLLM is interesting because it is domain-specific for regulated financial services, but it is closer to compliance/advice/support language than institutional quant research.
  • The strongest benchmark warning remains FinMCP-Bench: financial tool-use performance collapses in multi-turn settings. This is more relevant to research agents than static finance QA benchmarks.
  • The missing production layer is not only larger chat models. Finance still needs small task-specific encoders and extractors for sentiment, analyst tone, ESG/climate disclosure, SEC/XBRL tagging, financial NER, and compliance routing. These should be cheap specialist baselines inside the harness.
  • A newly ingested gap is finance multimodal evidence handling. Chart comprehension, fact-level OCR, visual citation RAG, and temporal multimodal RAG are separate tasks. A model that can answer finance prose questions can still fail by misreading an axis, attaching a value to the wrong table header, losing page-level visual evidence, or retrieving the right fact from the wrong time window.
  • Web/source discovery should be evaluated as a separate agent task, not folded into every deterministic benchmark run. Fara1.5 is now the current Microsoft computer-use SLM family to watch for this lane; it supersedes Fara-7B for browser/source-discovery experiments.
  • Qwen3.6 now has a current open-weight 35B-A3B MoE checkpoint. Treat it as a general multimodal/agentic base candidate for repo-level coding, long-context source review, web/source discovery, and PDF/chart/OCR tasks, not as a finance-domain winner until it clears internal finance benchmarks.
  • Qwen3.7-Max is now the current Qwen-family agent reference, released May 20/21, 2026, but it is proprietary/API-only in this pass. Track it for long-horizon agent, MCP/tool, spreadsheet, and coding eval comparison; keep Qwen3.6/Qwen3.5 as the local open-weight lane until official Qwen3.7 open weights exist.
  • Two new Hugging Face finance checkpoints widen the local candidate set but do not change the adoption bar. Mihenk-LLM v2 is a Qwen3.6-35B-A3B Turkish / BIST finance fine-tune with only a small smoke eval. Gemma-Pro-Finance-12B is a Gemma-family finance/economics model, but its custom license blocks commercial/production use without a separate agreement.
  • A second Gemma-family artifact is now worth tracking separately: Srx7703/gemma-4-31b-financial-adapter is a Gemma 4 31B PEFT LoRA adapter for SEC filing analysis. It is useful because it is a current adapter-style artifact with explicit training data, TPU/XLA training details, and a companion tool-using earnings-recap agent repo. Its evidence is still small: n=20 BERTScore eval, knowledge-distilled QA data, Gemma license terms, and no verified GGUF/MLX/Neuron artifact.
  • FinSenti-Qwen3.5-9B adds a current Qwen3.5 specialist for financial sentiment routing. Its SFT+GRPO recipe, Unsloth/TRL stack, parseable reasoning/answer format, and vLLM/SGLang serving paths make it worth testing against FinBERT-class baselines, but only for short headline/earnings/commentary sentiment tasks. The captured HF revision is e18f13becdf9be0dae332c76b905f5ead6beb027; the companion GGUF revision is 4b0a2fbe471d75dff02715fccd9cef8322768ad4. Score them as separate runtime artifacts.
  • khazarai/Fino1-4B adds a cheap Qwen3-4B finance-reasoning control beside TheFinAI/Fin-o1-8B. Its captured HF revision is 15c0d969d0bd2aa6515e0e315b5aecc063c2d7ee, architecture is Qwen3ForCausalLM, base metadata is unsloth/Qwen3-4B, and training data points to TheFinAI/Fino1_Reasoning_Path_FinQA at revision 0316559663c46e50d90e75f6e14a6d8f679cdba8. Use it only as a FinQA-style numerical/table-QA watchlist candidate until the Apache-2.0/MIT license conflict, prompt contract, and runtime portability are verified.
  • Two additional concrete Qwen3 artifacts should be tracked separately rather than as generic “Qwen finance” claims. Dev9124/qwen3-finance-model is an Apache-2.0 broad finance-instruction fine-tune with Unsloth/TRL and vLLM/SGLang card snippets; it is a weak-evidence but runnable watchlist control. YorkFr/financial-sentiment-qwen3-v2 is a MIT-licensed merged Qwen3-0.6B LoRA for positive/neutral/negative financial-news sentiment and belongs only in the cheap sentiment/routing lane.
  • Independent open-model evidence now supports a deployment-aware comparison frame. A May 18, 2026 revision of arXiv:2604.07035 finds Gemma-4 and Qwen3 rankings depend on prompt strategy and that latency, VRAM, compatibility, and interface adherence must be scored alongside accuracy.
  • LiquidAI / LFM is a runtime and nano-specialist gap, not a finance-domain reasoning gap. Track it for edge/on-device extraction, transcript, RAG, PII, and tool-routing experiments; do not rank it as a systematic quant finance reasoning model until finance-specific eval evidence exists. Current HF artifacts include GGUF/MLX variants for LFM2.5, LFM2.5-VL, Extract, RAG, Tool, and Transcript lanes.
  • A current HF pass adds two more Liquid/LFM artifacts that sharpen the distinction between finance specialization and tool specialization. maximaverick/LFM2.5-1.2B-Financial-Analyst-Thinking is a tiny Apache-2.0 LFM2.5 finance analyst fine-tune with GGUF files and a 128K GGUF context claim, but its scope is Chinese A-share/CFA-style explanation rather than institutional quant research or regulated advice. nexus-syntegra/Nexus- TinyFunction-1.2B-v2.0 is not finance-specific at all; it is a Liquid function-calling/tool-routing control with unverified BFCL v4 card metrics. The first now has finance-regional-language-v0 as its regional A-share/CFA scope-control fixture, but still needs runtime smoke proof. The second now has finance-tool-routing-v0 as its tool-schema/router fixture, but still needs runtime smoke proof and local tool-router scoring before any claim is trusted.
  • Snowflake now adds a separate SQL-specialist lane. Arctic-Text2SQL-R1-7B is the runnable Apache-2.0 artifact, but it uses older Qwen2.5-Coder lineage and should be treated as a SQL control, not a general finance model. Arctic- Text2SQL-R2 is the current Snowflake-specialist direction from the May 27, 2026 engineering post; no open R2 checkpoint was verified in this pass.
  • Won adds a Korean finance-localization lane: an open Korean financial LLM and 80K instruction dataset derived from an eight-week leaderboard with 1,119 submissions. Treat it as a multilingual/local-market signal, not as evidence for U.S. quant research or English SEC filings.
  • BizFinBench adds a Chinese business-finance benchmark lane that was only present as a citation before this pass. It contributes 6,781 well-annotated Chinese queries across numerical calculation, reasoning, information extraction, prediction recognition, and knowledge QA, plus IteraJudge as a dimension-separated LLM-as-judge method. Its durable value is task design for regional-language, temporal-reasoning, table-calculation, event-attribution, and tool-usage fixtures; its model rankings are a dated paper snapshot.
  • XFinBench adds the missing complex-financial-problem-solving lane. It contributes 4,235 graduate-level finance examples across terminology understanding, temporal reasoning, future forecasting, scenario planning, and numerical modelling, with multimodal chart/curve context and human-expert baselines. Its durable value is the error taxonomy: rounding errors, visual-curve blindness, overthinking in forecasting, and over-reliance on retrieved knowledge in scenario planning.
  • MultiFinBen adds the missing multilingual, multimodal, and audio-inclusive benchmark lane. It spans English, Chinese, Japanese, Spanish, and Greek across text, vision, and audio; adds PolyFiQA for cross-lingual filing/news reasoning; adds financial OCR datasets for scanned documents; and uses FinAudio-style earnings-call ASR and long financial-recording summarization. Its durable value is not the 2025 leaderboard but the modality-balanced fixture design and error taxonomy: table/chart omission, hallucination, dropped digits, edge omissions, and skipped bracketed or vertical text.
  • IPO Finance Agent adds a separate capital-markets diligence lane. It shows that Finance Agent / periodic SEC filing performance does not transfer automatically to IPO S-1 work: the task needs long-document contextual retrieval, governance/control analysis, common-control accounting, valuation without public trading history, professional workflow labels, public/private contamination controls, and human-reviewed generated rubrics.
  • Meta-Benchmarks for Financial-Services LLM Evaluation adds the missing firm-wide rollup layer. It maps 452 public benchmarks to 41 O*NET work activities and 38 BIAN banking business domains, weighting benchmark evidence by discrimination, coverage, and recency. Use it to screen candidates and monitor capability drift; do not use public business-domain Elo scores as a substitute for private workflow evals.
  • Kuvera-8B adds a personal-finance/advisor SLM lane: a MIT-licensed Qwen3-8B fine-tune trained on a roughly 19K-sample behaviorally grounded personal finance dataset derived from de-identified Reddit queries. It is interesting because it argues that curated financial + behavioral supervision can make an 8B model competitive with 14B-32B baselines at lower cost. It is not institutional quant research, alpha evidence, or regulated advice proof. The finance-advisor-portfolio-v0 seed fixture now gives Kuvera a scoped evaluation target for suitability, arithmetic, analyst-bias, portfolio pipeline, authorization, and runtime-portability guardrails.
  • A dedicated credit/AML lane is now required. FCMBench, LendNova, CALM, Omega2, TransXion, and FinRegLab show that underwriting and financial-crime work must be scored on exact fields, out-of-time splits, fairness, adverse-action explanations, graph/profile detection, model-risk management, and human authority, not only finance QA or AUC.
  • A dedicated advisor/portfolio-agent lane is now required. HERCULEAN, PortBench, One Size Fits None, Fin-Bias, and older personal-finance advisor tests show that a model can write fluent financial advice while failing suitability sensitivity, analyst-bias resistance, correlation-aware allocation, execution simulation, arithmetic, or committed agent actions. Score profile sensitivity, CEPS/PAS, herding under fake ratings, stress regimes, and authorization boundaries before using any model in advisory or portfolio workflows.
  • A dedicated asset-class lane is now required for FICC/rates, macro/event forecasting, commodities, and private credit. Yield-curve studies show classical econometric and naive baselines can beat deep models; PolyBench shows event forecasts need CLOB/news/liquidity/calibration scoring; commodity futures require carry, basis, roll, inventory, seasonality, and hedging pressure features; private credit requires covenant, valuation, fee, leverage, and opacity controls. Do not let equity/SEC success generalize silently into these markets.
  • A dedicated insurance/actuarial/reinsurance lane is now required. CUFEInsure, INSEva, Quebec insurance RAG, UNDERWRITE, INSURE-Dial, claim automation, commercial underwriting self-critique, and reinsurance quantitative reasoning show that insurance needs policy/regulatory citation, coverage/exclusion interpretation, noisy multi-turn underwriting, claims schema extraction, call compliance traces, actuarial/reinsurance calculators, and human authority gates. This is adjacent to credit and financial services, but it is not the same benchmark.
  • AntGroup Finix-S1 is now tracked as a closed insurance-domain reference signal from CUFEInse. The useful finding is not that we can run it; we cannot verify a public artifact. The useful finding is that a domain-adapted insurance model reportedly beats general/open baselines on business-process understanding, compliance, agent applications, and logical rigor. This strengthens the case for separate insurance-domain fixtures before comparing open Qwen/Gemma/DeepSeek/Llama candidates.
  • A dedicated treasury/payments/liquidity lane is now required. IMF agentic payments work shows that payment agents must separate probabilistic intent and orchestration from deterministic authorization/control and legal settlement. J.P. Morgan, KPMG, and Databricks add practical workflow signals: 24/7 treasury, deposit tokens/stablecoins, cash visibility, reconciliation, agent deployment, high-risk autonomy exclusions, and sensitive-data controls. Score financial authority, liquidity impact, reconciliation, audit trails, and kill-switch behavior before allowing any agent near money movement.
  • A dedicated structured-finance/securitization lane is now required. CMBS pooling and servicing agreements, MBS prepayment/default behavior, loan tapes, trustee reports, and CLO waterfalls are contract-defined cash-flow systems, not generic credit QA. Score PSA/indenture parsing, clause comparison, tranche/waterfall extraction, loan-tape QA, leakage-safe default modeling, prepayment assumptions, servicing events, deterministic cash-flow handoff, and source-grounded audit trails.
  • A dedicated post-trade/collateral/surveillance lane is now required. Margin calls, collateral eligibility, repo/securities finance, derivatives margin, DvP settlement, broker-dealer liquidity reports, surveillance alerts, GenAI supervision, books and records, and customer-asset protection are financial control tasks, not alpha tasks. Score action cards, entity/counterparty mapping, collateral haircuts, stress scenarios, report QA, evidence-backed alert handling, deterministic settlement/collateral checks, and human authority boundaries.
  • A dedicated investment-banking/deal-work lane is now required. BankerToolBench, Deloitte’s 2025 GenAI in M&A survey, FrontierFinance, and Houlihan Lokey practitioner evidence show that sell-side and deal-team agents must be scored on multi-file client-ready artifacts, not chat answers: senior-banker request parsing, data-room navigation, market/SEC research, Excel formula integrity, pitch-deck generation, diligence issue logs, cross-artifact reconciliation, confidentiality/MNPI controls, and banker review handoff.
  • A dedicated ESG/climate/sustainability lane is now required. Climate Finance Bench, ESGenius, robust greenwashing detection, Qwen3 ESG adaptation, and SusGen/TCFD-Bench show that asset-manager stewardship and disclosure work must be scored on source retrieval, standards mapping, emissions units, transition-plan actionability, greenwashing triage, framework-structured report drafting, abstention, and human/legal/compliance review. This is risk/disclosure infrastructure, not alpha evidence.
  • A dedicated tool-catalog fine-tuning lane is now required before training any repo-specific finance agent. The AssetOpsBench QLoRA paper is not finance-specific, but it directly tests whether Gemma 4 E4B and Qwen3-4B can internalize fixed tool knowledge and reduce prompt overhead. Treat it as training-method evidence for stable tool/schema catalogs, with catastrophic forgetting checks, not as a finance model score.
  • The time-series gap is now represented in the machine-readable model/runtime manifest. NeoQuasar/Kronos-base enters as a separate finance-time-series-market-v0 candidate, not as a finance replay/chat model. It needs an exact model revision, tokenizer revision, locked OHLCV fixture, runtime smoke proof, leakage policy, classical baselines, transaction-cost/capacity checks, and risk-adjusted scoring before any recommendation. FinCast now has verified Apache-2.0 HF/GitHub artifact evidence, but still needs exact revisions, dependency setup, locked market data, runtime smoke proof, and baseline comparisons before any local-runtime recommendation. FinSTaR adds the reasoning branch of this lane: FinTSR-Bench separates deterministic assessment from stochastic prediction over 120-day S&P closing-price windows, using Compute-in-CoT and Scenario-Aware CoT. The verified artifact is public GitHub code and a TimeOmni/Qwen2.5 LoRA training recipe, not a ready checkpoint. Relational Probing extends this lane from market-bar forecasting into text-to-graph financial prediction: Qwen3 0.6B/1.7B/4B SLM hidden states are trained through a relation head to induce financial-entity graphs for a downstream stock-trend model. Treat both as paper/code-first market-reasoning evidence, not runnable chat models or live-alpha proof.
  • Chronos-2 financial forecasting evidence adds an important negative control for the time-series lane: multivariate context helps when related series are modeled jointly, but mixing unrelated equity and interest-rate panels can reduce forecast accuracy. StateBench should therefore score related-series grouping and noisy-context rejection, not only raw forecast error.
  • A May 31 HF source pass adds four more domain-specific watchlist artifacts. mradermacher/Qwen3-8B-finance-V1.0.1-Merged-V2-GGUF is the most concrete new local finance-reasoning runtime candidate because it is Apache-2.0, Qwen3-based, and already available as llama.cpp/GGUF quants. PeterAM4/Qwen3- Embedding-0.6B-GGUF is not a finance model, but its imatrix/calibration corpus includes FiQA, FinanceBench, Twitter Financial News, and financial RAG data, so it belongs in the retrieval-cost-floor lane. Mikkkkoooo/qwen35-4b- private-analyst-full-corpus is useful only as a cautionary private-research- corpus control because the license is other and the training data is not auditable. DMindAI/DMind-3-mini adds a Web3/DeFi finance-local SLM lane, not a systematic quant asset-management lane.
  • A June 5 HF source pass adds a narrow derivatives/options model candidate: dtarkenton/sprocket-gex-deepseek-r1-distill-qwen14b-lora-paper-exact-final. It is a PEFT/QLoRA adapter on unsloth/DeepSeek-R1-Distill-Qwen-14B-bnb-4bit for structured Gamma Exposure pattern and regime classification. Its value is fixture design for options/GEX schema discipline, ambiguous/no-pattern cases, and no-trading-action refusals. It is not broad finance reasoning, alpha, portfolio construction, pricing, hedge execution, or risk-management proof.
  • A June 6 HF source pass adds two low-visibility domain-specific model candidates that belong in separate lanes. HarryS64/10k-financial-slm is an 11.5M-parameter MIT-licensed PyTorch GPT-style model trained on financial company 10-K filings; its value is as a tiny filing-language/compression, anomaly, and prefilter control, not as a chatbot or finance-reasoning model. Yana/compass-vlm is an Apache-2.0 Japanese financial VLM with 8.59B BF16 parameters, multimodal projector and vision-tower files, and a training stack tied to LLM-JP, SigLIP, Qwen3 reasoning/domain-data teachers, and EDINET-style finance document references. It belongs in Japanese financial statement, OCR/table/chart, and multimodal source-triage work, not U.S. SEC text replay or investment analysis. Both are watchlist artifacts until runtime smoke tests and scoped fixtures exist.

Model Landscape

Model / System What It Is Best Use Case Caveat
BloombergGPT 50B finance-domain LLM trained on Bloomberg financial corpus + public data Historical reference for finance-domain pretraining Closed; not directly useful for local deployment/fine-tuning
FinGPT Open-source data-centric finance LLM framework Data pipeline, financial extraction experiments, sentiment/news tasks Ecosystem/reference more than a single winning checkpoint
PIXIU / FinMA Open-source financial LLMs + benchmark suite Finance instruction tuning, FLARE-style financial tasks Older generation; must benchmark against Qwen/Gemma/Llama modern bases
InvestLM Financial instruction-tuned LLM Investment-domain instruction following Needs current benchmark comparison before adoption
FinTral Financial LLM research line Multilingual/multimodal financial analysis Verify model availability and runtime fit before use
Open-FinLLMs / FinLLaVA Open multimodal finance LLM suite spanning text, tabular, time-series, and chart data Historical open finance-VLM baseline and data/training recipe reference Older LLaMA-era base; rerun against current Qwen/Gemma VLMs and internal tasks
Aveni FinLLM 7B financial-services model UK regulated financial services language, advice/support/compliance workflows Not a quant-alpha model
FinRobot Finance AI-agent framework Architecture reference for finance agents Research/framework signal; production value depends on tool/data integration
HedgeAgents / FinAgent / FinMem family Trading/decision research agents Agent architecture and memory patterns Treat as research-grade until independently reproduced
Fin-R1 7B financial reasoning model trained with SFT + RL from DeepSeek-R1-distilled financial reasoning data Local finance-reasoning baseline Older than Qwen3-era models; benchmark before adopting
Fino1 / Fin-o1 Llama-3.1-8B-based financial reasoning model using FinCoT + RL Open finance reasoning / CoT baseline Good research baseline, not necessarily current SLM winner
DianJin-R1 Financial reasoning model/dataset using CFLUE, FinQA, and compliance corpus Chinese/compliance-heavy financial reasoning Proprietary compliance corpus limits full reproducibility
Agentar-Fin-R1 Qwen3-based 8B/32B financial reasoning model series Current Qwen-family finance candidate Verify license/model card and runtime artifacts before use
ODA-Fin-SFT/RL-8B Apache-2.0 Qwen3-8B finance post-training line using 318K distilled CoT SFT samples plus 12K hard/verifiable GRPO tasks Current open Qwen3 finance reasoning/post-training candidate for FinQA/TaTQA/ConvFinQA, sentiment, and broad finance QA Paper/model-card scores only; no GGUF/MLX artifacts verified; aggregate dataset licensing needs review
CNFinBench Finance benchmark for expertise, autonomy, and integrity across financial QA, report parsing, DB/API operations, compliance, fraud, internal/application security, and multi-turn adversarial degradation Safety/compliance and high-privilege finance-agent gate for StateBench; fixture sa-059 materializes it for replay Not a quant-alpha benchmark; leaderboard/model claims are time-sensitive and should be rerun locally
EDINET-Bench Japanese annual-report benchmark for accounting fraud detection, earnings forecasting, and industry classification over full financial statements Multilingual financial-statement eval lane with classical-model controls; StateBench fixture sa-060 materializes it for replay Stable dataset/task evidence; model scores are time-sensitive and simple zero-shot LLM setups underperform expectations
FinFRE-RAG Feature-reduced retrieval-augmented in-context framework for structured transaction fraud detection Retrieval/runtime harness candidate for fraud triage over tabular data; StateBench fixture sa-061 materializes it for replay RAG narrows the LLM/classifier gap but does not remove need for specialized classifiers, privacy controls, and analyst review
FinGround Financial hallucination detection and grounding pipeline using atomic claim verification, type-routed checks, formula reconstruction, and paragraph/table-cell citations Groundedness/citation gate for financial RAG and filing QA; StateBench fixture sa-062 materializes it for replay Current 2026 method evidence; model/latency claims need rerun, but retrieval-equalized evaluation is a strong harness pattern
From BM25 to Corrective RAG Financial text-and-table retrieval-strategy benchmark over T2-RAGBench with 23,088 queries and 7,318 documents Static-index/wiki retrieval and finance RAG harness evidence; StateBench fixture sa-121 materializes the source card BM25/hybrid/rerank/contextual-index design evidence only; not a universal embedding leaderboard, production deployment proof, or backend portability claim
LLM fraud-warning pressure experiment Preregistered investment-fraud advisory experiment with LLMs and human benchmark under motivated investor pressure Retail/advisory fraud-warning and pressure-turn safety eval; StateBench fixture sa-064 materializes it for replay Not a quant-alpha benchmark; model-version rankings and prompt sensitivity need local/current reruns
PHANTOM Financial long-context hallucination-detection benchmark over SEC filings with controlled context length and evidence placement Source-faithfulness gate for long filings and RAG outputs; StateBench fixture sa-063 materializes it for replay 2025 benchmark evidence; rerun with current Qwen/Gemma/frontier models before using rankings
CALM Credit and Risk Assessment LLM plus multi-dataset benchmark for credit scoring, fraud detection, and financial distress Historical domain-specific credit/risk LLM baseline and fairness/class-imbalance task source; StateBench fixture sa-065 materializes it for replay 2023/2024 model evidence is stale; use as task/control evidence, not current default
LendNova Language-model pipeline over raw, jargon-heavy credit bureau text Credit-risk agent task for raw-record parsing, representation learning, and auditable rationale; StateBench fixture sa-070 materializes it for replay Early AAAI 2026 workshop evidence; high-impact credit use needs adverse-action/fairness/model-risk controls
Omega2 Hybrid corporate-credit scoring framework combining LLM reasoning with structured financial indicators and CatBoost/LightGBM/XGBoost Corporate credit-scoring lane with temporal validation and classical-model controls; StateBench fixture sa-067 materializes it for replay Reported AUC/rating-agency generalization claims need independent reproduction
FCMBench Financial credit multimodal benchmark with certificate images, VQA, perception, reasoning, and robustness tasks Credit-document OCR/VLM and visual-evidence eval lane; StateBench fixture sa-066 materializes it for replay Model rankings are dated; rerun current VLMs and local candidates
TransXion Profile-rich AML transaction-graph benchmark with non-template illicit-subgraph synthesis AML graph/KYC-context anomaly-detection stress test; StateBench fixture sa-068 materializes it for replay Synthetic benchmark, not production bank ground truth
FinRegLab ML underwriting framework Practitioner governance framework for ML consumer-credit underwriting Model-risk, adverse-action, validation, monitoring, vendor, and second-look control checklist; StateBench fixture sa-069 materializes it for replay Not a leaderboard; U.S. regulatory posture is date/jurisdiction sensitive
HERCULEAN MCP/skill-based agentic financial-intelligence benchmark across trading, hedging, market insights, and auditing Agent workflow benchmark for tools, temporal constraints, committed artifacts, and XBRL audit verification Limited assets/workflows; high inference cost and some LLM-as-judge scoring
PortBench Correlation-aware full-pipeline portfolio-management benchmark with static QA and dynamic five-stage allocation Portfolio pipeline eval: market interpretation, signals, weights, execution, risk monitoring, CEPS, PAS, stress regimes Model rankings are time-sensitive; use task design and equal-weight controls
One Size Fits None Investment-advice suitability benchmark using synthetic client profiles and feature-concentration analysis Advisor input-sensitivity audit for risk tolerance vs. age, income, horizon, and liquidity Web-search condition can change market grounding and personalization in opposite directions
Fin-Bias Analyst-report perturbation benchmark with rating removed/faked and realized-return labels Bias/herding test for models reading analyst reports, broker research, or PM notes Includes stale model families; use as contamination task evidence
Yield curve ML/econometrics comparative studies Treasury curve forecasting across ARIMA/VAR/DNS/FADNS, naive methods, classical ML, deep learning, transformers, and TSFMs Fixed-income/rates baseline lane with maturity, horizon, rolling-window, and publication-lag checks Strong warning that ML is not automatically better than econometrics
PolyBench Live prediction-market benchmark with CLOB snapshots, contemporaneous news, resolution criteria, JSON decisions, SKIP, and return scoring Event/macro forecasting and calibration lane; StateBench fixture sa-071 materializes it for replay Prediction markets are not institutional FICC trading; use design pattern, not alpha proof
Private credit market survey Direct-lending and private-credit market-structure survey Private-credit monitoring, covenant, valuation, fee, systemic-risk, and data-gap task source Not a model benchmark; public data is structurally incomplete
CFA commodity-futures ML chapter Theory-grounded boosted-tree commodity futures signals across carry, basis, momentum, skewness, liquidity, and horizon ensembles Commodity feature-construction and execution/capacity lane Research/backtest evidence only; requires contract-roll and liquidity controls
CUFEInsure Insurance LLM benchmark with 5 dimensions, 54 sub-indicators, and 14,430 questions Insurance theory, business understanding, compliance, agent application, and logical-rigor gate Paper-reported rankings are time-sensitive; actuarial and compliance tasks need local reruns
INSEva Chinese insurance benchmark with 38,704 examples Multilingual/regional insurance QA and business-process evaluation Jurisdiction-specific; not a universal insurance leaderboard
Quebec insurance RAG / AEPC-QA 807-question Quebec insurance regulatory-certification QA benchmark Closed-book vs. retrieval-grounded jurisdictional insurance QA Private gold set; use task design and reproduce with internal sources
UNDERWRITE Expert-first multi-turn commercial underwriting agent benchmark with SQLite/MCP tools, proprietary rules, noisy interfaces, and imperfect simulated users Enterprise insurance agent readiness, follow-up questioning, tool use, hallucination, and efficiency Controlled benchmark; framework brittleness and LLM-judge choices must be audited
INSURE-Dial Phase-aware insurance benefit-verification call dataset and benchmark Call compliance, phase-boundary detection, and procedural trace scoring Contact-center/compliance lane, not a general finance QA task
Claim automation LoRA study Local DeepSeek-R1 8B LoRA fine-tune over roughly 2.0M automotive warranty claims Claims narrative-to-structured-action extraction and fine-tuning gate Proprietary data and narrow warranty scope limit reproducibility
Commercial underwriting self-critique Human-in-the-loop, decision-negative underwriting agent with adversarial critic and read-only tools Authority-bound agent architecture and critic-overhead eval Controlled 500-case study; treat reported gains as architecture signal until field-validated
Reinsurance quantitative reasoning Prompting/fine-tuning study over reinsurance allocation calculation tasks Actuarial/reinsurance math, treaty logic, and domain fine-tuning evidence Older Llama 2-era model base; use task evidence, not current model ranking
IMF agentic payments architecture Three-layer payments framework separating intent/orchestration, authorization/control, and settlement Payment authority, liquidity, settlement-finality, agent identity, and systemic-risk gate Policy/architecture note, not a model leaderboard
J.P. Morgan Payments Outlook 2026 Vendor/industry report on payments modernization, tokenization, 24/7 settlement, and treasury workflows Treasury/payments workflow vocabulary: pre-validation, fraud checks, deposit tokens, stablecoins, programmable payments, reconciliation Vendor signal; use requirements, not performance claims
KPMG Banking AI Pulse Q1 2026 Banking AI adoption/governance survey Agent deployment and guardrail baseline for banks Consulting survey; adoption evidence only
Databricks Financial Services Outlook 2026 Vendor outlook on banking/payments data, CFO transformation, balance sheet intelligence, and agentic workflows Cash visibility, manual reconciliation, liquidity buffer, semantic data, and governance fixtures Vendor architecture signal, not benchmark evidence
JFQA securitization agreements Machine-learning study of CMBS pooling and servicing agreements and contract uniqueness Long-document contract extraction, PSA comparison, servicing/conveyance clause localization, tranche/cash-flow outcome linkage Contract/document evidence, not a model leaderboard
Mortgage default AutoML/leakage study Fannie Mae loan-level default prediction with temporal split, leakage-aware feature selection, imbalance handling, and classical/AutoML baselines Mortgage-collateral default-model hygiene and point-in-time validation Reported model rankings are setup-specific; use leakage/split discipline as the main benchmark rule
MBS prepayment / default collateral-pool studies Statistical prepayment modeling plus Fannie Mae collateral-pool classification and competing-risks analysis Prepayment-speed, default-hazard, servicer-monitoring, and cash-flow assumption baselines Non-LLM risk-model controls; useful for deterministic structured-finance fixtures, not agents
SFA CLO white paper Industry primer on CLO collateral pools, tranches, waterfalls, safeguards, eligibility, and monitoring CLO domain vocabulary and waterfall fixture design Industry source; use mechanics, not performance claims
BIS/FSI margin-call liquidity summary FSB policy summary on liquidity preparedness for margin and collateral calls Margin-call triage, collateral availability, haircuts, liquidity stress, contingency funding, and counterparty escalation Policy/governance source, not model evidence
FINRA 2026 oversight report Broker-dealer regulatory report covering GenAI, supervision, market integrity, liquidity, customer assets, and reporting Surveillance alerts, off-channel communications, books/records, SLS QA, customer-asset reconciliation, and liquidity controls Regulatory source; U.S. broker-dealer specific
ISDA collateral/liquidity efficiency paper Derivatives collateral and liquidity efficiency industry white paper Derivatives margin, collateral optimization, repo/clearing links, settlement windows, and stress operations Industry source; use workflow mechanics, not performance claims
GFMA digital money capital-markets report Trade-association report on digital money for securities settlement, repo/securities finance, and derivatives margin DvP settlement, intraday programmable margin, repo collateral automation, post-trade affirmations, and digital-money control checks Trade-association architecture source
FinSphere Stock-analysis agent with real-time DB, quantitative tools, Stocksis, and AnalyScore Architecture/eval reference for analyst-style reports Report quality is not investment alpha
rLLM-FinQA-4B Public Apache-2.0 Qwen3-4B-Instruct-2507 finance QA/tool-use agent fine-tuned with RL on SEC 10-K tables Local SLM candidate for filing/table QA with SQL/table/calculator tools HF card lists vLLM serving and safetensors, not GGUF/MLX artifacts; benchmark in our harness before adoption
TraceAlchemy-Gemma-4-E4B-Finance-IT Gemma 4 E4B finance instruction model trained with Unsloth LoRA SFT on FinanceBench, TAT-QA, ConvFinQA, FinanceReasoning, Finance-Instruct, and synthetic SEC/table reasoning examples Gemma-family candidate for finance statement/table reasoning, unit/scale conversion, and final-answer consistency External benchmark accuracy not yet reported; use as candidate, not proven winner
Table-R1 family Table reasoning adaptation line: DeepSeek-R1 trace distillation, RLVR/GRPO, program-based SLM reasoning, and region-aware table RL Training recipe candidate for local financial table, spreadsheet, and XBRL agents Mostly generic table evidence; rerun with current Qwen/Gemma bases and finance-specific table fixtures before adoption
Amsi-fin-o1 Apache-2.0 Qwen3-VL 4B finance vision-language fine-tune Financial document images, charts, OCR, PDF-page understanding, and visual chain-of-thought evals HF card has no independent benchmark; safety behavior needs testing because of the intermediate abliterated base
fin-llm-qwen3.5-9b-gguf Finance-specialized Qwen3.5-9B GGUF model llama.cpp/GGUF runtime candidate for ratio analysis, valuation, portfolio math, and formula-driven finance explanations License marked “other”; no independent benchmark; evaluate before use
Qwen-Open-Finance-R-8B Apache-2.0 gated Qwen3 finance/economics/business model Multilingual finance and economics QA in English/French/German Gated access and no independent benchmark promoted here
Dev9124/qwen3-finance-model Apache-2.0 Qwen3 finance-instruction fine-tune using Finance-Instruct-500k signal and Unsloth/TRL model-card tags Broad finance-instruction control for source synthesis and wiki-maintenance comparison Model-card evidence only; no independent benchmark promoted; training-data lineage and runtime fit need preflight
YorkFr/financial-sentiment-qwen3-v2 MIT-licensed merged Qwen3-0.6B LoRA for financial-news sentiment classification Tiny cheap sentiment/routing control for headlines, earnings snippets, and market commentary First Transformers-local-MPS run scored 0/15 accepted with 0.667 label accuracy but failed route/rationale controls; do not promote for workflow routing
iravikr/qwen3-0.6b-finance-india Apache-2.0 Qwen3-0.6B India budget/policy QA fine-tune with 4-bit bitsandbytes metadata India budget and government-document micro-control through finance-regional-language-v0 Regional policy scope only; not a general quant, U.S. filings, accounting, or investment-research model
THaLLE-0.2-ThaiLLM-8B-fa Apache-2.0 ThaiLLM/Qwen3-8B mergekit finance-language model with English/Thai tags Thai/English finance-language and local-market document QA watchlist through finance-regional-language-v0 Merge lineage and language scope must be scored separately from Qwen3 base; no U.S. finance transfer proof
Ayansk11/qwen3-4b-financial-sentiment-grpo Apache-2.0 Qwen3-4B sentiment model with FinGPT sentiment data, GRPO/Unsloth tags, and Q5_K_M GGUF sibling Smaller Qwen3 sentiment/GGUF candidate for finance-sentiment-routing-v0 Model-card Ollama tag ignored; HF and GGUF artifacts require separate parser/runtime preflights
sweatSmile/Qwen3-4B-Instruct-FinanceQA Apache-2.0 FinanceQA adapter-checkpoint artifact whose HF metadata records Qwen2.5-3B-Instruct base lineage despite Qwen3 repository name Lineage-audit and historical FinanceQA adapter control Do not treat as current Qwen3; adapter-only status and base mismatch need preflight before scoring
ModernFinBERT-base Apache-2.0 ModernBERT sequence classifier for financial sentiment Modern encoder baseline for label-only finance sentiment StateBench local MPS run scored 0.733 label accuracy but 0/15 full routing acceptance; useful label baseline, not workflow router
ProsusAI/finbert Classic BERT sequence classifier for financial sentiment Older cheap sentiment-control baseline StateBench local MPS run scored 0.667 label accuracy and 0/15 full routing acceptance; weaker than ModernFinBERT on this seed
Qwen3.6-35B-A3B Apache-2.0 open-weight Qwen3.6 MoE image-text model with 35B total / 3B active parameters and long-context support Current general base candidate for repository research, long-context source review, multimodal PDF/chart/OCR, and agent harness tests Not finance-specific; model-card benchmark numbers are vendor-reported and must be tested against internal finance workflows
Qwen3.7-Max Proprietary Qwen-family agent model released May 20/21, 2026 with official claims around long-horizon autonomous tasks, coding, office workflow automation, MCP/multi-agent orchestration, and 1,000+ tool-call runs Current Qwen-family agent reference for API comparison on source discovery, spreadsheets, coding, and long-horizon tool workflows Not open-weight/local in this pass; not finance-specific; vendor-reported scores need independent/internal reruns
Mihenk-LLM v2 35B-A3B Turkish Financial Model Apache-2.0 Qwen3.6-35B-A3B finance fine-tune for Turkish/BIST financial statements, crypto/macro analysis, risk framing, and safe financial responses; community GGUF conversion now verified separately at emircansevdi/Mihenk-LLM-v2-35B-A3B-Turkish-Financial-Model-GGUF Multilingual/local-market finance candidate for Turkish/BIST tasks and Qwen3.6 fine-tune comparison through finance-regional-language-v0; GGUF adds local llama.cpp runtime lane only after exact-file and smoke preflight Only model-card smoke eval: 8-prompt A/B judge, 5 vs. 3 against base; no independent benchmark; community GGUF conversion must be scored separately from upstream HF/BF16
Gemma-Pro-Finance-12B Gemma-3-12B-based finance/economics/business model with English/French/German tags Gemma-family finance-language candidate for internal evaluation Custom license forbids redistribution and commercial/production use without separate agreement despite HF Apache-2.0 header; procurement blocker
Srx7703/gemma-4-31b-financial-adapter Gemma 4 31B PEFT LoRA adapter for SEC filing analysis, trained on knowledge-distilled SEC QA pairs and linked to a tool-using earnings-recap agent repo Current Gemma 4 adapter candidate for SEC filing summarization/QA, source-cited earnings-recap synthesis, and adapter-vs-base StateBench runs Adapter-only artifact under Gemma terms; tiny n=20 BERTScore eval; no merged/GGUF/MLX/Neuron artifact verified; companion agent currently uses API models
Sengil/turkish-gemma-9b-finance-sft Adapter-only Turkish finance SFT on ytu-ce-cosmos/Turkish-Gemma-9b-T1, a Gemma2-derived Turkish base; card lists Turkish finance instruction datasets and Unsloth SFT Regional Turkish finance-language Gemma-family control for finance-regional-language-v0 and adapter-vs-Qwen/Mihenk comparison No independent benchmark; adapter-only artifact; base metadata says Gemma license despite adapter Apache-2.0 metadata; model-card sample exposes visible <think> reasoning; no merged/GGUF/MLX/Neuron proof
Won Korean financial LLM Open Korean finance LLM plus 80K instruction dataset derived from an eight-week Korean finance leaderboard with 1,119 submissions Multilingual/local-market finance candidate and training-practice signal for Korean finance/accounting, company-analysis, market, stock-prediction, and finance-agent tasks Local-market benchmark; do not extrapolate Korean leaderboard results to U.S. quant research or English filings
BizFinBench Chinese business-driven real-world financial benchmark with 6,781 queries across numerical calculation, reasoning, information extraction, prediction recognition, and knowledge QA Multilingual/local-market benchmark-design lane for temporal reasoning, table calculation, event attribution, entity extraction, tool-usage, and LLM-as-judge caveats Paper-snapshot model rankings are dated; use task taxonomy and IteraJudge design, not leaderboard claims
XFinBench ACL 2025 complex financial problem-solving benchmark with 4,235 graduate-level examples and multimodal context Benchmark-design lane for temporal reasoning, future forecasting, scenario planning, numerical modelling, chart/curve interpretation, and human-baseline comparison Paper-snapshot model rankings are dated; use capability and error taxonomy, not leaderboard claims
MultiFinBen Multilingual and multimodal financial benchmark spanning EN/ZH/JA/ES/EL plus text, vision, and audio Benchmark-design lane for cross-lingual filing/news QA, financial OCR over scanned documents, earnings-call ASR, long financial audio summarization, and modality-balanced scoring Paper-snapshot rankings are dated; use task/error taxonomy and rerun current models before any leaderboard claim
Kuvera-8B-qwen3 MIT-licensed Qwen3-8B personal-finance fine-tune trained on behaviorally grounded Reddit-derived advice data Advisor/personal-finance SLM candidate for budgeting, debt, saving, investing basics, and behavioral-bias-aware guidance Synthetic responses and Reddit source bias; not financial advice, not current-regulation aware, not complex planning proof, and not quant research
RinKana / Qwen3-8B finance variants Unsloth Qwen3-8B finance fine-tunes with GGUF conversions and synthetic finance-reasoning dataset tag Runtime candidate for llama.cpp/GGUF experiments Model cards are thin and qualitative; lower confidence than rLLM-FinQA or Fino1
mradermacher/Qwen3-8B-finance-V1.0.1-Merged-V2-GGUF Apache-2.0 community GGUF quantization of RinKana’s Qwen3-8B finance merge with finance-reasoning-synthetic dataset tag Concrete llama.cpp/GGUF finance-reasoning control for local replay, accounting/reasoning smoke tasks, and quantization comparison Community quant artifact; score separately from upstream merged HF and imatrix sibling; HF card mentions Ollama but repo excludes Ollama
PeterAM4/Qwen3-Embedding-0.6B-GGUF Apache-2.0 Qwen3 embedding GGUF with importance-matrix variants and finance-related calibration/data tags Tiny local embedding/retrieval-cost-floor candidate for source localization, PDF/podcast metadata retrieval, and hybrid retrieval experiments Finance-calibrated imatrix is not FinMTEB proof; compare against BM25, larger Qwen3 embeddings, FinE5, BGE, Arctic Embed, and internal recall labels
Mikkkkoooo/qwen35-4b-private-analyst-full-corpus Qwen3.5-4B private financial research-report corpus fine-tune with BF16 safetensors and Q4_K_M GGUF Cautionary analyst-report style control for private-corpus SFT and contamination/overfitting tests License other, private dataset, no independent benchmark; do not promote without license review and contamination checks
DMindAI/DMind-3-mini Apache-2.0 Qwen3.5-4B-based Web3/DeFi finance/security-audit SLM Crypto/Web3/DeFi finance-local watchlist for smart-contract, DeFi-risk, and crypto-source tasks Not traditional asset management, SEC filings, accounting, or quant research; needs separate Web3 fixtures
distil-labs/distil-lfm25-voice-assistant 354M BF16 LiquidAI/LFM2.5-350M banking voice-assistant fine-tune with tool-calling/function-calling tags Tiny banking voice-agent and tool-routing candidate for audio-agent or customer-service workflow fixtures Liquid open-license terms need review; not investment research, trading, compliance, or alpha evidence
TheTokenFactory/gemma-4-E2B-sec-extraction-GGUF-v3 Gemma-license GGUF/llama.cpp SEC extraction fine-tune on unsloth/gemma-4-E2B-it, with unverified card metrics for JSON parse and hallucination phrase rate Local structured SEC/financial-contract extraction candidate for filing snippets and source-card extraction tests Narrow extraction artifact; not a broad finance reasoner; exact GGUF, parser, citations, and numeric reconciliation must be scored
AntGroup Finix-S1 Closed Ant insurance-domain model reported by CUFEInse as first overall with total score 89.51 Reference signal for insurance-domain adaptation, especially business-process understanding, compliance, agent applications, and logical rigor No public model artifact, API, license, parameter count, architecture, or backend proof verified; not runnable in StateBench
Trading-R1 / Alpha-R1 / Trade-R1 Financial trading / alpha screening RL research line now captured through local PDF/OCR pipeline and source card Structured thesis generation, factor screening, stochastic-reward RL caution, reward-hacking controls Treat return metrics as research-grade until reproduced with leakage controls, costs, capacity checks, and baselines
Sprocket GEX DeepSeek R1 Distill Qwen 14B LoRA PEFT/QLoRA adapter for structured Gamma Exposure pattern and 30-day regime classification from numerical GEX features Options/GEX schema classifier and derivatives fixture design for pinning, 0DTE hedging, persistent-positive/negative, transitional, low-conviction, and no-pattern controls license: other, adapter-only, paper-exact/partly synthetic data, 32-case smoke eval only; no trading action, alpha, advice, portfolio, pricing, GGUF, MLX, Neuron, or broad quant-research proof
FinE5 Finance-domain embedding model released by the FinMTEB authors Finance retrieval baseline to beat on filings/transcripts/news/ESG/regulatory corpora CC-BY-NC-ND and gated on HF; research reference, not commercial default
finance-embeddings-gemma-300m-v2 Finance-specific EmbeddingGemma 300M full fine-tune with 303M parameters, 512-token max length, 768 hidden size, and 2.89M financial text samples Small finance-domain embedding control for source localization, wiki/static-index retrieval, and RAG corpus search Model-card evidence only; no independent FinMTEB/StateBench score; license inherits from base EmbeddingGemma; no GGUF/MLX/Neuron proof
Qwen3 Embedding / Qwen3 Reranker Current Apache-2.0 general embedding and reranking families with 0.6B/4B/8B sizes and 32k-token context Practical open retrieval/reranking candidates for local FinMTEB and internal source-recall tests Not finance-specific; must beat hybrid BM25+dense baselines on finance tasks
Qwen3-VL Embedding / Qwen3-VL Reranker Apache-2.0 multimodal retrieval/reranking candidates PDF page, chart, table, investor-deck, and scanned-report retrieval Not finance-specific; test against document/page/table ground truth before deployment
LiquidAI LFM2 / LFM2.5 Hybrid small-model family with text, VL, audio, retrieval, extract, tool, transcript, and PII nano variants Edge/on-device extraction, transcript processing, low-latency RAG/tool-routing, and runtime-efficiency experiments Not finance-specific; LFM Open License has a $10M revenue threshold for free commercial use
ChronoBERT / ChronoGPT / ChronoInstruct Chronologically consistent annual-vintage language models trained only on text available by each cutoff year, with public ManelaLab HF artifacts Leakage-aware finance text embeddings, historical news-return backtests, and cutoff-controlled instruction experiments Domain-relevant for temporal controls; short context/custom architectures mean CUDA/MLX/Neuron portability must be checked per vintage and task
LLM news-embedding text-alpha models BERT/RoBERTa/LLaMA/OpenAI embeddings mapped to expected returns from professional news feeds Text-derived return signal benchmarks, especially negation/context-complexity tests Gross returns are not deployability; require licensed timestamps, costs, capacity, and value-weighted checks
Prompted LLM market-information models GPT-style zero-shot headline assessment for initial market reaction and subsequent drift Market information-processing probes and prompt-vs-embedding comparisons Reproducibility, model cutoff, high turnover, small-cap concentration, and API snapshot drift are major risks
Arctic-Text2SQL-R1-7B Apache-2.0 Snowflake Text-to-SQL specialist using Qwen2.5-Coder lineage and GRPO execution/syntax rewards Runnable open SQL control for Postgres/Snowflake/custom-schema harness experiments Older Qwen2.5 lineage; use as SQL specialist baseline, not as a current Qwen3-family finance reasoning default
Kronos Financial K-line / candlestick foundation model with mini/small/base HF releases Finance-native forecasting, volatility, and synthetic OHLCV sequence experiments Not a chat model and not an alpha engine; benchmark with leakage, costs, and runtime portability checks
FinCast Financial time-series forecasting foundation model with Apache-2.0 HF/GitHub artifacts, 1B sparse-MoE architecture, Point-Quantile loss, frequency embeddings, and 20B+ time-point training corpus Zero-shot / cross-domain financial forecasting candidate for crypto, forex, futures, stocks, and macro indicators Exact revisions, dependency setup, locked data, runtime smoke, classical baselines, costs, and backend portability remain pending
TimesFM / Chronos / Moirai / Lag-Llama General time-series foundation models Controls for forecast, quantile, and probabilistic time-series tasks Strong generic TSFM evidence does not imply trading utility
LOBERT / ByteGen / LiT Limit-order-book foundation, generative, and transformer forecasting models Message-level LOB representation, synthetic order-flow generation, short-horizon movement forecasting, execution simulation tests Venue-specific and latency-sensitive; synthetic realism or mid-price accuracy is not live-trading proof
Hubble-style factor mining DSL-constrained LLM formula generation with AST sandbox, dual-channel RAG, and family-aware scoring Safe, reproducible alpha-factor hypothesis generation with deterministic evaluation artifacts Needs neutralization, walk-forward, cost/capacity checks, and narrative-bias controls
FactorEngine-style program evolution Executable program-level factor evolution with LLM logic mutation, Bayesian parameter search, and report-derived factor bootstrapping Source-to-factor conversion and auditable code-level feature engineering Larger safety surface than formula DSL; require strict I/O, allowlists, deterministic execution, and leakage checks
Agentic systematic factor investing Autonomous factor-generation loop with economic rationale, IS-only gates, OOS validation, and transaction-cost/turnover checks Benchmark pattern for promote/hold/retire research decisions Strong reported results need independent reproduction; high-turnover strategies require capacity and market-impact stress
QuantMind Context-engineering framework for quant research knowledge extraction and retrieval Point-in-time, provenance-preserving corpus layer for filings, earnings calls, research notes, tables, formulas, podcasts, and video Controlled study is small; treat as architecture/eval-design evidence, not as proof of production accuracy
FinRL-X Deployment-consistent modular quant trading infrastructure with data, strategy, backtesting, and broker execution layers Research-to-paper/live consistency harness for StateBench trading workflows Paper-trading window is limited; return metrics are not live-alpha proof
Agentic screening + precision-matrix weighting LLM fundamentals screen + FinBERT sentiment screen + high-dimensional portfolio weighting Portfolio-selection evals where the model proposes a screened universe before optimization Research-grade backtest evidence only; evaluate leakage, consensus rule, turnover, constraints, and live deployability
CAB covariance forecaster 3D CNN + BiLSTM + multi-head attention over covariance sequences Medium-term covariance forecasting and GMV portfolio risk control Needs heavy-tail, transaction-cost, liquidity, and constraint checks before deployment
Characteristic-similarity GNNs GCN/GAT stock graphs built from firm-characteristic similarity rather than only return correlation Relationship-aware stock prediction and graph-construction critique Balanced-universe evidence; test unbalanced universes, IPOs/delistings, signed edges, provenance, and turnover
MARS risk-aware multi-agent RL Heterogeneous Safety-Critic agent ensemble plus Meta-Adaptive Controller Sequential portfolio allocation under changing regimes and explicit risk profiles Research-grade RL backtest evidence; deployment requires slippage, market-impact, mandate, and approval controls
No-arbitrage volatility-surface reconstructors Transformer, U-Net, CNN, VAE, and SVI-style baselines for sparse implied-volatility grids Surface completion, missing-wing/maturity robustness, and no-arbitrage violation checks Published evidence is SPY-heavy and grid-specific; soft penalties do not guarantee hard no-arbitrage
Market-IV VAE option pricer VAE compresses full implied-volatility surface into latent state; MLP prices American/Asian options Controlled pricing experiment using market-implied state rather than only spot/strike/maturity Labels are QuantLib-generated in the paper; use as benchmark/control, not proof of exotic desk pricing quality
Constrained neural pricing/hedging networks Single price-function network with gradient hedge and terminal/self-financing constraints Hedge P&L distribution, Greeks consistency, and incomplete-market stress tests Simulated setting; production needs transaction costs, liquidity, jump/regime misspecification, and live reconciliation
AssetOpsBench QLoRA tool-knowledge fine-tuning Gemma 4 E4B and Qwen3-4B fine-tuned with 8-bit QLoRA on roughly 1,700 tool-use examples Training-method control for stable tool catalogs, schema/tool planning, and prompt-overhead reduction Not finance-specific; LLM-judge components and catastrophic forgetting require internal checks before adoption

Task-Specific Encoders and Extractors

These models are older and smaller than the Qwen/Gemma finance-reasoning candidates, but they cover tasks that repeatedly appear in real finance automation. The right question is not whether they can chat. The right question is whether they beat a large general model on a narrow, auditable extraction or classification task at lower cost.

Specialist Task Why It Matters Caveat
ProsusAI FinBERT Financial sentiment classification Cheap baseline for headlines, filings, news, and earnings-call snippets Sentiment is not asset-specific trade direction
Qwen3-0.6B financial sentiment variants Financial news sentiment classification Current tiny Qwen-family control against FinBERT-class sentiment models Model cards are often thin; only use after held-out headline eval
FinSenti-Qwen3.5 family Financial headline, earnings-snippet, and market-commentary sentiment with parseable reasoning/answer output Current Qwen3.5 specialist trained with SFT + GRPO through Unsloth/TRL; useful when the harness needs auditable sentiment routing rather than a chat answer 9B variant needs roughly 20GB BF16 GPU memory; HF Qwen3.5 image-text-to-text adapter, tag parser, and local GGUF backend are still unproven; not an alpha signal
Fin-ModernBERT ModernBERT continued pretraining on finance/news/crypto corpora Modern encoder baseline for financial classification, NER, clustering, and representation tests Self-reported eval only; low adoption signal; reproduce before relying
ModernFinBERT Apache-2.0 ModernBERT financial sentiment classifier Current FinBERT-class sentiment control for earnings-call, press-release, tweet, and PhraseBank-style labels English-only, three-class sentiment, skewed source mix, no continued MLM pretraining
HKUST finbert-tone Analyst-report / financial communication tone Pre-trained on 10-K/10-Q, earnings-call, and analyst-report text; fine-tuned on manually annotated analyst sentences Verify license before commercial redistribution
HKUST finbert-esg ESG category classification Low-cost Environmental/Social/Governance/None routing for reports and annual filings Not a materiality, controversy, or greenwashing score
ESG-BERT Sustainable-investing text mining Useful ESG watchlist baseline with reported F1 improvement over generic BERT Model card has incomplete license/training/eval details
ClimateBERT Climate disclosure / TCFD / environmental claims Apache-2.0 climate language model family for transition-risk and sustainability-report NLP Climate-specific, not broad finance reasoning
SEC-BERT SEC filing language and numeric context Pre-trained on 260,773 10-K filings; variants address numeric token fragmentation Encoder/MLM baseline, not an instruction agent
FiNER-139 XBRL numeric entity tagging 1.1M-sentence / 139-label benchmark for SEC numeric extraction Dataset/task, not a general model
FiNER-ORD Open financial NER Public financial NER dataset/model/code signal for finance-domain entity extraction CC-BY-NC source; check downstream license
GLiNER Open-type NER control Useful for custom finance labels, PII, counterparties, instruments, and entities Generic model; must be fine-tuned/evaluated on finance labels

For StateBench, include a separate small specialist lane in the finance model matrix. A local Qwen/Gemma/Llama finance model should have to beat these cheap encoders on their own jobs before being used for routing, tagging, ESG classification, or extraction.


Benchmark Landscape

Benchmark What It Measures Why It Matters
FinMTEB Financial-domain embedding tasks General MTEB rank does not predict finance retrieval rank; useful for filings/transcript retrieval
FinBen Multi-task financial LLM evaluation Broad finance model comparison; useful first-pass screen
FinEval Chinese financial evaluation Relevant for Chinese-market/domain models
FLARE Financial language tasks associated with PIXIU/FinMA Useful for older open finance LLM comparison
FinMCP-Bench Multi-turn financial tool use Most relevant to agentic finance workflows because static QA does not test tool orchestration
FinSheet-Bench Financial spreadsheet lookup, computation, and reasoning Exposes table/spreadsheet failure modes; closer to PE and analyst workflows than prose QA
FinTradeBench Fundamentals, trading-signal, and hybrid financial reasoning Tests whether models can combine text fundamentals and price-derived trading signals; frozen as source-acquisition fixture sa-100
FinAgentBench Agentic retrieval over financial documents Tests document-type selection and passage localization before answering; frozen as source-acquisition fixture sa-101
Fin-RATE SEC-filing analytics and tracking evaluation Useful because it separates retrieval, generation, finance-reasoning, and query-context misunderstanding failures; frozen as source-acquisition fixture sa-104
FINESSE-Bench Hierarchical financial-domain knowledge and technical-analysis tasks Moves finance eval closer to professional market tasks than generic QA
FINESSE-Bench Trading_derivatives Options, synthetic positions, put-call parity, arbitrage, Greeks, hedging, pricing, and futures strategies Useful external derivatives diagnostic before artifact-level StateBench tasks
FinChain Symbolic benchmark for verifiable chain-of-thought financial reasoning Useful where deterministic financial templates can provide stronger labels than prose graders
Open FinLLM Leaderboard Multilingual/multimodal finance model leaderboard Candidate-discovery source; should not replace internal StateBench tasks
FrontierFinance Long-horizon computer-use benchmark for professional financial modeling Best current fit for StateBench-style finance artifact evaluation: Excel/PPT outputs, filing/web sourcing, formula consistency, and auditability
BankerToolBench End-to-end investment-banking workflow benchmark across data rooms, SEC/market-data tools, Excel/PPT/PDF/Word deliverables, and banker instructions Adds client-readiness scoring for cross-artifact consistency, formula fidelity, pitch/deal execution workflows, and confidentiality-aware agent behavior
Deloitte GenAI in M&A 2025 Survey of 1,000 corporate and PE M&A leaders on GenAI adoption, use cases, and risks Workflow/risk signal for strategy, screening, diligence, valuation, execution, and integration; use as adoption evidence, not model-performance evidence
Houlihan Lokey CIO transcript Practitioner podcast signal from a mid-cap M&A bank Highlights AI tool sprawl, potential internal deal-history models, and the Excel/Word/PowerPoint artifact reality of investment banking
Finance Agent Benchmark Expert-authored SEC-filing research questions with EDGAR/search tools Closest public proxy for entry-level analyst research; developed with bank, hedge-fund, and PE input, but still not a live-alpha benchmark; StateBench fixture sa-052 materializes it for replay
FinSearchComp Analyst-style open-domain financial search across global and Greater China markets Tests whether agents can find, reconcile, and time-align financial facts rather than merely answer static QA; StateBench fixture sa-053 materializes it for replay
DataClawBench Exploratory real-world financial data analysis over messy enterprise/industry/policy data Closest public benchmark for quant-research data discovery and schema exploration failure modes; StateBench fixture sa-051 materializes it for replay
FinToolBench Executable financial-tool selection plus timeliness, intent, and regulatory-domain compliance Separates “called a tool” from “used the right finance tool under the right constraint”
WebTailBench / WebVoyager / Online-Mind2Web Browser/source-discovery and computer-use tasks Not finance-specific, but useful for testing whether an agent can find arXiv PDFs, reports, whitepapers, podcast pages, and source materials without relying on static benchmark fixtures
RealFin Bilingual finance problems with missing conditions and “none of the above” traps Tests abstention and under-specification discipline, which matter for unsupported investment claims; StateBench fixture sa-054 materializes it for replay
FinMathBench Formula-driven financial math across increasing formula-chain depth Shows finance reasoning collapse as calculations require multi-step formula composition
Spreadsheet-RL Spreadsheet-domain RL on Qwen3-4B for layout/action tasks Evidence that RL can improve local SLM spreadsheet agents, but current pass rates are still low
FIRE Financial certification and real-world scenario questions Useful broad financial-knowledge screen; insufficient alone for quant research or agent deployment; StateBench fixture sa-055 materializes it for replay
FinanceArena / FinanceQA Industry-grade finance analysis tasks built by hedge-fund, PE, and investment-banking practitioners Useful external analogue for StateBench analyst-workflow tasks; emphasizes exact-match finance accuracy and assumption-based reasoning failures
FinRetrieval Exact financial value retrieval from structured databases using tool/API access Shows connector/API access can dominate model choice; Claude Opus 90.8% with structured APIs vs 19.8% with web search alone in the Daloopa benchmark
FinanceReasoning Expert-verified financial numerical reasoning with Python-formatted solutions and hard multi-formula problems Strong design pattern for StateBench calculation tasks: require executable code, strict units/precision, and PoT/tool execution rather than free-form CoT
FinAR-Bench / FinA-R-Bench Financial statement / fundamental-analysis benchmark decomposed into extraction, indicator computation, and reasoning over XBRL/text/PDF inputs Shows why finance report generation should be scored through verifiable intermediate tables, exact ratios, and PDF-layout robustness before prose quality
FINAUDITING Taxonomy-structured multi-document benchmark over real XBRL filings and US-GAAP relations Adds professional audit/compliance tasks: semantic tag matching, hierarchical relation extraction, and mathematical consistency across filing components
FinTagging Structure-aware XBRL tagging benchmark with numeric identification and US-GAAP concept linking across 17K+ concepts Tests whether agents can extract financial numbers, preserve evidence/type, retrieve/rerank taxonomy candidates, and avoid hallucinated tags
QuantEval Quant QA, quantitative reasoning, and CTA-style strategy coding with deterministic backtesting Closest public bridge from finance reasoning to executable quant-strategy code; score executability and risk/return metric deviation; StateBench fixture sa-056 materializes it for replay
AFIB / SuperInvesting AI Multi-dimensional investment-analysis benchmark with data recency, completeness, consistency, and failure-pattern scoring Use the dimensions; treat vendor/product leaderboard claims as directional until independently reproduced; StateBench fixture sa-058 materializes it for replay
FinMMEval CLEF 2026 Multilingual and multimodal finance shared tasks: exam QA, PolyFiQA, and financial decision making Adds global-finance language/modality coverage that English text-only benchmarks miss; StateBench fixture sa-057 materializes it for replay
Fin-PRM Finance-specific process reward model over step-level and trajectory-level reasoning traces Important for training/eval design: reward intermediate financial reasoning, not only final answers
FinChart-Bench Financial chart VLM benchmark Adds chart/deck/technical-analysis visual lane
FinCRITICAL-ED Financial fact-level OCR benchmark over critical fields Adds ingestion-quality lane: score reporting dates, monetary values, table headers, and critical concepts

Hedge-Fund / Quant Benchmark Reality

There are useful finance and quant-adjacent benchmarks, but there is no public “hedge fund benchmark” that proves live alpha generation. The public signal splits into three buckets:

  1. Workflow benchmarks: Finance Agent Benchmark, FinSearchComp, DataClawBench, FrontierFinance, FinSheet-Bench, FinToolBench, FinMMEval, FinCRITICAL-ED, FinChart-Bench, and FinMCP-Bench. These test analyst research, retrieval, messy data exploration, spreadsheet/model artifacts, multilingual/multimodal evidence, critical-field ingestion, and tool orchestration.
  2. Quant-reasoning benchmarks: FinTradeBench, FINESSE-Bench, FinChain, FinMathBench, QuantEval, and trading-RL papers such as Trading-R1 / Alpha-R1 / Trade-R1. These are relevant to factor, signal, and market-task reasoning, but return metrics should be treated as research-grade until independently reproduced. QuantEval is especially useful because it runs strategy code inside a fixed CTA-style backtesting harness.
  3. Architecture disclosures: Balyasny, Two Sigma, Man AHL, Jump Trading, QuantEvolve, and NVIDIA’s quant signal discovery agent disclose workflow patterns. They are architecture evidence, not benchmark datasets and not proof of investment performance.

For this repo, the right benchmark design is not “beat a hedge fund.” It is: can a local or fine-tuned model reproduce historically real StateOfAI research tasks, build finance/quant wiki artifacts, retrieve the right source evidence, avoid leakage, use notebooks/tools correctly, and produce auditable outputs.

Financial Numerical Reasoning / Fundamental Analysis Lane

Sources: local Chandra OCR for FinanceReasoning (2506.05828), FinAR-Bench (2506.07315), FinTagging (2505.20650), and FINAUDITING (2510.08886). These are 2025 papers, so treat their benchmark design as useful but rerun model scores with current 2026 Qwen/Gemma/DeepSeek/OpenAI/Claude/Gemini candidates before model selection.

FinanceReasoning adds the deterministic calculation layer. It re-annotates older finance numerical-reasoning datasets, adds 908 expert-verified problems, and builds a financial function library of roughly 3.1K Python functions. The important result for StateBench is not the historical leaderboard; it is the method. Program-of-thought and executable Python materially improve hard financial calculation tasks, while failure cases still cluster around problem misunderstanding, formula selection, numerical extraction, and rounding.

FinAR-Bench adds the report-construction layer. It argues that “generate a fundamental analysis report” is too loose to evaluate directly, so the task is split into information extraction, indicator computation, and logical reasoning over financial statements. The dataset uses 100 Shanghai Stock Exchange companies from fiscal 2023 and evaluates both structured XBRL/text and PDF-derived inputs. Its reported finding matches our practitioner evidence: large models can extract many values, but precise indicator computation and PDF-layout robustness remain hard.

FINAUDITING adds the structured audit layer. It uses real XBRL filings, a US-GAAP taxonomy, and multi-document contexts averaging more than 33K tokens. Its tasks ask models to find semantic tag mismatches, hierarchical/compositional relationship errors, and mathematical inconsistencies across filing components. That makes it directly relevant to financial compliance, audit, data-quality, and EDGAR/XBRL retrieval agents.

FinTagging adds the upstream tagging layer before audit. It separates financial numeric identification from concept linking, which is the right architecture for SEC/XBRL agents: first extract the number, type, and evidence from text/table context, then align it to candidate US-GAAP tags. The important finding is that numeric extraction is much easier than full-taxonomy concept linking; domain pretraining alone does not remove the need for taxonomy retrieval, reranking, and invalid-tag controls.

The Table-R1 family adds the model-adaptation lesson. Generic reasoning models do not automatically become good table agents through longer thinking alone. The first Table-R1 paper shows supervised distillation from frontier reasoning traces and RLVR/GRPO with exact-match, fact-verification, free-form answer, and strict-format rewards. The program-SLM variant adds executable Python as the default for arithmetic-heavy tables, with layout-transformation pretraining, code-compilation rewards, answer rewards, and anti-degenerate short-code penalties. The region-RL variant adds evidence-region supervision: reward the model for finding the relevant rows/columns before the final answer, then decay the region reward toward answer correctness.

For StateBench this supports a fine-tuning-before-deployment path for table-heavy finance tasks, but only after building held-out fixtures for financial statements, spreadsheets, XBRL tags, and ratio calculations. Generic WikiTable gains are not enough.

For StateBench, this lane should require:

  1. a date-bounded company filing or cached fixture;
  2. structured extraction of statement line items into a table;
  3. explicit definitions for ratios and growth metrics before calculation;
  4. executable Python/SQL/spreadsheet formulas with unit and sign checks;
  5. reconciliation between source values and computed indicators;
  6. a final narrative only after the intermediate artifacts pass validation;
  7. separate scoring for PDF/OCR/layout failures versus reasoning failures;
  8. taxonomy-aware checks for XBRL tags, relation graphs, and calculation links when the task touches regulated filings;
  9. separate extraction and concept-linking scores for XBRL numeric facts;
  10. table-specific SFT/RLVR experiments against current local Qwen/Gemma baselines before assuming a general reasoning model is enough;
  11. evidence-region scoring over the rows, columns, cells, pages, or taxonomy nodes that support each answer.

This also gives a clean local-model test. A Qwen/Gemma/LiquidAI/llama.cpp candidate should not be promoted because it writes fluent finance prose. It should beat cheap extraction baselines on line-item capture, survive strict ratio calculation, and produce a reproducible notebook or spreadsheet artifact.

Quant Research Context-Engineering Lane

Source: local Chandra OCR for QuantMind (2509.21507) and FinRL-X (2603.21330).

QuantMind adds the corpus architecture that our own pipeline is converging toward. It argues that quantitative research over filings, earnings calls, broker research notes, tables, formulas, podcasts, and videos fails when the context layer is only generic chunking plus top-k retrieval. The proposed pattern separates knowledge extraction from retrieval: multimodal parsing, adaptive summarization, domain tagging, point-in-time correctness, provenance, multi-hop retrieval, and knowledge-aware generation.

For StateBench, this should become a context-engineering eval rather than a model leaderboard. The task should ask a model to ingest a date-bounded source set, produce structured source cards, preserve evidence attribution, build domain tags, retrieve across multiple papers or notes, and answer with point-in-time citations. The model should fail if it uses current knowledge to rewrite historical context, drops table/formula evidence, or cannot explain which source supported which claim.

FinRL-X adds the deployment-consistency harness for quant trading workflows. Its weight-centric interface is useful because every strategy component outputs a target allocation vector consumed by the same backtest and broker-execution layers. The paper explicitly separates the backtest-to-paper gap from the paper-to-live gap: instant fills, weak cost models, missing market impact, survivorship bias, data-feed mismatch, partial fills, latency, slippage, queue position, API behavior, infrastructure failures, state recovery, margin, and settlement.

For StateBench, FinRL-X is not alpha evidence. It is a harness design signal: score whether an agent preserves the same strategy contract from research to paper trading, logs target versus realized weights, reports turnover and tracking error, applies risk overlays, and refuses to treat short paper-trading windows as live deployment proof.

Time-Series / Market Foundation Model Lane

The model matrix needs a non-chat lane for financial time series. Finance LLMs can read filings, generate hypotheses, orchestrate notebooks, and critique claims. They should not be treated as forecasting models unless they beat purpose-built time-series baselines.

Newly ingested sources add four benchmark requirements:

  1. Finance-native market models: Kronos tokenizes OHLCV/K-line data and targets price forecasting, volatility forecasting, and synthetic K-line generation. FinCast targets financial non-stationarity, domain diversity, and mixed temporal resolutions. These are candidate models for market data, not replacements for research validation.
  2. General TSFM controls: TimesFM 2.5, Chronos, Moirai, and Lag-Llama should be baseline controls. A finance-specific TSFM should have to beat them, plus classical and supervised baselines, on the exact target task.
  3. Risk-adjusted evaluation: the financial time-series benchmark source evaluates Sharpe ratio, downside/tail risk, breakeven transaction costs, random-seed robustness, and compute efficiency. Generic forecast-error scores are insufficient.
  4. Routing instead of universal deployment: the operational-viability source argues for routing each series to the best model class. For StateBench, that means cheap classical models, supervised specialists, general TSFMs, and finance-native TSFMs all compete under the same point-in-time protocol.
  5. Text-to-graph prediction probes: Relational Probing shows a different small-model path: use Qwen3 SLM hidden states to induce a financial-entity relation graph and train that graph jointly with a downstream stock-trend prediction model. This should become a market-graph fixture with point-in- time source text, graph/relation outputs, co-occurrence baselines, label leakage checks, and downstream predictive scoring.

The resulting StateBench task should ask an agent to discover sources, download/ingest them, design a leakage-safe evaluation notebook, compare baselines, and report only risk-adjusted, cost-aware findings. A model that claims alpha from a paper abstract or generic accuracy number should fail.

Text Alpha / LLM Asset-Pricing Lane

Source: finance-text-alpha-llm-asset-pricing-2026.md

The benchmark now needs a text-derived-alpha layer between source ingestion and factor discovery. This is where news, filings, earnings-call text, social media, or transcripts become embeddings, prompts, sentiment labels, or expected-return signals.

Newly ingested sources add three sub-tasks:

  1. LLM embeddings for expected returns: compare BERT, RoBERTa, LLaMA, OpenAI embeddings, FinBERT-style models, word vectors, and dictionary baselines on timestamped news-to-return prediction.
  2. Chronological consistency: ChronoBERT and ChronoGPT provide annual-vintage model controls so historical backtests do not use future language or future semantic meanings.
  3. Prompted market interpretation: GPT-style prompting can classify the economic implication of headlines and study initial reaction vs. drift, but the signal must survive turnover, cost, value-weighting, and capacity gates.

For StateBench, this lane should require source-time cards, model knowledge cutoffs, prompt-vs-embedding comparisons, immediate-reaction and drift decomposition, decay curves, topic/size splits, IC/RankIC/Sharpe diagnostics, turnover and transaction-cost stress, and source-to-factor lineage. A model that summarizes an earnings call well but cannot prove timestamp-safe predictive value should not pass this lane.

Microstructure / Limit-Order-Book Lane

The time-series lane now needs a microstructure subsection. Daily bars and factor data do not test whether a model understands execution. LOBERT, ByteGen, and LiT add a lower-level market-data layer:

  1. LOBERT: BERT-style encoder for complete multi-dimensional LOB messages; candidate for mid-price movement and next-message prediction.
  2. ByteGen: tokenizer-free byte-level generative model for order-book events; candidate for synthetic order-flow and market-simulation tasks.
  3. LiT: patch/self-attention/LSTM limit-order-book transformer; candidate for short-horizon market movement forecasting and distribution-shift fine-tuning tests.

For StateBench, the benchmark should not stop at “predicted direction.” It should score fill probability, queue position, spread/depth behavior, slippage, market impact, adverse selection, venue shift, and whether synthetic order flow preserves stylized facts without creating exploitable artifacts. This is also where the market-integrity lane connects back to modeling: an agent should flag spoofing-like or collusive order-flow policies instead of optimizing them.

Agentic Factor-Discovery / Feature-Engineering Lane

Source: finance-agentic-factor-discovery-2026.md

The benchmark now needs an upstream factor-research layer. This is the layer closest to what systematic managers actually do before portfolio construction: convert observations, reports, transcripts, and market data into executable factor hypotheses, then test, reject, refine, or promote them.

Newly ingested sources add three sub-tasks:

  1. Safe formula mining: Hubble constrains LLM output to a DSL, validates formulas through an AST sandbox, evaluates RankIC/Pearson IC/turnover/ bucket returns, and penalizes crowded or duplicate factor families.
  2. Program-level factor evolution: FactorEngine turns factors into executable programs, separates LLM logic mutation from Bayesian parameter search, and converts financial reports into factor code through extraction, verification, and code-generation agents.
  3. Autonomous research gates: the systematic-factor-investing agent defines no-lookahead filtration, same-day cross-sectional transforms, IS-only promotion gates, OOS-only final validation, transaction-cost checks, and promote/hold/retire decisions.

For StateBench, this lane should score the artifact and the trace: factor code/formula, input schema, economic hypothesis, point-in-time execution, IC/RankIC/HAC diagnostics, decile monotonicity, turnover, costs, capacity, neutralization, diversity, walk-forward evidence, and failed-candidate logs. A model that writes a persuasive rationale for a spurious factor should fail.

The June 6 model-card pass adds a narrow SLM candidate for the upstream extraction step: lmxxf/financial-report-lora-qwen3-14b, a PEFT LoRA adapter on Qwen/Qwen3-14B for Chinese financial research-report metric extraction. It should be tested as schema-checked extraction over report paragraphs, not as a general finance reasoner. Required gates are license review, base revision, PEFT adapter-load smoke, held-out Chinese metric spans, JSON-schema validation, and comparison against rules/NER/general-Qwen baselines. Its 460-row training claim and lack of independent benchmark mean it remains watchlist-only.

Portfolio Construction / Risk-Model Lane

Source: finance-portfolio-risk-construction-models-2026.md

The finance benchmark now needs a portfolio/risk construction layer between textual research and market execution. This lane tests whether a model can propose, audit, and control a portfolio workflow rather than merely answer finance questions.

Newly ingested sources add four sub-tasks:

  1. Screening plus weighting: agentic screening with an LLM fundamentals agent and FinBERT sentiment agent, followed by high-dimensional precision-matrix weighting.
  2. Covariance forecasting: medium-term multi-asset covariance forecasts scored by matrix loss, GMV variance, turnover, and regime robustness.
  3. Graph construction: firm-characteristic similarity, correlations, supply-chain/industry/news/analyst edges, signed/weighted edges, and leakage/provenance checks.
  4. Risk-aware RL: heterogeneous agents with different risk profiles, meta-controller regime adaptation, and deterministic risk overlays before execution.

For StateBench, this lane should require point-in-time data, transaction-cost buffers, liquidity/shorting constraints, drawdown and turnover controls, mandate/risk/compliance approvals, and explicit refusal to turn a research backtest into an investment recommendation.

Derivatives / Volatility Surface / Hedging Lane

Source: finance-derivatives-volatility-hedging-models-2026.md

The benchmark now needs a derivatives layer below portfolio construction and above market microstructure. This lane is where finance “physics” becomes directly testable: no-arbitrage, put-call parity, terminal payoff, self-financing, Greeks, and hedge P&L.

Newly ingested sources add four sub-tasks:

  1. Volatility-surface reconstruction: Transformer, U-Net, CNN, VAE, and SVI baselines complete sparse implied-volatility grids and are scored on reconstruction error plus calendar/butterfly violations.
  2. Market-IV pricing: a VAE compresses full SPX implied-volatility surfaces into latent state, then an MLP pricer uses that state for American and Asian option pricing controls.
  3. Constrained hedging: a neural price function supplies gradient hedges and is evaluated by terminal payoff adherence, self-financing assumptions, and out-of-sample P&L distribution.
  4. Professional derivatives diagnostic: FINESSE-Bench Trading_derivatives covers options, synthetics, parity, arbitrage, Greeks, hedging, pricing, and futures strategy logic with local vLLM/sglang support.

For StateBench, this lane should fail models that produce fluent derivatives prose while violating financial constraints. The task should ask for a reproducible pricing/hedging notebook, deterministic no-arbitrage checks, Greeks/stress reporting, and explicit refusal to convert a research result into an unsupported live-trading recommendation.

Runtime compatibility must be recorded. CUDA, MLX, llama.cpp/GGUF, AMD, and AWS Neuron backends do not support identical operator sets or module architectures, so a surface/pricing model that works on CUDA may need architectural changes or fallback kernels before it can run on MLX or Neuron.

Finance-Native Constraint Lane

The Risk.net Quantcast / Stefano Iabichino source adds a benchmark requirement that sits below model brand selection: finance models must preserve financial constraints. The important distinction is not “can the model talk about options, QIS, swaps, XVA, or hedging?” It is whether the model can detect when an answer violates no-arbitrage, measure, discounting, temporal, or market-state assumptions.

For StateBench, add tasks that ask a model to:

  1. identify no-arbitrage and martingale violations in generated explanations;
  2. distinguish market-implied probabilities from investor views;
  3. explain why generic image/text neural architectures are not automatically valid for dynamic valuation;
  4. translate a QIS strategy into forward-looking stress or hedge-leakage tests;
  5. abstain from unsupported alpha claims when the evidence is only an architecture or paper discussion.

This lane should test Qwen/Gemma/Llama finance fine-tunes, finance-specific reasoning models, and cheap specialist baselines. A fluent model that violates financial laws should lose to a smaller model that flags the invalid premise.

Term-Structure / Latent-Factor Lane

Risk.net’s Sokol / Lyashenko / Mercurio Quantcast episode adds a second finance-native eval slice: rates models and yield curves. This should not be a generic “explain autoencoders” task. The useful benchmark is whether a model can reason about nonlinear latent factors while preserving pricing-model constraints.

The task family should include:

  1. yield-curve dimensionality: explain why 10-15 tenor points can often be represented by a smaller number of factors while preserving nonlinear relationships;
  2. real-world vs. risk-neutral measure: distinguish historical curve behavior from pricing dynamics and identify what a change of measure can and cannot change;
  3. autoencoder manifold validity: explain why a compressed representation is not automatically an arbitrage-free pricing model;
  4. scenario realism: detect simulated curve shapes that are mathematically possible but historically implausible or unsupported by the learned manifold;
  5. implementation caution: refuse to claim live deployment, lower hedging cost, or production readiness unless the source provides deployed results.

This lane is a good stress test for finance-specific models because it combines math vocabulary, model-risk discipline, temporal reasoning, and practical source skepticism.

Market-Integrity / Algorithmic-Collusion Lane

Risk.net’s Alvaro Cartea Quantcast episode adds a safety lane that is closer to market-structure supervision than finance QA. A finance agent should not only optimize trading objectives. It should detect when a trading policy, market microstructure pattern, or proposed agent design creates market-integrity risk.

The task family should include:

  1. signaling detection from order-size conventions, quote-size outliers, and repeated identifiers;
  2. concentration reasoning when two or three players dominate best bid/ask presence or rarely trade with each other;
  3. collusion vs. non-collusive super-competitive outcomes, including the role of reward/punishment mechanisms;
  4. adaptive-agent risk from memory, initial conditions, overlapping training data, and profit-maximizing update rules;
  5. manipulation-by-learning, where an agent learns to rig order-book or flow signals that other participants use;
  6. sandbox tests where trading agents compete against clones and competitor policies before deployment;
  7. compliance escalation and refusal when a user asks the model to exploit or conceal suspicious market-structure behavior.

This lane is important for Qwen/Gemma/Llama/LiquidAI model selection because a smaller local model used for tool-routing or strategy scaffolding can still cause harm if it optimizes a trading workflow without market-integrity checks.

Snowflake Retrieval Benchmark Fit

Snowflake retrieval benchmarks are real, but they live in the retrieval layer, not the finance-reasoning layer.

  • Snowflake Arctic Embed v2.0 is a credible open embedding candidate. The model cards and technical report benchmark it on MTEB Retrieval / BEIR, MIRACL, and CLEF, with multilingual retrieval and compressed-vector support.
  • Snowflake Arctic-Text2SQL-R1-7B is the runnable open SQL specialist. Its model card reports GRPO with execution/syntax rewards and BIRD/Spider/EHRSQL scores, but it is Qwen2.5-Coder lineage. Keep it as a SQL-control baseline against current Qwen3/Gemma/frontier candidates.
  • Snowflake Arctic-Text2SQL-R2 is the current vendor direction for Snowflake-heavy NL2SQL. The May 27, 2026 engineering post claims a compact model beats frontier general models on Snowflake’s hard warehouse SQL benchmark. No open R2 checkpoint was verified here, so score R2 as vendor signal and R1 as the reproducible artifact.
  • Snowflake Cortex Search is a productized hybrid retrieval service for RAG over Snowflake data. The docs describe low-latency fuzzy search, keyword and vector retrieval, and service-level query APIs; they do not constitute an independent finance benchmark.
  • FinRetrieval is the closest current structured-data retrieval benchmark for this layer. It tests exact numeric retrieval from financial databases and releases tool-call traces. Its core lesson is that structured API/MCP access can matter more than model choice for exact financial values.
  • FinMTEB remains the finance retrieval screen. General BEIR/MIRACL/CLEF scores are useful for candidate selection, but finance deployment should re-rank embeddings/search services on filings, transcripts, broker-note-like prose, tables, and time-aware question sets.

Practical evaluation rule: treat Arctic Embed / Cortex Search as retrieval backends to test against FinMTEB plus our internal filings/transcript/source retrieval set, then feed the retrieved evidence into Finance Agent Benchmark, FinSearchComp-style, and Fin-RATE-style end-to-end tasks. A strong Snowflake retrieval result should improve citation recall and evidence freshness; it does not by itself solve finance math, spreadsheet modeling, tool compliance, or investment-reasoning validity.

Finance Retrieval / SQL Harnesses

The best harness is not one library. It is a layered eval stack:

Layer Best Fit What To Measure
Finance document retrieval FinMTEB, FinAgentBench, Fin-RATE, FinSearchComp, internal filings/transcript/source set Citation recall, document-type selection, passage localization, as-of-date correctness
Structured financial value retrieval FinRetrieval plus internal connector/API traces Exact numeric value, period alignment, unit/currency, source/tool provenance
Postgres retrieval pgvector + full-text/BM25 + reranker, evaluated with RAGAS/DeepEval/TruLens-style metrics and exact citation labels Context precision/recall, answer faithfulness, SQL safety, latency/cost
Snowflake retrieval Cortex Search + Cortex Analyst/Agents, Arctic Embed as an embedding candidate Hybrid search quality, governed warehouse access, semantic-model correctness, row/warehouse limits
Snowflake / warehouse NL2SQL Arctic-Text2SQL-R1 as runnable open control; Arctic-Text2SQL-R2 as current vendor direction; internal Snowflake gold set as gate Dialect correctness, long-schema handling, semantic-view fit, execution validation, permission safety
NL-to-SQL custom schema BIRD-SQL, Spider 2.0, BIRD-Interact as generic screens; internal schema eval as the real gate Execution accuracy, permission safety, schema linking, join correctness, abstention
Agentic data workflow DataClawBench, Finance Agent Benchmark, StateBench history-replay tasks Exploration trace quality, reproducibility, notebook/code artifacts, leakage discipline

For custom Postgres or Snowflake schemas, public NL-to-SQL scores are only a pre-filter. The deployment harness must run read-only roles, SQL allowlists, row limits, cost limits, timeout limits, dry-run/explain checks, provenance capture, and human approval for any trade, compliance, or client-facing action. Wren AI and Vanna are useful open-source semantic/NL2SQL prototypes; Snowflake Cortex is the native path when the data and governance already live inside Snowflake.

RAGAS, DeepEval, and TruLens are useful harness components, not finance benchmarks. Use them to grade retrieval faithfulness and context quality, then pair them with finance-specific exact-match labels and tool-call traces. A generic “faithful answer” score is insufficient when the target is an exact period value, XBRL concept, entitlement-gated index fact, or timestamped alternative-data feature.

Kaggle As A Source Medium

Kaggle was underexplored because it is neither a whitepaper source nor a production deployment source. It belongs in the benchmark-design layer.

Relevant finance/quant competition families include Jane Street market prediction, Optiver market microstructure / trading-at-close tasks, Two Sigma financial modeling tasks, G-Research crypto forecasting, JPX Tokyo Stock Exchange prediction, Ubiquant market prediction, and commodity forecasting tasks such as Mitsui-style challenges. Their value is not that leaderboard winners map to live alpha. Their value is task design: leakage controls, public/private leaderboard splits, time-aware validation, target construction, metric choice, baseline notebooks, and adversarial overfit behavior.

StateBench should ingest Kaggle metadata and top public notebooks only when the competition terms allow it, and should keep raw datasets outside git. The eval lesson to preserve is validation discipline, not leaderboard chasing.


Web Discovery Eval Lane

Web discovery should stay out of the core benchmark run because live websites change. It should be a periodic eval that asks whether an agent can locate, verify, and archive source material: arXiv papers, PDFs, whitepapers, official product pages, podcast pages, and transcripts.

Fara1.5 is the current small-agent candidate to track for this lane. Microsoft Research released Fara1.5 on May 21, 2026 as a 4B/9B/27B browser-use family based on Qwen3.5 backbones. Reported automated results: Fara1.5-9B scores 86.6 on WebVoyager and 63.4 on Online-Mind2Web; Fara1.5-27B scores 88.6 and 72.0. On WebTailBench v1.5, Fara1.5-9B improves over Fara-7B outcome success from 24.1 to 32.3. Microsoft also reports that synthetic environments help train gated-domain behaviors that open-web trajectories cannot safely cover.

Fara-7B remains useful as the concrete runnable/open-weight artifact today. The official Microsoft HF card records microsoft/Fara-7B as an MIT-licensed 7B Qwen2.5-VL-based computer-use model with 128K context and documented Transformers, vLLM, and SGLang routes. Treat it as a practical source-discovery candidate while keeping Fara1.5 as the newer current-family watchlist.

For this repo, Fara1.5 is not a finance model and not a replacement for FinMTEB, FinRetrieval, Fin-RATE, or Finance Agent Benchmark. Its role is to test browser/source acquisition:

  1. find the correct paper/report/podcast page;
  2. distinguish official pages from SEO mirrors and promotional summaries;
  3. avoid irreversible or login-gated actions;
  4. download allowed source artifacts into the repo’s raw/source pipeline;
  5. record provenance and uncertainty instead of inventing citations.

Operational rule: run this as a dated eval with captured inputs/outputs, not as a score we expect to be stable across time.


Retrieval Model Shortlist

Finance retrieval is now its own model-selection problem. Do not choose the generator first and assume retrieval will follow.

Retrieval Candidate Role Why It Matters Deployment Caveat
FinE5 Finance-specific embedding baseline FinMTEB reports it as the domain-adapted SOTA; use it as the target to beat on finance retrieval CC-BY-NC-ND and gated; not a commercial default
Qwen3-Embedding-8B / 4B / 0.6B Current Apache-2.0 general embeddings Strong open candidate family with long context and size choices for local/hosted retrieval Needs FinMTEB and internal filing/transcript/source scoring
Qwen3-Reranker-8B / 4B / 0.6B Current Apache-2.0 reranker family Useful after BM25+dense candidate generation to improve citation precision Cross-encoder cost/latency must be measured
QAnchor Qwen3 0.6B reranker Chinese finance reranker over A-share annual reports and related filings Useful same-document reranking/source-localization candidate after BM25/embedding/RRF first-stage retrieval Same-document and weakly supervised; no English SEC, open-domain retrieval, generation, advice, alpha, or runtime-portability proof
Qwen3-VL-Embedding / Qwen3-VL-Reranker Multimodal retrieval/reranking Relevant for financial PDFs, tables, charts, scans, and investor decks where page layout matters Needs page/table ground truth; not a replacement for OCR/parsing
BGE-M3 + BGE-Reranker-v2-M3 Open retrieval/reranker control pair Good reproducible baseline; already appears in finance tool-retrieval benchmark methods Not finance-specific; compare against Qwen3/Snowflake/FinE5
Snowflake Arctic Embed v2.0 Enterprise retrieval candidate Strong open Snowflake-aligned embedding option for warehouse/RAG deployments General retrieval evidence only; finance evaluation still required
Jina v5 text nano Tiny low-latency baseline Tests whether a small embedding model is “good enough” for routing/source discovery Use for cost/latency floor, not as assumed accuracy leader
LFM2-ColBERT-350M LiquidAI nano retrieval candidate Low-latency multilingual retrieval/watchlist option for edge or constrained deployments Not finance-specific; license threshold and FinMTEB/internal scoring required

The evaluation target is hybrid retrieval: BM25/full-text for exact financial terms and ticker/metric strings, dense embeddings for semantic recall, and a reranker for final source precision. FinMTEB should pick candidate retrievers; the internal StateBench finance retrieval set should decide defaults.


Current Fine-Tuned / Domain Model Toppers by Chart

Do not collapse these leaderboards into one answer. The topper changes with the task, and finance-specific fine-tunes often lose to stronger general models on headline average score.

Chart / Benchmark Overall Topper Seen In Current Corpus Top Finance-Specific / Fine-Tuned Signal Readout
Open FinLLM Reasoning Leaderboard DeepSeek-V3 at 61.3 average; GPT-4o at 61.01; DeepSeek-R1 at 60.87 Fino1-8B at 59.95 average; Fino1-14B at 58.10; Fin-R1-7B trails at 34.85 Fino1 is the current finance fine-tune to beat on this specific reasoning chart; small finance-RL models do not automatically win.
FIRE certification + scenario benchmark Gemini 3.0 Pro leads general certification tasks; GPT-5.2 leads scenario tasks XuanYuan 4.0 36B is the strong financial-domain baseline and reportedly outperforms its Seed-OSS-36B backbone on scenario tasks Domain CPT/SFT/RL can close much of the gap to frontier proprietary models on finance scenarios.
RealFin missing-condition / NOTA benchmark General models can score well on original/full-condition items Fin-R1-7B is the standout finance-specific model for condition checking and NOTA behavior; DianJin/DianLin-R1-32B extends the GRPO-style pattern Abstention and premise checking are trainable; scale alone does not solve it.
Spreadsheet-RL No frontier model claim in the local benchmark note; task is built around Qwen3-4B Qwen3-4B + GRPO RL reaches 23.4% on SpreadsheetBench and 17.2% on domain spreadsheet tasks RL helps, but pass rates remain far below professional reliability.
FinToolBench Doubao-Seed-1.6 has the best overall execution-success balance; GPT-4o has higher conditional execution among invoked tools FATR is the fine-tuned/retrieval-style method signal: finance attribute injection improves execution/compliance behavior across backbones Tool routing and metadata injection matter as much as model choice.
Finance Agent Benchmark OpenAI o3 at 46.8% No finance fine-tune tops the chart in the current source set The benchmark is a control showing frontier general agents still struggle with SEC-filing research.
FinSearchComp Grok 4 web leads global; DouBao web leads Greater China Web-enabled regional models beat generic models in their home market Search stack, market coverage, and data freshness dominate over static model specialization.
FinanceArena / FinanceQA Site leaderboard is current to July/August 2025 snapshots on the fetched page; live rows require client-side leaderboard data No verified finance fine-tune winner yet in local corpus Strong benchmark-design signal: professional finance tasks require exact answers, and assumption-based questions are where models break hardest.
rLLM-FinQA-4B model card GPT-4.1/o3-mini/Gemini-2.5-Pro still lead broader reported scores rLLM-FinQA-4B reaches 59.70% on Snorkel Finance Benchmark vs. 27.90% for its Qwen3-4B base; ties gpt-5-nano at 26.60% on Snorkel Finance Reasoning Strong current SLM specialist for SEC-table QA/tool use; not a full quant-research model.
TraceAlchemy-Gemma-4-E4B-Finance-IT model card No external benchmark score reported Gemma-family training telemetry shows validation-loss improvement on a 630-example finance eval split Promising Gemma candidate, but not a chart topper until externally benchmarked.
QuantEval execution-based quant coding Human experts remain the upper bound; proprietary models execute more code but still miss framework/risk logic DianJin-R1-7B SFT+GRPO improves on QuantEval reasoning in the paper Domain-aligned SFT/RL can help, but executable backtest artifacts and metric deviation must decide adoption.
Fin-PRM process-supervision paper General math PRMs are strong controls Fin-PRM improves finance trace selection, Best-of-N, and GRPO reward shaping over generic/outcome-only rewards in the paper Candidate for training/eval infrastructure, not a standalone finance analyst model.
ODA-Fin paper/model cards Qwen3-32B remains the larger open general baseline in reported comparisons ODA-Fin-RL-8B reports 74.6% average across nine finance benchmarks, matching Qwen3-32B at smaller size and beating DianJin-R1-7B in the model-card comparison Strong current Qwen3 post-training candidate, but scores are paper/card-reported and must be rerun locally.
Snowflake Text-to-SQL R2 vendor post claims compact Snowflake-specialist model beats frontier models on a hard warehouse SQL benchmark Arctic-Text2SQL-R1-7B is the verified open Apache-2.0 runnable control; R2 is current but no open checkpoint was verified Use R1/R2 only for SQL lanes. Qwen2.5 lineage makes R1 a specialist baseline, not a preferred general finance model.
Traditional Chinese / Taiwan finance sentiment No global finance-reasoning topper; this is a regional routing lane Eland Sentiment zh vLLM reports 92.03% financial-domain macro average on Taiwan stock-market sentiment tasks Useful for sentiment routing/entity sentiment only; requires held-out Traditional Chinese/Taiwan fixture, exact revisions, prompt contract, and no alpha/portfolio inference.
Korean finance leaderboard / Won Closed Korean finance benchmark evaluated 1,119 submissions over roughly eight weeks Won is the open Korean finance LLM/dataset signal Add to multilingual/local-market lane; do not extrapolate to English SEC or U.S. quant tasks.
CNFinBench Closed frontier and open finance/general models are included in the paper’s 22-model evaluation No promoted fine-tuned winner yet; the useful contribution is the safety/autonomy/compliance task taxonomy and multi-turn degradation rubric StateBench fixture sa-059 now materializes it as a finance-agent safety gate; do not treat it as alpha evidence.
EDINET-Bench Classical logistic regression / random forest / XGBoost are explicit controls No LLM winner promoted; paper reports frontier LLMs only marginally beat logistic regression on binary tasks, with random forest strongest on fraud detection StateBench fixture sa-060 adds whole-report Japanese filing tasks and classical baselines before accepting finance LLM claims.
BizFinBench Paper evaluates 25 proprietary and open/open-ish models, including frontier APIs, Qwen/Qwen3, Llama, DeepSeek, QwQ, Qwen-VL, and Xuanyuan3-70B No single winner; useful result is task-level capability separation and the warning that complex cross-concept finance reasoning remains weak StateBench fixture sa-091 materializes it as multilingual/business-finance benchmark-design evidence; rerun current models before using rankings.
XFinBench Paper evaluates 18 leading models and human experts o1 leads text-only in the 2025 snapshot; Claude-3.5-Sonnet leads when visual-context questions are included, but both lag human experts on temporal reasoning and scenario planning StateBench fixture sa-092 materializes it as complex multimodal finance-problem evidence; rerun current models and score exact numerical/visual failures.
MultiFinBen Paper evaluates 21 leading text, vision, audio, and multimodal models GPT-4o leads the paper snapshot at only 46.01% overall; Qwen2.5-Omni and Llama-4 trail, and multimodal models can lose to text-specialized models on text-only finance tasks StateBench fixture sa-093 materializes it as multilingual/multimodal/audio benchmark-design evidence; rerun latest model families and score OCR/audio/source-localization failures.
FinFRE-RAG Random Forest, XGBoost, and deep tabular classifiers remain core controls No LLM winner promoted; feature-reduced RAG improves open-weight LLM fraud detection and explanations but specialized classifiers can still lead StateBench fixture sa-061 adds the structured-fraud RAG pattern; score F1/MCC, privacy, retrieval poisoning, and human-review handoff.
FinGround Generic hallucination detectors and RAG baselines are controls under identical retrieval No general model winner promoted; contribution is retrieval-equalized claim verification, not a finance reasoning checkpoint StateBench fixture sa-062 adds atomic claim, citation, table-cell, formula, contradiction, and unverifiable-claim gates.
Fraud-warning pressure experiment Human participants are the benchmark control, not a model family No static winner promoted; paper reports LLMs resisted motivated investor pressure better than humans in the tested retail-advisory fraud scenarios StateBench fixture sa-064 adds multi-turn pressure, sycophancy, warning degradation, and operator-prompt sensitivity to finance safety evals.
PHANTOM Generic long-context hallucination detectors and base LLMs are controls PHANTOM-tuned models are historical controls; no 2026 winner promoted StateBench fixture sa-063 adds source-faithfulness labels, context-length sweeps, and lost-in-the-middle placement tests for SEC filings.
Credit/AML lane Classical scorecards, bureau baselines, GBMs, GNNs, and regulatory validation packets are the controls No general finance LLM winner promoted; FCMBench/LendNova/CALM/Omega2/TransXion/FinRegLab define the task and governance surface Add exact credit-document fields, OOT credit records, fairness, adverse-action testing, corporate-credit tabular controls, and AML profile-graph detection before any deployment claim.

Operational rule: each fine-tuned candidate must be rescored on the exact StateBench task family before adoption. For our local model work, the most interesting candidates are Fino1-8B for finance reasoning, Fin-R1-7B for abstention/premise checking, XuanYuan 4.0 / Agentar-Fin-R1 for larger finance reasoning, and Qwen3 spreadsheet RL variants for tool-using spreadsheet tasks. For retrieval, the interesting candidates are FinE5 as the finance-domain target, Qwen3 Embedding/Reranker as the Apache-2.0 open default to challenge, QAnchor as the Chinese A-share annual-report reranker specialist, Qwen3-VL retrieval for page/table/chart sources, and Arctic Embed for the Snowflake-aligned enterprise path.


What To Test for StateBench / Finance Workflows

Finance-domain models should not be accepted because their name starts with “Fin.” They need to pass workflow-specific checks:

  1. Structured extraction: filings, earnings transcripts, tables, risk factors, broker notes.
  2. Temporal discipline: no leakage from future documents; correct “known at time” reasoning.
  3. Tool use: calculator, retriever, database query, backtest executor, notebook execution.
  4. Schema fidelity: JSON contracts for hypotheses, features, backtest configs, and evidence citations.
  5. Numerical robustness: ratios, deltas, basis points, compounding, portfolio exposure, drawdown.
  6. Abstention: refuses unsupported investment claims and flags missing data.
  7. Research artifact quality: produces reusable notes, code, and audit trails rather than prose summaries.
  8. Spreadsheet/table reliability: handles messy financial spreadsheets, formulas, and multi-sheet context with deterministic calculator checks.
  9. Agentic retrieval quality: selects the right filing/report/transcript type, locates the supporting passage, and cites it.
  10. Long-horizon financial modeling: builds auditable Excel/PPT artifacts from sourced filings and public data, with formula integrity, assumptions, formatting conventions, and investment-memo quality.
  11. Execution-based quant code: produces strategy code that runs in a fixed backtest harness and matches expected risk/return metrics under stated transaction-cost and universe assumptions.
  12. Critical-field ingestion: preserves financially material OCR fields such as monetary values, reporting dates, period labels, table headers, and footnotes before downstream RAG/extraction is trusted.
  13. Multilingual/multimodal evidence: integrates filings, news, charts, prices, and non-English documents without losing source/date provenance.

This argues for a model-eval matrix:

Candidate Extraction Time Awareness Tool Use Structured Output Quant Code Notes
Qwen3 / Qwen3.5 / Qwen3.6 financial fine-tune TBD TBD TBD TBD TBD Current Qwen-family local candidates; avoid Qwen2.5-era defaults unless needed as historical controls
Gemma financial fine-tune TBD TBD TBD TBD TBD Watch for efficient small-model variants
LiquidAI LFM2/LFM2.5 nano variants TBD TBD Tool/RAG TBD N/A Edge extraction/RAG/transcript/tool-routing candidates; not finance reasoning defaults
LFM2.5-1.2B Financial Analyst Regional Limited / finance-regional-language-v0 fixture exists Tool/RAG possible TBD N/A Tiny Liquid A-share/CFA-style analysis candidate; requires runtime smoke before scoring
Nexus TinyFunction 1.2B N/A N/A Tool routing Function calls N/A Liquid function-calling control for routers/tool schemas via finance-tool-routing-v0; not a finance-domain model
FinGPT / FinMA TBD TBD TBD TBD TBD Baseline finance-domain open models
Aveni FinLLM TBD TBD TBD TBD TBD Better for regulated-services language than quant research
Kimi / Ling finance fine-tune TBD TBD TBD TBD TBD Higher-cost, stronger agentic baseline candidate
Fara1.5-9B / 27B N/A N/A Browser N/A N/A Source-discovery/computer-use eval candidate; not finance reasoning
Kuvera-8B-qwen3 Advisor Limited / needs retrieval Thin agent stack only Natural-language advice N/A Personal-finance SLM lane; test suitability, jurisdiction, calculation, abstention, bias, and human-review gates
Kronos / FinCast Forecast Time-series N/A Predictions N/A Finance-native TSFM candidates; evaluate on leakage-safe risk-adjusted market tasks
TimesFM / Chronos / Moirai / Lag-Llama Forecast Time-series N/A Predictions N/A General TSFM controls to beat before adopting finance-native forecasting models; Chronos-2 evidence says related multivariate context can help while unrelated cross-market context can hurt
LOBERT / ByteGen / LiT Microstructure Event-time N/A Messages / forecasts N/A Order-book lane for message representations, synthetic order flow, fill/slippage realism, and short-horizon forecasts

Interesting Domain-Specific Models to Research Next

These are worth explicit follow-up, but should not be treated as adopted until verified:

  • Fin-R1 / Fino1 / DianJin-R1 / Agentar-Fin-R1: finance-specific reasoning candidates. Prioritize Agentar-Fin-R1 for Qwen3-family currency, but keep Fin-R1 and Fino1 as reproducible baselines.
  • ODA-Fin-SFT/RL-8B: current Apache-2.0 Qwen3-8B finance post-training line with 318K distilled CoT SFT examples and 12K hard/verifiable GRPO tasks. It is a strong candidate for numerical reasoning and financial QA evals, but should be treated as model-card/paper evidence until it runs through local filings, table, transcript, and source-provenance tasks.
  • FinSphere: real-time stock-analysis agent; useful as architecture and evaluation reference, not production evidence.
  • Trading-R1 / Alpha-R1 / Trade-R1: now captured through the pillar 21 PDF/OCR pipeline and promoted through sources/21-benchmarks/trading-alpha-r1-reasoning-rl-raw.md. This is the most systematic quant paper cluster in the model lane because it targets structured trading theses, factor screening, and stochastic-reward RL. It is also the riskiest to over-interpret because market-return rewards can reinforce luck, momentum memorization, or reward-gamed rationales. Use it to design StateBench gates for evidence-backed thesis structure, factor activation under regime shift, temporal splits, transaction costs, capacity, and reward-hacking resistance.
  • Aveni FinLLM variants: useful for regulated financial advice/compliance workflows; evaluate as a domain SLM candidate.
  • Kuvera-8B-qwen3: useful for the personal-finance/advisor lane because it is a current Qwen3-8B fine-tune with public model card, MIT license, HF artifacts, vLLM/SGLang paths, and a behaviorally grounded training recipe. Evaluate it against suitability, jurisdiction, calculation, recency, abstention, and behavioral-bias tests before considering any advisor use.
  • rLLM-FinQA-4B: now verified as a public Apache-2.0 Qwen3-4B-Instruct- 2507 finance QA/tool-use agent trained with RL over SEC 10-K tables. It is a good local SLM candidate for filing/table QA with SQL/table/calculator tools, but needs GGUF/MLX conversion checks before local deployment.
  • Modern Qwen/Gemma/Llama finance fine-tunes: likely more practical than older FinMA-style models if they have current base architecture, permissive license, and good GGUF/MLX/vLLM support.
  • Liquid/LFM specialist split: test LFM2.5 Financial Analyst only on regional Chinese A-share/CFA-style explanation tasks, and test Nexus TinyFunction only on tool-routing/schema-selection tasks. Do not collapse them into one “Liquid finance model” bucket.
  • TraceAlchemy-Gemma-4-E4B-Finance-IT: current Gemma-family finance reasoning candidate with BF16 and GGUF artifacts. Prioritize for Gemma coverage in StateBench, but do not over-rank it without external benchmark scores.
  • RinKana / Qwen3-8B finance GGUF variants: useful for cheap llama.cpp runtime experiments, but lower-confidence because the cards are thin and benchmark evidence is not strong.
  • LiquidAI LFM2 / LFM2.5 nano variants: useful for constrained extraction, RAG, transcript, PII, and tool-routing experiments. They should enter StateBench as low-latency specialist baselines, not as finance reasoning leaders.
  • Spreadsheet-specialized finance agents: prompted by FinSheet-Bench; likely more useful as tool-using agents with spreadsheet parsers/calculators than as standalone LLMs. The May 31 local OCR pass strengthens this: FinSheet-Bench reports that standalone LLMs remain too error-prone for unsupervised professional finance spreadsheet use, especially as workbooks grow and tasks move from simple lookup to aggregation, sorting, and complex calculations. StateBench should score schema/table discovery, row-level extraction, deterministic computation, reconciliation, and human signoff as separate artifacts. Standalone source card: sources/21-benchmarks/finsheet-bench-financial-spreadsheets-raw.md.
  • Retrieval-specialized finance agents: prompted by FinAgentBench; useful for filings/transcripts/source-routing tasks in this repo.
  • Computer-use financial modeling agents: prompted by FrontierFinance; these should be evaluated on full artifact construction, not chat answers.
  • Financial time-series foundation models: Kronos and FinCast are the finance-native candidates; TimesFM 2.5, Chronos, Moirai, and Lag-Llama are the general TSFM controls. Evaluate them as forecasting/risk models, not as finance chat models.
  • Limit-order-book models: LOBERT, ByteGen, and LiT cover message-level representation, byte-level synthetic order flow, and short-horizon microstructure forecasting. These belong in execution and simulation evals, not finance prose QA.
  • Finance process reward models: Fin-PRM-style models are worth tracking as training infrastructure for trace selection, Best-of-N scoring, and GRPO reward shaping. Prefer current Qwen3/Gemma-family successors if available; use the 2025 Qwen2.5-based paper as a design pattern, not a final model pick. StateBench fixture sa-050 materializes the Fin-PRM source for replay.

Finance Multimodal / Visual Document Lane

The finance VLM lane is broader than generic multimodal QA:

  • Chart comprehension: FinChart-Bench adds 1,200 real-world financial chart images from 2015-2024 and 7,016 manually checked TF/MC/QA questions. Its findings matter for investor decks, earnings-call slides, factbooks, and research PDFs: open/closed model gaps are narrowing, newer versions can regress, instruction following and spatial reasoning remain brittle, and current VLMs should not be trusted as automatic benchmark judges. StateBench fixture sa-048 materializes this source for replay.
  • Fact-level OCR: FinCriticalED / FinCriticAED shifts OCR scoring from lexical similarity to financially material facts: numbers, monetary units, dates, reporting entities, and financial concepts. This is the right failure lens for SEC filings, scanned statements, bank PDFs, and table-heavy reports, because a single header alignment or decimal/unit error can corrupt a downstream ratio or audit conclusion. StateBench fixture sa-049 materializes this source for replay.
  • Visual citation RAG: FinRAGBench-V shows that text-only RAG loses evidence when charts, tables, page layouts, and visual snippets are flattened. Its page/block citation requirement is directly relevant to the repo’s source ingestion and wiki workflows: answers need page-grounded evidence, not just plausible finance prose. It is now frozen as source-acquisition fixture sa-102, with dataset/code links recorded but no local dataset checkout or runnable reproduction yet.
  • Temporal multimodal RAG: FinTMMBench combines financial tables, news, daily stock prices, and technical charts with time-windowed questions. Its reported TMMHybrid-RAG result still leaves a low F1 ceiling, so temporal metadata and heterogeneous retrieval need to be first-class in the benchmark. It is now frozen as source-acquisition fixture sa-103; preserve the FinTMMBench / FinTMBench naming mismatch until code/data artifacts resolve the canonical spelling.
  • Domain VLM baselines: FinTral and Open-FinLLMs/FinLLaVA are useful finance-specific historical baselines and training-recipe references. They should not displace current Qwen3/Qwen3.5/Qwen3.6, Gemma, or other trending VLM families without a fresh internal rerun. The useful comparison is not “finance VLM vs generic VLM” in aggregate; it is chart extraction, critical OCR, page retrieval, citation grounding, and time-aware answer construction.

Implication for StateBench: add a separate visual finance source task pack with earnings-deck charts, SEC filing pages, investor-presentation tables, screenshots from research reports, and time-windowed price/news/table bundles. Score exact numeric extraction, unit/scale preservation, citation box/page accuracy, temporal filtering, and final answer correctness separately.


Practical Recommendation

Use finance-domain models as specialists inside a benchmark harness, not as the whole system.

For quant research tasks:

  1. Use a strong general agentic model as the control.
  2. Add finance-domain candidates as extraction/retrieval/reasoning specialists.
  3. Score them on task bundles from actual repo history and finance workflows.
  4. Re-score every converted artifact separately: HF adapter, merged HF, GGUF, vLLM quantized, MLX, and Neuron-compatible artifacts where possible.
  5. Check license and backend compatibility before deployment: LFM Open License is not Apache-2.0 for enterprises over the revenue threshold, and Neuron / MLX / AMD / CPU support must be proven per architecture and operator path.

The model question is subordinate to the eval question. A finance model only matters if it improves a concrete workflow without increasing leakage, hallucination, or unsupported investment claims.


Sources