See also (wiki): wiki/quant-asset-management-ai.md · wiki/ai-model-evaluation-benchmarks.md
See also (operations): finance-quant-ai-objective-coverage-audit-2026.md · finance-quant-source-ingestion-queue-2026.md · finance model/runtime bakeoff · finance source-discovery eval · finance replay scorecard · finance retrieval/SQL harness · finance time-series market suite
Source ledger: sources/06-industry-verticals/finance-quant-ai-coverage-audit-2026-raw.md
Executive Summary
The finance/quant corpus is no longer missing the obvious source categories. It now covers:
- public hedge-fund and asset-manager workflow signals;
- institutional quant process controls;
- official talent, hiring, hackathon, and university-partnership leakage;
- practitioner podcasts and video source discovery;
- financial data connector and MCP-style retrieval ecosystems;
- alternative-data and Kaggle/competition artifacts;
- finance-domain models, SLMs, retrieval models, and SQL specialists;
- Snowflake/Postgres/NL2SQL harness choices;
- allocator governance, professional standards, and human-in-the-loop controls;
- asset-class-specific benchmark lanes for FICC, commodities, private credit, structured finance, derivatives, post-trade, ESG, insurance, payments, deal work, and advisor/portfolio agents.
The remaining gap is now scored execution and governed data-agent harnesses,
not source awareness. The current 118-task finance-replay-v0 StateBench suite is
materialized and validated, so the next missing layer is running candidate
models/backends against it and recording comparable model/runtime scorecards.
Current dry-run rows prove manifest and scorecard plumbing only; they leave
pass-rate columns blank by design. The oracle expected-patch control now passes
118/118, proving the current suite/check ceiling without claiming model
performance.
The concrete run order and phase gates now live in
finance-model-runtime-bakeoff-v0.
The latest local verification on 2026-06-02 resolves the bakeoff queue’s
preflight-file gap: all 27 phase-0 preflight records now exist, including the
paper-first Relational Probing Qwen3 SLM market-graph method, and every target
suite count matches its manifest. That is readiness plumbing, not model
performance; scored runs remain at 0/27 ready because endpoint, license,
runtime, cost, corpus, parser, backend, and Neuron compile evidence is still
missing per artifact.
For retrieval, Postgres, Snowflake, and custom-schema NL2SQL, the first static
answer-artifact scorecard row now exists, proving the 46-task fixture and
scorecard path. The missing layer is still a separate model/backend and
database-backed harness with preflight records, semantic-layer revisions,
executed SQL, retrieved evidence, RBAC checks, cost caps, and refusal gates.
Until those runs exist, model recommendations remain research-backed
hypotheses rather than measured choices.
Covered Lanes
| Lane | Current Evidence | Status |
|---|---|---|
| institutional quant workflow | Acadian PDF/OCR, Acadian podcast transcripts, Zhe Chen interview, Man Group, Two Sigma, AQR, HRT, Balyasny, Renaissance official careers/PDF signal, QFBench executable quant-code benchmark signal | Covered for source synthesis; 118-task replay suite now materialized |
| Agentic factor discovery | Hubble, FactorEngine, Beyond Prompting, FinRL-X | Covered as research-loop architecture and guardrail evidence; needs deterministic factor-discovery harness and live data/cost/capacity proof before any alpha claim |
| Top hedge funds and quant shops | Balyasny, Man AHL, Two Sigma, Bridgewater, Schonfeld, Millennium, D. E. Shaw, Tower, WorldQuant, XTX, CFM, Squarepoint, PDT, Jump, Citadel, Jane Street, HRT, Voleon, G-Research, QRT, Winton, AIMA allocator/industry survey evidence, agentic-finance market-structure survey | Covered as public evidence; no public alpha proof. Jane Street now adds first-party industrial-ML constraints; AIMA now adds 150-manager / 18-investor GenAI adoption, DDQ, governance, and AI-washing controls; agentic-finance survey adds system-level stability/accountability evaluation requirements |
| Official social/talent leakage | Millennium hackathons, Jump AI/ML page, Schonfeld training lab, Tower governance article, QRT Labs, Bridgewater AIA pages, Jane Street ML/performance pages | Covered as workflow/control evidence |
| Podcasts/audio/video | Acadian, Two Sigma/TWIML, Man Group/Tech Talks Daily, Risk.net Quantcast, Curious Quant, The Derivative, Top Traders, Flirting with Models, Chat With Traders, Quant/Financial Engineering | Covered for promoted extracts and queue; not exhaustive |
| Vendor/product signals | QuantConnect, Bloomberg/FactSet/LSEG/S&P/AlphaSense, Anthropic finance templates, Snowflake, Databricks, NVIDIA, JPMorgan, Mastercard | Covered as architecture/product evidence |
| Official talent/social leakage | Millennium, Jump, Schonfeld, Tower, QRT, Jane Street, Bridgewater and other official firm/careers/social leads | Covered as an executable finance-talent-leakage-v0 seed fixture plus source cards; not part of finance replay |
| Finance benchmarks | FinMTEB, FinMCP-Bench, FinanceBench, FinDER, FinTMMBench, FinRAGBench-V, FinanceArena, FinSheet, FinTrade, FinAgentBench, FinToolBench, FinSearchComp, FrontierFinance, CNFinBench and others | Covered for benchmark design |
| Advisor / portfolio agents | HERCULEAN, PortBench, One Size Fits None, Fin-Bias, Kuvera, personal-finance advisor tasks | Covered as executable finance-advisor-portfolio-v0 seed fixture plus source cards |
| Kaggle / competition artifacts | Finance Kaggle source ledger, JPX/Kaggle PDF evidence, MLE-Bench, CoMind, MLE-Dojo, and finance-kaggle-artifact-v0 |
Covered as source-discovery plus executable artifact-digestion seed fixture |
| Retrieval and SQL | FinRetrieval, Snowflake Cortex, Arctic Embed, Arctic-Text2SQL R1/R2, pgvector, pgvectorscale, Wren, Vanna, Defog SQL-Eval, BIRD, Spider 2.0, BIRD-Interact, BEAVER, EnterpriseMem | Covered for candidate selection; StateBench suite blueprint now exists; needs database fixtures, internal gold set, and task-level scores |
| Current model candidates | Qwen3/Qwen3.5/Qwen3.6, Qwen3.7-Max API reference, Gemma candidates, LiquidAI LFM runtime candidates, Fin-R1, Fino1, DianJin-R1, Agentar-Fin-R1, ODA-Fin, rLLM-FinQA, TraceAlchemy Gemma, Kuvera, CPA-Qwen3, FinGuard, Won, Kronos | Covered as watchlist; needs local/API/task-specific scores |
| Financial time-series / market models | Kronos, FinCast, TimesFM, Chronos, Moirai, Lag-Llama, LOBERT, ByteGen, LiT | Covered as a separate executable seed fixture and Kronos preflight; not part of finance replay |
What Is Still Missing
1. Historical StateBench Replay Scoring
The user-requested benchmark should test whether models would have helped build this repo and research corpus over time. The first executable fixture layer now exists, but it has not yet been run across candidate models and runtimes.
The first finance replay fixture manifest now lives at
statebench/suites/finance-replay-v0.md
with a machine-readable companion at
statebench/suites/finance-replay-v0.json.
It now contains 118 materialized tasks: 80 source-acquisition tasks, 15
wiki-maintenance tasks, and 23 reflect-audit tasks. Local validation passed for
JSON parsing, task listing, check compilation, and full expected-patch replay.
Dry-run result rows and a 118/118 oracle expected-patch control now exist in
finance-replay-v0-scorecard,
but it is still not a scored model/backend run.
QFBench also now has a separate executable seed fixture:
finance-quant-code-v0.
That fixture scores a static answer artifact 8/8 for executable quant-code
benchmark discipline and now also runs a tiny local code-execution smoke 3/3
over option payoff, drawdown/turnover, and factor-rank tasks. It covers
source-boundary, Docker/numerical-verifier contract, memo-overclaim rejection,
production-boundary, baseline comparison, task scope, QuantEval bridge, runtime
trace, and deterministic verifier wiring. It is intentionally not a live
QFBench run or a claim that any model can produce profitable quant code.
The suite now tests:
- find and download stable papers, reports, PDFs, and podcast/audio sources;
- ingest PDFs through the existing OCR pipeline;
- ingest podcasts through
scripts/podcast_mine.py; - extract source cards with URL, date, owner, claim type, and promotion tier;
- update the correct research pillar and wiki page;
- preserve caveats around alpha, performance, licensing, and authority;
- pass
validate_research.pyandgit diff --check.
The next evaluation step is to run this suite against a small model/backend matrix and record not only pass/fail, but also cost, latency, context length, artifact format, backend, quantization, compatibility constraints, and failure mode.
2. Model/Runtime Scorecard
The repo has a strong candidate list, but it does not yet have comparable scores across model/runtime choices.
The first scorecard should compare:
- frontier API controls for source discovery and synthesis;
- Qwen3/Qwen3.5/Qwen3.6 local candidates;
- Gemma finance candidates;
- LiquidAI LFM2/LFM2.5 nano/runtime candidates;
- finance-specific models such as Fino1, Fin-R1, ODA-Fin, rLLM-FinQA, and TraceAlchemy Gemma;
- retrieval models such as FinE5, Qwen3 Embedding/Reranker, BGE, Arctic Embed, Jina nano, and LFM2-ColBERT;
- SQL specialists such as Arctic-Text2SQL-R1 and future R2 availability;
- market/time-series specialists such as Kronos only in the dedicated
finance-time-series-market-v0lane after leakage-safe fixtures and trading utility metrics exist.
Backend compatibility must be recorded per task: MLX, llama.cpp/GGUF, CUDA, AWS Neuron/Inferentia, AMD/ROCm, and API. Neuron compatibility is only relevant when serving/training on AWS Inferentia/Trainium-class infrastructure.
3. Direct Official Social Archival
The new talent-leakage lane promotes official firm pages and official job listings. It still does not fully archive official LinkedIn posts because third-party social pages are brittle and often require browser/session access.
Promotion rule remains:
- official firm/careers/university pages can be findings;
- named practitioner transcripts can be findings;
- third-party LinkedIn commentary is source discovery only;
- official LinkedIn posts can strengthen a claim only when URL/date/author and primary-source backing are preserved.
4. More Top-Fund Negative Evidence
The corpus is clear that public sources do not prove genAI alpha. It can still improve by collecting negative or skeptical material from quant shops:
- HRT-style warnings about benchmarks not matching trading reality;
- Winton-style selection-bias and interpretability commentary;
- AQR/systematic quant overfitting, small-data, low-signal-to-noise, and domain-knowledge controls;
- allocator due-diligence questions that expose weak AI programs.
This matters because StateBench should reward models that avoid overclaiming, not models that generate the most exciting investment narrative.
The AQR “Can Machines Learn Finance?” PDF is now locally captured and OCR’d in pillar 06. HRT’s official trading-AI, data, and benchmark articles now add the market-making version of the same negative-evidence lane. Together they should be used to turn skepticism into concrete StateBench checks: does the model notice that adding predictors is not the same as adding independent target observations, that low signal-to-noise makes flexible models fragile, that prediction / optimization / execution are separate tasks, that data needs relevance / uniqueness / lookahead / sample-size / noise / provenance gates, and that economic theory and human expertise are controls rather than optional commentary?
5. Source-Ingestion Robustness Scorecards
The web-discovery lane should stay an eval, not a benchmark run each time. The suite now has stable source-acquisition fixtures, but they still need a scorecard that separates stable replay from live-web/source-discovery behavior:
- arXiv papers;
- official PDFs;
- whitepapers/reports;
- podcast pages and RSS entries;
- official firm pages;
- official conference/session pages;
- Kaggle competition pages and notebooks.
Success is not “browse the current web perfectly.” Success is finding the official item, downloading allowed artifacts, recording provenance, and abstaining when the source is not verifiable. Live web can refresh the fixture queue, but only the frozen task fixtures should count as comparable benchmark scores.
The operational queue for this work now lives at finance-quant-source-ingestion-queue-2026.md. It separates promotable sources from discovery-only social material and ties each source class to the existing PDF, podcast/audio, source-card, and StateBench pipelines.
6. Finance Retrieval and SQL Execution Harness
The retrieval/NL2SQL decision guide now has an operational StateBench suite
blueprint at
statebench/suites/finance-retrieval-sql-harness-v0.md.
This is deliberately separate from finance-replay-v0.
The first seed fixture contract now lives at
statebench/fixtures/finance-retrieval-sql/v0/.
It contains a Postgres-first schema, a small retrieval corpus, forty-six gold
tasks, and a Snowflake semantic-view contract. A static answer-artifact control
now records 46/46 in
finance-retrieval-sql-v0-scorecard.md.
It is not a scored model/backend run yet.
The missing build work is:
- freeze representative document-RAG fixtures from existing PDFs, transcripts, source cards, and wiki evidence;
- implement a runner for the
finance-retrieval-sql-v0Postgres fixture and gold tasks; - test the Snowflake-compatible schema manifest and semantic-view contract;
- adapt Defog SQL-Eval style execution scoring for local Postgres and future Snowflake runs;
- add retrieval recall, citation, RBAC, cost, stale-data, and unsafe-action refusal checks before any production-like deployment test.
Priority Next Work
- Run the 118-task
finance-replay-v0suite against a small bakeoff matrix: frontier API control, one Qwen3-family local/API candidate, one Gemma-family candidate, one LiquidAI runtime candidate, and one finance-specific candidate such as rLLM-FinQA or ODA-Fin. - Run
finance-quant-code-v0against frontier, Qwen, and Gemma candidates, then graduate the tiny execution smoke into a larger code-execution harness only after answer-artifact and local-verifier controls stay clean. - Create a model/runtime result table with task, model, backend, quantization, context length, cost, latency, pass/fail, and failure mode.
- Add a stable source-discovery fixture manifest for the next queue of papers,
PDFs, whitepapers, podcasts, official firm pages, and Kaggle pages. The
live/non-comparable backlog is now staged in
finance-source-discovery-eval-v0and should be promoted into frozen source-acquisition tasks only after evidence packs are captured. - Build the first database-backed retrieval/SQL fixture slice from
finance-retrieval-sql-harness-v0: Postgres first, Snowflake semantic-view contract second, external model bakeoff third. - Convert repeated winning behavior into training pairs only after the eval proves the target behavior and after Unsloth, LLaMA-Factory, AWS native, and MLX training paths are compared for cost, feature fit, and backend limits.
- Add Neuron preflight records only for model artifacts that would actually be hosted on AWS Inferentia/Trainium; do not treat Neuron as a generic model requirement.
Bottom Line
The corpus now answers “what are we missing?” at the research-map level and the
first replay suite exists. The remaining gap is measured execution: run models
through finance-replay-v0, build the database-backed retrieval/SQL harness,
capture model/runtime scorecards, and only then decide whether fine-tuning,
conversion, or AWS-hosted deployment is worth the cost.