← Industry Verticals 🕐 9 min read
Industry Verticals

Finance and Quant AI Coverage Audit

The finance/quant corpus is no longer missing the obvious source categories.

See also (wiki): wiki/quant-asset-management-ai.md · wiki/ai-model-evaluation-benchmarks.md

See also (operations): finance-quant-ai-objective-coverage-audit-2026.md · finance-quant-source-ingestion-queue-2026.md · finance model/runtime bakeoff · finance source-discovery eval · finance replay scorecard · finance retrieval/SQL harness · finance time-series market suite

Source ledger: sources/06-industry-verticals/finance-quant-ai-coverage-audit-2026-raw.md


Executive Summary

The finance/quant corpus is no longer missing the obvious source categories. It now covers:

  • public hedge-fund and asset-manager workflow signals;
  • institutional quant process controls;
  • official talent, hiring, hackathon, and university-partnership leakage;
  • practitioner podcasts and video source discovery;
  • financial data connector and MCP-style retrieval ecosystems;
  • alternative-data and Kaggle/competition artifacts;
  • finance-domain models, SLMs, retrieval models, and SQL specialists;
  • Snowflake/Postgres/NL2SQL harness choices;
  • allocator governance, professional standards, and human-in-the-loop controls;
  • asset-class-specific benchmark lanes for FICC, commodities, private credit, structured finance, derivatives, post-trade, ESG, insurance, payments, deal work, and advisor/portfolio agents.

The remaining gap is now scored execution and governed data-agent harnesses, not source awareness. The current 118-task finance-replay-v0 StateBench suite is materialized and validated, so the next missing layer is running candidate models/backends against it and recording comparable model/runtime scorecards. Current dry-run rows prove manifest and scorecard plumbing only; they leave pass-rate columns blank by design. The oracle expected-patch control now passes 118/118, proving the current suite/check ceiling without claiming model performance. The concrete run order and phase gates now live in finance-model-runtime-bakeoff-v0. The latest local verification on 2026-06-02 resolves the bakeoff queue’s preflight-file gap: all 27 phase-0 preflight records now exist, including the paper-first Relational Probing Qwen3 SLM market-graph method, and every target suite count matches its manifest. That is readiness plumbing, not model performance; scored runs remain at 0/27 ready because endpoint, license, runtime, cost, corpus, parser, backend, and Neuron compile evidence is still missing per artifact. For retrieval, Postgres, Snowflake, and custom-schema NL2SQL, the first static answer-artifact scorecard row now exists, proving the 46-task fixture and scorecard path. The missing layer is still a separate model/backend and database-backed harness with preflight records, semantic-layer revisions, executed SQL, retrieved evidence, RBAC checks, cost caps, and refusal gates. Until those runs exist, model recommendations remain research-backed hypotheses rather than measured choices.


Covered Lanes

Lane Current Evidence Status
institutional quant workflow Acadian PDF/OCR, Acadian podcast transcripts, Zhe Chen interview, Man Group, Two Sigma, AQR, HRT, Balyasny, Renaissance official careers/PDF signal, QFBench executable quant-code benchmark signal Covered for source synthesis; 118-task replay suite now materialized
Agentic factor discovery Hubble, FactorEngine, Beyond Prompting, FinRL-X Covered as research-loop architecture and guardrail evidence; needs deterministic factor-discovery harness and live data/cost/capacity proof before any alpha claim
Top hedge funds and quant shops Balyasny, Man AHL, Two Sigma, Bridgewater, Schonfeld, Millennium, D. E. Shaw, Tower, WorldQuant, XTX, CFM, Squarepoint, PDT, Jump, Citadel, Jane Street, HRT, Voleon, G-Research, QRT, Winton, AIMA allocator/industry survey evidence, agentic-finance market-structure survey Covered as public evidence; no public alpha proof. Jane Street now adds first-party industrial-ML constraints; AIMA now adds 150-manager / 18-investor GenAI adoption, DDQ, governance, and AI-washing controls; agentic-finance survey adds system-level stability/accountability evaluation requirements
Official social/talent leakage Millennium hackathons, Jump AI/ML page, Schonfeld training lab, Tower governance article, QRT Labs, Bridgewater AIA pages, Jane Street ML/performance pages Covered as workflow/control evidence
Podcasts/audio/video Acadian, Two Sigma/TWIML, Man Group/Tech Talks Daily, Risk.net Quantcast, Curious Quant, The Derivative, Top Traders, Flirting with Models, Chat With Traders, Quant/Financial Engineering Covered for promoted extracts and queue; not exhaustive
Vendor/product signals QuantConnect, Bloomberg/FactSet/LSEG/S&P/AlphaSense, Anthropic finance templates, Snowflake, Databricks, NVIDIA, JPMorgan, Mastercard Covered as architecture/product evidence
Official talent/social leakage Millennium, Jump, Schonfeld, Tower, QRT, Jane Street, Bridgewater and other official firm/careers/social leads Covered as an executable finance-talent-leakage-v0 seed fixture plus source cards; not part of finance replay
Finance benchmarks FinMTEB, FinMCP-Bench, FinanceBench, FinDER, FinTMMBench, FinRAGBench-V, FinanceArena, FinSheet, FinTrade, FinAgentBench, FinToolBench, FinSearchComp, FrontierFinance, CNFinBench and others Covered for benchmark design
Advisor / portfolio agents HERCULEAN, PortBench, One Size Fits None, Fin-Bias, Kuvera, personal-finance advisor tasks Covered as executable finance-advisor-portfolio-v0 seed fixture plus source cards
Kaggle / competition artifacts Finance Kaggle source ledger, JPX/Kaggle PDF evidence, MLE-Bench, CoMind, MLE-Dojo, and finance-kaggle-artifact-v0 Covered as source-discovery plus executable artifact-digestion seed fixture
Retrieval and SQL FinRetrieval, Snowflake Cortex, Arctic Embed, Arctic-Text2SQL R1/R2, pgvector, pgvectorscale, Wren, Vanna, Defog SQL-Eval, BIRD, Spider 2.0, BIRD-Interact, BEAVER, EnterpriseMem Covered for candidate selection; StateBench suite blueprint now exists; needs database fixtures, internal gold set, and task-level scores
Current model candidates Qwen3/Qwen3.5/Qwen3.6, Qwen3.7-Max API reference, Gemma candidates, LiquidAI LFM runtime candidates, Fin-R1, Fino1, DianJin-R1, Agentar-Fin-R1, ODA-Fin, rLLM-FinQA, TraceAlchemy Gemma, Kuvera, CPA-Qwen3, FinGuard, Won, Kronos Covered as watchlist; needs local/API/task-specific scores
Financial time-series / market models Kronos, FinCast, TimesFM, Chronos, Moirai, Lag-Llama, LOBERT, ByteGen, LiT Covered as a separate executable seed fixture and Kronos preflight; not part of finance replay

What Is Still Missing

1. Historical StateBench Replay Scoring

The user-requested benchmark should test whether models would have helped build this repo and research corpus over time. The first executable fixture layer now exists, but it has not yet been run across candidate models and runtimes.

The first finance replay fixture manifest now lives at statebench/suites/finance-replay-v0.md with a machine-readable companion at statebench/suites/finance-replay-v0.json. It now contains 118 materialized tasks: 80 source-acquisition tasks, 15 wiki-maintenance tasks, and 23 reflect-audit tasks. Local validation passed for JSON parsing, task listing, check compilation, and full expected-patch replay. Dry-run result rows and a 118/118 oracle expected-patch control now exist in finance-replay-v0-scorecard, but it is still not a scored model/backend run.

QFBench also now has a separate executable seed fixture: finance-quant-code-v0. That fixture scores a static answer artifact 8/8 for executable quant-code benchmark discipline and now also runs a tiny local code-execution smoke 3/3 over option payoff, drawdown/turnover, and factor-rank tasks. It covers source-boundary, Docker/numerical-verifier contract, memo-overclaim rejection, production-boundary, baseline comparison, task scope, QuantEval bridge, runtime trace, and deterministic verifier wiring. It is intentionally not a live QFBench run or a claim that any model can produce profitable quant code.

The suite now tests:

  1. find and download stable papers, reports, PDFs, and podcast/audio sources;
  2. ingest PDFs through the existing OCR pipeline;
  3. ingest podcasts through scripts/podcast_mine.py;
  4. extract source cards with URL, date, owner, claim type, and promotion tier;
  5. update the correct research pillar and wiki page;
  6. preserve caveats around alpha, performance, licensing, and authority;
  7. pass validate_research.py and git diff --check.

The next evaluation step is to run this suite against a small model/backend matrix and record not only pass/fail, but also cost, latency, context length, artifact format, backend, quantization, compatibility constraints, and failure mode.

2. Model/Runtime Scorecard

The repo has a strong candidate list, but it does not yet have comparable scores across model/runtime choices.

The first scorecard should compare:

  • frontier API controls for source discovery and synthesis;
  • Qwen3/Qwen3.5/Qwen3.6 local candidates;
  • Gemma finance candidates;
  • LiquidAI LFM2/LFM2.5 nano/runtime candidates;
  • finance-specific models such as Fino1, Fin-R1, ODA-Fin, rLLM-FinQA, and TraceAlchemy Gemma;
  • retrieval models such as FinE5, Qwen3 Embedding/Reranker, BGE, Arctic Embed, Jina nano, and LFM2-ColBERT;
  • SQL specialists such as Arctic-Text2SQL-R1 and future R2 availability;
  • market/time-series specialists such as Kronos only in the dedicated finance-time-series-market-v0 lane after leakage-safe fixtures and trading utility metrics exist.

Backend compatibility must be recorded per task: MLX, llama.cpp/GGUF, CUDA, AWS Neuron/Inferentia, AMD/ROCm, and API. Neuron compatibility is only relevant when serving/training on AWS Inferentia/Trainium-class infrastructure.

3. Direct Official Social Archival

The new talent-leakage lane promotes official firm pages and official job listings. It still does not fully archive official LinkedIn posts because third-party social pages are brittle and often require browser/session access.

Promotion rule remains:

  • official firm/careers/university pages can be findings;
  • named practitioner transcripts can be findings;
  • third-party LinkedIn commentary is source discovery only;
  • official LinkedIn posts can strengthen a claim only when URL/date/author and primary-source backing are preserved.

4. More Top-Fund Negative Evidence

The corpus is clear that public sources do not prove genAI alpha. It can still improve by collecting negative or skeptical material from quant shops:

  • HRT-style warnings about benchmarks not matching trading reality;
  • Winton-style selection-bias and interpretability commentary;
  • AQR/systematic quant overfitting, small-data, low-signal-to-noise, and domain-knowledge controls;
  • allocator due-diligence questions that expose weak AI programs.

This matters because StateBench should reward models that avoid overclaiming, not models that generate the most exciting investment narrative.

The AQR “Can Machines Learn Finance?” PDF is now locally captured and OCR’d in pillar 06. HRT’s official trading-AI, data, and benchmark articles now add the market-making version of the same negative-evidence lane. Together they should be used to turn skepticism into concrete StateBench checks: does the model notice that adding predictors is not the same as adding independent target observations, that low signal-to-noise makes flexible models fragile, that prediction / optimization / execution are separate tasks, that data needs relevance / uniqueness / lookahead / sample-size / noise / provenance gates, and that economic theory and human expertise are controls rather than optional commentary?

5. Source-Ingestion Robustness Scorecards

The web-discovery lane should stay an eval, not a benchmark run each time. The suite now has stable source-acquisition fixtures, but they still need a scorecard that separates stable replay from live-web/source-discovery behavior:

  • arXiv papers;
  • official PDFs;
  • whitepapers/reports;
  • podcast pages and RSS entries;
  • official firm pages;
  • official conference/session pages;
  • Kaggle competition pages and notebooks.

Success is not “browse the current web perfectly.” Success is finding the official item, downloading allowed artifacts, recording provenance, and abstaining when the source is not verifiable. Live web can refresh the fixture queue, but only the frozen task fixtures should count as comparable benchmark scores.

The operational queue for this work now lives at finance-quant-source-ingestion-queue-2026.md. It separates promotable sources from discovery-only social material and ties each source class to the existing PDF, podcast/audio, source-card, and StateBench pipelines.

6. Finance Retrieval and SQL Execution Harness

The retrieval/NL2SQL decision guide now has an operational StateBench suite blueprint at statebench/suites/finance-retrieval-sql-harness-v0.md. This is deliberately separate from finance-replay-v0.

The first seed fixture contract now lives at statebench/fixtures/finance-retrieval-sql/v0/. It contains a Postgres-first schema, a small retrieval corpus, forty-six gold tasks, and a Snowflake semantic-view contract. A static answer-artifact control now records 46/46 in finance-retrieval-sql-v0-scorecard.md. It is not a scored model/backend run yet.

The missing build work is:

  1. freeze representative document-RAG fixtures from existing PDFs, transcripts, source cards, and wiki evidence;
  2. implement a runner for the finance-retrieval-sql-v0 Postgres fixture and gold tasks;
  3. test the Snowflake-compatible schema manifest and semantic-view contract;
  4. adapt Defog SQL-Eval style execution scoring for local Postgres and future Snowflake runs;
  5. add retrieval recall, citation, RBAC, cost, stale-data, and unsafe-action refusal checks before any production-like deployment test.

Priority Next Work

  1. Run the 118-task finance-replay-v0 suite against a small bakeoff matrix: frontier API control, one Qwen3-family local/API candidate, one Gemma-family candidate, one LiquidAI runtime candidate, and one finance-specific candidate such as rLLM-FinQA or ODA-Fin.
  2. Run finance-quant-code-v0 against frontier, Qwen, and Gemma candidates, then graduate the tiny execution smoke into a larger code-execution harness only after answer-artifact and local-verifier controls stay clean.
  3. Create a model/runtime result table with task, model, backend, quantization, context length, cost, latency, pass/fail, and failure mode.
  4. Add a stable source-discovery fixture manifest for the next queue of papers, PDFs, whitepapers, podcasts, official firm pages, and Kaggle pages. The live/non-comparable backlog is now staged in finance-source-discovery-eval-v0 and should be promoted into frozen source-acquisition tasks only after evidence packs are captured.
  5. Build the first database-backed retrieval/SQL fixture slice from finance-retrieval-sql-harness-v0: Postgres first, Snowflake semantic-view contract second, external model bakeoff third.
  6. Convert repeated winning behavior into training pairs only after the eval proves the target behavior and after Unsloth, LLaMA-Factory, AWS native, and MLX training paths are compared for cost, feature fit, and backend limits.
  7. Add Neuron preflight records only for model artifacts that would actually be hosted on AWS Inferentia/Trainium; do not treat Neuron as a generic model requirement.

Bottom Line

The corpus now answers “what are we missing?” at the research-map level and the first replay suite exists. The remaining gap is measured execution: run models through finance-replay-v0, build the database-backed retrieval/SQL harness, capture model/runtime scorecards, and only then decide whether fine-tuning, conversion, or AWS-hosted deployment is worth the cost.