← Industry Verticals 🕐 17 min read
Industry Verticals

Finance/Quant AI Objective Coverage Audit

> **Document role:** internal completion audit for the active finance/quant AI

See also: finance-quant-ai-coverage-audit-2026.md · finance-quant-source-ingestion-queue-2026.md · statebench-model-candidate-registry-2026.md · finance-model-runtime-bakeoff-v0 · finance-source-discovery-eval-v0 · finance-sentiment-routing-v0

Document role: internal completion audit for the active finance/quant AI research goal. It maps the requested end state to current repo evidence, distinguishes proven coverage from remaining work, and keeps the scope from collapsing into only source collection or only benchmark plumbing.


Objective Requirements

The active objective decomposes into eight concrete requirements:

  1. identify what is missing in finance and quant AI across industry, papers, whitepapers, reports, podcasts, articles, and official social/talent signals;
  2. determine what is working and what is not working at systematic quant shops;
  3. capture leakage from podcasts, articles, official pages, job listings, and LinkedIn/company social without overclaiming alpha;
  4. ingest valuable sources through the repo’s existing PDF, podcast, source-card, wiki, and StateBench pipelines;
  5. separate live web/source discovery from deterministic benchmark scoring;
  6. identify interesting domain-specific models and SLMs not previously tracked;
  7. record backend/runtime limits across MLX, llama.cpp/GGUF, CUDA, AMD/ROCm, API, SageMaker, Bedrock, and AWS Neuron/Inferentia/Trainium;
  8. prove the work with executable checks, scorecards, or explicit evidence tables rather than loose notes.

Coverage Matrix

Requirement Current Evidence Status Remaining Gap
Industry/whitepaper/research-paper map finance-quant-ai-coverage-audit-2026.md, asset-management-quant-ai-gap-map-2026.md, finance-domain-models-benchmarks-2026.md, finance-ai-benchmark-cluster-2026.md, pillar 06/21 raw ledgers and OCR outputs including BlueFin, WorkstreamBench, and Deep FinResearch Bench Covered at source-awareness level Need scored model/runtime runs before calling any model/backend better for this repo
What is working at systematic quant shops top-hedge-fund-ai-agents-failure-modes-2026.md, buy-side-quant-ai-practitioner-signals-2026.md, quant-ai-infrastructure-moats-2026.md, finance-quant-audio-video-deep-dive-2026.md Covered as public workflow evidence No independent public proof of genAI alpha; keep claims at workflow/platform/governance level
What is not working Two Sigma temporal leakage / overfitting notes, HRT benchmark and data-provenance warnings, AQR small-data/low-signal OCR, Acadian model-capacity/overfitting controls, Winton selection-bias/signal-discipline commentary, FrontierFinance spreadsheet/state failures, BlueFin/WorkstreamBench workbook-agent failures, Deep FinResearch professional-report failures, CFA/Aon governance controls, U.S. Senate HSGAC hedge-fund AI/ML governance report, Balyasny factor-workflow caveats Covered as negative-evidence lane with materialized guardrail checks through ra-123 plus current benchmark/source cards Run real model/backend candidates against the guardrail checks and compare failure modes; add frozen tasks for workbook/report quality and regulatory audit-trail/disclosure quality if these become target workflows
Podcasts/audio/video leakage Acadian, Jane Street, Risk.net Quantcast, Curious Quant, The Derivative, Man Group, Tech Talks Daily, Better System Trader queue, scripts/podcast_sources.yaml Covered for promoted extracts and queue Continue targeted, not bulk, audio archival; reject weak host commentary
Articles / official pages / official social, vendor products, and talent leakage Millennium, Jump, Schonfeld, Tower, WorldQuant, Jane Street, HRT, XTX, G-Research, QRT, Winton, Bridgewater, Man AHL, Acadian, QuantConnect, terminal incumbents, finance-agent connector vendors, AI-native workflow platforms including CurrentAI (sa-122), plus sa-073, sa-087 ScaleDown finance-workflow rejection/watchlist, finance-talent-leakage-v0 static-control scorecard at 10/10, finance-vendor-product-signal-v0 fixture gates, and frontier-api-vendor-product-signal-control-preflight.json Covered at evidence-classification level Official LinkedIn/company-social capture remains brittle; require URL/date/account/archive or primary-source backing before promotion; product claims still need independent measurement before ROI/alpha promotion; vendor-product frontier/API row still needs cost estimate, cost cap, and endpoint/API route before scoring
Existing ingestion pipelines scripts/ingest_pdfs.py, scripts/podcast_mine.py, source ledgers under sources/, pillar 13 audio/source outputs, StateBench source-acquisition tasks, and the finance-tool-routing-v0 static-control scorecard at 10/10; the tool-routing runner now has a preflight-gated endpoint path Covered operationally Need continued validation that new source batches do not leave large binaries in git; real tool-router model rows still need endpoint/runtime smoke and approved preflight
Live source discovery as eval, not benchmark finance-source-discovery-eval-v0.md, source_discovery_eval.py, fixture-template coverage for sa-003, sa-006, sa-007, sa-012, sa-013, sa-023, sa-028, sa-072, sa-073, sa-080, sa-082, sa-084, sa-085, and sa-086, plus sa-081 as a quiet-firm official careers/PDF replay fixture Covered and executable; validator now reports 12 fixture families with materialized examples, including vendor-product signal classification, SQL benchmark-quality / dialect-portability discovery, HSGAC regulatory negative evidence, and BankerToolBench deal-work benchmark capture Materialize more frozen tasks only after evidence packs are captured
Domain-specific models and SLMs finance-domain-models-benchmarks-2026.md, statebench-model-candidate-registry-2026.md, finance-model-runtime-candidates-v0.json, finance-model-runtime-bakeoff-v0.md, finance-accounting-compliance-v0.md, finance-regional-language-v0-scorecard.md, finance-time-series-market-v0-scorecard.md, finance-quant-code-v0-scorecard.md, finance-treasury-payments-liquidity-v0-scorecard.md, finance-agentic-factor-discovery-v0-scorecard.md, finance-portfolio-risk-construction-v0-scorecard.md, Fin-o1, CPA-Qwen3, FinGuard, Amsi-fin MLX VLM, Fara-7B/Fara1.5 browser/source-discovery, ScaleDown task-specific SLM/context-compression (sa-123), Kronos time-series market model, regional-qwen3-finance-slm-watchlist-raw.md, current-finance-slm-watchlist-2026-05-31-raw.md, and kronos-financial-time-series-foundation-model-raw.md source cards Covered as watchlist/manifest/run queue with broader accounting/compliance, multimodal, browser/source-discovery, regional-language, India-policy, small-sentiment, local GGUF finance-reasoning, tiny embedding/retrieval, context compression, private-corpus cautionary, Web3/DeFi, financial time-series/market-model, treasury/payments/liquidity, agentic factor discovery, portfolio/risk construction, and executable quant-code lanes; accounting/compliance seed fixture scores 10/10, regional-language scores 12/12, time-series/market scores 14/14, treasury/payments/liquidity scores 12/12, agentic factor discovery scores 12/12, portfolio/risk construction scores 12/12, quant-code answer fixture scores 8/8, and quant-code execution smoke scores 3/3 static controls; treasury/payments, factor-discovery, and portfolio/risk lanes now each have three manifest candidate artifacts for future scoring No approved model/backend preflights or scored model runs yet; model rankings remain hypotheses
Retrieval, SQL, Snowflake, Postgres, custom schema finance-retrieval-sql-harness-v0.md, finance-retrieval-sql-v0-scorecard.md, finance-retrieval-nl2sql-harnesses-2026.md, sql-benchmark-quality-dialect-portability-raw.md, 46-task fixture, runner, static answer-artifact score summary, postgres-preflight-local-20260530.json, and snowflake-semantic-contract-local-20260530.json Executable seed plus durable static-control scorecard exists; local Postgres preflight records exact environment blocker; Snowflake semantic contract now passes local fixture-alignment checks; sa-084 adds module-diagnostic and dialect-portability source-acquisition coverage for NL2SQLBench/PARROT Postgres execution proof blocked locally by missing driver; real Snowflake Cortex/warehouse/RBAC/cost proof not run; no approved model/backend/preflight score yet; harness still needs module-level schema-selection/query-revision/dialect-portability scoring beyond the seed static fixture
Sentiment/routing specialists finance-sentiment-routing-v0.md, finance-sentiment-routing-v0-scorecard.md, finance-measured-runtime-evidence-ledger-2026.md, 15-task synthetic held-out fixture, YorkFr, ModernFinBERT, ProsusAI FinBERT, and FinSenti manifest targets, full-routing and label-only score rows Executable seed plus durable static-control scorecard plus real local model/backend rows exist; label-only mode now separates classifier usefulness from workflow-router acceptance; measured ledger records ModernFinBERT as current label-only baseline at 11/15 and all real local sentiment models at 0/15 full-routing acceptance YorkFr and label-only encoders still fail full routing controls; need a stronger routing candidate such as FinSenti or Nexus TinyFunction after adapter/backend proof
Backend/runtime compatibility Candidate manifest, preflight schema, suite-run preflight gates, retrieval backend preflight gate, model portability docs, a validated finance-model-runtime-bakeoff-v0 queue with 28 preflight records and 0/28 ready, and finance-scored-run-readiness-2026-06-02.md Covered as controls Every concrete model/backend still needs artifact-specific preflight before scoring; nearest rows are frontier API cost-cap approval, vendor-product frontier/API cost/route approval, CPA-Qwen3 endpoint smoke, Kuvera advisor endpoint smoke, Nexus tool-routing endpoint smoke, Qwen3.6/Gemma endpoint smoke, or regional-language endpoint smoke
Comparable model/runtime scores finance-replay-v0 118 tasks, finance-replay-v0-scorecard, dry-run scorecard rows, 118/118 full oracle expected-patch control, retrieval/SQL answer scorer, static retrieval/SQL summary, finance-quant-code-v0 answer scorer and execution smoke, regional-language/time-series static scorecards, and candidate manifest command generation Infrastructure, metadata dry-run rows, current full-suite oracle ceiling, negative-evidence guardrails, vendor-platform source fixture, quiet-firm official-careers fixture, Balyasny factor-agent caveat fixture, QFBench executable quant-code source fixture, quant-code static controls, regional/market-model static controls, and current model-card follow-up source tasks now exist Run approved model/backend candidates and record non-oracle scorecards

Evidence Strength

Evidence Class What It Proves What It Does Not Prove
Raw source ledgers, PDFs, OCR outputs, podcast transcripts, and source cards The repo can preserve provenance for papers, whitepapers, reports, practitioner media, and official pages That the source is high value, complete, or current after retrieval
Pillar 06 and 21 synthesis docs The research map now covers the major finance/quant source lanes, vendor signals, practitioner leaks, benchmarks, and domain-model candidates That any model/backend can reproduce the work or that public sources prove live investment alpha
finance-source-discovery-eval-v0 and fixture-template coverage Live discovery is separated from deterministic benchmark scoring and has 12 promotable frozen task patterns: podcast RSS, official firm/careers, official social, negative evidence, paper/PDF, Kaggle, SQL/retrieval, SQL benchmark-quality/dialect portability, domain model/runtime, vendor/customer architecture, vendor-product signal classification, and finance workflow vendor platforms. Current fixture-coverage validation passes for all 12 families; sa-087 now demonstrates rejection/watchlist discipline when a vendor does not prove finance workflow scope. That live web results are stable or comparable across time
finance-tool-routing-v0, 10 tasks, answer scorer, and static-control summary The repo has an executable policy for routing finance/quant source work through PDF ingestion, podcast mining, source cards, StateBench fixture creation, retrieval/SQL refusal, paid-API refusal, runtime policy, backend portability, official-social boundaries, and git safety That any function-calling model or router is reliable, production-safe, or competent at finance reasoning
finance-talent-leakage-v0, 10 tasks, answer scorer, durable static-control summary, and scorecard The repo has an executable policy for promoting official firm/careers/university/social evidence while rejecting unverified recruiter summaries, reposts, informal social commentary, and alpha/autonomy overclaims; the static-control row records 10/10 That live LinkedIn coverage is complete, that a firm has production AI beyond official evidence, that talent signals prove alpha, or that a source-discovery/browser model can find and archive new pages
finance-vendor-product-signal-v0, 12 tasks, answer scorer, durable static-control summary, scorecard, endpoint runner path, and frontier/API control preflight The repo has an executable policy for promoting finance AI product architecture while rejecting generic marketing, unverified metrics, podcast anecdotes, terminal-replacement claims, and alpha/ROI overclaims; the static-control row records 12/12, the runner can score a preflight-approved endpoint, and frontier-api-vendor-product-signal-control is now a concrete pending candidate That any vendor product is recommended, independently measured, proven to improve returns/productivity, or that the frontier/API candidate has cost, route, endpoint, or scored-run approval
finance-treasury-payments-liquidity-v0, 12 tasks, answer scorer, durable static-control summary, and scorecard The repo has an executable policy for treasury, payments, liquidity, reconciliation, settlement-finality, controlled agentic money-movement, kill-switch, audit-trail, and runtime-portability claims; the static-control row records 12/12 That any model can move money, that a payment rail is endorsed, that settlement is legally final, that sanctions/fraud controls are sufficient, or that a model/backend has been approved for this suite
finance-kaggle-artifact-v0, 10 tasks, answer scorer, durable static-control summary, and scorecard The repo has an executable policy for treating finance Kaggle/competition pages as benchmark-design evidence: official metadata, metric/split, leakage, baseline, notebook status, MLE-harness, community-artifact timing, git-boundary, and not-alpha checks; the static-control row records 10/10 That a leaderboard result proves finance alpha, production deployment, execution quality, autonomous investment authority, or model/backend competence
finance-replay-v0, 118 materialized tasks, dry-run rows, 118/118 oracle expected-patch control, and scorecard plumbing The repo has a deterministic replay surface for source acquisition, wiki maintenance, overclaim repair, run metadata aggregation, and executable check ceiling That any candidate model has passed it
Frontier/API cost-estimation preflight statebench.runners.finance_replay_cost_estimate estimates the 118-task replay control at $7.51 before any paid API call, recorded in statebench/results/finance-replay-v0-cost-estimate-frontier-api-20260601-expanded118.json That spending is approved, that actual tokenization/output/retries will match the estimate, or that the frontier model will pass
finance-regional-language-v0, 12 tasks, answer scorer, durable static-control summary, and scorecard The repo has an executable policy for regional finance-language specialists: A-share/CFA, Turkish/BIST, Thai/English, India budget-policy, regulated-advice refusal, translation integrity, benchmark-only overclaim, backend portability, exact-revision gates, and unverified social/vendor claims; the static-control row records 12/12 That any regional SLM transfers to U.S. finance, SEC filing analysis, accounting compliance, institutional quant research, investment advice, live trading, or alpha generation
finance-time-series-market-v0, 14 tasks, answer scorer, durable static-control summary, and scorecard The repo has an executable policy for financial time-series and market foundation models: source discovery, runnable-artifact separation, point-in-time leakage controls, forecast metrics, related/unrelated multivariate context, trading utility, transaction costs, capacity, and runtime portability; the static-control row records 14/14 That Kronos, FinCast, TimesFM, Chronos, Moirai, Lag-Llama, LOBERT, ByteGen, LiT, or any market model has forecasting skill, trading utility, live alpha, or backend portability
BankerToolBench deal-work source fixture sa-086 freezes the investment-banking/deal-work source ledger around BankerToolBench: data rooms, market data, SEC filings, Excel models, PowerPoint decks, PDF/Word reports, multi-file deliverables, banker rubrics, client-readiness caveats, and cross-artifact consistency That buy-side quant alpha, live trading, or production client readiness is proven
finance-quant-code-v0, 11 tasks, answer scorer, execution smoke, sample-pass artifacts, scorecard, and candidate-manifest target coverage The repo has a seed harness for QFBench/QuantEval-style executable quant-code discipline: source-boundary, Docker/numerical-verifier contract, memo-overclaim rejection, production-boundary, baseline comparison, task scope, QuantEval bridge, runtime trace, option payoff, drawdown/turnover, and factor-rank verifier wiring. Static controls record 8/8 answer-artifact and 3/3 local execution-smoke rows. That QFBench itself has been run, that any model can generate correct quant code, or that a benchmark pass proves live alpha
finance-agentic-factor-discovery-v0, 12 tasks, answer scorer, durable static-control summary, and scorecard The repo has an executable policy for source-to-factor conversion, DSL/AST safety, unsafe program rejection, point-in-time evaluation, factor diagnostics, implementation realism, diversity/crowding, neutralization/attribution, walk-forward promote/hold/retire decisions, narrative-bias rejection, and runtime portability; the static-control row records 12/12 That any factor has live alpha, that any model can generate valid factor code, that a backtest is production-ready, or that a model/backend is approved for this suite
finance-portfolio-risk-construction-v0, 12 tasks, answer scorer, durable static-control summary, and scorecard The repo has an executable policy for portfolio screen audits, covariance experiment design, graph construction critique, constrained weight construction, action-first RL policy, risk-aware agent controls, inverse-RL limits, backtest skepticism, and runtime portability; the static-control row records 12/12 That any portfolio is recommended, that an optimizer/RL policy is production-ready, that backtest metrics prove allocation quality, or that a model/backend is approved for this suite
U.S. Senate hedge-fund AI/ML report source card and OCR The repo has a government/regulatory source for inconsistent AI/ML definitions, human-review ambiguity, vague client disclosures, testing/review process gaps, overfitting/backtesting limits, and audit-trail/version-control recommendations across Citadel, Renaissance Technologies, Bridgewater, AI Capital Management, Numerai, and WorldQuant evidence That any named fund has proven alpha, that current 2026 deployment status is known, or that regulatory recommendations are implemented
finance-retrieval-sql-harness-v0, 46 tasks, answer scorer, static-control summary, local Postgres preflight artifact, and local Snowflake semantic-contract artifact The repo has a seed harness for finance retrieval, NL2SQL, Postgres, and Snowflake-style semantic contracts; the answer-artifact scorecard path works end-to-end; local Postgres readiness is blocked specifically on missing Python driver; the Snowflake semantic-view YAML aligns with the fixture schema/corpus/controls That a real model/backend, Postgres/Snowflake backend, pgvector/Cortex route, warehouse execution trace, cost cap, or production RBAC layer has been scored
finance-sentiment-routing-v0, 15 tasks, answer scorer, static-control summary, local MPS classifier rows, parser-contract row, and measured-runtime ledger The repo has a held-out synthetic scorer for headline/earnings/commentary sentiment, routing, abstention, and no-alpha overclaim checks; ModernFinBERT, ProsusAI FinBERT, and YorkFr have real local MPS label-only rows; ModernFinBERT leads label-only at 11/15; all real local sentiment models score 0/15 on full routing; a Qwen3 GRPO parser-contract row proves output parsing can score 15/15 when the answer contract is satisfied That any sentiment classifier is ready for full workflow routing; parser-contract success is not model-generation proof; label-only rows do not prove source acquisition, SQL, accounting, portfolio, or alpha competence
Model candidate manifest, bakeoff queue, preflight schema, accounting/compliance fixture, advisor/portfolio fixture, regional-language fixture, time-series seed fixture, quant-code fixture, and talent/social leakage seed fixture Candidate families, runtimes, backend constraints, run phases, hardware-specific gates, and first accounting/compliance/advisor/regional/time-series/quant-code/talent-leakage guardrail tasks are recorded, including newly promoted Fin-o1 Qwen3 reasoning, CPA-Qwen3 accounting, Kuvera personal-finance, FinGuard compliance-guard, Amsi-fin MLX VLM, Fara-7B source-discovery, Kronos market-model candidates, and a self-driving-portfolio controlled-delegation task for IPS/script/peer-review/human-approval boundaries. The advisor/portfolio static control records 13/13, regional-language records 12/12, time-series/market records 14/14, and quant-code records 8/8 answer-artifact plus 3/3 execution-smoke controls. That Qwen, Gemma, Liquid, Fara, Kronos, Kuvera, finance specialists, MLX, GGUF, CUDA, AMD, API, or Neuron are approved for a given task

What We Are Still Missing

The remaining gap is measured execution, not source awareness. The corpus now has enough finance/quant industry, model, benchmark, podcast, report, and official-practitioner evidence to start answering the strategic question. What is still missing is proof that specific models and backends can do the work better, cheaper, or more reliably.

The highest-value unfinished items are:

  1. First scored model finance-replay-v0 run. Dry-run rows now prove manifest and scorecard aggregation, and the full oracle expected-patch control now proves 118/118 suite/check executability across source acquisition, wiki maintenance, and reflect-audit tasks. The current suite adds ra-123 for Balyasny multi-strat factor-agent caveats and sa-083 for QFBench executable quant-code benchmark evidence. The refreshed frontier/API cost estimate is $7.512744; next run a frontier API control after explicit cost_cap, or one current open/local candidate after endpoint smoke. A fresh qwen3-6-35b-a3b-hf localhost smoke artifact shows localhost:8000 is not currently serving the model, so the Qwen route is still blocked on backend_memory_check. Summarize with the scorecard runner as non-oracle model/backend rows.
  2. First approved model preflight. Fill one preflight record for a candidate that can actually run in the current environment; do not approve Neuron until there is real Inferentia/Trainium compile evidence. The current readiness note ranks the first viable paths: frontier API is blocked only on cost_cap; cpa-qwen3-8b-v0 has all source checks passed but needs an approved endpoint; Kuvera now has a preflight-gated advisor/portfolio endpoint runner path but still needs runtime smoke; Qwen3.6/Gemma need endpoint-memory smoke; regional SLMs now have a preflight-gated endpoint runner path but still need endpoint smoke and preflight approval.
  3. First quant-code model/backend row. The new finance-quant-code-v0 seed fixture records the QFBench/QuantEval execution discipline, scores a static answer artifact 8/8, and runs a tiny deterministic execution smoke 3/3. The missing layer is a real model/backend row, then a larger code-execution harness with realistic frozen inputs, task-local data, verifier, oracle/reference solution, repeated runs, one-shot baseline, tool traces, cost, latency, and failure-mode capture.
  4. First retrieval/SQL backend score. The static answer-artifact control now records 46/46 in statebench/results/finance-retrieval-sql-v0/sample-pass-artifact.summary.json. The missing layer is a preflight-approved model/backend row, followed by Postgres and Snowflake proof. Local Postgres preflight is captured and currently fails only on missing driver availability (psycopg, psycopg2, or pg8000), so the next environment step is installing a driver and pointing the harness at a disposable database. The Snowflake semantic contract now has a passing local fixture-alignment artifact, but the missing Snowflake layer is still real Cortex/warehouse execution with RBAC, traces, and cost metadata.
  5. Sentiment/routing next candidate. The YorkFr, ModernFinBERT, and ProsusAI FinBERT local MPS rows exist. ModernFinBERT leads label-only acceptance at 11/15, while YorkFr and ProsusAI accept 10/15; all score 0/15 accepted on full workflow routing. The qwen3-grpo-parser-contract row scores 15/15 as a parser/output-contract smoke, not as a model-quality result. Next compare FinSenti after the Qwen3.5/FinSenti adapter or GGUF backend is proven.
  6. Official social hardening. Convert official LinkedIn/company-social discoveries into source cards only when the account, URL, date, and archive/screenshot or primary-source backing are preserved. The finance-talent-leakage-v0 static-control scorecard now records 10/10 for the promotion-policy fixture, and sa-073 makes this a first-class source-discovery family rather than a buried subcase of firm/careers pages. The remaining work is scoring real browser/source-discovery agents and adding new frozen tasks only when official evidence packs are captured. 6a. Vendor product-signal scoring. The new finance-vendor-product-signal-v0 fixture covers QuantConnect-style research pipelines, assistant teams, incumbent terminal AI, governed connector ecosystems, AI-native workflow platforms, named-customer quote caveats, podcast source-discovery routing, terminal-replacement overclaims, vendor metrics, MCP/tool surfaces, and wiki promotion policy. sa-082 now makes this a materialized source-acquisition fixture family in finance-source-discovery-eval-v0. The remaining work is scoring real models on it and only expanding it when a new vendor pattern appears. The runner now has a preflight-gated endpoint path, and the dedicated frontier-api-vendor-product-signal-control preflight exists, but this lane still needs cost estimate/cap and endpoint/API-route evidence before a non-static model/backend row can be produced. 6b. Agentic factor-discovery scoring. The new finance-agentic-factor-discovery-v0 fixture covers Hubble/FactorEngine style source-to-factor conversion, safe DSL/AST execution, point-in-time validation, diagnostics, turnover/cost/capacity controls, diversity/crowding, neutralization, walk-forward promote/hold/retire decisions, narrative-bias rejection, and runtime portability. The remaining work is a preflight-approved model row plus a deterministic toy factor-code harness with point-in-time data, reference results, cost/capacity checks, and failure traces. 6c. Portfolio/risk-construction scoring. The new finance-portfolio-risk-construction-v0 fixture covers screened-universe audits, covariance/risk experiment design, graph-construction critique, constrained weights, action-first RL policy, risk-aware multi-agent controls, inverse-RL evidence limits, backtest skepticism, and runtime portability. The remaining work is a preflight-approved model row plus deterministic covariance/weighting and graph-construction harnesses with point-in-time data, reference outputs, turnover/cost/capacity checks, and survivorship-bias controls.
  7. Negative-evidence task expansion. The first expansion is now materialized in ra-119 through ra-123: Acadian process overclaim, Two Sigma temporal leakage, Winton selection bias, cross-firm autonomous agent-alpha collapse, and Balyasny factor-agent workflow caveats. sa-085 adds the Senate/HSGAC regulatory source-acquisition counterpart for system-definition consistency, human-review boundary, test cadence, version control, disclosure adequacy, and audit-trail preservation. The remaining work is scoring real candidates and adding more tasks only when new practitioner or regulatory evidence introduces a distinct failure mode.
  8. Domain SLM bakeoff. Score at least one current Qwen-family artifact, one Gemma-family artifact, one Liquid task specialist, one finance reasoning specialist, one accounting/compliance specialist on finance-accounting-compliance-v0, one advisor/personal-finance specialist on finance-advisor-portfolio-v0, one treasury/payments/liquidity row on finance-treasury-payments-liquidity-v0, one agentic factor-discovery row on finance-agentic-factor-discovery-v0, one portfolio/risk-construction row on finance-portfolio-risk-construction-v0, one multimodal finance VLM, one browser/source-discovery agent such as Fara-7B or Fara1.5, one regional/language finance SLM where a matching fixture exists, one local GGUF finance-reasoning artifact such as Qwen3-8B finance GGUF, one retrieval/reranking specialist or tiny embedding artifact on the appropriate suite, and one finance-native time-series market model such as Kronos only after finance-time-series-market-v0 has locked data, leakage controls, baselines, transaction-cost/capacity assumptions, and runtime smoke proof.
  9. Podcast and article quality triage. Keep adding only named-practitioner, source-backed media to findings; vendor/product podcasts and social posts stay as discovery unless they expose architecture, process controls, failure modes, or named customer evidence.
  10. Kaggle and competition harness extraction. Preserve competition metrics, split rules, leakage controls, kernels/notebooks, and local graders as eval design evidence; do not treat competition rankings as finance-alpha proof. The finance-kaggle-artifact-v0 static-control scorecard now records 10/10 for official metadata, metric/split, leakage, baseline, notebook, MLE-harness, community-timing, git-boundary, and not-alpha checks. The remaining work is scoring real browser/source-discovery agents and building larger executable competition-style harnesses only when frozen task-local data, reference solutions, leakage controls, runtime traces, and git exclusions are present.
  11. Treasury/payments execution harness. The finance-treasury-payments-liquidity-v0 static-control suite now covers the regulated money-movement policy boundary. The remaining work is a preflight-approved model row plus synthetic mandate/invoice/statement/ledger fixtures and deterministic tool traces for authorization, sanctions/fraud, settlement, reconciliation, kill-switch, and audit controls.
  12. Investment-banking/deal-work execution harness. BankerToolBench is now frozen as source-acquisition fixture sa-086, but the repo still needs a deterministic StateBench task family for senior-banker request parsing, data-room provenance, market/SEC research, Excel formula integrity, deck/memo generation, cross-artifact reconciliation, confidentiality/MNPI, and human-review handoff.

Current Interpretation

What top finance and quant shops appear to be doing publicly:

  • using AI to accelerate research, coding, document analysis, feature discovery, data exploration, and analyst/PM workflows;
  • building governed internal platforms around data lineage, entitlement checks, notebooks, backtests, logs, and evaluation gates;
  • using ML at industrial scale for market prediction, execution, risk, storage, and low-latency deployment;
  • treating broad chat assistants as shallow unless embedded in firm-specific tools, permissions, and workflows.

What is not proven or not working reliably:

  • public sources do not independently prove genAI-driven live alpha;
  • general agents can widen the hypothesis funnel and worsen false discovery;
  • LLMs introduce historical knowledge-cutoff and temporal-leakage risks;
  • finance artifacts can look plausible while formulas, state, citations, or execution traces are wrong;
  • vendor product pages are architecture/product signals unless named customers or reproducible measurements are present.

Completion Standard

The objective should not be marked complete until current repo evidence proves:

  1. promoted source batches have raw source records or explicit reject/watchlist decisions;
  2. source-discovery categories have frozen StateBench fixture coverage;
  3. domain-model candidates have runnable manifest entries and preflight gates;
  4. at least one scored model/runtime run exists for finance-replay-v0;
  5. at least one scored retrieval/SQL answer-artifact run exists, and at least one model/backend retrieval/SQL row exists with backend metadata and preflight validation;
  6. docs and wiki pages distinguish workflow evidence from alpha/performance evidence;
  7. validation commands for the touched runners, manifests, and fixtures pass.

Until those scorecards exist, the correct status is: research map and ingestion control plane are strong; model/runtime recommendations remain unmeasured.

Next Verification Commands

Use these checks before claiming another layer is complete:

python3 -m statebench.runners.source_discovery_eval \
  --queue statebench/suites/finance-source-discovery-eval-v0.json \
  --validate-only

python3 -m statebench.runners.source_discovery_eval \
  --queue statebench/suites/finance-source-discovery-eval-v0.json \
  --section fixture-coverage \
  --format markdown

python3 -m statebench.runners.finance_retrieval_sql_fixture \
  --fixture-dir statebench/fixtures/finance-retrieval-sql/v0 \
  --answers-file statebench/fixtures/finance-retrieval-sql/v0/sample_answers/pass.json \
  --format markdown

python3 -m statebench.runners.finance_retrieval_sql_fixture \
  --fixture-dir statebench/fixtures/finance-retrieval-sql/v0 \
  --snowflake-contract-check \
  --format markdown

python3 -m statebench.runners.model_candidate_manifest \
  --manifest statebench/suites/finance-model-runtime-candidates-v0.json \
  --format markdown

python3 -m statebench.runners.suite_runner \
  --suite-id finance-replay-v0-oracle-full-20260601-expanded118 \
  --suite-file statebench/suites/finance-replay-v0.json \
  --backend frontier-api \
  --model oracle-expected-patch \
  --policy overwrite \
  --oracle-expected-patch