See also: finance-quant-ai-coverage-audit-2026.md · finance-quant-source-ingestion-queue-2026.md · statebench-model-candidate-registry-2026.md · finance-model-runtime-bakeoff-v0 · finance-source-discovery-eval-v0 · finance-sentiment-routing-v0
Document role: internal completion audit for the active finance/quant AI research goal. It maps the requested end state to current repo evidence, distinguishes proven coverage from remaining work, and keeps the scope from collapsing into only source collection or only benchmark plumbing.
Objective Requirements
The active objective decomposes into eight concrete requirements:
- identify what is missing in finance and quant AI across industry, papers, whitepapers, reports, podcasts, articles, and official social/talent signals;
- determine what is working and what is not working at systematic quant shops;
- capture leakage from podcasts, articles, official pages, job listings, and LinkedIn/company social without overclaiming alpha;
- ingest valuable sources through the repo’s existing PDF, podcast, source-card, wiki, and StateBench pipelines;
- separate live web/source discovery from deterministic benchmark scoring;
- identify interesting domain-specific models and SLMs not previously tracked;
- record backend/runtime limits across MLX, llama.cpp/GGUF, CUDA, AMD/ROCm, API, SageMaker, Bedrock, and AWS Neuron/Inferentia/Trainium;
- prove the work with executable checks, scorecards, or explicit evidence tables rather than loose notes.
Coverage Matrix
| Requirement | Current Evidence | Status | Remaining Gap |
|---|---|---|---|
| Industry/whitepaper/research-paper map | finance-quant-ai-coverage-audit-2026.md, asset-management-quant-ai-gap-map-2026.md, finance-domain-models-benchmarks-2026.md, finance-ai-benchmark-cluster-2026.md, pillar 06/21 raw ledgers and OCR outputs including BlueFin, WorkstreamBench, and Deep FinResearch Bench | Covered at source-awareness level | Need scored model/runtime runs before calling any model/backend better for this repo |
| What is working at systematic quant shops | top-hedge-fund-ai-agents-failure-modes-2026.md, buy-side-quant-ai-practitioner-signals-2026.md, quant-ai-infrastructure-moats-2026.md, finance-quant-audio-video-deep-dive-2026.md | Covered as public workflow evidence | No independent public proof of genAI alpha; keep claims at workflow/platform/governance level |
| What is not working | Two Sigma temporal leakage / overfitting notes, HRT benchmark and data-provenance warnings, AQR small-data/low-signal OCR, Acadian model-capacity/overfitting controls, Winton selection-bias/signal-discipline commentary, FrontierFinance spreadsheet/state failures, BlueFin/WorkstreamBench workbook-agent failures, Deep FinResearch professional-report failures, CFA/Aon governance controls, U.S. Senate HSGAC hedge-fund AI/ML governance report, Balyasny factor-workflow caveats | Covered as negative-evidence lane with materialized guardrail checks through ra-123 plus current benchmark/source cards |
Run real model/backend candidates against the guardrail checks and compare failure modes; add frozen tasks for workbook/report quality and regulatory audit-trail/disclosure quality if these become target workflows |
| Podcasts/audio/video leakage | Acadian, Jane Street, Risk.net Quantcast, Curious Quant, The Derivative, Man Group, Tech Talks Daily, Better System Trader queue, scripts/podcast_sources.yaml |
Covered for promoted extracts and queue | Continue targeted, not bulk, audio archival; reject weak host commentary |
| Articles / official pages / official social, vendor products, and talent leakage | Millennium, Jump, Schonfeld, Tower, WorldQuant, Jane Street, HRT, XTX, G-Research, QRT, Winton, Bridgewater, Man AHL, Acadian, QuantConnect, terminal incumbents, finance-agent connector vendors, AI-native workflow platforms including CurrentAI (sa-122), plus sa-073, sa-087 ScaleDown finance-workflow rejection/watchlist, finance-talent-leakage-v0 static-control scorecard at 10/10, finance-vendor-product-signal-v0 fixture gates, and frontier-api-vendor-product-signal-control-preflight.json |
Covered at evidence-classification level | Official LinkedIn/company-social capture remains brittle; require URL/date/account/archive or primary-source backing before promotion; product claims still need independent measurement before ROI/alpha promotion; vendor-product frontier/API row still needs cost estimate, cost cap, and endpoint/API route before scoring |
| Existing ingestion pipelines | scripts/ingest_pdfs.py, scripts/podcast_mine.py, source ledgers under sources/, pillar 13 audio/source outputs, StateBench source-acquisition tasks, and the finance-tool-routing-v0 static-control scorecard at 10/10; the tool-routing runner now has a preflight-gated endpoint path |
Covered operationally | Need continued validation that new source batches do not leave large binaries in git; real tool-router model rows still need endpoint/runtime smoke and approved preflight |
| Live source discovery as eval, not benchmark | finance-source-discovery-eval-v0.md, source_discovery_eval.py, fixture-template coverage for sa-003, sa-006, sa-007, sa-012, sa-013, sa-023, sa-028, sa-072, sa-073, sa-080, sa-082, sa-084, sa-085, and sa-086, plus sa-081 as a quiet-firm official careers/PDF replay fixture |
Covered and executable; validator now reports 12 fixture families with materialized examples, including vendor-product signal classification, SQL benchmark-quality / dialect-portability discovery, HSGAC regulatory negative evidence, and BankerToolBench deal-work benchmark capture | Materialize more frozen tasks only after evidence packs are captured |
| Domain-specific models and SLMs | finance-domain-models-benchmarks-2026.md, statebench-model-candidate-registry-2026.md, finance-model-runtime-candidates-v0.json, finance-model-runtime-bakeoff-v0.md, finance-accounting-compliance-v0.md, finance-regional-language-v0-scorecard.md, finance-time-series-market-v0-scorecard.md, finance-quant-code-v0-scorecard.md, finance-treasury-payments-liquidity-v0-scorecard.md, finance-agentic-factor-discovery-v0-scorecard.md, finance-portfolio-risk-construction-v0-scorecard.md, Fin-o1, CPA-Qwen3, FinGuard, Amsi-fin MLX VLM, Fara-7B/Fara1.5 browser/source-discovery, ScaleDown task-specific SLM/context-compression (sa-123), Kronos time-series market model, regional-qwen3-finance-slm-watchlist-raw.md, current-finance-slm-watchlist-2026-05-31-raw.md, and kronos-financial-time-series-foundation-model-raw.md source cards |
Covered as watchlist/manifest/run queue with broader accounting/compliance, multimodal, browser/source-discovery, regional-language, India-policy, small-sentiment, local GGUF finance-reasoning, tiny embedding/retrieval, context compression, private-corpus cautionary, Web3/DeFi, financial time-series/market-model, treasury/payments/liquidity, agentic factor discovery, portfolio/risk construction, and executable quant-code lanes; accounting/compliance seed fixture scores 10/10, regional-language scores 12/12, time-series/market scores 14/14, treasury/payments/liquidity scores 12/12, agentic factor discovery scores 12/12, portfolio/risk construction scores 12/12, quant-code answer fixture scores 8/8, and quant-code execution smoke scores 3/3 static controls; treasury/payments, factor-discovery, and portfolio/risk lanes now each have three manifest candidate artifacts for future scoring |
No approved model/backend preflights or scored model runs yet; model rankings remain hypotheses |
| Retrieval, SQL, Snowflake, Postgres, custom schema | finance-retrieval-sql-harness-v0.md, finance-retrieval-sql-v0-scorecard.md, finance-retrieval-nl2sql-harnesses-2026.md, sql-benchmark-quality-dialect-portability-raw.md, 46-task fixture, runner, static answer-artifact score summary, postgres-preflight-local-20260530.json, and snowflake-semantic-contract-local-20260530.json |
Executable seed plus durable static-control scorecard exists; local Postgres preflight records exact environment blocker; Snowflake semantic contract now passes local fixture-alignment checks; sa-084 adds module-diagnostic and dialect-portability source-acquisition coverage for NL2SQLBench/PARROT |
Postgres execution proof blocked locally by missing driver; real Snowflake Cortex/warehouse/RBAC/cost proof not run; no approved model/backend/preflight score yet; harness still needs module-level schema-selection/query-revision/dialect-portability scoring beyond the seed static fixture |
| Sentiment/routing specialists | finance-sentiment-routing-v0.md, finance-sentiment-routing-v0-scorecard.md, finance-measured-runtime-evidence-ledger-2026.md, 15-task synthetic held-out fixture, YorkFr, ModernFinBERT, ProsusAI FinBERT, and FinSenti manifest targets, full-routing and label-only score rows | Executable seed plus durable static-control scorecard plus real local model/backend rows exist; label-only mode now separates classifier usefulness from workflow-router acceptance; measured ledger records ModernFinBERT as current label-only baseline at 11/15 and all real local sentiment models at 0/15 full-routing acceptance |
YorkFr and label-only encoders still fail full routing controls; need a stronger routing candidate such as FinSenti or Nexus TinyFunction after adapter/backend proof |
| Backend/runtime compatibility | Candidate manifest, preflight schema, suite-run preflight gates, retrieval backend preflight gate, model portability docs, a validated finance-model-runtime-bakeoff-v0 queue with 28 preflight records and 0/28 ready, and finance-scored-run-readiness-2026-06-02.md |
Covered as controls | Every concrete model/backend still needs artifact-specific preflight before scoring; nearest rows are frontier API cost-cap approval, vendor-product frontier/API cost/route approval, CPA-Qwen3 endpoint smoke, Kuvera advisor endpoint smoke, Nexus tool-routing endpoint smoke, Qwen3.6/Gemma endpoint smoke, or regional-language endpoint smoke |
| Comparable model/runtime scores | finance-replay-v0 118 tasks, finance-replay-v0-scorecard, dry-run scorecard rows, 118/118 full oracle expected-patch control, retrieval/SQL answer scorer, static retrieval/SQL summary, finance-quant-code-v0 answer scorer and execution smoke, regional-language/time-series static scorecards, and candidate manifest command generation |
Infrastructure, metadata dry-run rows, current full-suite oracle ceiling, negative-evidence guardrails, vendor-platform source fixture, quiet-firm official-careers fixture, Balyasny factor-agent caveat fixture, QFBench executable quant-code source fixture, quant-code static controls, regional/market-model static controls, and current model-card follow-up source tasks now exist | Run approved model/backend candidates and record non-oracle scorecards |
Evidence Strength
| Evidence Class | What It Proves | What It Does Not Prove |
|---|---|---|
| Raw source ledgers, PDFs, OCR outputs, podcast transcripts, and source cards | The repo can preserve provenance for papers, whitepapers, reports, practitioner media, and official pages | That the source is high value, complete, or current after retrieval |
| Pillar 06 and 21 synthesis docs | The research map now covers the major finance/quant source lanes, vendor signals, practitioner leaks, benchmarks, and domain-model candidates | That any model/backend can reproduce the work or that public sources prove live investment alpha |
finance-source-discovery-eval-v0 and fixture-template coverage |
Live discovery is separated from deterministic benchmark scoring and has 12 promotable frozen task patterns: podcast RSS, official firm/careers, official social, negative evidence, paper/PDF, Kaggle, SQL/retrieval, SQL benchmark-quality/dialect portability, domain model/runtime, vendor/customer architecture, vendor-product signal classification, and finance workflow vendor platforms. Current fixture-coverage validation passes for all 12 families; sa-087 now demonstrates rejection/watchlist discipline when a vendor does not prove finance workflow scope. |
That live web results are stable or comparable across time |
finance-tool-routing-v0, 10 tasks, answer scorer, and static-control summary |
The repo has an executable policy for routing finance/quant source work through PDF ingestion, podcast mining, source cards, StateBench fixture creation, retrieval/SQL refusal, paid-API refusal, runtime policy, backend portability, official-social boundaries, and git safety | That any function-calling model or router is reliable, production-safe, or competent at finance reasoning |
finance-talent-leakage-v0, 10 tasks, answer scorer, durable static-control summary, and scorecard |
The repo has an executable policy for promoting official firm/careers/university/social evidence while rejecting unverified recruiter summaries, reposts, informal social commentary, and alpha/autonomy overclaims; the static-control row records 10/10 |
That live LinkedIn coverage is complete, that a firm has production AI beyond official evidence, that talent signals prove alpha, or that a source-discovery/browser model can find and archive new pages |
finance-vendor-product-signal-v0, 12 tasks, answer scorer, durable static-control summary, scorecard, endpoint runner path, and frontier/API control preflight |
The repo has an executable policy for promoting finance AI product architecture while rejecting generic marketing, unverified metrics, podcast anecdotes, terminal-replacement claims, and alpha/ROI overclaims; the static-control row records 12/12, the runner can score a preflight-approved endpoint, and frontier-api-vendor-product-signal-control is now a concrete pending candidate |
That any vendor product is recommended, independently measured, proven to improve returns/productivity, or that the frontier/API candidate has cost, route, endpoint, or scored-run approval |
finance-treasury-payments-liquidity-v0, 12 tasks, answer scorer, durable static-control summary, and scorecard |
The repo has an executable policy for treasury, payments, liquidity, reconciliation, settlement-finality, controlled agentic money-movement, kill-switch, audit-trail, and runtime-portability claims; the static-control row records 12/12 |
That any model can move money, that a payment rail is endorsed, that settlement is legally final, that sanctions/fraud controls are sufficient, or that a model/backend has been approved for this suite |
finance-kaggle-artifact-v0, 10 tasks, answer scorer, durable static-control summary, and scorecard |
The repo has an executable policy for treating finance Kaggle/competition pages as benchmark-design evidence: official metadata, metric/split, leakage, baseline, notebook status, MLE-harness, community-artifact timing, git-boundary, and not-alpha checks; the static-control row records 10/10 |
That a leaderboard result proves finance alpha, production deployment, execution quality, autonomous investment authority, or model/backend competence |
finance-replay-v0, 118 materialized tasks, dry-run rows, 118/118 oracle expected-patch control, and scorecard plumbing |
The repo has a deterministic replay surface for source acquisition, wiki maintenance, overclaim repair, run metadata aggregation, and executable check ceiling | That any candidate model has passed it |
| Frontier/API cost-estimation preflight | statebench.runners.finance_replay_cost_estimate estimates the 118-task replay control at $7.51 before any paid API call, recorded in statebench/results/finance-replay-v0-cost-estimate-frontier-api-20260601-expanded118.json |
That spending is approved, that actual tokenization/output/retries will match the estimate, or that the frontier model will pass |
finance-regional-language-v0, 12 tasks, answer scorer, durable static-control summary, and scorecard |
The repo has an executable policy for regional finance-language specialists: A-share/CFA, Turkish/BIST, Thai/English, India budget-policy, regulated-advice refusal, translation integrity, benchmark-only overclaim, backend portability, exact-revision gates, and unverified social/vendor claims; the static-control row records 12/12 |
That any regional SLM transfers to U.S. finance, SEC filing analysis, accounting compliance, institutional quant research, investment advice, live trading, or alpha generation |
finance-time-series-market-v0, 14 tasks, answer scorer, durable static-control summary, and scorecard |
The repo has an executable policy for financial time-series and market foundation models: source discovery, runnable-artifact separation, point-in-time leakage controls, forecast metrics, related/unrelated multivariate context, trading utility, transaction costs, capacity, and runtime portability; the static-control row records 14/14 |
That Kronos, FinCast, TimesFM, Chronos, Moirai, Lag-Llama, LOBERT, ByteGen, LiT, or any market model has forecasting skill, trading utility, live alpha, or backend portability |
| BankerToolBench deal-work source fixture | sa-086 freezes the investment-banking/deal-work source ledger around BankerToolBench: data rooms, market data, SEC filings, Excel models, PowerPoint decks, PDF/Word reports, multi-file deliverables, banker rubrics, client-readiness caveats, and cross-artifact consistency |
That buy-side quant alpha, live trading, or production client readiness is proven |
finance-quant-code-v0, 11 tasks, answer scorer, execution smoke, sample-pass artifacts, scorecard, and candidate-manifest target coverage |
The repo has a seed harness for QFBench/QuantEval-style executable quant-code discipline: source-boundary, Docker/numerical-verifier contract, memo-overclaim rejection, production-boundary, baseline comparison, task scope, QuantEval bridge, runtime trace, option payoff, drawdown/turnover, and factor-rank verifier wiring. Static controls record 8/8 answer-artifact and 3/3 local execution-smoke rows. |
That QFBench itself has been run, that any model can generate correct quant code, or that a benchmark pass proves live alpha |
finance-agentic-factor-discovery-v0, 12 tasks, answer scorer, durable static-control summary, and scorecard |
The repo has an executable policy for source-to-factor conversion, DSL/AST safety, unsafe program rejection, point-in-time evaluation, factor diagnostics, implementation realism, diversity/crowding, neutralization/attribution, walk-forward promote/hold/retire decisions, narrative-bias rejection, and runtime portability; the static-control row records 12/12 |
That any factor has live alpha, that any model can generate valid factor code, that a backtest is production-ready, or that a model/backend is approved for this suite |
finance-portfolio-risk-construction-v0, 12 tasks, answer scorer, durable static-control summary, and scorecard |
The repo has an executable policy for portfolio screen audits, covariance experiment design, graph construction critique, constrained weight construction, action-first RL policy, risk-aware agent controls, inverse-RL limits, backtest skepticism, and runtime portability; the static-control row records 12/12 |
That any portfolio is recommended, that an optimizer/RL policy is production-ready, that backtest metrics prove allocation quality, or that a model/backend is approved for this suite |
| U.S. Senate hedge-fund AI/ML report source card and OCR | The repo has a government/regulatory source for inconsistent AI/ML definitions, human-review ambiguity, vague client disclosures, testing/review process gaps, overfitting/backtesting limits, and audit-trail/version-control recommendations across Citadel, Renaissance Technologies, Bridgewater, AI Capital Management, Numerai, and WorldQuant evidence | That any named fund has proven alpha, that current 2026 deployment status is known, or that regulatory recommendations are implemented |
finance-retrieval-sql-harness-v0, 46 tasks, answer scorer, static-control summary, local Postgres preflight artifact, and local Snowflake semantic-contract artifact |
The repo has a seed harness for finance retrieval, NL2SQL, Postgres, and Snowflake-style semantic contracts; the answer-artifact scorecard path works end-to-end; local Postgres readiness is blocked specifically on missing Python driver; the Snowflake semantic-view YAML aligns with the fixture schema/corpus/controls | That a real model/backend, Postgres/Snowflake backend, pgvector/Cortex route, warehouse execution trace, cost cap, or production RBAC layer has been scored |
finance-sentiment-routing-v0, 15 tasks, answer scorer, static-control summary, local MPS classifier rows, parser-contract row, and measured-runtime ledger |
The repo has a held-out synthetic scorer for headline/earnings/commentary sentiment, routing, abstention, and no-alpha overclaim checks; ModernFinBERT, ProsusAI FinBERT, and YorkFr have real local MPS label-only rows; ModernFinBERT leads label-only at 11/15; all real local sentiment models score 0/15 on full routing; a Qwen3 GRPO parser-contract row proves output parsing can score 15/15 when the answer contract is satisfied |
That any sentiment classifier is ready for full workflow routing; parser-contract success is not model-generation proof; label-only rows do not prove source acquisition, SQL, accounting, portfolio, or alpha competence |
| Model candidate manifest, bakeoff queue, preflight schema, accounting/compliance fixture, advisor/portfolio fixture, regional-language fixture, time-series seed fixture, quant-code fixture, and talent/social leakage seed fixture | Candidate families, runtimes, backend constraints, run phases, hardware-specific gates, and first accounting/compliance/advisor/regional/time-series/quant-code/talent-leakage guardrail tasks are recorded, including newly promoted Fin-o1 Qwen3 reasoning, CPA-Qwen3 accounting, Kuvera personal-finance, FinGuard compliance-guard, Amsi-fin MLX VLM, Fara-7B source-discovery, Kronos market-model candidates, and a self-driving-portfolio controlled-delegation task for IPS/script/peer-review/human-approval boundaries. The advisor/portfolio static control records 13/13, regional-language records 12/12, time-series/market records 14/14, and quant-code records 8/8 answer-artifact plus 3/3 execution-smoke controls. |
That Qwen, Gemma, Liquid, Fara, Kronos, Kuvera, finance specialists, MLX, GGUF, CUDA, AMD, API, or Neuron are approved for a given task |
What We Are Still Missing
The remaining gap is measured execution, not source awareness. The corpus now has enough finance/quant industry, model, benchmark, podcast, report, and official-practitioner evidence to start answering the strategic question. What is still missing is proof that specific models and backends can do the work better, cheaper, or more reliably.
The highest-value unfinished items are:
- First scored model
finance-replay-v0run. Dry-run rows now prove manifest and scorecard aggregation, and the full oracle expected-patch control now proves118/118suite/check executability across source acquisition, wiki maintenance, and reflect-audit tasks. The current suite addsra-123for Balyasny multi-strat factor-agent caveats andsa-083for QFBench executable quant-code benchmark evidence. The refreshed frontier/API cost estimate is$7.512744; next run a frontier API control after explicitcost_cap, or one current open/local candidate after endpoint smoke. A freshqwen3-6-35b-a3b-hflocalhost smoke artifact showslocalhost:8000is not currently serving the model, so the Qwen route is still blocked onbackend_memory_check. Summarize with the scorecard runner as non-oracle model/backend rows. - First approved model preflight. Fill one preflight record for a candidate
that can actually run in the current environment; do not approve Neuron
until there is real Inferentia/Trainium compile evidence.
The current readiness note ranks the first viable paths: frontier API is
blocked only on
cost_cap;cpa-qwen3-8b-v0has all source checks passed but needs an approved endpoint; Kuvera now has a preflight-gated advisor/portfolio endpoint runner path but still needs runtime smoke; Qwen3.6/Gemma need endpoint-memory smoke; regional SLMs now have a preflight-gated endpoint runner path but still need endpoint smoke and preflight approval. - First quant-code model/backend row. The new
finance-quant-code-v0seed fixture records the QFBench/QuantEval execution discipline, scores a static answer artifact8/8, and runs a tiny deterministic execution smoke3/3. The missing layer is a real model/backend row, then a larger code-execution harness with realistic frozen inputs, task-local data, verifier, oracle/reference solution, repeated runs, one-shot baseline, tool traces, cost, latency, and failure-mode capture. - First retrieval/SQL backend score. The static answer-artifact control now
records
46/46instatebench/results/finance-retrieval-sql-v0/sample-pass-artifact.summary.json. The missing layer is a preflight-approved model/backend row, followed by Postgres and Snowflake proof. Local Postgres preflight is captured and currently fails only on missing driver availability (psycopg,psycopg2, orpg8000), so the next environment step is installing a driver and pointing the harness at a disposable database. The Snowflake semantic contract now has a passing local fixture-alignment artifact, but the missing Snowflake layer is still real Cortex/warehouse execution with RBAC, traces, and cost metadata. - Sentiment/routing next candidate. The YorkFr, ModernFinBERT, and
ProsusAI FinBERT local MPS rows exist. ModernFinBERT leads label-only
acceptance at
11/15, while YorkFr and ProsusAI accept10/15; all score0/15accepted on full workflow routing. Theqwen3-grpo-parser-contractrow scores15/15as a parser/output-contract smoke, not as a model-quality result. Next compare FinSenti after the Qwen3.5/FinSenti adapter or GGUF backend is proven. - Official social hardening. Convert official LinkedIn/company-social
discoveries into source cards only when the account, URL, date, and
archive/screenshot or primary-source backing are preserved. The
finance-talent-leakage-v0static-control scorecard now records10/10for the promotion-policy fixture, andsa-073makes this a first-class source-discovery family rather than a buried subcase of firm/careers pages. The remaining work is scoring real browser/source-discovery agents and adding new frozen tasks only when official evidence packs are captured. 6a. Vendor product-signal scoring. The newfinance-vendor-product-signal-v0fixture covers QuantConnect-style research pipelines, assistant teams, incumbent terminal AI, governed connector ecosystems, AI-native workflow platforms, named-customer quote caveats, podcast source-discovery routing, terminal-replacement overclaims, vendor metrics, MCP/tool surfaces, and wiki promotion policy.sa-082now makes this a materialized source-acquisition fixture family infinance-source-discovery-eval-v0. The remaining work is scoring real models on it and only expanding it when a new vendor pattern appears. The runner now has a preflight-gated endpoint path, and the dedicatedfrontier-api-vendor-product-signal-controlpreflight exists, but this lane still needs cost estimate/cap and endpoint/API-route evidence before a non-static model/backend row can be produced. 6b. Agentic factor-discovery scoring. The newfinance-agentic-factor-discovery-v0fixture covers Hubble/FactorEngine style source-to-factor conversion, safe DSL/AST execution, point-in-time validation, diagnostics, turnover/cost/capacity controls, diversity/crowding, neutralization, walk-forward promote/hold/retire decisions, narrative-bias rejection, and runtime portability. The remaining work is a preflight-approved model row plus a deterministic toy factor-code harness with point-in-time data, reference results, cost/capacity checks, and failure traces. 6c. Portfolio/risk-construction scoring. The newfinance-portfolio-risk-construction-v0fixture covers screened-universe audits, covariance/risk experiment design, graph-construction critique, constrained weights, action-first RL policy, risk-aware multi-agent controls, inverse-RL evidence limits, backtest skepticism, and runtime portability. The remaining work is a preflight-approved model row plus deterministic covariance/weighting and graph-construction harnesses with point-in-time data, reference outputs, turnover/cost/capacity checks, and survivorship-bias controls. - Negative-evidence task expansion. The first expansion is now
materialized in
ra-119throughra-123: Acadian process overclaim, Two Sigma temporal leakage, Winton selection bias, cross-firm autonomous agent-alpha collapse, and Balyasny factor-agent workflow caveats.sa-085adds the Senate/HSGAC regulatory source-acquisition counterpart for system-definition consistency, human-review boundary, test cadence, version control, disclosure adequacy, and audit-trail preservation. The remaining work is scoring real candidates and adding more tasks only when new practitioner or regulatory evidence introduces a distinct failure mode. - Domain SLM bakeoff. Score at least one current Qwen-family artifact, one
Gemma-family artifact, one Liquid task specialist, one finance reasoning
specialist, one accounting/compliance specialist on
finance-accounting-compliance-v0, one advisor/personal-finance specialist onfinance-advisor-portfolio-v0, one treasury/payments/liquidity row onfinance-treasury-payments-liquidity-v0, one agentic factor-discovery row onfinance-agentic-factor-discovery-v0, one portfolio/risk-construction row onfinance-portfolio-risk-construction-v0, one multimodal finance VLM, one browser/source-discovery agent such as Fara-7B or Fara1.5, one regional/language finance SLM where a matching fixture exists, one local GGUF finance-reasoning artifact such as Qwen3-8B finance GGUF, one retrieval/reranking specialist or tiny embedding artifact on the appropriate suite, and one finance-native time-series market model such as Kronos only afterfinance-time-series-market-v0has locked data, leakage controls, baselines, transaction-cost/capacity assumptions, and runtime smoke proof. - Podcast and article quality triage. Keep adding only named-practitioner, source-backed media to findings; vendor/product podcasts and social posts stay as discovery unless they expose architecture, process controls, failure modes, or named customer evidence.
- Kaggle and competition harness extraction. Preserve competition metrics,
split rules, leakage controls, kernels/notebooks, and local graders as eval
design evidence; do not treat competition rankings as finance-alpha proof.
The
finance-kaggle-artifact-v0static-control scorecard now records10/10for official metadata, metric/split, leakage, baseline, notebook, MLE-harness, community-timing, git-boundary, and not-alpha checks. The remaining work is scoring real browser/source-discovery agents and building larger executable competition-style harnesses only when frozen task-local data, reference solutions, leakage controls, runtime traces, and git exclusions are present. - Treasury/payments execution harness. The
finance-treasury-payments-liquidity-v0static-control suite now covers the regulated money-movement policy boundary. The remaining work is a preflight-approved model row plus synthetic mandate/invoice/statement/ledger fixtures and deterministic tool traces for authorization, sanctions/fraud, settlement, reconciliation, kill-switch, and audit controls. - Investment-banking/deal-work execution harness. BankerToolBench is now
frozen as source-acquisition fixture
sa-086, but the repo still needs a deterministic StateBench task family for senior-banker request parsing, data-room provenance, market/SEC research, Excel formula integrity, deck/memo generation, cross-artifact reconciliation, confidentiality/MNPI, and human-review handoff.
Current Interpretation
What top finance and quant shops appear to be doing publicly:
- using AI to accelerate research, coding, document analysis, feature discovery, data exploration, and analyst/PM workflows;
- building governed internal platforms around data lineage, entitlement checks, notebooks, backtests, logs, and evaluation gates;
- using ML at industrial scale for market prediction, execution, risk, storage, and low-latency deployment;
- treating broad chat assistants as shallow unless embedded in firm-specific tools, permissions, and workflows.
What is not proven or not working reliably:
- public sources do not independently prove genAI-driven live alpha;
- general agents can widen the hypothesis funnel and worsen false discovery;
- LLMs introduce historical knowledge-cutoff and temporal-leakage risks;
- finance artifacts can look plausible while formulas, state, citations, or execution traces are wrong;
- vendor product pages are architecture/product signals unless named customers or reproducible measurements are present.
Completion Standard
The objective should not be marked complete until current repo evidence proves:
- promoted source batches have raw source records or explicit reject/watchlist decisions;
- source-discovery categories have frozen StateBench fixture coverage;
- domain-model candidates have runnable manifest entries and preflight gates;
- at least one scored model/runtime run exists for
finance-replay-v0; - at least one scored retrieval/SQL answer-artifact run exists, and at least one model/backend retrieval/SQL row exists with backend metadata and preflight validation;
- docs and wiki pages distinguish workflow evidence from alpha/performance evidence;
- validation commands for the touched runners, manifests, and fixtures pass.
Until those scorecards exist, the correct status is: research map and ingestion control plane are strong; model/runtime recommendations remain unmeasured.
Next Verification Commands
Use these checks before claiming another layer is complete:
python3 -m statebench.runners.source_discovery_eval \
--queue statebench/suites/finance-source-discovery-eval-v0.json \
--validate-only
python3 -m statebench.runners.source_discovery_eval \
--queue statebench/suites/finance-source-discovery-eval-v0.json \
--section fixture-coverage \
--format markdown
python3 -m statebench.runners.finance_retrieval_sql_fixture \
--fixture-dir statebench/fixtures/finance-retrieval-sql/v0 \
--answers-file statebench/fixtures/finance-retrieval-sql/v0/sample_answers/pass.json \
--format markdown
python3 -m statebench.runners.finance_retrieval_sql_fixture \
--fixture-dir statebench/fixtures/finance-retrieval-sql/v0 \
--snowflake-contract-check \
--format markdown
python3 -m statebench.runners.model_candidate_manifest \
--manifest statebench/suites/finance-model-runtime-candidates-v0.json \
--format markdown
python3 -m statebench.runners.suite_runner \
--suite-id finance-replay-v0-oracle-full-20260601-expanded118 \
--suite-file statebench/suites/finance-replay-v0.json \
--backend frontier-api \
--model oracle-expected-patch \
--policy overwrite \
--oracle-expected-patch