See also: finance-quant-ai-objective-coverage-audit-2026.md · finance-quant-ai-coverage-audit-2026.md · buy-side-quant-ai-practitioner-signals-2026.md · …/13-multimodal-sources/finance-quant-audio-video-deep-dive-2026.md
Source ledger: sources/06-industry-verticals/finance-quant-source-ingestion-queue-2026-raw.md
This queue converts the remaining finance/quant source gaps into ingestion work. It is deliberately stricter than a watchlist: an item belongs here only when it can be captured through an existing repo pipeline or converted into a StateBench source-discovery fixture.
Promotion Rules
| Source Type | Promote To Findings? | Required Evidence |
|---|---|---|
| arXiv paper / academic PDF | Yes | URL, title, date, PDF in research/*/raw, OCR output, source card |
| Official firm/report PDF | Yes | URL, publisher, date, PDF/OCR, vendor/conflict caveat |
| Named practitioner podcast/video | Yes, selectively | Episode URL, publication date, guest, transcript or local ASR, claim caveat |
| Official firm page / careers page | Yes | URL, date retrieved, firm owner, exact workflow/control claim |
| Official LinkedIn/company social post | Conditional | URL, account, date, author/account, screenshot/export where possible, primary-source backing |
| Third-party LinkedIn/commentary | No | Discovery only unless it points to a primary source |
| Kaggle competition/notebook/discussion | Conditional | Official competition metadata first; notebooks/discussions as practitioner signals only |
| Vendor product marketing | Conditional | Named customer, architecture, metric, or failure mode; otherwise product-signal only |
Priority Queue
| Priority | Item | Why It Matters | Pipeline | Promotion Gate |
|---|---|---|---|---|
| 1 | Frontier/API and local model score runs over finance-replay-v0 |
Coverage audit says the main remaining gap is measured execution, not source awareness | python3 -m statebench.runners.suite_runner --suite-file statebench/suites/finance-replay-v0.json ... then python3 -m statebench.runners.scorecard ... |
Manifest has artifact/backend metadata; pass/fail is from task checks, not dry-run |
| 2 | Official top-fund social and careers archival | Official LinkedIn/job pages leak platform, workflow, governance, and skill demand, but current capture is brittle | Manual/browser capture to source ledger; promote only stable official pages into research notes | URL/date/account preserved; no third-party inference |
| 3 | Negative/skeptical quant-shop material | StateBench should reward overclaim resistance and leakage/overfitting skepticism | Official pages/PDFs into research/06-industry-verticals/raw; podcast transcripts through pillar 13 |
Names concrete failure mode: overfitting, leakage, benchmark mismatch, interpretability, capacity, or governance |
| 4 | Risk.net Quantcast paper-linked episodes | Strong bridge between quant papers and practitioner vocabulary | scripts/podcast_mine.py --show risk-net-quantcast --episode <episode> --archive-audio |
Transcript or linked paper archived; claims tied to paper/source |
| 5 | Acadian/Man/Two Sigma/AQR/Balyasny adjacent appearances | systematic-manager direct workflow leakage is the highest-value audio/video category; Bloomberg Odd Lots now adds a verified Balyasny/multi-strat RSS lane | scripts/podcast_mine.py targeted episode runs; official transcript/PDF preferred where available |
Named practitioner, workflow boundary, no alpha overclaim |
| 6 | Kaggle finance competition artifacts | Underexplored medium for leakage controls, metrics, splits, local graders, and notebook archetypes; first stable PDFs now OCR’d in pillar 21 | Create source cards; archive official competition pages and high-signal public notebooks as discovery metadata; store only small PDFs/metadata in git | Official competition metadata promoted; notebooks/discussions remain practitioner signal |
| 7 | Vendor/customer conference media with finance guests | Useful only when a named customer gives architecture or failure-mode detail | scripts/podcast_mine.py for YouTube/session audio, or manual source card for official session pages |
Named customer, concrete workflow, not only product demo |
| 8 | Official PDFs and reports discovered by web/source eval | Stable target for source-acquisition tasks | Download to research/<pillar>/raw, then scripts/.venv-tts/bin/python3 scripts/ingest_pdfs.py --pillar <NN> |
OCR complete, metadata sidecar/source ledger created |
| 9 | Podcast/RSS feeds with durable transcripts | Source-discovery and vocabulary layer | Add to scripts/podcast_sources.yaml, then targeted scripts/podcast_mine.py --show <slug> --limit <n> --archive-audio |
Feed verified; audio archived or transcript preserved; weak episodes rejected |
| 10 | Official data/platform connector docs | Needed for finance retrieval and SQL harnesses | Source card plus research note under finance/retrieval pillars | Documents actual connector, permissions, audit, schema, or cost boundary |
| 11 | SQL benchmark-quality and module-diagnostic papers | Public SQL leaderboard scores can be unstable when annotations or gold SQL are wrong | Download PDF to research/21-benchmarks/raw, OCR with Chandra, add source card and harness note |
Captures annotation-quality caveat, module boundary, cost/latency metric, dialect-portability failure, and internal-gold-query implication |
| 12 | MultiFinBen-style multilingual/multimodal finance sources | Finance eval coverage now needs text, scanned-document OCR, and audio treated as distinct modalities | Ingest PDFs through pillar 21; use podcast/audio pipeline only for raw episode media; add source-acquisition fixture | Publish date and retrieval date recorded; model rankings treated as dated; task/error taxonomy promoted |
Recently Captured Benchmark Sources
| Source | Date | Captured Artifacts | Promotion |
|---|---|---|---|
| BizFinBench | 2025-05-26 | sa-091, arXiv PDF metadata, Chandra OCR, source card |
Chinese business-finance task taxonomy and IteraJudge evidence |
| XFinBench | July 2025 | sa-092, ACL PDF metadata, Chandra OCR, source card |
Complex finance problem-solving, temporal/scenario reasoning, and visual-context error evidence |
| MultiFinBen | 2025-06-16; v3 2025-10-11 | sa-093, arXiv PDF metadata, Chandra OCR, source card |
Multilingual text, financial OCR, and financial audio benchmark-design evidence |
| AIMA The Long-Short / CFM quant episode | 2026-03-25 page; June 2025 conversation | sa-110, official AIMA page, Apple-verified Acast RSS feed, aima-long-short podcast registry entry, local audio sidecar, WhisperX-MLX transcript, promoted pillar 13 note |
Transcript-backed systematic-investing practitioner vocabulary; not AI-agent deployment or performance evidence |
| Voleon ML-first investment process | No visible firm-page publication date; official pages retrieved 2026-06-05 | sa-115, official firm page, official careers/job pages, standalone source card |
ML-first systematic-investment process, financial prediction, trading-operations supervision, execution/data-pipeline constraints; not LLM-agent autonomy or alpha proof |
| D. E. Shaw machine-teaching optimizer evidence | 2023 article; investment-management page retrieved 2026-06-05 with March 1, 2026 firm context | sa-117, official Machine Teaching article, official investment-management page, standalone source card |
Human-machine optimizer workflow evidence: explicit inputs, constraints, sensitivity, transaction costs, correlation/tail/common-investor risk, and human intervention; not GenAI-agent or alpha proof |
| Point72 / Cubist systematic and Market Intelligence workflow evidence | Cubist statistics as of 2026-01-01; Quant Academy article 2025-01-06; official pages retrieved 2026-06-05 | sa-118, official Cubist, Market Intelligence, Investment Services, and Cubist Quant Academy pages, standalone source card |
Systematic ML return-prediction workflow, compliant alternative-data research products, risk/portfolio/execution support, and full trade lifecycle; not LLM-agent, autonomous-trading, or alpha proof |
| Citadel EQR and Citadel Securities systematic ML pipeline evidence | EQR articles published 2026-02-20 and 2026-04-28; official pages retrieved 2026-06-05 | sa-119, official Citadel quantitative research, EQR, and Citadel Securities quantitative research pages, standalone source card |
Full-pipeline systematic research evidence: alternative data, AI/ML, forecasting, portfolio construction, optimization, market impact, execution, simulation, and feedback; not LLM-agent, autonomous-trading, or alpha proof |
| AQR Can Machines Learn Finance skeptical ML evidence | AQR page published 2019-06-07; PDF/OCR retrieved 2026-05-29; official pages rechecked 2026-06-05 | sa-120, official AQR ML page, official journal page, PDF metadata, Chandra OCR, standalone source card |
Small-data / low-signal-to-noise return-prediction control, overfit/out-of-sample/economic-theory/human-expertise guardrails; not GenAI-agent, current model benchmark, alpha proof, or runtime/backend proof |
| G-Research production LLM code-review pattern | Official article published 2026-05-13; retrieved 2026-06-06 | sa-130, official G-Research engineering article, standalone source card |
Quant-firm production LLM control pattern: untrusted-model boundary, source-of-truth rules index, Pydantic/JSON structured output limits, two-pass recall/precision evaluation, bounded repair, provider abstraction, cost telemetry, and non-blocking PR comments; not alpha proof or runtime/backend proof |
| BigFinanceBench workflow-grounded financial-research benchmark | arXiv published 2026-06-02; retrieved 2026-06-06 | sa-127, arXiv PDF metadata, Chandra OCR, arXiv API metadata, standalone source card |
928 auditable financial-research tasks, derivation-level rubrics, source/period/accounting-definition/calculation trace scoring; public subset/harness access and paper model scores require local reproduction |
| Hedge-Bench hedge-fund analyst reasoning benchmark | arXiv posted 2026-06-02; retrieved 2026-06-06 | sa-126, arXiv PDF metadata, Chandra OCR, GitHub repository metadata, standalone source card |
Open-ended hedge-fund analyst reasoning, expert-move coverage, source-file citations, counter-evidence, ambiguity, and hallucination checks; public model scores require local reproduction and LLM-judge caveat |
Pipeline Commands
PDF/report ingestion:
scripts/.venv-tts/bin/python3 scripts/ingest_pdfs.py --pillar 06
scripts/.venv-tts/bin/python3 scripts/ingest_pdfs.py --pillar 21
Raw PDF files captured under research/**/raw are local source material and
are ignored for new ingests. Commit the OCR output (*.chandra.txt or chapter
chunks), *.pdf.meta.md sidecars, source-card markdown, and any StateBench
fixture entries that prove the source was captured. Some older PDFs are already
tracked in git; do not use that precedent for new finance/quant captures.
Podcast/audio ingestion:
scripts/.venv-tts/bin/python3 scripts/podcast_mine.py \
--show risk-net-quantcast \
--limit 3 \
--archive-audio
Archived podcast and conference audio/video stays local under
sources/13-multimodal-sources/. Commit only the .meta.md sidecar, ASR
transcript in research/13-multimodal-sources/**/raw, promoted episode note,
and source-discovery or benchmark metadata.
Finance replay model scoring:
python3 -m statebench.runners.suite_runner \
--suite-file statebench/suites/finance-replay-v0.json \
--suite-id finance-replay-v0-<artifact>-<backend>-YYYYMMDD \
--backend <frontier-api-or-openai-compatible-url> \
--model <model-id> \
--policy resume \
--artifact-id <exact-artifact-id> \
--artifact-format <api|hf|gguf|mlx|neuron_compiled|awq|gptq|fp8> \
--artifact-precision <api|bf16|fp16|q4_k_m|fp8|...> \
--weights-revision <revision-or-snapshot> \
--backend-hardware <api|apple_m3_max|cuda_h100|aws_inf2|...> \
--backend-constraint "record backend/operator/context/license limits" \
--context-length <effective-context-tokens> \
--preflight-record statebench/preflights/<artifact>-preflight.json
Scorecard summary:
python3 -m statebench.runners.scorecard \
--results-root statebench/results \
--suite-prefix finance-replay-v0 \
--format markdown
What To Capture In Each Source Card
Every promoted source should preserve:
- URL and retrieval date;
- publication date;
- publisher/owner;
- named speaker/author where applicable;
- source type: paper, PDF, podcast, official page, job listing, social post, conference session, competition, repository, or model card;
- claim type: workflow, model, benchmark, governance, data connector, execution, alpha/performance, hiring/talent, or failure mode;
- credibility tier and conflict caveat;
- ingestion path: PDF OCR, podcast ASR, source card, StateBench fixture, or source-discovery-only;
- promotion decision: finding, source-discovery, watchlist, or reject.
Stable Source-Discovery Fixture Ideas
These should become frozen source-acquisition tasks rather than live-web leaderboard runs:
- Find an arXiv finance benchmark paper and record its PDF, title, authors, date, and benchmark purpose.
- Find a named quant podcast episode, archive the audio/transcript, and extract only workflow claims tied to a guest.
- Find an official hedge-fund AI/careers page and classify whether it proves workflow, tooling, hiring demand, governance, or only marketing.
- Find a Kaggle finance competition and extract metric, public/private split, leakage controls, and official data schema.
- Find a small stable Kaggle-related PDF or paper, download it to
research/21-benchmarks/raw, OCR it withscripts/ingest_pdfs.py --pillar 21, and record why the source is eval-design evidence rather than alpha evidence. - Find a vendor finance-agent product page and separate product claims from named-customer architecture evidence.
- Find a Snowflake/Postgres/NL2SQL source and extract dialect, schema, execution, cost, and permission constraints.
- Find an NL2SQL benchmark-quality paper and record whether it changes the StateBench SQL harness, the public-leaderboard caveat, or the internal gold query audit process.
The vendor/product promotion boundary is now executable through
finance-vendor-product-signal-v0.
Use it when a source-discovery model tries to promote QuantConnect-style
assistant teams, finance-agent connector pages, incumbent terminal AI, AI-native
workflow vendors, named-customer quotes, podcast mentions, or terminal
replacement claims.
The first live-source eval queue is now tracked in
statebench/suites/finance-source-discovery-eval-v0.md.
It records a Better System Trader RSS dry run and marks the discovered episodes
as source-discovery until transcript review proves they contain useful
workflow or failure-mode evidence.
Validate that queue before promoting any live discovery batch:
python3 -m statebench.runners.source_discovery_eval \
--queue statebench/suites/finance-source-discovery-eval-v0.json \
--format markdown
Generate frozen source-acquisition task templates from the queue with:
python3 -m statebench.runners.source_discovery_eval \
--queue statebench/suites/finance-source-discovery-eval-v0.json \
--section fixture-templates \
--format markdown
Check which template families already have source-acquisition task examples:
python3 -m statebench.runners.source_discovery_eval \
--queue statebench/suites/finance-source-discovery-eval-v0.json \
--section fixture-coverage \
--format markdown
The eval now tracks twelve fixture families, matching the current backlog: podcast RSS episode discovery, official firm/careers page discovery, official social-post discovery, negative-evidence discovery, paper/PDF discovery, Kaggle competition discovery, SQL/retrieval source discovery, domain model/runtime discovery, vendor/customer architecture discovery, vendor-product signal classification, finance workflow vendor platform discovery, and SQL benchmark-quality / dialect-portability discovery. This closes the earlier gap where LinkedIn/social evidence, negative-evidence, model/runtime, vendor/product, and SQL benchmark-quality batches were listed in the backlog but were not first-class frozen-fixture templates.
Newly Captured Gap Items
-
BizFinBench business-driven financial benchmark - arXiv
2505.19457, submitted 2025-05-26 and retrieved 2026-06-05. Captured through the pillar 21 PDF/OCR pipeline asresearch/21-benchmarks/raw/bizfinbench-business-driven-financial-benchmark-2505.19457.chandra.txtand promoted throughsources/21-benchmarks/bizfinbench-business-driven-financial-benchmark-raw.md. Promotion: multilingual/business-finance benchmark-design evidence, not current model-ranking proof. It strengthens the regional-language, numerical-calculation, financial-time-reasoning, event-attribution, tool-usage, and LLM-as-judge lanes by requiring task-level decomposition, Chinese/local-market scope labels, and reruns before trusting paper-snapshot model rankings. -
XFinBench complex financial problem solving - ACL Findings 2025, published July 2025 and retrieved 2026-06-05. Captured through the pillar 21 PDF/OCR pipeline as
research/21-benchmarks/raw/xfinbench-complex-financial-problem-solving-2025.findings-acl.457.chandra.txtand promoted throughsources/21-benchmarks/xfinbench-complex-financial-problem-solving-raw.md. -
AIMA The Long-Short / CFM quant episode - official AIMA episode page published 2026-03-25 and retrieved 2026-06-05. Captured as
sources/06-industry-verticals/aima-long-short-cfm-quant-2026-raw.mdwith Apple lookup verification of the Acast RSS feed and a newscripts/podcast_sources.yamlshow entryaima-long-short. Audio was archived locally with a.meta.mdsidecar, raw transcript was written toresearch/13-multimodal-sources/aima-long-short/raw/rss-8026b27a9d70.txt, and a promoted extract was added atresearch/13-multimodal-sources/aima-long-short/2026-03-25-philip-seager-cfm-systematic-multistrategy.md. Promotion: transcript-backed systematic-investing practitioner vocabulary. Do not use as current AI-agent deployment, alpha, allocator-suitability, or model/runtime evidence. Promotion: complex multimodal financial problem-solving benchmark-design evidence, not current model-ranking proof. It strengthens temporal reasoning, future forecasting, scenario planning, numerical modelling, visual chart/curve interpretation, knowledge-augmentation, and human-expert-baseline gates. -
LLM stock forecasting from a hedge-fund perspective - arXiv
2605.05211, submitted 2026-04-10 and retrieved 2026-06-02. Captured through the pillar 21 PDF/OCR pipeline asresearch/21-benchmarks/raw/llm-stock-forecasting-hedge-fund-perspective-2605.05211.chandra.txt. Promotion: benchmark-design evidence for text-alpha and trading-agent guardrails. It strengthens the negative-evidence lane by requiring full market-cycle evaluation, non-LLM baselines, trading-aligned metrics, leakage controls, and illiquidity/capacity checks before any LLM stock forecasting claim is treated as useful. -
Balyasny / OpenAI AI research engine - OpenAI customer case study, published 2026-03-06 and retrieved 2026-06-05. Captured through
sources/06-industry-verticals/balyasny-openai-ai-research-engine-raw.mdand materialized as source-acquisition fixturesa-098. Promotion: production-bar finance research-agent architecture evidence, not alpha proof or audited ROI. It strengthens the vendor-product and research-agent lanes by requiring pre-deployment model evaluation, scoped tools/data access, traceable reasoning paths, compliance guardrails, workflow-specific agents, and human decision-support boundaries before vendor speedup claims are trusted. -
Man Group / Anthropic / AlphaGPT partnership - official Man Group announcement and PDF mirror, published 2026-02-11 and retrieved 2026-06-05. Captured through
sources/06-industry-verticals/man-anthropic-alphagpt-partnership-raw.mdand materialized as source-acquisition fixturesa-099. Promotion: systematic-manager frontier-lab partnership and workflow-integration evidence, not investment-performance evidence. It strengthens the systematic quant practitioner lane by requiring agents to distinguish Claude / Claude Skills / Claude Code / AlphaGPT proprietary workflow signals from runnable local, MLX, llama.cpp, SageMaker, Bedrock, or Neuron artifacts. -
FinSheet-Bench financial spreadsheets - arXiv
2603.07316, published 2026-03-07 and retrieved 2026-05-31. Captured through the pillar 21 PDF/OCR pipeline and promoted throughsources/21-benchmarks/finsheet-bench-financial-spreadsheets-raw.md. Promotion: benchmark-design evidence for private-markets diligence, spreadsheet extraction, and deterministic-computation guardrails. It strengthens the negative-evidence lane by showing that standalone frontier models remain too error-prone for unsupervised professional spreadsheet use, especially on irregular layouts, fund dividers, multi-line headers, aggregation, sorting, and complex calculations. -
Agentic AI in Finance survey - arXiv
2604.21672, published 2026-04-23 and retrieved 2026-05-28. Captured through the pillar 06 PDF/OCR pipeline and promoted throughsources/06-industry-verticals/agentic-ai-finance-survey-2026-raw.md. Promotion: benchmark-design and governance evidence for multi-agent trading, portfolio, risk, compliance, market-structure, and systemic-risk tasks. It strengthens the negative-evidence lane by requiring system-level stability, liquidity, market-impact, shock-recovery, accountability, audit-trail, and continuous-validation checks before autonomous finance-agent claims are trusted. -
U.S. Senate HSGAC hedge-fund AI/ML report - June 2024 staff report, retrieved 2026-06-02. Captured through the pillar 06 PDF/OCR pipeline as
research/06-industry-verticals/raw/us-senate-hedge-funds-ai-ml-2024.chandra.txtand promoted throughsources/06-industry-verticals/us-senate-hedge-funds-ai-ml-2024-raw.md. Promotion: government/regulatory evidence for hedge-fund AI/ML governance, not alpha evidence. It is now frozen as source-acquisition fixturesa-085. It strengthens the negative-evidence lane by requiring consistent system definitions, human-review boundaries, testing/review cadence, client-disclosure adequacy, version control, audit trails, overfitting/backtesting skepticism, and systemic-risk checks for herding, manipulation, and market stability. -
SQL benchmark quality and dialect portability - NL2SQLBench / arXiv
2604.16493plus PARROT / arXiv2509.23338, promoted on 2026-06-02 throughsources/21-benchmarks/sql-benchmark-quality-dialect-portability-raw.mdand source-acquisition fixturesa-084. Promotion: retrieval/NL2SQL harness design evidence, not finance-alpha evidence. It strengthens the Snowflake/Postgres/custom-schema lane by requiring module-level schema-selection, SQL-generation, query-revision, dialect-portability, execution-result, cost/latency, permission, semantic-layer, provenance, abstention, and internal-gold-query checks instead of trusting generic BIRD/Spider scores. -
BankerToolBench investment-banking workflows - arXiv
2604.11304, retrieved 2026-05-29 and promoted throughsources/21-benchmarks/finance-investment-banking-deal-work-benchmarks-2026-raw.md. Promotion: benchmark-design evidence for deal-work and investment-banking agents, not buy-side alpha evidence. It is now frozen as source-acquisition fixturesa-086. It strengthens the multi-artifact workflow lane by requiring senior-banker request parsing, data-room provenance, market/SEC research, Excel/PPT/PDF/Word/CSV consistency, spreadsheet formula fidelity, confidentiality/compliance controls, and human-review handoff before any client-readiness claim. -
The Fund AI Pod funds-industry AI interviews - verified on 2026-06-05 through the official site, Apple Podcasts, and executable RSS feed
https://anchor.fm/s/106e86050/podcast/rss, promoted throughsources/13-multimodal-sources/the-fund-ai-pod-raw.mdand added toscripts/podcast_sources.yaml. Highest-value leads are Shu Bai on hedge-fund research and thesis challenge, Pat Starling on FactSet AI Foundry / APIs versus MCPs / auditability, Alex Benke on Ridgeline system-of-record agents, Alex Dunegan on FX/currency workflows, and Simmons & Simmons on governed legal AI. Archive audio/transcripts before promoting claims; sponsor, host, and LinkedIn metadata stay discovery-only. The Shu Bai episode was archived/transcribed and promoted on 2026-06-05 as practitioner workflow evidence for hedge-fund research/thesis-challenge patterns. The Pat Starling episode was archived/transcribed and promoted on 2026-06-05 as vendor/platform architecture evidence for FactSet AI Foundry, MCP/API/feed delivery, auditability, grounded answers, model routing, and buy-side analyst research workflows. Remaining episodes stay source-discovery until transcript review. -
scaledown.ai finance workflow platform signal - queued on 2026-06-05 as source-acquisition fixture
sa-087, then resolved on 2026-06-06 throughsources/06-industry-verticals/scaledown-ai-finance-workflow-platform-raw.md. Promotion decision: keep as finance-workflow source-discovery/watchlist, not a pillar 06 product finding. Official ScaleDown docs reviewed on 2026-06-06 prove context engineering, task-specific SLMs, compression, extraction, AST/BM25 pruning, semantic retrieval, and span offsets, but do not name a concrete finance workflow, data/source connector, investment artifact, finance governance control, auditability boundary, access-control boundary, or named finance customer architecture. Valid promotion is instead the technicalsa-123context-engineering source card. Do not promote third- party commentary, social mentions, unsourced logos, inferred deployment state, alpha, ROI, investment-performance, or autonomous capital-allocation claims. -
CurrentAI proactive investment-agent product signal - official page retrieved 2026-06-06 and promoted through
sources/06-industry-verticals/currentai-proactive-investment-agents-raw.md; materialized as source-acquisition fixturesa-122. Promotion: finance workflow-vendor evidence for investment-team background agents, financial model updates, scenario analysis, value-chain read-throughs, earnings previews, and management-meeting model-impact analysis. Evidence boundary: product-surface signal only. Do not use as alpha, autonomous-trading, customer-adoption, audited-productivity, model-superiority, or backend- portability evidence. -
ScaleDown context-engineering / task-specific SLM signal - official docs retrieved 2026-06-06 and promoted through
sources/21-benchmarks/scaledown-context-engineering-slm-raw.md; materialized as source-acquisition fixturesa-123. Promotion: SLM/context infrastructure evidence for task-specific context extraction, prompt compression, AST/BM25 code-context pruning, local-embedding/FAISS semantic retrieval, preservation controls for domain-specific terms, and offset/confidence-bearing extraction. Evidence boundary: not finance outcome evidence and not MLX/GGUF/CUDA/ROCm/SageMaker/Bedrock/Neuron compatibility evidence unless separate source-card proof exists. -
Hudson River Trading AI, benchmark, and data guardrail source - recorded on 2026-06-05 through
sources/06-industry-verticals/hrt-ai-benchmark-data-source-raw.mdand materialized as source-acquisition fixturesa-088. Promotion: first-party negative-evidence and benchmark-design source for trading-agent, quant-code, alternative-data, and benchmark-mismatch checks. HRT’s official articles preserve three durable guardrails: trading AI should be decomposed into prediction, optimization, execution, risk, and regulatory / market- participant constraints; ML papers should be judged for simplicity, reproducibility, and generality rather than leaderboard gains alone; and data sources should be tested for relevance, uniqueness, leakage/lookahead, sample size, noise, and provenance. The Bloomberg Odd Lots / Omny episode “How Hudson River Trading Actually Uses AI” was captured through the podcast pipeline on 2026-06-05 asresearch/13-multimodal-sources/odd-lots/raw/rss-0562c5190e8c.txtwith an excluded local MP3 and committed audio sidecar. It is promoted only as named practitioner workflow and guardrail evidence for short-horizon market-data ML, infrastructure, operational risk, regulatory controls, and LLM backtest contamination checks. -
Acadian 2026 Investor Forum systematic-manager scale source - recorded on 2026-06-05 through
sources/06-industry-verticals/acadian-2026-investor-forum-systematic-scale-raw.mdand materialized as source-acquisition fixturesa-089. Promotion: current official public systematic-manager context, not AI-agent or alpha proof. The page records $196B of client assets as of 2026-03-31, an advanced alpha engine analyzing 65,000+ securities daily with hundreds of unique signals, more than 100 investment experts, and a five-year benchmark-outperformance claim that is explicitly limited to eligible composite AUM. Its most useful benchmark lesson is the authority boundary: Acadian separates discretionary AUM from model advisory assets where it does not have trading authority. Agents should preserve those dates, caveats, and authority boundaries rather than flattening the page into a generic “AI alpha” claim. -
Trading-R1 / Alpha-R1 / Trade-R1 trading-reasoning RL cluster - captured on 2026-06-05 through pillar 21 PDF/OCR ingestion and promoted through
sources/21-benchmarks/trading-alpha-r1-reasoning-rl-raw.md; materialized as source-acquisition fixturesa-090. Promotion: systematic quant benchmark-design evidence for structured trading theses, context-aware alpha screening, stochastic market rewards, RAG-grounded process verification, and reward- hacking checks. The local PDFs remain ignored; the committed artifacts are Chandra OCR text, PDF metadata sidecars, the source card, and the frozen fixture. Do not promote reported return/drawdown results as live alpha or autonomous trading proof without code, data, transaction-cost, capacity, baseline, and cross-regime reproduction.
Current Interpretation
The next missing work is not more indiscriminate scraping. The corpus needs fewer but stronger captures:
- source cards for official firm/social/careers pages that currently exist only as brittle web signals;
- audio archives/transcripts for the highest-value quant practitioner episodes;
- OCR for newly discovered PDFs/reports;
- frozen source-discovery fixtures to test whether models can find stable papers, reports, podcasts, and official pages;
- scored StateBench runs proving which models and backends can actually do the ingestion, wiki-maintenance, and overclaim-repair work.
This keeps web discovery as an eval and prevents time-sensitive browsing from contaminating the deterministic finance replay benchmark.