← Findings 🕐 90 min read
Findings

Hedge Fund Research-Machine Leakage Tape

This tape tracks public, legally available clues about how quant and hedge-fund research machines work. It is not a portfolio, alpha, or investment recommendation ledger.


Generated: 2026-06-11

This tape tracks public, legally available clues about how quant and hedge-fund research machines work. It is not a portfolio, alpha, or investment recommendation ledger.

Ranked Signals

Fund Date Leaked clue Implied data/source surface Control Confidence Source
ScaleDown Finance Workflow Promotion Rejection 2026-06-06 ScaleDown is useful negative evidence for the ingestion loop: official docs show context compression, extraction, classification, retrieval, code-context pruning, span offsets, and confidence scores, but no named finance workflow, connector, portfolio artifact, finance customer, governance boundary, or investment-research deployment. generic context documents; compressed prompts; span offsets; confidence scores; local embeddings; FAISS indexes; AST/BM25 code context finance-specific source requirement; official customer/workflow evidence requirement; connector or entitlement proof requirement; artifact/source preservation; backend-specific runtime caveat; no-alpha/no-finance-deployment caveat HIGH sources/06-industry-verticals/scaledown-ai-finance-workflow-platform-raw.md
Voleon 2026-06-05 Voleon exposes an ML-first investment-process boundary: flexible statistical models for financial prediction, research-to-production handoff, trading-operations supervision, production trading systems, execution engineering, and data pipelines. financial prediction datasets; production trading data; research data pipelines; trading-operations feedback academic research discipline; scalability and risk-management caveats; trading-operations supervision; production engineering boundary HIGH sources/06-industry-verticals/voleon-ml-first-investment-process-raw.md
Point72 / Cubist 2026-06-05 Point72/Cubist exposes the compliant alternative-data research-product layer: data sourcing with investment teams, Compliance, and external partners; petabyte-scale AI/ML tooling; risk, portfolio construction, execution, and post-trade analysis. publicly available data; alternative datasets; petabytes of data; market data; post-trade analysis data Compliance review; data nuance and bias checks; real-world verification; permitted-use checks HIGH sources/06-industry-verticals/point72-cubist-market-intelligence-raw.md
Citadel Securities / Systematic Trading 2026-06-05 Citadel Securities leaks the market-maker research loop: observe patterns, form mathematically precise hypotheses, use proprietary tools and large-scale datasets to identify predictive signals, combine signals with market-impact models and risk factors, deploy automated strategies, read market feedback, and iterate. proprietary research tools; large-scale datasets; predictive market signals; market-impact models; risk factors; market feedback hypothesis precision; data-supported strategy gate; market-impact and risk-factor checks; simulation and compute boundary; no-LLM-agent caveat HIGH sources/06-industry-verticals/citadel-eqr-systematic-ml-pipeline-raw.md
Kaggle Quant Sponsors 2026-05-31 Finance Kaggle artifacts expose the public evaluation mechanics around quant work: hidden targets, public/private splits, time-aware validation, purged/embargoed splits, metric design, notebook timing, and local grader/submission artifacts. competition datasets; public/private leaderboards; rules pages; notebooks; discussion timestamps; local grader artifacts public/private split; time-aware validation; purging and embargoing; solution-leakage checks; not alpha proof caveat HIGH research/21-benchmarks/finance-kaggle-competition-artifacts-2026.md
Two Sigma / 2026 AI Outlook 2026-05-29 Two Sigma’s official 2026 outlook leaks the operating shift from idea generation to idea evaluation: company-aware tools, raw-data-to-feature compression, agentic AI under governance, and explicit concern that more hypotheses/backtests worsen overfitting and knowledge-cutoff leakage. raw research data; company context; feature stores; hypothesis logs; backtest artifacts; pretrained-model knowledge cutoffs idea-evaluation gates; timestamped data; overfitting resistance; leakage controls; human research discipline HIGH sources/06-industry-verticals/two-sigma-2026-ai-investment-management-outlook-raw.md
Jump Trading 2026-05-29 Jump describes custom foundation models and LLM agents integrated across tools and data, API/HPC serving, 50+ daily AI tools, 75%+ weekly firmwide LLM usage, and a sub-24-hour model-to-trader feedback loop. petabyte-scale data; simulation outputs; trading/research/core-infrastructure data; tool usage telemetry high-stakes adversarial-environment discipline; trader feedback loop; API/HPC serving boundary HIGH sources/06-industry-verticals/jump-trading-ai-ml-workflow-signal-raw.md
Jane Street Industrial ML Constraints 2026-05-29 Jane Street’s official pages leak the industrial ML constraint stack: noisy market data, ultra-low latency, regime shifts, nonstationarity, self-impact, distribution identification, microsecond inference, high-throughput market data, custom CUDA/hardware, determinism, and tail-event measurement. market data; returns and options data; market-relevant Twitter streams; trading system telemetry; high-throughput multicast messages; ML training data noise and nonstationarity caveat; self-impact control; low-latency deployment constraint; determinism check; tail-event measurement; no LLM-agent autonomy caveat HIGH sources/06-industry-verticals/jane-street-industrial-ml-constraints-raw.md
Tower Research Capital 2026-05-26 Tower frames capital-markets AI as structured environments with knowledge graphs, agent harnesses, tool/API/data controls, observability, token budgets, execution limits, ROI measurement, and zero-trust permissions. knowledge graphs; permissioned tool APIs; feedback-loop traces; benchmark KPIs; ROI telemetry observability; token budgeting; execution limits; zero-trust permissions; ROI measurement HIGH sources/06-industry-verticals/tower-ai-agent-harness-governance-raw.md
G-Research 2026-05-13 G-Research exposes a production LLM reliability pattern: treat model output as unverified, validate against a source-of-truth rules index, require JSON/Pydantic structure, split recall and precision checks, bound repair attempts, abstract providers, and log cost. source-of-truth rules index; diffs; expected findings; provider cost telemetry; structured JSON outputs Pydantic validation; two-pass recall/precision gates; bounded JSON repair; non-blocking human review; cost telemetry HIGH sources/06-industry-verticals/g-research-production-llm-code-review-raw.md
Schonfeld Strategic Advisors 2026-05-12 Schonfeld’s Fundamental Equity AI Lab trains PMs and analysts to automate earnings prep, idea generation, document analysis, inbox triage, and Excel workflows using proprietary internal systems and frontier-model partnerships. earnings materials; documents; email/inbox streams; spreadsheets; proprietary internal systems structured pilots; downside evaluation; rollout gates HIGH sources/06-industry-verticals/buy-side-quant-ai-practitioner-signals-2026-raw.md
Citadel EQR 2026-04-28 Citadel EQR publicly describes a systematic pipeline from real-world observation and dataset identification to forecasts, portfolio construction, optimization, market impact, execution, testing, and market feedback. proprietary datasets; alternative data; market feedback; liquidity and execution data; simulation artifacts testing before deployment; rapid market feedback; prototype simulation; discussion and refinement HIGH sources/06-industry-verticals/buy-side-quant-ai-practitioner-signals-2026-raw.md
Jane Street Official ML Blog Ledger 2026-04-22 Jane Street’s official ML blog ledger leaks the research-culture control surface: model inspection, mechanistic interpretability puzzles, sequential-model inductive-bias reasoning, market information propagation, large low-signal dataset handling, and non-Python ML tooling all matter before a finance model is trusted. neural-network weights; model specifications; market data; large low-signal labels; dataset shuffle artifacts; sequence-model experiments; market information propagation visualizations; Kaggle market-prediction artifacts full model-spec inspection; dependency tracing; SAT/ILP-style solver trace; brute-force insufficiency check; positional-encoding theory review; low signal-to-noise caveat; competition-versus-production boundary HIGH sources/06-industry-verticals/jane-street-official-ml-blog-ledger-raw.md
JPMorgan Chase Internal Agent Platform 2026-04-13 JPMorgan leaks the enterprise-finance agent platform pattern: model choice is secondary to controlled connectivity into data, technology, and process estates; internal adoption scales through reusable components, gap surveillance, central triage, data-lineage control, and top-down redesign of cross-team processes. internal enterprise data; process telemetry; user-found gap reports; platform usage traces; data lineage records; cross-team workflow records data privacy boundary; data-lineage control; central gap surveillance; shared-platform triage; bottom-up and top-down adoption split; process redesign gate HIGH research/13-multimodal-sources/beyond-the-pilot/2026-04-13-what-30k-jpmorgan-ai-agents-taught-me.md
Balyasny Asset Management 2026-03-06 Balyasny exposes a production research-agent operating model: centralized Applied AI, internal model benchmarks, scoped tools, proprietary data access, traceable reasoning, compliance guardrails, and team-specific agent customization. proprietary financial data; internal benchmarks; central-bank speeches; merger-arbitrage deal data; team-specific research corpora internal benchmarks; scoped data/tool access; traceable reasoning paths; compliance guardrails; human decision support boundary HIGH sources/06-industry-verticals/balyasny-openai-ai-research-engine-raw.md
Man Group / AHL 2026-02-11 Man discloses a frontier-lab partnership where Claude sits between raw datasets and investment decisions through code execution, workflow-described Skills, portfolio-team synthesis, Claude Code productivity, and proprietary AlphaGPT idea-generation tooling. raw datasets; financial risk models; ad hoc communications; textual information; automated reports; agentic AI outputs provider dependency disclosure; decision-support framing; alpha-claim caveat; runtime portability caveat HIGH sources/06-industry-verticals/man-anthropic-alphagpt-partnership-raw.md
MSCI IndexAI 2026-02 MSCI exposes the licensed-data connector control plane: entitlement-grounded index data access through MSCI ONE, ChatGPT, Claude, and client systems, with required tool-call provenance and warnings for unsupported connector answers. MSCI index data; index constituents; weights; factor exposures; ESG metrics; methodology definitions entitlement checks; tool-call references; connector access controls; usage audit; abstain/warn on missing provenance HIGH research/06-industry-verticals/msci-indexai-insights-user-guide-2026.md
XTX Markets 2026-01-28 XTX exposes the industrial-ML substrate: price forecasts over 50,000+ instruments, 25,000 GPUs, 650 PB usable storage, and TernFS built because ML research storage outgrew general filesystems. market data across 50,000+ instruments; ML research datasets; filesystem snapshots; cluster data-access traces; price-forecast artifacts immutable files; snapshots; partial-write safety; external permission boundary HIGH sources/06-industry-verticals/quant-ai-infrastructure-moats-2026-raw.md
Two Sigma 2025-06-17 LLMs sit upstream in equities feature forecasting; observed real-world events such as hiring signs can become cross-company features. hiring/labor traces; multimodal observations; company-level feature stores point-in-time temporal validation; overfitting checks HIGH research/13-multimodal-sources/twiml/2025-06-17-llms-for-equities-feature-forecasting-at-two-sigma.md
Jane Street 2025-03-10 Jane Street frames research as exploration, data collection, modeling, and productionization under low-data/high-noise market conditions. market data; hand-collected datasets; ticker/split/error checks; production system telemetry out-of-sample belief discipline; data-quality checks; simple baseline before deep model; production gate HIGH research/13-multimodal-sources/jane-street-signals-and-threads/2025-03-10-finding-signal-in-the-noise.md
Citadel / Citadel Securities AI-Fluent Intern Pipeline 2026-06-09 Citadel and Citadel Securities leak an AI-fluent talent-pipeline control surface: a 2026 intern class selected from more than 115,900 applicants is screened for AI fluency, adaptability, judgment, and STEM strength, given access to the same AI tools as full-time employees, assigned high-impact projects, and evaluated through regular reviews and leadership presentations. intern project artifacts; employee AI-tool environments; recruiting assessment signals; performance-review notes; leadership presentation materials; quant and trading training context media-source caveat; intern-versus-full-time authority separation; Citadel LLC versus Citadel Securities boundary; tool access does not imply production trading authority; project-output human review; AI fluency versus alpha proof separation; no specific model or dataset disclosure MEDIUM sources/06-industry-verticals/citadel-ai-fluent-intern-talent-pipeline-raw.md
Eland Traditional Chinese Finance Sentiment 2026-06-06 Eland leaks a regional sentiment-runtime lane: Traditional Chinese Taiwan finance sentiment needs entity, opinion, stance, and overall labels, explicit vLLM system-prompt injection, separated merged/LoRA/GGUF artifacts, revision capture, and held-out local-market replay before promotion. Taiwan stock-market forum text; Taiwan finance news text; Traditional Chinese sentiment labels; entity-sentiment examples; opinion-sentiment examples; vLLM merged weights; LoRA adapter; GGUF variants system-prompt contract; exact revision capture requirement; artifact-form separation; held-out Taiwan finance fixture; model-card metric reproduction gate; runtime portability caveat MEDIUM sources/21-benchmarks/eland-sentiment-zh-vllm-raw.md
AntGroup Finix-S1 Insurance Domain Model 2026-06-06 Finix-S1 leaks the closed-domain-model reference lane: a non-open Ant insurance model reportedly leads CUFEInse on insurance knowledge, industry understanding, security/compliance, agent applications, and rigorous reasoning, but the useful signal is the missing fixture class around insurance law, product knowledge, underwriting, claims, servicing, compliant marketing, and deterministic actuarial handoff. CUFEInse benchmark items; insurance theory questions; insurance industry understanding tasks; insurance security/compliance tasks; insurance agent application tasks; policy clause examples; underwriting and claims scenarios closed-model availability caveat; public artifact absence check; benchmark-report provenance; runtime/API contract verification before scoring; policy/citation/calculation fixture requirement; human authority over underwriting and claims; no-Ollama/no-runtime inference MEDIUM sources/21-benchmarks/antgroup-finix-s1-insurance-domain-model-raw.md
D. E. Shaw 2026-06-05 D. E. Shaw describes optimizers in systematic and discretionary investment contexts as human-machine tools that force PMs and traders to quantify assumptions, evaluate tradeoffs, and confront cognitive bias. portfolio constraints; optimization inputs; shared data and technology resources; hypothesis and validation artifacts comparative human-machine review; assumption explicitness; statistically robust hypothesis testing MEDIUM sources/06-industry-verticals/de-shaw-machine-teaching-optimizer-raw.md
Acadian Asset Management / Investor Forum 2026-06-05 Acadian’s 2026 investor forum leaks the scale and boundary conditions of a live systematic manager: an alpha engine analyzing 65,000+ securities daily with hundreds of signals, cross-asset coverage, investment-expert staffing, and explicit model-advisory assets without trading authority. 65,000+ security universe; hundreds of signal inputs; equities/fixed-income/alternatives coverage; client-asset composites; model-advisory asset records; benchmark performance composites as-of-date discipline; composite-AUM caveats; discretionary vs model-advisory separation; no causal AI-alpha attribution; firm-owned performance caveat MEDIUM sources/06-industry-verticals/acadian-2026-investor-forum-systematic-scale-raw.md
FinTradeBench Fundamentals-Trading Signal Reasoning 2026-06-03 FinTradeBench leaks a useful research-control split: finance agents must reconcile regulatory-filing fundamentals with price/volume-derived trading signals while preserving historical windows, data vintages, golden indicators, and the boundary between benchmark reasoning and live trading authority. regulatory filings; company fundamentals; historical price data; volume dynamics; NASDAQ-100 company windows; golden key indicators golden indicator citation; historical-window checks; fundamentals versus trading-signal separation; RAG degradation checks; no-live-trading caveat MEDIUM sources/21-benchmarks/fintradebench-financial-reasoning-raw.md
SQL Benchmark Quality / Dialect Portability Controls 2026-06-02 The SQL benchmark-quality card leaks why finance agents need module-level SQL diagnostics: schema selection, candidate generation, query revision, dialect translation, execution results, permission labels, ambiguity correction, and gold-query audits must be scored separately from public BIRD/Spider leaderboard numbers. NL2SQLBench traces; BIRD development data; ScienceBench development data; PARROT SQL translation pairs; 22 production-grade database systems; finance schema contracts; gold SQL and result sets Correct/Incorrect/Error Rate split; Pass@k; gold SQL annotation audit; execution-first metric; dialect-portability tests; semantic-view/schema contract; RBAC/entitlement labels MEDIUM sources/21-benchmarks/sql-benchmark-quality-dialect-portability-raw.md
Hedge-Bench Analyst-Agent Benchmark 2026-06-02 Hedge-Bench leaks the target shape of hedge-fund analyst agents: closed document packs, professional reasoning traces, clear investment positions, strongest counter-evidence, ambiguity reconciliation, source-file citations, and hallucination penalties. closed document packs; professional analyst traces; valuation materials; M&A materials; competitive-positioning evidence; risk documents; source-file citations closed materials; deterministic task tests; LLM-judge caveat; citation grounding; unsupported-claim penalties; reproduction requirement MEDIUM sources/21-benchmarks/hedge-bench-financial-reasoning-agents-raw.md
BigFinanceBench Research-Agent Benchmark 2026-06-02 BigFinanceBench leaks the auditable financial-research workflow: entity identification, source selection, line-item retrieval, accounting adjustment, formula construction, calculation, and final synthesis are separately rubric-scored from visible tool traces. public filings; EDGAR search results; market data; web sources; line items; accounting definitions; calculation traces; rubric criteria point-weighted rubrics; visible trajectory grading; source-date correctness; formula/calculation consistency; partial-credit failure localization; LLM-judge caveat MEDIUM sources/21-benchmarks/bigfinancebench-workflow-grounded-financial-research-agents-raw.md
QFBench Quantitative Finance Agents 2026-06-01 QFBench leaks the executable quant-agent task shape: agents must write and run quantitative-finance code in Docker sandboxes over task-local data, then pass strict numerical tests for options, risk, factor models, event studies, market microstructure, volatility, rates, portfolio risk, and SEC-event workflows. task-local datasets; instruction.md files; task.toml metadata; Python financial libraries; pytest verifiers; reference solutions; leaderboard run traces binary pass/fail tests; strict numerical tolerances; three independent runs; oracle solution precheck; solution-hint leakage controls; live-leaderboard caveat MEDIUM sources/21-benchmarks/qfbench-quantitative-finance-agent-benchmark-raw.md
Renaissance Technologies 2026-05-31 Renaissance leaks quiet-firm infrastructure through official careers material: research/data-processing software, technical models for predicting and trading markets, C++/Rust, low-level CPU/GPU, distributed computing, compiler, and research-infrastructure programming. research data-processing systems; statistical model artifacts; trading algorithm infrastructure; high-performance compute traces official-site authority check; anti-impersonation warning; careers-source classification; no-agent-claim boundary MEDIUM sources/06-industry-verticals/renaissance-technologies-official-careers-signal-raw.md
Local Finance SLM Runtime Watchlist 2026-05-31 The finance SLM watchlist leaks the local-runtime control surface: finance reasoning, embedding, analyst-report style, SEC extraction, bookkeeping, sentiment, tool-routing, banking voice, DeFi, and regional finance models must be scored by artifact provenance, license, quantization, runtime support, context window, dataset contamination, and task-specific fixture fit. Hugging Face model cards; GGUF metadata; finance-reasoning synthetic data; financial RAG embedding data; private analyst-report corpus; SEC contract extraction instructions; bookkeeping/accounting datasets; sentiment datasets license review; contamination check; quantization-separated scoring; runtime support verification; no-Ollama policy; benchmark replay before promotion; task-specific fixture matching; card-metric verification MEDIUM sources/21-benchmarks/current-finance-slm-watchlist-2026-05-31-raw.md
Kronos K-Line Market Foundation Model 2026-05-31 Kronos leaks the artifact shape for finance-native market models: OHLCV/amount/timestamp schemas are tokenized into K-line sequences for autoregressive forecasting, volatility forecasting, and synthetic sequence generation, with backend and tokenizer revisions needing separate promotion gates. OHLC columns; volume and amount fields; timestamped K-line records; 45 global exchange histories; 12B+ claimed K-line records; HF model cards; GitHub repository exact model revision pin; tokenizer revision pin; OHLCV schema check; point-in-time splits; classical baseline comparison; transaction-cost/slippage/capacity caveat; backend-specific smoke tests MEDIUM sources/21-benchmarks/kronos-financial-time-series-foundation-model-raw.md
Fara-7B Browser Source-Discovery Agent 2026-05-31 Fara-7B leaks a runnable source-discovery lane for this corpus: an open-weight browser/computer-use model can be evaluated on finding official arXiv, report, whitepaper, vendor, podcast, transcript, and PDF pages while preserving URL, date, author, source authority, and stop conditions. official arXiv pages; paper PDFs; vendor documentation; whitepapers; podcast episode pages; transcript URLs; browser screenshots; web trajectories MIT license check; artifact revision pin; primary-source versus mirror distinction; URL/date/author preservation; sandboxed browser execution; stop before login/payment/personal data; backend-specific runtime caveat MEDIUM sources/21-benchmarks/fara7b-browser-source-discovery-raw.md
ModernFinBERT / ProsusAI Sentiment Baselines 2026-05-30 The FinBERT baseline card leaks the sentiment-routing boundary: label-only classifiers can reach usable sentiment accuracy while still failing route, rationale, and full-workflow acceptance, so research agents need separate label-only, route, abstention, and explanation score lanes. Financial PhraseBank; financial reasoning aggregate data; 15-task finance sentiment routing fixture; classifier label maps; local Transformers/MPS run summaries explicit label map; label accuracy separated from route accuracy; abstention accuracy; full-fixture acceptance gate; local runtime trace; model-card metric caveat MEDIUM sources/21-benchmarks/modernfinbert-prosusai-finbert-sentiment-baselines-raw.md
rLLM-FinQA SEC Table Tool Agent 2026-05-29 rLLM-FinQA leaks the tool-use shape for SEC table QA: a finance model is trained around 10-K questions, SQL queries, table lookup, calculator calls, company tables, and GRPO-style post-training rather than free-form financial chat alone. SEC 10-K filings; 5,110 QA pairs; 207 companies; 6,923 filing tables; SQL query traces; table lookup evidence; calculator-tool outputs; Snorkel Finance score reports tool-use provenance check; table-cell grounding; calculator replay; model-card metric caveat; vLLM-only runtime note; no GGUF/MLX verification; local benchmark replay before promotion MEDIUM sources/21-benchmarks/rllm-finqa-4b-model-card-raw.md
WorldQuant 2026-05-29 WorldQuant leaks the competition-medium research funnel: BRAIN/IQC participation, 263,000+ public alpha submissions, AI as a research partner for scanning papers, generating hypotheses, running simulations, refining strategies, and shifting advantage toward better question framing. WorldQuant BRAIN alpha submissions; IQC participant metadata; research papers; simulation logs; strategy-refinement artifacts; competition statistics competition vs production boundary; sponsor/date/count preservation; participant-signal caveat; no deployable-alpha inference MEDIUM sources/06-industry-verticals/worldquant-ai-iqc-talent-signal-raw.md
Treasury Payments and Liquidity Control Lane 2026-05-29 The treasury/payments ledger leaks the authority split for agentic money movement: probabilistic agents can propose payment, liquidity, fraud, reconciliation, and treasury actions only when deterministic authorization, settlement/legal-finality, agent identity, audit, kill-switch, and human override layers stay separate. payment instructions; treasury cash-flow records; liquidity forecasts; fraud and sanctions signals; settlement records; deposit-token/stablecoin/CBDC rails; banking agent survey responses; reconciliation records three-layer decision/authorization/settlement split; mandate-based authorization; Know-Your-Agent identity; programmable payment controls; audit trail; tiered human-in-the-loop; kill switch; manual override; legal-finality caveat MEDIUM sources/21-benchmarks/finance-treasury-payments-liquidity-benchmarks-2026-raw.md
Structured Finance and Securitization Controls 2026-05-29 The structured-finance ledger leaks a contract-defined cash-flow machine: agents can extract PSA clauses, collateral-pool fields, servicing language, loan-tape mappings, prepayment/default drivers, tranche waterfalls, eligibility rules, triggers, subordination, excess spread, and CLO monitoring fields only if deterministic waterfall and legal controls remain authoritative. pooling and servicing agreements; CMBS contract text; Fannie Mae loan performance data; mortgage prepayment histories; loan tapes; CLO white papers; tranche cash-flow assumptions section/page source localization; strict temporal mortgage-default split; class-imbalance review; deterministic waterfall engine; prepayment/default model baseline; legal/servicing authority check; cash-flow outcome linkage MEDIUM sources/21-benchmarks/finance-structured-securitization-benchmarks-2026-raw.md
Squarepoint / PDT 2026-05-29 Squarepoint and PDT expose the industrial research loop: systematic research and trading integrated through technology, ML-based quantitative models, automated strategies, backtests, and rigorous research/testing before live deployment. historical market data; backtest outputs; quantitative model artifacts; automated trading system telemetry Dream/Experiment/Validate/Repeat loop; rigorous testing before live deployment; historical backtest discipline MEDIUM sources/06-industry-verticals/buy-side-quant-ai-practitioner-signals-2026-raw.md
QRT Labs 2026-05-29 QRT Labs exposes a frontier-AI talent and research pipeline with Imperial, Cambridge, and Oxford, explicitly including foundation AI models, agentic systems, HPC, cybersecurity, hardware design, and mathematical modelling. academic research outputs; foundation-model research artifacts; HPC research environments; early-career researcher project traces university partnership boundary; research-phase framing; talent/source-discovery classification MEDIUM sources/06-industry-verticals/quant-ai-talent-hiring-leakage-2026-raw.md
Post-Trade Collateral and Surveillance Controls 2026-05-29 The post-trade ledger leaks the control layer around trades: finance agents can triage margin calls, collateral eligibility, repo/securities-lending mappings, settlement checks, surveillance alerts, records, liquidity reports, and reconciliations, but cannot move collateral or alter regulatory/customer-asset records without deterministic controls. margin and collateral call records; collateral inventories; repo and securities-finance records; derivatives margin data; FINRA supervision reports; off-channel communications; customer-asset reconciliation records; settlement and DvP records early-warning indicators; contingency funding plan; stress-test scenario design; collateral transferability and settlement check; external-party reconciliation; books-and-records supervision; human/legal/risk escalation MEDIUM sources/21-benchmarks/finance-post-trade-collateral-surveillance-benchmarks-2026-raw.md
Portfolio and Risk-Construction Model Lane 2026-05-29 The portfolio/risk-construction ledger leaks the full decision pipeline beyond stock picking: agents may screen fundamentals and news, forecast covariance, build graph relationships, propose weights, and adapt risk profiles, but promotion depends on point-in-time features, covariance loss, drawdown, turnover, liquidity, frictions, and risk-overlay evidence. S&P 500 fundamentals; financial news sentiment; ETF covariance histories; CRSP equities; Open Asset Pricing characteristics; dynamic firm-similarity graphs; DJIA and Hang Seng market regimes point-in-time feature gate; classical covariance baseline comparison; GMV variance and turnover; dynamic graph leakage check; drawdown and volatility review; transaction-cost/liquidity/shorting caveat; human execution authority MEDIUM sources/21-benchmarks/finance-portfolio-risk-construction-models-2026-raw.md
PortBench Portfolio-Management Benchmark 2026-05-29 PortBench leaks a portfolio-agent evaluation stack: market interpretation, signal generation, weight optimization, execution simulation, and risk monitoring must be scored for correlation reasoning, cross-stage error propagation, stress regimes, turnover, slippage, commission, and drawdown. six-asset-class histories; 183 instruments; correlation matrices; stress-regime data; portfolio weights; execution simulation outputs; risk-monitoring traces dual-layer correlation scoring; CEPS; profile alignment; Sharpe/Sortino/Calmar; maximum drawdown; slippage/commission; stress gate MEDIUM sources/21-benchmarks/portbench-portfolio-management-raw.md
Millennium 2026-05-29 Millennium leaks an internal AI-enablement operating model: a global AI advisory group, foundational and custom tools, end-user partnership, AI training events, and hackathon prototypes spanning agentic AI, RAG, MCP, and workflow automation. internal workflow context; AI tool/resource catalogs; hackathon submissions; RAG corpora; MCP connector candidates; end-user feedback official firm-source boundary; prototype vs production separation; training/advisory framing; no investment-performance inference MEDIUM sources/06-industry-verticals/millennium-ai-advisory-hackathon-raw.md
Kaggle Quant Replay Harnesses 2026-05-29 The replay artifact leaks how quant competitions become agent-evaluation work: recover canonical rules, metrics, kernels, discussions, public/private split semantics, leakage controls, local grading, submission artifacts, and iterative experiment traces rather than asking static QA questions. Jane Street competition artifacts; Optiver competition artifacts; Two Sigma news competition artifacts; G-Research crypto competition artifacts; rules pages; kernels; discussion timestamps; local grader artifacts public/private split preservation; canonical metric check; local grader reproduction; leaderboard-to-production caveat; capacity/slippage/regime caveat MEDIUM sources/21-benchmarks/finance-kaggle-competition-artifacts-replay-raw.md
Insurance Actuarial and Reinsurance Controls 2026-05-29 The insurance ledger leaks a regulated-adjacent finance workflow: agents can route underwriting, claims, policy QA, compliance calls, actuarial calculations, and reinsurance treaty reasoning, but binding decisions require rule engines, regulatory-policy retrieval, deterministic calculators, schema/state-machine guards, adversarial critique, and human authorization. CUFEInsure questions; Chinese insurance benchmark examples; historical warranty claims; Quebec insurance certification questions; underwriting manuals; SQLite underwriting backends; insurance call transcripts; reinsurance allocation training material rule-engine check; retrieval citation and source localization; schema validation; state-machine guard conditions; adversarial self-critique; deterministic actuarial calculator; human final authorization MEDIUM sources/21-benchmarks/finance-insurance-actuarial-benchmarks-2026-raw.md
HERCULEAN Agentic Finance Benchmark 2026-05-29 HERCULEAN leaks the environment contract for financial agents: trading, hedging, market-insights, and auditing workflows expose observations, tools, actions, temporal constraints, output schemas, and evaluation criteria through skill/MCP environments, and agents must commit artifacts rather than end with prose. MCP environment observations; tool traces; workflow state; XBRL facts; market-insight report artifacts; trading and hedging outputs return/Sharpe/drawdown scoring; structured report rubrics; deterministic XBRL fact verification; persistent memory disabled; built-in web search disabled; cost trace MEDIUM sources/21-benchmarks/herculean-agentic-financial-intelligence-raw.md
Gemma SEC Filing Adapter 2026-05-29 The Gemma SEC adapter leaks a current finance fine-tuning pattern: small knowledge-distilled SEC QA pairs, LoRA adapter training, held-out semantic-score reporting, and a companion tool-using earnings-recap agent over EDGAR, market data, web search, charts, citations, and factuality judging. 381 SEC summaries; 1,060 knowledge-distilled QA pairs; 10-K filings; 10-Q filings; 8-K filings; 69 S&P 500 companies; EDGAR data; market data; web search results; chart artifacts base-versus-adapter comparison; held-out evaluation caveat; BERTScore-not-factuality caveat; Gemma license review; adapter/base separation; no merged/GGUF/MLX/Neuron verification; StateBench reproduction gate MEDIUM sources/21-benchmarks/gemma-4-sec-filing-adapter-model-card-raw.md
Finance Text Alpha / LLM Asset Pricing 2026-05-29 The text-alpha ledger leaks the point-in-time language-model control surface: stock-return text signals depend on licensed news feeds, headlines, earnings calls, chronological model cutoffs, timestamp windows, return-label construction, immediate-reaction versus drift decomposition, turnover, transaction costs, and small-cap/capacity concentration. Refinitiv Thomson Reuters news; Dow Jones Newswire headlines; CRSP returns; Datastream-EIKON international stocks; earnings-call text; OpenAI embeddings; ChronoBERT/ChronoGPT vintages strict timestamp cutoff; chronologically consistent model vintage; event-level split; IC/RankIC evaluation; value-weighted versus equal-weighted comparison; turnover and transaction-cost gate; full-cycle regime requirement MEDIUM sources/21-benchmarks/finance-text-alpha-llm-asset-pricing-2026-raw.md
Finance Safety Fraud Compliance Control Ledger 2026-05-29 The safety/fraud/compliance ledger leaks the failure modes that finance-agent demos hide: multi-turn compliance decay, transactional-intent escalation, structured-fraud class imbalance, motivated investor pressure, Japanese whole-filing tasks, atomic claim verification, table-cell operand grounding, and long-context filing hallucination all need explicit negative-evidence gates. CNFinBench safety tasks; transaction-fraud datasets; investment fraud scenarios; EDINET annual reports; SEC filings; financial tables; retrieved evidence paragraphs; atomic claim labels; long-context filing spans multi-turn attack testing; risk-interception checks; precision/recall and MCC over accuracy; analyst-in-loop thresholding; retrieval-equalized evaluation; paragraph/table-cell citations; lost-in-the-middle placement tests; human legal/compliance review MEDIUM sources/21-benchmarks/finance-safety-fraud-compliance-2026-raw.md
Finance Retrieval and NL2SQL Harness Ledger 2026-05-29 The retrieval/NL2SQL ledger leaks the data-access substrate for research agents: finance answers depend on semantic views, RBAC, verified SQL examples, hybrid lexical/dense retrieval, reranking, table-aware evidence, citation localization, temporal metadata, cost/latency tracing, and ambiguity-correction loops rather than a generic vector store. Snowflake semantic views; Cortex Analyst verified queries; Cortex Search indexes; Postgres vectors; SQL DDL; business metric documentation; SEC filings; earnings reports; financial tables; visual report pages RBAC and warehouse permissions; gold SQL and verified-query regression; hybrid BM25/dense baseline; cross-encoder reranking; citation precision/recall; temporal/as-of-date checks; token/cost/latency reporting; annotation-error review MEDIUM sources/21-benchmarks/finance-retrieval-nl2sql-harnesses-2026-raw.md
Finance Retrieval / NL2SQL Harness Stack 2026-05-29 The finance retrieval/NL2SQL stack leaks the deployment harness shape: Snowflake, Postgres, OSS SQL agents, vector retrieval, semantic layers, traces, cost, latency, entitlement, data-vintage, citation, dialect, and abstention checks belong in one scored bakeoff rather than scattered demo claims. Cortex Analyst/Search/Agents traces; Snowflake Semantic Views; Postgres pgvector rows; pgvectorscale indexes; Wren MDL; Vanna training pairs; Defog SQL-Eval tasks; FinanceBench/FinDER/T2-RAGBench-style evidence verified query evaluations; RBAC checks; generated SQL execution; filtered recall; issuer/fund/date/permission filters; citation precision; dialect portability; abstention reason; audit trace MEDIUM sources/21-benchmarks/finance-retrieval-nl2sql-harness-stack-raw.md
Finance Credit AML Risk Control Ledger 2026-05-29 The credit/AML ledger leaks the regulated-risk control surface for finance agents: credit-document VLMs, raw bureau text, corporate credit scoring, AML transaction graphs, fairness metrics, adverse-action explanations, reject inference, out-of-time validation, profile-aware anomaly detection, and independent model-risk challenge all belong in the harness before any agent touches credit or AML decisions. credit certificate images; credit-bureau records; commercial credit ratings; corporate financials; KYC profiles; transaction graphs; fraud/default labels; adverse-action reason codes; FinRegLab framework out-of-time validation; class-imbalance metrics; fairness metrics; reject inference review; specific and accurate adverse-action reasons; independent effective challenge; distribution-shift monitoring; human compliance authority MEDIUM sources/21-benchmarks/finance-credit-aml-risk-models-2026-raw.md
Fin-Bias Analyst-Herding Benchmark 2026-05-29 Fin-Bias leaks the opinion-contamination control surface for research agents: 8,868 Yahoo Finance analyst reports are tested with analyst ratings present, removed, and replaced by fake ratings, then scored against realized stock returns and 60-day cumulative abnormal returns rather than analyst opinion. Yahoo Finance analyst reports; Bullish/Neutral/Bearish ratings; fake-rating perturbations; MPQA subjectivity lexicon; realized stock returns; 60-day cumulative abnormal returns Herding Score; realized-return labels; quantile-based ground truth; rating-present/removed/fake conditions; analyst-label imbalance caveat MEDIUM sources/21-benchmarks/fin-bias-analyst-herding-raw.md
FICC Macro Commodities and Private Credit Controls 2026-05-29 The FICC/private-credit ledger leaks the non-equity research boundary: rates and macro agents must handle curve state, maturity structure, publication lags, order-book/liquidity simulation, commodity rolls, storage and hedging pressure, capacity/crowding, private-credit opacity, stale valuations, covenant monitoring, fees, and leverage channels. Treasury constant-maturity yields; LSEG zero-coupon Treasury yields; macro and financial indicators; Polymarket CLOB snapshots; commodity futures data; Bloomberg commodity features; private-credit AUM and direct-lending evidence sliding/expanding window validation; publication-lag discipline; classical econometric baseline; calibration and abstention scoring; liquidity/slippage simulation; capacity and crowding checks; stale-valuation/systemic-risk caveat MEDIUM sources/21-benchmarks/finance-ficc-macro-private-credit-benchmarks-2026-raw.md
ESG Climate Sustainability Evidence Benchmarks 2026-05-29 The ESG/climate benchmark cluster leaks the stewardship evidence layer: sustainability agents must retrieve complete report pages, tables, figures, units, dates, fiscal years, scope categories, standards taxonomies, actionability labels, TCFD/ISSB/GRI/SASB/CDP provenance, and human/legal review gates. sustainability reports; ESG PDFs; IPCC reports; GRI standards; SASB standards; IFRS/ISSB standards; TCFD reports; CDP disclosures page/table/figure provenance; hybrid retrieval and reranking; commitment/plan/implemented action labels; evidence coverage check; abstention/review gate; regulatory disclosure caveat MEDIUM sources/21-benchmarks/finance-esg-climate-sustainability-benchmarks-2026-raw.md
Derivatives Volatility and Hedging Controls 2026-05-29 The derivatives ledger leaks a finance-physics gate for agentic pricing work: volatility surface completion, option pricing, Greeks, and hedging must preserve no-arbitrage, put-call parity, calendar and butterfly constraints, terminal payoff, self-financing behavior, and hedge P&L distributions. SPY option surfaces from 2008-2025; SPX option quotes; arbitrage-free volatility surfaces; QuantLib option labels; simulated incomplete-market paths; FINESSE derivatives QA tasks calendar-spread no-arbitrage; butterfly-convexity penalty; put-call parity; terminal payoff embedding; self-financing loss; out-of-sample hedge P&L mean/std/tail quantiles; transaction-cost/liquidity/market-impact caveat MEDIUM sources/21-benchmarks/finance-derivatives-volatility-hedging-models-2026-raw.md
CPA-Qwen3 Accounting Compliance Candidate 2026-05-29 CPA-Qwen3 leaks the accounting/compliance specialist preflight lane: CPA, GAAP, IFRS, tax, audit-risk, regulatory-interpretation, and professional-skepticism claims must be separated from model-card persona text and reproduced against jurisdiction, citation, calculation, refusal, and human-review tasks. Finance-Instruct-500k dataset reference; Qwen3-8B safetensors; CPA task prompts; GAAP/IFRS/tax-code references; audit-risk scenarios; regulatory-compliance prompts; community GGUF quantization metadata Apache-2.0 license review; upstream-vs-GGUF artifact separation; jurisdiction boundary; citation requirement; calculation replay; professional-skepticism/refusal gate; human-review escalation MEDIUM sources/21-benchmarks/cpa-qwen3-accounting-compliance-model-raw.md
BlueFin Spreadsheet-Agent Benchmark 2026-05-29 BlueFin leaks the concrete artifact layer that buy-side and deal teams cannot replace with prose QA: professional finance spreadsheet agents must synthesize, manipulate, and comprehend workbooks while preserving formula behavior, dynamic recalculation, formatting, reviewability, and stakeholder utility. professional finance spreadsheets; workbook formulas; spreadsheet formatting; granular rubric criteria; expert-human judge validation 3,225 granular rubric criteria; expert-human judge validation; dynamic correctness checks; formula fidelity checks; negative-evidence caveat MEDIUM sources/21-benchmarks/bluefin-financial-spreadsheet-agents-raw.md
BankerToolBench Deal-Work Agents 2026-05-29 The deal-work benchmark ledger leaks the investment-banking research machine around data rooms, SEC filings, market data platforms, Excel models, PowerPoint decks, Word/PDF reports, senior-banker requests, veteran-banker rubrics, and cross-artifact client readiness. data rooms; SEC filings; market data platforms; Excel financial models; PowerPoint pitch decks; Word/PDF reports; internal deal history; banker notes veteran-banker rubrics; cross-artifact consistency checks; spreadsheet formula fidelity; source/tool grounding; confidentiality and compliance review; human sign-off MEDIUM sources/21-benchmarks/finance-investment-banking-deal-work-benchmarks-2026-raw.md
Agentic Factor Discovery Frameworks 2026-05-29 The agentic factor-discovery ledger leaks the quant-research promotion gate: LLMs can generate factor formulas or executable programs, but candidates must pass safe execution, point-in-time data handling, no-lookahead checks, IC/RankIC/ICIR diagnostics, turnover/cost/capacity review, and deterministic evaluator promotion before a narrative is trusted. U.S. equity panels; A-share Qlib-style data; CRSP daily equities; financial reports; factor libraries; operator trees; program-level factor code; research artifact logs AST sandbox; structural security checks; complexity limits; semantic validation; RankIC and ICIR scoring; OOS holdout; transaction-cost review; capacity and market-impact caveat MEDIUM sources/21-benchmarks/finance-agentic-factor-discovery-2026-raw.md
FinGuard Financial Compliance Guard 2026-05-28 FinGuard leaks the regulation-grounded guardrail lane for financial agents: compliance detection should be induced from regulatory documents, scored at query and response level, adapted to institution-specific policies, and separated from generic safety refusal taxonomies. Chinese financial regulations; institution-specific policy documents; FinGuard-Bench labels; query-level compliance labels; response-level compliance labels; synthetic grounded training data; self-play reinforcement traces expert query-level labels; expert response-level labels; jurisdiction boundary; generic guard-model comparison; policy-document adaptation test; reported-performance reproduction gate; no open-weight verification caveat MEDIUM sources/21-benchmarks/finguard-financial-compliance-guard-model-raw.md
Bridgewater AIA Labs / Agentic Research Platform Hiring 2026-05-28 Bridgewater’s AIA Labs job listing leaks the productization layer around its artificial-investor program: a team led by Greg Jensen is hiring to evolve and scale an AI research platform for systematic investment research, explicitly naming agentic development tooling, LLMs, agents, tool use, harnesses, orchestration, and scientist/investor workflows. official job listing text; AI research platform requirements; agentic tooling requirements; LLM and tool-use harnesses; scientist and investor workflow needs; compensation and talent-market signal job-post source caveat; hiring intent versus deployed system separation; productization versus alpha attribution separation; agentic tooling does not imply autonomous capital allocation; official-company listing provenance; compensation as talent signal not performance proof MEDIUM sources/06-industry-verticals/buy-side-quant-ai-practitioner-signals-2026-raw.md
Snowflake Arctic Text2SQL Finance Harness 2026-05-27 Snowflake Arctic leaks the warehouse-specialist SQL lane for finance agents: open R1 is a runnable Qwen2.5-Coder GRPO SQL artifact, while R2 signals Snowflake-dominant mid-training, realistic SQL, collision-resistant execution rewards, long-context schema reasoning, and internal-hard-benchmark specialization. Snowflake schemas; semantic views; generated SQL; BIRD/Spider/EHRSQL/ScienceBenchmark tasks; warehouse execution results; internal hard Text-to-SQL benchmark traces read-only SQL; schema allowlist; semantic-view versioning; row and cost limits; execution-result comparison; RBAC tests; refusal behavior; vendor-reported R2 caveat MEDIUM sources/21-benchmarks/snowflake-arctic-text2sql-model-cards-raw.md
Discretionary Hedge-Fund PM Workflow 2026-05-27 A discretionary PM workflow leak: AI helps gather facts, organize filings/conference/expert/social/podcast evidence, translate source material into financial-model assumptions, and run thesis-monitoring agents that seek confirming and contradicting evidence. filings; industry conferences; alternative datasets; expert-network transcripts; YouTube; Reddit; podcasts; public AI-company interviews fiduciary accountability; human PM decision boundary; source-to-model traceability; contradiction capture MEDIUM research/13-multimodal-sources/the-fund-ai-pod/2026-05-27-shu-bai-hedge-fund-ai-research-workflow.md
DataClawBench Exploratory Financial Data Analysis 2026-05-27 DataClawBench leaks the messy-data research-agent harness: agents must discover sources and schemas without prior cleanup, reason across enterprise, industry, and policy files, handle 2.06M noisy records, missing values, naming mismatches, disclosure lags, rounding conventions, and intermediate milestone failures. enterprise files; industry files; policy-domain files; local database records; web search results; financial think-tank consulting scenarios; intermediate milestone traces read-only container workspace; 1200 second timeout; process-trace scoring; execution-efficiency scoring; milestone diagnostics; source/schema/noise uncertainty caveat MEDIUM sources/21-benchmarks/dataclawbench-exploratory-financial-data-analysis-raw.md
FinSTaR Financial Time-Series Reasoning 2026-05-24 FinSTaR leaks the time-series reasoning split finance agents need: deterministic assessment tasks should compute observable drawdown, volatility, trend, support/resistance, correlation, and relative-performance quantities, while stochastic prediction tasks require explicit horizons, scenarios, OOD stock/period splits, and classical baseline comparisons. S&P 500 daily closing prices; 120-day price windows; OOD stock splits; OOD period splits; raw stock data; baseline model outputs deterministic versus stochastic task split; explicit horizon control; Compute-in-CoT check; Scenario-Aware CoT caveat; OOD stock/period splits; classical and TSFM baseline comparison MEDIUM sources/21-benchmarks/finstar-financial-time-series-reasoning-raw.md
Top Traders Unplugged / QIS 2026-05-23 Baltas leaks the QIS product boundary: systematic wrappers express macro views, but capacity, crowding, market depth, desk execution risk, slippage ownership, and client-visible product form determine whether a strategy can be implemented. QIS index definitions; macro view inputs; market depth/liquidity data; crowding indicators; execution research; client slippage data capacity checks; market-depth checks; crowding checks; slippage ownership classification; execution research MEDIUM research/13-multimodal-sources/top-traders-unplugged/2026-05-23-nick-baltas-qis-trend-following-chaos.md
WorkstreamBench Spreadsheet Workflow Agents 2026-05-21 WorkstreamBench leaks the evaluation shape for complete finance spreadsheet workstreams: financial modeling, forecasting, scenario analysis, and full spreadsheet artifact creation need separate scoring for accuracy, formulas, format, readability, revision ease, and degradation across chained calculations. spreadsheet workstreams; financial models; forecasting workbooks; scenario-analysis sheets; formula chains; format/readability criteria accuracy scoring; formula scoring; format/readability scoring; dynamic correctness checks; professional finance standard caveat MEDIUM sources/21-benchmarks/workstreambench-finance-spreadsheet-agents-raw.md
Welton Investment Partners / Trend Following 2026-05-21 Welton leaks the managed-futures evaluation gate: systematic strategies need a definable edge, definable process, explicit exceptional-risk handling, shared tracking/execution/evaluation infrastructure, stress-correlation checks, holding-power analysis, flow-driven inference tests, and edge half-life/crowding review as AI and faster data compress signal windows. liquid market prices; trend and macro signals; volatility budgets; stress-period returns; correlation matrices; flow-driven market impulses; option skew; edge-decay observations definable edge statement; definable process statement; exceptional-risk statement; stress-period correlation check; holding-power evaluation; source exclusivity/crowding check; AI-assisted decay caveat MEDIUM research/13-multimodal-sources/the-derivative/2026-05-21-patrick-welton-trend-following-process.md
Two Sigma / Agentic Systems 2026-05-20 Two Sigma’s CAIS preview leaks its agent-systems mental model: value sits in compound AI systems around models, with harnesses, tool design, schema compliance, execution-trace policy checks, retrieval/memory limits, routing, operational reports, and security defenses. execution traces; tool schemas; PDFs; spreadsheets; slide decks; retrieval memory; production deployment reports; coding-agent traces benchmarking; schema compliance; runtime safety enforcement; security/privacy review; trace-level checks; contamination controls MEDIUM sources/06-industry-verticals/two-sigma-cais-agentic-systems-raw.md
FactSet AI Foundry 2026-05-20 FactSet exposes the finance-data connector control plane: MCP/API/feed/cloud channels, helper services for context budgeting, deterministic specs, internal model gateways, eval frameworks, lineage/provenance, source links, and grounded-answer constraints for analyst research workflows. financials; models; transcripts; news; filings; expert-network calls; management meetings; sell-side research; conference notes; podcasts; YouTube; FactSet data lineage/provenance; extraction/enrichment/QA; source links; grounded-answer constraints; eval frameworks; human review; permissioned data access MEDIUM research/13-multimodal-sources/the-fund-ai-pod/2026-05-20-pat-starling-factset-ai-foundry.md
Flirting with Models / Crypto Market Making 2026-05-11 Gu leaks a crypto market-making control map: token launches face missing reference prices, thin order books, weak comparables, uncertain holder behavior, service-vs-prop incentives, venue/counterparty risk, organic-flow dependence, and adverse selection. order books; venue quality data; organic flow; token holder behavior; comparables; contract terms; inventory P&L; liquidation/ADL mechanics inventory-risk controls; wide/small initial quoting; counterparty-risk checks; manipulation/adverse-selection checks; TradFi analogy caveat MEDIUM research/13-multimodal-sources/flirting-with-models/2026-05-11-john-gu-crypto-market-making-cold-start.md
Top Traders Unplugged / Regime Adaptive Fund 2026-05-09 Dunne leaks a regime-adaptive portfolio construction surface: AI is treated as a macro data/policy distortion, while trend following becomes the systematic tactical layer over strategic equity, bond, and gold exposure, with risk budgets sized by volatility/correlation and portfolio equity beta allowed to swing from above 1 to negative. AI capex and data-center indicators; earnings concentration; labor-share data; inflation and rate-volatility data; trend-following exposures; equity/bond/gold/commodity/currency markets; volatility and correlation estimates; portfolio beta history volatility/correlation sizing; risk-budget decomposition; macro shock channel separation; equity-beta monitoring; historical regime analysis; wrapper/suitability caveat; performance-claim refusal MEDIUM research/13-multimodal-sources/top-traders-unplugged/2026-05-09-alan-dunne-ai-inflation-regime-adaptive-portfolio.md
Hedge-Fund Stock-Forecasting Overclaim Controls 2026-05-08 The hedge-fund stock-forecasting critique leaks the promotion gate for text-alpha claims: sentiment, reports, earnings calls, relationships, price tokenization, and multi-agent trading systems must survive full-cycle regimes, timestamp cutoffs, event grouping, costs, slippage, liquidity, capacity, and abstention outside validated states. financial news; social media; financial reports; earnings calls; cross-asset relationships; price series; event timestamps full-market-cycle requirement; forward-only timestamp cutoff; same-event fold grouping; P&L/Sharpe/drawdown/turnover metrics; transaction-cost/slippage/liquidity/capacity gate; regime abstention MEDIUM sources/21-benchmarks/llm-stock-forecasting-hedge-fund-perspective-raw.md
Chronos-2 Multivariate Financial Forecasting 2026-05-08 Chronos-2 leaks a time-series control that generic finance agents usually miss: related-series grouping can improve forecasts, but mixing unrelated equities and rates can degrade accuracy, so more market context is not automatically better. Magnificent-7 equity panels; U.S. Treasury rate series; combined equity/rates panels; rolling monthly windows from 2000 through 2025; forecast horizons; input-window variants rolling monthly evaluation; RMSE; MAPE; univariate baseline; related-series positive control; unrelated-series negative control; classical baseline requirement MEDIUM sources/21-benchmarks/chronos2-multivariate-financial-forecasting-raw.md
Aspect Capital 2026-05-07 Aspect leaks the managed-futures research-control stack: weak trend edges become useful only through multi-timeframe modeling, market-capacity checks, preprocessing and signal mapping, risk overlays, crisis-duration expectations, and proof that new models or markets add real diversification rather than correlated variants. trend signals; market capacity data; market correlation matrices; alternative-market price histories; portfolio risk data; model overlap diagnostics; crisis-period path data capacity and crowding checks; stress-time correlation review; market price-discovery eligibility; weak-edge aggregation test; dynamic sizing review; incremental diversification check; investor livability caveat MEDIUM research/13-multimodal-sources/the-derivative/2026-05-07-christopher-reeve-aspect-crisis-alpha-trend.md
FinSafetyBench Financial Safety Red-Team 2026-05-01 FinSafetyBench leaks the financial-agent refusal surface: finance assistants must reject or redirect requests tied to financial crimes, unethical behavior, compliance violations, implicit manipulation, bilingual English-Chinese attacks, and real-world case-grounded red-team prompts rather than only optimize helpfulness. real-world financial crime cases; financial ethics standards; English red-team prompts; Chinese red-team prompts; attack-setting variants; compliance-violation categories 14-subcategory safety taxonomy; three attack settings; direct and implicit manipulation split; jurisdiction/language caveat; prompt-defense limitation caveat; rerun requirement for exact policy corpus MEDIUM sources/21-benchmarks/finsafetybench-financial-safety-redteam-raw.md
Cambridge CCAF Financial Services AI Survey 2026-04-28 Cambridge CCAF leaks the financial-services deployment risk landscape around agentic AI: 52% of industry respondents are piloting or beyond, investment research and trading/portfolio intelligence are active use cases, but provider concentration, data quality, hallucinations, loss of oversight, market-abuse risk, explainability gaps, and weak bias monitoring remain unresolved. financial-services survey responses; investment-research workflows; trading and portfolio-intelligence workflows; foundation-model provider usage; cloud provider usage; risk and compliance responses; regulator responses self-selection bias caveat; regulator-industry gap check; provider concentration check; human oversight risk review; market-abuse risk review; explainability and bias-monitoring gate MEDIUM sources/06-industry-verticals/cambridge-ccaf-2026-ai-financial-services-raw.md
RealFin Missing-Premise Finance Reasoning 2026-04-26 RealFin leaks the abstention layer finance agents need: compare full-condition questions with condition-missing variants, require None-of-the-Above rejection, detect underdetermined premises, and treat over-commit behavior as a safety failure rather than a reasoning success. English CFA-style source questions; Chinese CPA-style source questions; full-condition prompts; condition-missing prompts; NOTA answer sets; bilingual failure traces full-condition versus missing-condition split; NOTA formulation; 15-model comparison; over-commit failure analysis; implicit-assumption detection; accuracy-not-sufficient caveat MEDIUM sources/21-benchmarks/realfin-missing-premise-finance-reasoning-raw.md
IMF / Stacklok Finance Agent Control Architecture 2026-04-24 IMF/Stacklok leak the institutional control split for finance agents: probabilistic agents can express intent and orchestrate workflows, but authorization/control and settlement/execution need deterministic identity, mandates, limits, approvals, audit trails, revocation, and tiered human-in-the-loop gates; MCP adoption is still mostly pilot or limited production in financial services. mandate documents; agent identity records; authorization logs; payment or execution intents; MCP server inventories; tool permission records; audit trails; financial-services MCP survey responses architectural separation of decision and execution; mandate-based authorization; programmable controls; audit trail; revocation path; tiered HITL; vendor-survey caveat MEDIUM sources/06-industry-verticals/investment-agentic-ai-infrastructure-2026-raw.md
FinGround Atomic Financial Claim Verifier 2026-04-23 FinGround leaks the downstream verification layer finance research agents need: hybrid retrieval over text and tables, query-complexity routing, atomic claim decomposition, financial claim-type routing, formula reconstruction, table-cell operand retrieval, paragraph/table-cell citations, and flag-only mode for high-stakes contexts. SEC filings; financial document text; financial tables; table cells; formula templates; retrieved evidence identifiers contradicted versus unverifiable verdicts; formula reconstruction; table-cell citation precision; retrieval-equalized evaluation; confidence thresholds; human verification for material numerical claims MEDIUM sources/21-benchmarks/finground-financial-hallucination-grounding-raw.md
Agentic Finance Systemic-Risk Survey 2026-04-23 The agentic-finance survey leaks the market-structure control layer: component accuracy is insufficient because adaptive multi-agent systems can change liquidity, volatility, market depth, shock recovery, price discovery, and concentration through correlated behavior. multi-agent interaction traces; market-depth metrics; liquidity and volatility series; execution-cost records; shock-recovery windows; tool-call and audit logs system-stability metrics; regime robustness; market-impact and execution-cost checks; audit trails; human review points; continuous validation MEDIUM sources/06-industry-verticals/agentic-ai-finance-survey-2026-raw.md
Deep FinResearch Report-Quality Benchmark 2026-04-22 Deep FinResearch Bench leaks the professional research-report quality gate: investment research agents must be scored on qualitative rigor, quantitative forecasting, valuation accuracy, claim credibility, verifiability, premise support, citation grounding, and analyst-signoff boundaries. financial research reports; valuation assumptions; forecast inputs; claims and citations; verifiability evidence; professional report rubrics qualitative rigor rubric; valuation accuracy check; forecasting discipline; claim-level support; verifiability check; negative-evidence caveat MEDIUM sources/21-benchmarks/deep-finresearch-bench-raw.md
CFA Institute / Aon Asset-Manager AI Governance 2026-04-22 CFA/Aon leak the allocator due-diligence control plane for manager AI: investment workflows move from data to feature engineering, modeling, explainability, decision-making, and deployment, while AI governance requires human-in-the-loop accountability, privacy/security protocols, responsible sourcing, steering committees, use-case oversight, audit trails, model tracking, version control, and performance monitoring. earnings calls; filings; sustainability reports; ESG analytics; credit-risk surveillance data; prices; economic indicators; trade information; manager AI survey responses human-in-the-loop policy; data privacy protocols; responsible data sourcing; AI steering committee oversight; information-security evaluation; audit trails; model version control; explainability and stability metrics MEDIUM sources/06-industry-verticals/cfa-aon-asset-manager-ai-governance-2026-raw.md
QRAFTI Agentic Quant Research 2026-04-20 QRAFTI leaks a concrete quant-research agent architecture: specialist factor/report/code/context agents use MCP-style tools over WRDS/CRSP/Compustat panels, RAG data catalogs, computation graphs, typed transforms, generated code audit trails, reflection, and trace-level factor replication checks. CRSP monthly stocks; Compustat annual files; WRDS identifiers; PERMNO/GVKEY/CUSIP mappings; factor libraries; RAG data catalogs; computation graphs row/month-count checks; cosine similarity to reference factors; quantile-rank comparison; portfolio-weight comparison; trace and computation-graph inspection; entitlement caveat MEDIUM sources/21-benchmarks/qrafti-agentic-quant-research-raw.md
TransXion AML Graph Benchmark 2026-04-17 TransXion leaks the AML investigation substrate: profile-rich transaction graphs require KYC context, persistent entity profiles, behavioral baselines, directed temporal multigraphs, non-template anomaly synthesis, illicit subgraphs, extreme-class-imbalance metrics, and investigator escalation packets. transaction graphs; KYC profiles; entity demographics; behavioral attributes; laundering labels; temporal transaction edges; illicit subgraphs profile-aware simulation caveat; 6:2:2 temporal split; AUC/AP/F1/PR-AUC/precision-at-K; class-imbalance gate; graph-baseline comparison; not compliance certification caveat MEDIUM sources/21-benchmarks/transxion-aml-graph-benchmark-raw.md
COMPASS-VLM Japanese Financial Documents 2026-04-16 COMPASS-VLM leaks a regional multimodal filing stack: Japanese financial-document agents combine SigLIP vision, LLM-JP text, Qwen reasoning teachers, OCR data, domain-specific QA, TAT-QA, ConvFinQA, FinQA, EDINET Bench, and Japanese public-source PDFs. Japanese financial PDFs; EDINET Bench material; Cabinet Office PDFs; Financial Services Agency PDFs; Ministry of Finance PDFs; TAT-QA; ConvFinQA; FinQA Apache-2.0 artifact check; dataset-license review; VLM loader verification; vision-tower/projector presence check; text-only comparison caveat; runtime portability caveat MEDIUM sources/21-benchmarks/compass-vlm-japanese-financial-raw.md
BlackRock / Aladdin Data Cloud 2026-04-14 BlackRock leaks the investment-data productization layer beneath AI workflows: Aladdin moves from analytics factory to intelligence factory by centralizing governed public/private-market data, Aladdin and non-Aladdin datasets, normalized/mapped investment data, whole-portfolio views, client data products, and Python/R access for thousands of business engineers. Aladdin data; non-Aladdin client data; public markets data; private markets data; normalized and mapped investment data; millions of nightly data files; on-demand reports; Python/R analysis traces governed platform boundary; data normalization and mapping; public/private data reconciliation; vendor-conference caveat; product architecture not alpha evidence; human analytics ownership MEDIUM research/13-multimodal-sources/snowflake-summit/2026-04-14-from-analytics-to-intelligence-blackrocks-journey-to-data-pr.md
TokenFactory Gemma SEC Extraction GGUF 2026-04-12 TokenFactory’s Gemma SEC extraction GGUF leaks a narrow local-structured-extraction lane: filing and contract snippets need parser contracts, canonical term-type checks, hallucination-phrase checks, exact GGUF file selection, citation reconciliation, numeric reconciliation, and license review before local SEC extraction is promoted. SEC filing snippets; financial contract extraction instructions; GGUF metadata; Gemma4 architecture metadata; JSON parser outputs; canonical term labels; hallucination phrase checks Gemma license review; exact GGUF file selection; JSON parser contract; canonical term-type metric replay; hallucination phrase metric replay; citation check; numeric reconciliation; no-Ollama policy MEDIUM sources/21-benchmarks/tokenfactory-gemma-sec-extraction-gguf-v3-raw.md
Relational Probing Financial Graph Signals 2026-04-11 Relational Probing leaks a source-to-graph research lane: small language models can encode financial entity relationships from text into induced graphs for downstream stock-trend or event-prediction models, but promotion needs point-in-time text, entity universes, graph baselines, downstream labels, leakage controls, and cost comparison against prompt extraction. point-in-time financial text; financial entity universes; language-model hidden states; induced relation graphs; stock-trend labels; event-prediction labels; co-occurrence baselines co-occurrence baseline; point-in-time split; graph-output audit; downstream prediction holdout; transaction-cost and capacity caveat; single-24GB-GPU feasibility check; no-checkpoint caveat MEDIUM sources/21-benchmarks/relational-probing-financial-prediction-qwen3-slm-raw.md
LLM Stock-Forecasting Guardrail 2026-04-10 The hedge-fund stock-forecasting review leaks the overclaim-control checklist: sentiment and news extraction are not trading systems unless they survive full-cycle regimes, timestamp cutoffs, event-level splits, leakage removal, P&L/Sharpe/drawdown/turnover/cost/slippage/liquidity/capacity evaluation, and abstention outside validated states. financial news; social media; earnings calls; financial reports; cross-asset relationships; price series; event timestamps; trading-cost data full-market-cycle caveat; strict timestamp cutoffs; same-event fold grouping; transaction-cost/slippage checks; capacity and liquidity limits; classical baseline comparison MEDIUM sources/21-benchmarks/llm-stock-forecasting-hedge-fund-perspective-raw.md
FinRegLab Consumer-Credit ML Governance 2026-04-09 FinRegLab leaks the consumer-credit model-risk checklist: underwriting agents need representative data by segment/time, target-rate analysis, reject-inference decisions, feature rationale, adverse-action reason codes, perturbation testing, out-of-time validation, swapset analysis, monitoring, vendor controls, and second-look policies. credit bureau data; cash-flow information; alternative data; thin-file/no-file borrower records; model features; adverse-action reasons; monitoring logs independent effective challenge; out-of-time validation; implementation consistency check; fairness/explainability review; perturbation reason-code test; U.S. policy date-sensitivity caveat MEDIUM sources/21-benchmarks/finreglab-ml-consumer-credit-underwriting-raw.md
FrontierFinance Financial-Modeling Agents 2026-04-07 FrontierFinance leaks the client-readiness failure mode for financial-modeling agents: agents can retrieve data and draft rationales quickly, but cell references, formula dependency chains, spreadsheet state tracking, and structural integrity still fail enough that rebuilds can be cheaper than repair. EDGAR filings; public-company data; Excel models; PPT deliverables; LBO models; DCF models; M&A models; three-statement models; lender models expert rubrics; formula dependency checks; cell-reference consistency; full-restart rate; human expert comparison; LLM-judge caveat; client-ready threshold MEDIUM sources/21-benchmarks/frontierfinance-2026-raw.md
Financial Text-and-Table RAG Retrieval Controls 2026-04-02 The text-and-table RAG benchmark leaks the retrieval layer finance agents need before synthesis: mixed SEC text/table QA must compare BM25, dense, hybrid, reranked, contextual, query-expansion, HyDE, and corrective retrieval while preserving issuer, reporting period, metric labels, and table evidence. T2-RAGBench; FinQA; ConvFinQA; TAT-DQA; 23,088 queries; 7,318 mixed text-and-table documents; SEC filings; annual reports; company/reporting-period metadata BM25 baseline; hybrid retrieval comparison; reranker trace requirement; candidate-pool depth disclosure; issuer and reporting-period metadata check; HyDE/multi-query caution for numeric tasks; citation localization scored separately MEDIUM sources/21-benchmarks/financial-text-table-rag-strategies-raw.md
Self-Driving Portfolio Architecture 2026-04-01 The self-driving portfolio paper leaks the institutional portfolio-agent control architecture: an Investment Policy Statement governs roughly 50 specialist agents, 20+ portfolio-construction methods, peer review, voting, CIO memo synthesis, forecast accountability, and script-delegated computation. Investment Policy Statement; capital-market assumptions; covariance matrices; asset-class histories; portfolio proposals; peer-review votes; forecast-vs-realized records IPS constraint gate; JSON output contracts; markdown audit artifacts; script-delegated computation; peer-review dissent preservation; L3/L4 not L5 autonomy boundary MEDIUM sources/06-industry-verticals/self-driving-portfolio-agentic-asset-management-raw.md
PolyBench Prediction-Market Forecasting 2026-04 PolyBench leaks an event-forecasting execution harness: prediction agents must use timestamp-locked market states, contemporaneous news, CLOB liquidity, abstention, structured decisions, confidence thresholds, slippage, and confidence-to-return calibration rather than answer-only accuracy. Polymarket events; point-in-time market snapshots; Central Limit Order Book states; contemporaneous news; official resolution criteria; Google News; Polymarket Gamma API no-future-data leakage gate; strict Bayesian baseline; structured JSON decision contract; confidence threshold 0.6; SKIP/abstention; CWR/APY/Sharpe; slippage and liquidity checks MEDIUM sources/21-benchmarks/polybench-prediction-market-forecasting-raw.md
One Size Fits None Suitability Benchmark 2026-04 One Size Fits None leaks the suitability-audit failure mode: LLM investment advice can collapse to one dominant feature, especially stated risk tolerance, while age, income, horizon, liquidity, product constraints, and diversification receive weak influence despite personalized-sounding rationales. 1,000 synthetic client profiles; 20 standardized investment products; Latin hypercube samples; web-search condition outputs; structured allocation outputs; surrogate model features feature concentration; HHI; weighted Jaccard similarity; diversification score; personalization score; round-number bias check MEDIUM sources/21-benchmarks/one-size-fits-none-investment-advice-raw.md
KPMG Asset Management / Private Equity Agent Survey 2026-04 KPMG’s asset-management/private-equity survey leaks the institutional agent-control posture: 39% actively deploying AI agents, only 2% orchestrating multiple agents across workflows, 77% using human validation of outputs, and 80% either preferring trusted-provider agents or disallowing autonomy for high-risk use cases. agent deployment survey responses; shared knowledge bases; unified dashboards; cross-functional workflow records; decision-routing records; data privacy and cyber-risk controls human validation of agent outputs; trusted-provider preference; high-risk autonomy prohibition; skills-gap caveat; scaling barrier caveat; survey-source caveat MEDIUM sources/06-industry-verticals/investment-agentic-ai-infrastructure-2026-raw.md
GFMA Digital-Money Capital-Markets Controls 2026-04 GFMA leaks the tokenized-settlement control surface for capital-markets agents: digital money use in securities settlement, repo/securities finance, and derivatives margin depends on atomic DvP, programmable margin calls, collateral eligibility, capital and haircut treatment, settlement finality, interoperability, privacy, and formal legal/risk/compliance sign-offs. tokenized deposit records; deposit token ledgers; wCBDC settlement records; stablecoin reserves and redemption data; repo collateral records; securities lending records; derivatives margin calls; ISO 20022 messages; DLT transaction records atomic DvP finality check; collateral eligibility and haircut review; capital treatment review; AML/KYC/CFT controls; settlement-finality legal review; privacy and data-standard checks; risk/compliance/legal/new-product sign-off MEDIUM research/21-benchmarks/raw/gfma-digital-money-capital-markets-2026.chandra.txt
Fraud-Warning Investor-Pressure Study 2026-04 The fraud-warning study leaks a safety harness for investment-advice agents: test neutral versus motivated investor framing, repeated pressure turns, complete conversation history, warning presence, warning intensity, endorsement reversal, and prompt-robustness checks before trusting fraud warnings. fraudulent investment scenarios; high-risk opportunity scenarios; AI advisory conversations; human benchmark responses; pressure-turn transcripts; judge-coded warning labels preregistration; IRB record; GPT-4o blind judge; Claude cross-judge validation; Cohen’s kappa reliability; minimum warning threshold; system-prompt robustness check MEDIUM sources/21-benchmarks/llm-fraud-warning-investor-pressure-raw.md
CFA / Aon Asset-Manager Oversight 2026-04 CFA and Aon expose the professional oversight workflow: data, feature engineering, modeling, explainability, decision-making, and deployment must be mapped to AI policy, privacy controls, sourcing rules, staff training, human review, output verification, and auditability. AI policies; data privacy protocols; responsible sourcing guidelines; staff training evidence; audit trails; RAG source-use traces; manager survey responses strict human-in-the-loop review; bias checks; explainability; audit trails; secondary validation; privacy/IP routing; deployment monitoring MEDIUM research/06-industry-verticals/cfa-aon-asset-manager-ai-governance-2026.md
BCG Global Asset Management AI-First Report 2026-04 BCG’s asset-management report leaks the target operating model vendors and managers are selling into: AI-first asset managers rebuild workflows around agentic stacks, shared model environments, orchestration layers, modular compute, unified data architecture, autonomous execution-flow prechecks, research coverage expansion, and governance embedded from the outset. BCG EXPAND asset-management benchmarking data; trade precheck records; order-management records; fill-monitoring data; investment-research coverage data; portfolio-construction and risk data; RFP and client-coverage workflows top-down ambition gate; governance-from-outset requirement; operating-model caveat; asset-class variation caveat; adoption-compression caveat; survey/projection caveat MEDIUM sources/06-industry-verticals/bcg-global-asset-management-ai-first-2026-raw.md
BCG AI-First Asset-Manager Target Model 2026-04 BCG leaks the board-level AI target model for asset managers: agentic workflows are expected to expand research coverage, client coverage, investment operations, trading flow, portfolio construction/risk, RFP response, and trade-error reduction, but the useful signal is the required platform stack of shared model environments, orchestration layers, modular compute, unified data architecture, governance, auditability, traceability, and accountability. asset-manager operating metrics; investment operations workflows; research coverage records; portfolio construction and risk data; trading flow data; RFP response artifacts; client coverage data; BCG BIFMA maturity data capacity-projection caveat; Sharpe-compression caveat; auditability; traceability; accountability; governance layer; unified data architecture requirement MEDIUM research/06-industry-verticals/bcg-global-asset-management-ai-first-2026.md
Versor Investments 2026-03-25 Agents are framed as junior researchers that read papers, implement ideas, structure conviction evals, and mine podcast/transcript sentiment. papers; podcasts; transcripts; factor ideas; model/code artifacts; eval logs structured evals; model/task fit checks MEDIUM sources/06-industry-verticals/versor-hedge-fund-huddle-agent-workflow-raw.md
Capital Fund Management 2026-03-25 CFM emphasizes scientific research discipline: statistical significance, luck-vs-skill separation, model clusters, decorrelation, and rejection of weak/correlated ideas. alternative data; strategy/model libraries; portfolio correlations; shared research/data platform statistical significance; decorrelation tests; Sharpe/risk framing; idea rejection MEDIUM research/13-multimodal-sources/aima-long-short/2026-03-25-philip-seager-cfm-systematic-multistrategy.md
CFM / AIMA Long-Short 2026-03-25 CFM leaks a systematic multi-strategy research-control stack: model-based research, real-time decision-making, portfolio construction, risk management, statistical significance, luck-vs-skill separation, strategy/model decorrelation, alternative data, shared research/data platforms, cloud storage, and compute scale. systematic strategy models; portfolio construction inputs; risk-management data; alternative data; shared research/data platforms; cloud storage; compute-scale experiment traces statistical significance; luck-vs-skill separation; strategy and model decorrelation; risk-management review; portfolio-construction gate; AI-agent overclaim caveat MEDIUM sources/06-industry-verticals/aima-long-short-cfm-quant-2026-raw.md
Agentic AI Portfolio Screening Framework 2026-03-25 The agentic screening paper leaks a three-layer portfolio research pattern: one LLM screens fundamentals, a FinBERT agent screens news, the agents intersect buy/sell decisions into a high-conviction subset, and a separate high-dimensional precision-matrix optimizer determines weights. S&P 500 universe; firm size; book-to-market ratio; twelve-month momentum; financial news articles; IBES analyst recommendations; WRDS-accessed characteristics; historical returns no-returns input to LLM-S; as-of-date firm characteristic queries; manual rule application; look-ahead bias prevention; sensible-screening condition; single-agent and human-analyst baselines; out-of-sample Sharpe evaluation; performance-claim reproduction caveat MEDIUM research/21-benchmarks/raw/agentic-ai-screening-portfolio-investment-2603.23300.chandra.txt
FinRL-X Deployment-Consistent Trading Stack 2026-03-21 FinRL-X leaks the research-to-execution control surface: quant agents should preserve one weight-centric interface across data ingestion, stock selection, allocation, timing, risk overlay, backtesting, broker-integrated paper trading, monitoring, and execution rather than treating signals, orders, and portfolio weights as interchangeable. market data; fundamental data; macro data; news text; LLM-derived sentiment signals; raw data snapshots; processed feature stores; broker execution traces; paper-trading fills shared trading calendar; persistent raw snapshots; processed-feature reproducibility; weight-vector interface contract; transaction-cost model; market-impact caveat; paper-to-live execution gap; operational resilience checks MEDIUM research/21-benchmarks/raw/finrl-x-ai-native-quant-trading-2603.21330.chandra.txt
Autonomous Factor Investing Agent Loop 2026-03-16 Beyond Prompting leaks an auditable alpha-research loop: a factor agent generates symbolic formulas from bounded primitives, executes deterministic panel construction, evaluates every candidate under a fixed metric protocol, promotes/holds/retires via explicit gates, and updates memory from prior outcomes. U.S. equity market data; raw price primitives; trading volume; liquidity variables; volatility states; candidate factor formulas; factor performance metrics; agent memory logs fixed variable universe; bounded expression complexity; strict no-look-ahead rules; pre-committed evaluation protocol; in-sample versus out-of-sample split; IC t-stat gate; long-short Sharpe gate; transaction-cost and turnover review MEDIUM research/21-benchmarks/raw/beyond-prompting-agentic-systematic-factor-investing-2603.14288.chandra.txt
FinToolBench Financial Tool Use 2026-03-09 FinToolBench leaks the financial tool-routing control layer: agents face 760 executable tools, 295 tool-required queries, live RapidAPI and AkShare interfaces, and can fail even after invocation through timeliness mismatch, intent mismatch, domain mismatch, stale data, or unsafe escalation from informational to transactional use. tool manifests; RapidAPI endpoints; AkShare interfaces; live financial APIs; tool execution logs; finance attributes for timeliness, intent, and regulatory domain Tool Invocation Rate; Conditional Execution Rate; Timeliness Mismatch Rate; Intent Mismatch Rate; Domain Mismatch Rate; permission and audit-log requirement MEDIUM sources/21-benchmarks/fintoolbench-financial-tool-use-raw.md
ODA-Fin Qwen3 Finance Post-Training 2026-03-07 ODA-Fin leaks the current finance post-training recipe: Qwen3-8B checkpoints are specialized with verified CoT SFT samples and a hard-but-verifiable GRPO subset, then scored across finance QA, sentiment, FOMC, Finova, FinEval, TaTQA, FinQA, and ConvFinQA, but every runtime artifact and training-data license must be separated before promotion. ODA-Fin-SFT-318K samples; ODA-Fin-RL-12K hard examples; Qwen3-235B-A22B-Thinking distillation outputs; FinEval; Finova; FinanceIQ; FOMC; Financial PhraseBank; Headlines; FinQA; TaTQA; ConvFinQA SFT vs RL checkpoint separation; training-data license review; artifact revision pin; held-out StateBench replay; backend-specific smoke test; reported-score caveat; cost/failure-mode metadata MEDIUM sources/21-benchmarks/oda-fin-qwen3-model-cards-raw.md
FinSheet-Bench Private-Markets Spreadsheet Controls 2026-03-07 FinSheet-Bench leaks the private-markets spreadsheet failure surface: portfolio-monitoring workbooks contain irregular layouts, fund dividers, multi-line headers, multi-sheet structures, finance conventions, row-level extraction, deterministic arithmetic, reconciliation, and abstention gates. private-equity portfolio-monitoring templates; synthetic financial spreadsheet data; multi-sheet workbooks; irregular headers; fund-boundary markers; source cells structure-discovery checks; row-level provenance; deterministic arithmetic outside the LLM; reconciliation against source cells; abstention or human-review gate MEDIUM sources/21-benchmarks/finsheet-bench-financial-spreadsheets-raw.md
EDINET-Bench Japanese Financial Statements 2026-03-05 EDINET-Bench leaks the whole-filing regional benchmark pattern: Japanese annual-report agents need EDINET API ingestion, text/table synthesis across full reports, accounting-fraud triage, earnings-direction prediction, industry classification, contamination controls, and classical model baselines. EDINET annual reports; Japanese financial statements; balance sheets; profit and loss statements; cash-flow statements; industry classifications; fraud labels logistic regression baseline; random forest baseline; ROC-AUC/MCC/accuracy; year-wise performance analysis; model-cutoff contamination caveat; auditor replacement caveat MEDIUM sources/21-benchmarks/edinet-bench-japanese-financial-statements-raw.md
FinRetrieval Connector Benchmark 2026-03 FinRetrieval leaks the connector advantage: exact financial-value retrieval depends more on structured data/API access and provenance than on generic model choice, with bundled evidence showing a large gap between structured API and web-search-only configurations. structured financial databases; filings; earnings releases; investor presentations; 4,000+ public-company source data; tool traces 500-question retrieval set; web-only vs structured-API comparison; source provenance requirement; vendor-origin caveat; internal-data reproduction requirement MEDIUM sources/21-benchmarks/finretrieval-2603.04403-raw.md
FinMCP-Bench Financial Tool Use 2026-03 FinMCP-Bench leaks the MCP reliability cliff for finance agents: 613 samples over 10 scenarios and 33 sub-scenarios use 65 real financial MCP servers, with single-tool, multi-tool, and multi-turn splits showing that multi-turn financial tool exact match can collapse to roughly 3%. 65 financial MCP servers; Qieman production-log-derived tool traces; real and synthetic user queries; single-tool samples; multi-tool samples; multi-turn samples single vs multi-tool split; multi-turn exact-match scoring; authorization caveat; parameter validation requirement; audit/cost/freshness limits; rollback requirement MEDIUM sources/21-benchmarks/finmcp-bench-financial-mcp-tool-use-raw.md
CoMind / MLE-Live Kaggle Agents 2026-02-27 CoMind leaks the modern competition-research agent loop: agents ingest timestamped Kaggle discussions, public kernels, datasets, model checkpoints, and dependency graphs; coordinators sample promising artifacts; analyzers score novelty/feasibility/effectiveness/efficiency; idea proposers maintain memory; coding agents run parallel notebook experiments; evaluators enforce official metrics. Kaggle task descriptions; competition datasets; public kernels; discussion threads; published datasets; model checkpoints; dependency graphs; submission.csv artifacts; leaderboard snapshots timestamped pre-deadline artifact access; public-dataset restriction; official metric grader; separate containers; 24-hour run budget; parallel-agent count disclosure; live leaderboard reporting caveat MEDIUM research/21-benchmarks/raw/comind-mle-live-kaggle-agents-2506.20640.chandra.txt
Chat With Traders / Systematic Backtest Reconciliation 2026-02-27 Mabe leaks the research-to-live control loop for systematic traders: migrate discretionary process in stages, automate position sizing/order staging/journaling before full autonomy, compare actual trades against the backtest, and explain missed fills, slippage, locate costs, stop logic, and bar-resolution optimism. backtest trades; actual live or paper trades; fills; slippage records; short locate costs; order journal; bar-resolution assumptions post-trade reconciliation; fill and slippage review; short-locate cost check; stop-logic caveat; bar-resolution optimism check; no-alpha-evidence caveat MEDIUM research/13-multimodal-sources/chat-with-traders/2026-02-27-dave-mabe-systematic-trading-backtested-confidence.md
FIRE Financial Intelligence Reasoning Controls 2026-02-25 FIRE leaks the business-workflow gap behind finance exams: professional certification scores do not prove real-world financial scenario competence, so agents need rubric-scored business problems across accounting, insurance, funds, futures, banking, securities, credit approval, marketing, fraud detection, and ROI-style operational contexts. CFA questions; CPA questions; FRM questions; CIA/CISA/CFP/CMA material; scenario-based business problems; closed-form reference answers; open-ended rubric answers; internal business activity traces certification-vs-scenario split; problem-wise rubrics; closed-form reference answers; business-value gap check; domain matrix coverage; internal workflow metric requirement; model-ranking staleness caveat MEDIUM sources/21-benchmarks/fire-financial-intelligence-reasoning-raw.md
Numerai 2026-02-23 Craib leaks Numerai’s crowdsourced-alpha control system: obfuscated 2,000-dimensional data produces 6,000-stock predictions, scored on subsequent 20-day residual returns and meta-model contribution, while staking, factor neutralization, MMC, crowding controls, liquidity/short-availability universe pruning, and agentic research scaffolds decide which external models can affect the fund. obfuscated equity feature matrix; 6,000-stock anonymized universe; 20-day residual returns; stake records; NMR/stablecoin staking data; meta-model predictions; Barra-style risk model factors; crowding indicators; short availability data; liquidity data; internet news for predictive LLM features data obfuscation; live-performance staking; overfitting penalty; residual correlation scoring; MMC/orthogonal-alpha gate; factor neutralization beyond Barra; crowding stress review; liquidity and short-availability filters; walk-forward cross-validation scaffold; performance-claim caveat MEDIUM research/13-multimodal-sources/flirting-with-models/2026-02-23-richard-craib-numerai-crowdsourced-alpha.md
Fin-RATE SEC Filing Tracking Benchmark 2026-02-14 Fin-RATE leaks the filing-analysis control plane behind research agents: detail reasoning, cross-entity comparison, and longitudinal company tracking fail in different ways and need separate entity, filing-period, as-of-date, retrieval, and generation diagnostics. SEC filings; company disclosures; filing types; reporting periods; entity comparison sets; retrieved contexts entity alignment check; reporting-period alignment check; as-of-date control; retrieval-versus-generation split; ground-truth-context versus RAG distinction MEDIUM sources/21-benchmarks/fin-rate-sec-filings-analytics-tracking-raw.md
Eva-4B Earnings-Call Evasion Detection 2026-02-04 Eva-4B/EvasionBench leaks a transcript-derived disclosure-quality lane: analyst-question and management-answer pairs can be classified as direct, intermediate, or fully evasive using large-scale earnings-call data, frontier-model consensus labels, human gold labels, and narrow domain classifiers. S&P Capital IQ earnings-call transcripts; 22.7 million Q&A pairs; analyst questions; management answers; 84K balanced training samples; 1K expert human-labeled gold samples; Hugging Face model-card metadata direct/intermediate/fully-evasive taxonomy; dual frontier-LLM annotation; three-judge majority vote; human inter-annotator agreement; gold evaluation set; subjective-label caveat; backend preflight requirement MEDIUM sources/21-benchmarks/eva4b-evasionbench-earnings-call-evasion-raw.md
FinMMEval Multilingual Multimodal Finance 2026-02 FinMMEval leaks the global finance evaluation surface: research agents need multilingual exam QA, PolyFiQA evidence grounding, SEC 10-K/10-Q excerpts, multilingual news, market reports, regulatory documents, investor communications, and a separate Buy/Hold/Sell decision sandbox with risk metrics. SEC filing excerpts; 10-K/10-Q excerpts; multilingual news; market reports; regulatory documents; investor communications; BTC and TSLA daily market contexts expert-annotated PolyFiQA split; inter-annotator agreement; evidence-supported QA; fixed-cadence decision evaluation; Sharpe and drawdown reporting; live-environment caveat MEDIUM sources/21-benchmarks/finmmeval-multilingual-multimodal-finance-raw.md
QuantEval Strategy-Code Benchmark 2026-01-16 QuantEval leaks a deterministic backtest-harness pattern: 1,575 samples span knowledge QA, quantitative reasoning, and 60 strategy-coding tasks, with CTA-style backtesting, fixed asset universe, cost model, standardized execution/risk constraints, and metric-deviation scoring. U.S. ETF histories; large-cap stock histories; CTA backtest configuration; strategy-code samples; expert annotations; domain-aligned training samples annualized return MAE; maximum drawdown MAE; Sharpe ratio MAE; RetDraw MAE; leakage prevention; expert review; deduplication MEDIUM sources/21-benchmarks/quanteval-quantitative-finance-tasks-raw.md
Trading-R1 Alpha-R1 Trade-R1 Reasoning Controls 2026-01-08 The R1 trading-reasoning cluster leaks the systematic-research reward problem: trading agents need evidence-grounded theses, point-in-time data, regime-aware factor activation, noisy-reward safeguards, triangular evidence/reasoning/decision consistency, and cost/capacity/regime checks before any profitable-looking result gets credit. heterogeneous financial data; market news; trading corpora; factor definitions; point-in-time market context; retrieved evidence point-in-time data gate; evidence-reasoning-decision consistency; semantic reward strategy; reward-hacking guard; transaction-cost and capacity check; runtime artifact caveat MEDIUM sources/21-benchmarks/trading-alpha-r1-reasoning-rl-raw.md
LendNova Credit Bureau Text Model 2026-01-01 LendNova leaks the raw credit-bureau text workflow: a language model turns jargon-heavy commercial bureau records into a credit story, joins application profile, credit activity, demand, repayment, stress, and static layers, then tests default and charge-off prediction against holdout, out-of-time, and industrial bureau baselines. anonymized commercial credit bureau records; approximately one million Canadian customers; raw credit bureau text; credit application profiles; credit activity; credit demand; repayment outcomes; financial stress fields; charge-off and delinquency labels train/validation/holdout split; out-of-time deployment simulation; February 1 2018 cutoff; September 2018 application anchor; industrial bureau baseline comparison; class-imbalance handling; confidential-data caveat MEDIUM sources/21-benchmarks/lendnova-credit-risk-language-model-raw.md
FCMBench Credit Multimodal Evidence 2026-01-01 FCMBench leaks the multimodal credit-document validation stack: financial agents must extract exact fields from privacy-safe certificates, connect IDs, addresses, bank accounts, dates, and income evidence across documents, and survive blur, reflection, occlusion, and capture-artifact perturbations before any credit workflow claim is trusted. 26 certificate types; 5,198 privacy-compliant images; 13,806 VQA samples; Chinese and English certificates; synthetic-to-physical document photos; IDs; addresses; bank accounts; income and date fields exact-match atomic field scoring; perception/reasoning/robustness task split; macro F1; normalized robustness ratios; blur/reflection/occlusion perturbation tests; semantic-paraphrase rejection; privacy-safe synthetic-document caveat MEDIUM sources/21-benchmarks/fcmbench-financial-credit-multimodal-raw.md
ThoughtLab / Grant Thornton Investment Firms 2025-12 The AI-powered investment-firm survey leaks the asset/wealth/hedge-fund operating backlog: data quality/access, roadmap gaps, skills, system integration, transparency, risk, regulation, and human oversight dominate adoption friction while current use clusters around document summarization, regulatory/tax monitoring, risk/fraud protection, portfolio support, trade settlement, code, custody, and records. advisor CRM notes; customer surveys; complex documents; regulatory and tax monitoring sources; risk and fraud data; portfolio-support data; trade-settlement records; custody and financial-statement records data-quality and access checks; third-party due diligence; testing and auditing; risk monitoring; regulatory dialogue; human oversight requirement MEDIUM sources/06-industry-verticals/grant-thornton-thoughtlab-ai-powered-investment-firm-2026-raw.md
Grant Thornton / ThoughtLab AI-Powered Investment Firm 2025-12 Grant Thornton/ThoughtLab leak the investment-firm maturity gap: leaders differ from starters through aligned AI/business strategy, AI-ready IT and data infrastructure, governance/risk/regulatory frameworks, future-of-work preparation, and agentic-era redesign, while blockers concentrate in conservative culture, data quality, roadmap gaps, skills, and employee resistance. 500-firm investment-management survey responses; BCG BIFMA maturity scores; front-office use-case adoption; middle/back-office use-case adoption; AI payback reports; data-quality and culture barrier responses sponsor-framing caveat; maturity-score segmentation; payback distribution review; culture/data-quality barrier tracking; governance/risk framework check; AI roadmap requirement MEDIUM research/06-industry-verticals/grant-thornton-thoughtlab-ai-powered-investment-firm-2026.md
FinFRE-RAG Structured Fraud Retrieval 2025-12 FinFRE-RAG leaks a privacy-sensitive structured-fraud workflow: feature importance reduction, normalized numeric features, categorical filtering, vectorized transactions, top-N similar exemplar retrieval, label-aware examples, natural-language serialization, risk scores, and analyst escalation gates. fraud transaction datasets; numeric transaction features; categorical transaction features; retrieved similar transactions; label-aware exemplars; reviewer notes Random Forest feature-importance baseline; F1/MCC/precision/recall; class-imbalance gate; human escalation; prompt-injection stress test; retrieval-poisoning stress test MEDIUM sources/21-benchmarks/finfre-rag-structured-fraud-detection-raw.md
CNFinBench Finance-Agent Safety 2025-12 CNFinBench leaks a high-privilege finance-agent safety harness: execution chains span requirement parsing, step planning, tool invocation, result verification, DB/API operations, long conversations, role adaptation, self-reflection, and adversarial multi-turn compliance attacks. financial system APIs; databases; A-share annual reports; loan applications; text segments; simulated chatbot logs; adversarial prompts HICS; risk-type deductions; multi-turn consistency tracking; severity-adjusted penalties; least-privilege checks; minimum-disclosure checks; dynamic leaderboard caveat MEDIUM sources/21-benchmarks/cnfinbench-finance-agent-safety-compliance-raw.md
Risknet Quantcast / UBS Finance-Native NN 2025-11-18 Risk net exposes a finance-native model-control lane: neural architectures should preserve asset-pricing semantics, no-arbitrage, numeraires, time/depth interpretation, Markovian activations, QIS/XVA stress states, and view-vs-market distinctions. market-implied structures; QIS definitions; XVA/hedging scenarios; investor views; portfolio loss states; pricing paths no-arbitrage checks; martingale/discounting assumptions; measure/time semantics; economic hallucination detection; stress-state localization MEDIUM research/13-multimodal-sources/risk-net-quantcast/2025-11-18-stefano-iabichino-finance-native-neural-networks.md
FinCRITICAL-ED Financial Fact OCR 2025-11-14 FinCRITICAL-ED leaks the critical-field ingestion layer: finance document pipelines must preserve numeric facts, temporal facts, monetary units, entities, financial concepts, values, signs, decimals, scale, currency symbols, row-column alignment, and table headers before any downstream analysis is trusted. financial statements; SEC filings; tax forms; securities transaction records; financial legal documents; page-level PDFs; table headers finance-domain annotators; inter-annotator agreement; senior financial expert review; version-controlled ground truth; Financial Fact Accuracy; deterministic exact-match override MEDIUM sources/21-benchmarks/fincriticaled-financial-fact-ocr-raw.md
Omega2 Corporate Credit Scoring 2025-11 Omega2 leaks the corporate-credit agent validation stack: LLM orchestration is wrapped around structured financial data, vector retrieval, knowledge graphs, typed ontologies, API layers, gradient-boosting models, temporal rolling-window validation, train-only transformations, calibration, SHAP, stability monitoring, and leakage exclusion. Kaggle corporate credit rating dataset; corporate credit ratings; financial ratios; reporting periods; rating-agency labels; knowledge graph entities; vector database records k=5 temporal folds; train-only imputation; future leakage exclusion; contemporaneous feature derivation; Brier and calibration metrics; Population Stability Index; SHAP and permutation importance; train-AUC overfit warning MEDIUM sources/21-benchmarks/omega2-corporate-credit-scoring-raw.md
Hudson River Trading 2025-10-31 HRT uses low-level market events for short-horizon prediction; neural nets produce plans/predictions while audited risk-checked layers act in markets. order-book events; market event data; petabyte-scale storage; GPU training; model serving and routing systems release checks; intraday sanity checks; numerical stability; regulatory trust; LLM contamination checks MEDIUM research/13-multimodal-sources/odd-lots/2025-10-31-how-hudson-river-trading-actually-uses-ai.md
MultiFinBen Multilingual Multimodal Finance Benchmark 2025-10-11 MultiFinBen leaks the cross-border multimodal evaluation surface: finance agents need multilingual text, scanned-document OCR, chart/table vision, financial audio ASR, earnings-call summarization, cross-lingual evidence integration, difficulty-aware task selection, and modality-balanced reporting. English finance sources; Chinese finance sources; Japanese finance sources; Spanish finance sources; Greek finance sources; scanned financial documents; earnings-call audio difficulty-aware selection; modality-balanced score reporting; manual factuality assessment; OCR omission checks; numeric misinterpretation checks; live leaderboard caveat MEDIUM sources/21-benchmarks/multifinben-multilingual-multimodal-financial-benchmark-raw.md
FinAgentBench Filing-Retrieval Agents 2025-10-03 FinAgentBench leaks the retrieval plumbing behind finance research assistants: before answer generation, an agent must select document type, localize the relevant passage or chunk, understand 10-K/10-Q/8-K/earnings-call/DEF 14A distinctions, and control context length, cost, and latency. 10-K filings; 10-Q filings; 8-K filings; earnings-call transcripts; DEF 14A proxy statements; S&P 500 disclosures document-type selection check; passage/chunk localization check; citation precision independent of fluency; count/version caveat; not full research-agent caveat MEDIUM sources/21-benchmarks/finagentbench-agentic-retrieval-raw.md
FinSearchComp Financial Search Reasoning 2025-09-16 FinSearchComp leaks the expert financial-search failure surface: agents must resolve official filings, disclosures, earnings releases, footnotes, GAAP/non-GAAP distinctions, corporate actions, fiscal calendars, renamed entities, partial/conflicting evidence, regional source coverage, and real-time data without stale or shallow search. official filings; corporate disclosures; earnings releases; footnotes; GAAP and non-GAAP metrics; corporate-action records; real-time data; regional financial sources date freshness scoring; source provenance; timezone and fiscal-calendar checks; unit and currency conversion checks; regional coverage audit; expert-human gap caveat MEDIUM sources/21-benchmarks/finsearchcomp-financial-search-reasoning-raw.md
AIMA Alternative-Investment AI Leaders 2025-09-16 AIMA leaks the allocator-facing governance surface: fund managers report broad GenAI use, rising front-office expectations, investor DDQ questions, policy/training/privacy/explainability concerns, and a need to avoid AI washing. fund-manager survey responses; institutional-investor interviews; DDQs; AI policies; training records; privacy and model-governance artifacts DDQ evidence; governance foundations; risk/limitation awareness; responsible training; right-model/right-use-case matching; AI-washing caveat MEDIUM sources/06-industry-verticals/aima-charting-course-ai-leaders-alternative-investment-2025-raw.md
FinRAGBench-V Visual Citation RAG 2025-09-09 FinRAGBench-V leaks the visual-citation contract for finance RAG: reports, statements, prospectuses, and research PDFs must preserve charts, tables, page layout, visual blocks, page-level citations, block-level citations, box bounds, and image crops instead of flattening all evidence to text. financial reports; financial statements; prospectuses; research reports; chart blocks; table blocks; page-layout evidence retrieval/generation/citation split; page-level citation metrics; block-level citation metrics; box-bounding evaluation; image-cropping evaluation; no local reproduction caveat MEDIUM sources/21-benchmarks/finragbench-v-visual-citation-rag-raw.md
INSEVA Chinese Insurance Benchmark 2025-09-04 INSEVA leaks a regulated-finance evaluation lane that broad finance QA misses: insurance agents need jurisdiction/language scope, product terminology, coverage and exclusion reasoning, claims-process knowledge, marketing-growth and service-dialogue safety, and explicit faithfulness-vs-completeness scoring before any policy or customer workflow is trusted. Chinese insurance professional exams; regulatory standards; real-world insurance business data; policy/product terminology; coverage and exclusion examples; claims-process dialogue; multi-turn service records jurisdiction/language scope check; faithfulness before completeness; multi-turn dialogue scoring; procedural-knowledge evaluation; unsupported coverage-statement rejection; human escalation for missing policy/regulatory evidence; leaderboard-stability caveat MEDIUM sources/21-benchmarks/inseva-chinese-insurance-benchmark-raw.md
FinCast Financial Time-Series Foundation Model 2025-08-27 FinCast leaks a sparse-MoE financial forecasting stack: regime shifts, domain diversity, variable frequencies, quantile loss, learnable frequency embeddings, and expert routing need to be evaluated separately from investable utility. 20B+ claimed time points; stocks; crypto; forex; futures; macroeconomic indicators; FinCast-Paper-test reproduction dataset; HF model card; GitHub code locked point-in-time data; classical/econometric baselines; supervised baseline comparison; quantile calibration; tail-risk checks; transaction-cost and slippage separation; exact HF/GitHub revision pin MEDIUM sources/21-benchmarks/fincast-financial-time-series-foundation-model-raw.md
FinTMMBench Temporal Multimodal Finance RAG 2025-08-03 FinTMMBench leaks the time-windowed multimodal retrieval problem: finance agents must route across tables, news, daily prices, and technical charts while preserving daily, weekly, monthly, quarterly, and annual windows, as-of periods, data vintages, and modality-specific failure attribution. NASDAQ-100 company data; financial tables; news articles; daily stock prices; technical charts; time-window metadata daily/weekly/monthly/quarterly/annual window checks; as-of period statement; data-vintage control; modality-routing failure split; temporal retrieval failure split; current/historical leakage prevention MEDIUM sources/21-benchmarks/fintmmbench-temporal-multimodal-rag-raw.md
Fin-PRM Financial Process Reward Model 2025-08 Fin-PRM leaks the process-supervision layer for finance reasoning: step-level traces, trajectory-level rewards, expert analyses, financial knowledge-base extraction, offline trace selection, Best-of-N inference, GRPO reward shaping, and test-time scaling can score intermediate reasoning rather than only final answers. CFLUE; CFLUE-Neo triplets; expert financial analyses; financial knowledge base; DeepSeek-R1 traces; Qwen auxiliary judgments step-level reward supervision; trajectory-level reward supervision; binary cross-entropy objective; finance-specific concept fidelity; runtime/deployment caveat; no trading proof caveat MEDIUM sources/21-benchmarks/fin-prm-process-reward-model-raw.md
Acadian Asset Management 2025-07-28 AI modules feed an existing systematic pipeline: pattern detection, analyst-bias correction, earnings-call Q&A extraction, expected-return forecasts, portfolio construction, and order routing. earnings-call Q&A; analyst estimates; newsflow; supplier/customer context; peer fundamentals; technical patterns domain-knowledge review; modular signal validation; out-of-sample discipline; transaction-cost constraints MEDIUM research/13-multimodal-sources/fear-greed-interviews/2025-07-28-acadian-zhe-chen-ai-investment-process.md
FinChart-Bench Financial Chart VLM 2025-07-14 FinChart-Bench leaks the visual-source ingestion gap for investment agents: earnings decks, Bloomberg charts, corporate slides, and financial websites require chart-type recognition, spatial reasoning, exact-answer discipline, manual ambiguity filtering, and skepticism toward VLMs as automated judges. financial chart images; company presentation slides; Bloomberg News charts; official corporate websites; earnings deck visuals; QA labels manual chart filtering; two rounds of human evaluation; QA correction; unambiguous single-token answers; exact-match scoring; VLM-judge unreliability caveat MEDIUM sources/21-benchmarks/finchart-bench-financial-chart-vlm-raw.md
XFinBench Complex Finance Reasoning Controls 2025-07 XFinBench leaks the complex-finance evaluation gap: research agents need to solve graduate-level terminology, temporal reasoning, future forecasting, scenario planning, numerical modeling, chart/curve interpretation, long-table context, and exact financial calculations rather than only retrieve prose. graduate-level finance textbooks; finance terminology bank; long financial tables; visual chart context; time-series-like finance context; financial calculation prompts; human-expert baseline answers human-expert baseline; exact numerical tolerance; temporal-reasoning error analysis; scenario-planning caveat; curve-position/intersection checks; retrieval-overreliance review; dated leaderboard caveat MEDIUM sources/21-benchmarks/xfinbench-complex-financial-problem-solving-raw.md
Balyasny Asset Management 2025-06-21 Central quant research supports PM pods through factor models, hedging, portfolio advisory, performance/risk/drawdown support, and execution research. factor libraries; pod/PM history; risk exposures; execution data; performance and drawdown logs pervasiveness; persistence; interpretability; diversified reproduction; point-in-time factor definitions MEDIUM research/13-multimodal-sources/odd-lots/2025-06-21-giuseppe-paleologo-quant-investing-multi-strat.md
FinanceReasoning Numerical-Control Benchmark 2025-06 FinanceReasoning leaks the numeric-control layer for financial research agents: answer quality depends on disambiguated questions, corrected ground truth, executable Python solutions, unit/sign/percentage precision, strict error bands, and a reviewed function library rather than fluent financial explanations. CodeFinQA; CodeTAT-QA; FinCode; FinanceMath; Investopedia articles; financial function library; expert-reviewed Python solutions; hybrid table/text contexts answer correction audit; question disambiguation audit; 0.2 percent error margin; unit/sign/decimal-place checks; expert review; program-executed numerical results; reasoner-versus-programmer comparison MEDIUM research/21-benchmarks/raw/finance-reasoning-2506.05828.chandra.txt
FinAR-Bench Fundamental Analysis Controls 2025-06 FinAR-Bench leaks the control surface for fundamental-analysis report agents: report generation must decompose into statement-data extraction, indicator computation, and logical reasoning over observed financial conditions rather than relying on fluent prose. Shanghai Stock Exchange annual reports; XBRL financial statements; PDF financial-statement pages; income statements; balance sheets; cash-flow statements; financial indicators; reasoning-condition tables ground-truth XBRL labels; Markdown table normalization; Hungarian assignment for header matching; relative numerical error scoring; PDF extraction backend comparison; LLM-judge caveat for reasoning; zero-tolerance precision boundary MEDIUM research/21-benchmarks/raw/finar-bench-fundamental-analysis-2506.07315.chandra.txt
BizFinBench Chinese Business-Finance Reasoning 2025-05-26 BizFinBench leaks a local-market business-finance task taxonomy: Chinese finance agents need numerical calculation, reasoning, information extraction, prediction recognition, knowledge QA, anomalous-event attribution, time reasoning, tool usage, entity recognition, and dimension-separated judging. Chinese financial queries; financial tables; business-context prompts; financial entities; event descriptions; time-reasoning tasks IteraJudge caveat; dimension-decoupled grading; time-reasoning emphasis; model-ranking staleness caveat; local-market scope boundary; no alpha proof caveat MEDIUM sources/21-benchmarks/bizfinbench-business-driven-financial-benchmark-raw.md
Finance Agent Benchmark Analyst Research 2025-05-20 Finance Agent Benchmark leaks the analyst-research harness: 537 expert-authored questions across nine financial task categories use recent documents, public/private/test splits, Google Search, EDGAR access, filing search, ParseHTML, RetrieveInformation, context-window management, up to 50 loop iterations, and cost/time scoring. SEC filings; EDGAR database; recent financial documents; public validation set; private validation set; test set; search results; retrieved HTML rubric-based evaluation; manual author review; contradiction rubric; LLM-as-judge caveat; public/private/test split; human expert baseline MEDIUM sources/21-benchmarks/finance-agent-benchmark-analyst-research-raw.md
ISDA Collateral Liquidity Agentic AI Controls 2025-05 ISDA’s collateral-liquidity paper leaks the derivatives collateral-agent control surface: agentic AI can recommend collateral, reconcile valuation disputes, forecast funding stress, and rebalance collateral portfolios only when eligibility schedules, ISDA terms, CDM data standards, DLT/tokenized-asset records, stress-event protocols, and legal/operational gates are explicit. margin calls; collateral inventories; eligibility schedules; haircuts and concentration limits; ISDA agreement terms; valuation models; repo and financing market data; CDM records; tokenized collateral records; stress-event protocols; liquidity forecasts collateral eligibility check; haircut and concentration-limit check; legal enforceability review; operational settlement check; stress-protocol consent; CDM data-standard contract; liquidity stress testing; human risk/legal/operations approval; no autonomous collateral movement MEDIUM research/21-benchmarks/raw/isda-collateral-liquidity-efficiency-derivatives-2025.chandra.txt
Risknet Quantcast / Term-Structure Autoencoders 2025-03-27 The term-structure episode leaks a finance-specific model gate: autoencoders can learn yield-curve manifolds, but pricing/hedging requires real-world vs risk-neutral measure discipline, no-arbitrage constraints, HJM-style conditions, scenario realism, and production-readiness refusal. yield curves; latent curve factors; pricing paths; XVA scenarios; asset-liability simulations; potential future exposure; expected shortfall no-arbitrage gate; risk-neutral constraint; measure discipline; overfitting control; scenario realism; data-quality caveat MEDIUM research/13-multimodal-sources/risk-net-quantcast/2025-03-27-sokol-lyashenko-mercurio-autoencoding-term-structure.md
Won / FINKRX Korean Finance LLM 2025-03-24 Won / FINKRX leaks the Korean local-market evaluation pattern: finance agents need domestic company analysis, financial markets, stock-price prediction tasks, financial-agent tasks, open-ended QA, Korean/English reasoning format controls, and an 80K-plus instruction dataset lineage. Korean finance benchmark submissions; Won-Instruct rows; domestic company-analysis prompts; financial-market questions; stock-price prediction questions; financial-agent tasks license check; artifact revision pin; tokenizer/runtime check; Korean finance eval rerun; local-market scope boundary; U.S. generalization caveat MEDIUM sources/21-benchmarks/won-korean-finance-llm-raw.md
Man Group / AHL 2025-02-26 ManGPT provides safe/audited LLM access; Alpha/Rosa and ArcticDB-style data infrastructure are the research substrate; Man Numeric uses ML in a material share of signal models. earnings-call transcripts; Reddit/meme-stock monitoring; level-three market data; sector-specific data; Python/proprietary code audited LLM access; source provenance; human-acceptable trade rationale MEDIUM research/13-multimodal-sources/tech-talks-daily/2025-02-26-man-group-data-driven-future.md
FinE5 / FinMTEB Finance Embedding Controls 2025-02-17 FinE5 leaks the finance-RAG retrieval gate: filings, annual reports, earnings-call transcripts, ESG reports, regulatory filings, and source-card retrieval need finance-domain embedding benchmarks, hybrid BM25+dense controls, citation localization, as-of-date checks, permissions, latency, cost, and license review before a research agent can trust retrieved context. financial news articles; corporate annual reports; 10-K filings; 10-Q filings; ESG reports; regulatory filings; earnings-call transcripts; finance source cards FinMTEB benchmark comparison; BM25 baseline; hybrid BM25+dense control; training/test overlap check; CC-BY-NC-ND license gate; gated-access review; source/page precision; latency and cost measurement MEDIUM sources/21-benchmarks/fine5-finance-embedding-model-card-raw.md
PHANTOM Long-Context Filing Hallucination Benchmark 2025 PHANTOM leaks the long-context failure surface for filing QA: SEC filing answers must be tested for source faithfulness across faithful, contradicted, misrepresented, and unverifiable answers while varying context length, evidence placement, filing type, and lost-in-the-middle sensitivity. EDGAR SEC filings; 10-K filings; 8-K filings; 497K filings; DEF 14A filings; long-context filing spans; hallucination labels human validation; faithful/contradicted/misrepresented/unverifiable labels; context-length stratification; evidence-placement stratification; filing-type stratification; not extrinsic factuality caveat MEDIUM sources/21-benchmarks/phantom-financial-long-context-hallucination-raw.md
Bridgewater Associates 2024-12 AIA Labs decomposes investment research into causal maps, data finding, coding, charts, critique, replanning, and subagent supervision. Bridgewater databases; causal maps; approved prior plans; research reports; generated code and charts causal reasoning over correlation; diagnosable oversight; approved-plan retrieval; editable code review MEDIUM research/13-multimodal-sources/aws-reinvent/2024-12-bridgewater-aia-labs-financial-services.md
MLE-Bench / Kaggle Agent Harness 2024-10 MLE-Bench leaks the competition-agent evaluation harness behind Kaggle-style quant replay: package datasets, generated experiment code, submission.csv artifacts, private leaderboard scoring, internet/hardware/runtime disclosures, and scaffold comparisons so agent performance is not confused with public-leaderboard overfitting. Kaggle competition datasets; generated code; submission.csv files; private leaderboard labels; runtime and hardware metadata; scaffold configuration private leaderboard preference; internet access disclosure; hardware disclosure; runtime reporting; scaffold reporting; historical score caveat MEDIUM sources/21-benchmarks/mle-bench-kaggle-agent-harness-raw.md
Open-FinLLMs Multimodal Finance Corpus 2024-08 Open-FinLLMs leaks the historical multimodal finance pretraining recipe: financial agents need text, tabular, time-series, and chart coverage, with separate corpus, instruction, and multimodal-tuning ledgers rather than a single finance-chat score. 52B-token financial corpus; financial reports; papers; market data; news and social media; historical stock prices; SEC filings; 573K financial instructions; 1.43M multimodal tuning pairs dataset-mixture documentation; financial-versus-general corpus ratio; task-family split; zero/few-shot/supervised separation; multimodal task separation; older-base-model caveat; current-baseline rerun requirement MEDIUM research/21-benchmarks/raw/open-finllms-multimodal-finance-2408.11878.chandra.txt
Risknet Quantcast / Market Integrity 2024-07-24 Cartea leaks the trading-agent safety surface: adaptive algorithms that update rules from market feedback can learn signaling, reward/punishment, clone interaction, initial-condition dependence, memory-based coordination, or signal rigging without explicit misconduct instructions. order-size patterns; quote concentration; market-maker interaction data; training data overlap; agent memory traces; sandbox clone interactions; order-book pressure signals concentration checks; reward/punishment analysis; memory and initial-condition review; clone-competition sandbox; legal/compliance escalation MEDIUM research/13-multimodal-sources/risk-net-quantcast/2024-07-24-alvaro-cartea-algorithmic-collusion-market-integrity.md
Senate Hedge-Fund AI/ML Review 2024-06 The Senate report exposes the governance failure modes across Citadel, Renaissance, Bridgewater, AI Capital, Numerai, and WorldQuant: inconsistent AI/ML terminology, varied human-review points, uneven testing cadence, high-level client disclosures, and need for versioning/audit trails. committee briefings; hedge-fund AI/ML descriptions; backtest evidence; audit trails; client disclosures; regulator input common definitions; risk assessments; standardized audits; version-control frameworks; trade audit trails; overfitting skepticism MEDIUM sources/06-industry-verticals/us-senate-hedge-funds-ai-ml-2024-raw.md
ChatGPT News-Drift Trading Friction Study 2024-05 The ChatGPT news-forecasting study leaks the implementation boundary for text-alpha agents: headline interpretation can align with immediate market reaction and short post-announcement drift, but usefulness is dominated by timestamping, overnight versus intraday entry, small-stock effects, negative-news asymmetry, turnover, transaction costs, liquidity, and partial rebalancing. firm-specific news headlines; GPT-4 news scores; overnight news timestamps; intraday news timestamps; opening and closing prices; firm and date fixed effects; RavenPack sentiment scores; market-cap filters strict news-release timestamping; overnight and intraday split; firm/date fixed effects; double-clustered standard errors; 5/10/20 bps round-trip cost scenarios; turnover and partial-rebalancing tests; small and illiquid stock exclusion MEDIUM research/21-benchmarks/raw/chatgpt-forecast-stock-price-movements-2304.07619.chandra.chunk_05-ChatGPT-and-Market-Information-Processing.txt
CALM Credit and Risk Assessment LLM 2024-02-18 CALM leaks an older but concrete credit/risk agent workbench: combine 9 datasets, 14K samples, 45K-plus instruction examples, tabular-to-text prompts, credit scoring, fraud detection, distress identification, claim analysis, fairness metrics, and expert-system baselines. German credit data; Australian credit data; Lending Club 2007-2018; Credit Card Fraud; ccFraud; Polish bankruptcy data; Taiwan Economic Journal data; PortoSeguro; Travel Insurance expert-system baseline; accuracy/F1/MCC; miss-rate reporting; Equal Opportunity Difference; Average Odds Difference; Disparate Impact; sensitive-attribute audit; class-imbalance caveat MEDIUM sources/21-benchmarks/calm-credit-risk-assessment-raw.md
SiC Tech / Brevan Howard SIG Backtesting 2023-12-11 Ren leaks the quant research operating stack: multi-asset and derivatives backtesting must model instrument/strategy/portfolio abstractions, convex payoff structuring, per-unit versus bps execution costs, exchange calendars, close auctions, half-days, FX/ETF timing, one-minute and irregular intraday bars, alternative-data entity mapping, and LLM tool APIs as clients with small tool surfaces, JSON schemas, low-latency orchestration, and human researcher-as-director authority. instrument master; prices and derivative pricing inputs; execution-cost schedules; exchange calendars; close auctions; holiday and half-day schedules; FX intraday snapshots; ETF NAV and underlying baskets; alternative time series; unstructured text/image data; OpenAPI specs; JSON tool responses; ticker lookup services transaction-cost model by instrument; calendar and half-day checks; point-in-time execution timing; entity mapping; unstructured-to-signal feedback loop; API/tool-count minimization; JSON schema validation; query privacy and enterprise/private-cloud controls; assumption surfacing and human preference capture; P&L and strategy-assumption review MEDIUM research/13-multimodal-sources/flirting-with-models/2023-12-11-bin-ren-text2quant-llm-backtesting-engine.md
Curious Quant / GenAI Quant Process 2023-06-29 Kollo’s GenAI leak is not alpha but process: richer NLP for topic delineation, semantic association graphs, investor-attention monitoring, thematic classification, financial-statement footnote review, and client/product explanation after quant validation. timestamped news; transcripts; filings; financial-statement footnotes; theme corpora; client/product narratives one-history warning; correlation-vs-causality separation; quant evidence before narrative; no alpha overclaim MEDIUM research/13-multimodal-sources/curious-quant/2023-06-29-generative-ai-quant-investment-process.md
Acadia Quaternion / Derivatives Model Validation 2023-06-12 Kienitz leaks the derivatives model-validation boundary for ML hedging: model-agnostic data-driven delta hedges, GMM conditional distributions, closed-form deltas/prices for difficult stochastic models, clean-data and outlier assumptions, synthetic-data caveats, regulated IMM explainability/stability gates, and AI-assisted vendor/database/XML integration that remains under human quant and model-risk control. option underlyings; riskless rates; Heston/Bates/rough-volatility paths; Gaussian mixture parameters; Monte Carlo outputs; synthetic VAE/GAN data; real market data; trade XML; vendor database feeds; QuantLib and ORE model code clean-data assumption disclosure; outlier review; Monte Carlo method review; distribution-cutoff caveat; explainability and stability checks; regulated IMM/model-risk review; source-code and documentation inspection; peer-review challenge; human quant validation MEDIUM research/13-multimodal-sources/quantspeak/2023-06-12-jorg-kienitz-ml-hedging-model-validation.md
Winton 2023-03-16 Winton leaks durable failure-mode controls: ML helps most where data is rich, slower strategies often need interpretability and simplicity, text/report-scale backtests can require ML, and selection bias is an organizational risk. high-volume market data; report-scale text corpora; backtest artifacts; idea-to-live research infrastructure selection-bias awareness; interpretability preference; data-richness threshold; live-implementation skepticism MEDIUM sources/06-industry-verticals/buy-side-quant-ai-practitioner-signals-2026-raw.md
HRT First-Party Benchmark Discipline 2022-05-16 HRT’s first-party articles leak the applied-trading negative-control stack: separate prediction, optimization, and execution; reject leaderboard-only ML claims; score simplicity, reproducibility, generality, data relevance, uniqueness, lookahead, sample size, noise, provenance, and backtest-to-live consistency. market data; alternative datasets; benchmark papers; random-seed and hyperparameter traces; data provenance records; backtest/live comparison artifacts prediction/optimization/execution separation; simplicity check; reproducibility check; generality check; lookahead/leakage check; sample-size and noise review MEDIUM sources/06-industry-verticals/hrt-ai-benchmark-data-source-raw.md
Acadian Asset Management / Trading Desk 2021-06-28 Acadian exposes the execution-control layer beneath systematic signals: traders collaborate with PM/research, understand order intent and urgency, run OMS/EMS analytics and TCA, manage venue toxicity and information leakage, and feed execution results back into research and portfolio construction. OMS/EMS analytics; market data; broker chats; blotter context; TCA; venue quality data; risk models; post-trade analysis human last-line-of-defense; asset-class-specific execution policy; information-leakage controls; market-impact checks; post-trade TCA MEDIUM research/13-multimodal-sources/acadian-behind-the-signals/2021-06-28-evolution-buy-side-trading.md
Acadian Asset Management / ML Research 2021-03-22 Acadian’s ML note leaks a disciplined systematic-research recipe: hypothesis-led feature selection, raw filings/fundamentals/text ingestion, nonlinear interaction discovery, random-forest-style baselines, train/validation/out-of-sample splits, interpretability, and humility when data disproves the story. raw filings; fundamentals; NLP/text metrics; company-linkage data; risk/event data; out-of-sample test sets train/validation/out-of-sample split; capacity/depth limits; simple baseline comparison; overfitting guardrails; Shapley/LIME-style interpretability MEDIUM research/13-multimodal-sources/acadian-behind-the-signals/2021-03-22-machine-learning-quant-investing.md
Curious Quant / RL Finance 2020-07-16 Halperin leaks the RL control framing: finance is an action/control problem where forecast accuracy matters only through allocation, hedge, execution, or risk action under explicit state/action/reward, uncertainty, interpretability, and data-availability constraints. portfolio states; order flow; trade data; aggregate prices; market-response data; reward/utility definitions; risk and uncertainty estimates interpretability gate; risk-vs-uncertainty distinction; state/action/reward explanation; mandate/leverage/turnover/drawdown overlays; data-availability check MEDIUM research/13-multimodal-sources/curious-quant/2020-07-16-igor-halperin-reinforcement-learning-finance.md
Curious Quant / Alternative-Data Research 2020-02-05 The alternative-data leak is a dataset-selection workflow: reject most candidate feeds, map data to fundamentals before returns, run perfect-foresight ceiling checks, match horizon/capacity, and require a mechanism before backtest enthusiasm. earnings-call transcripts; 10-K/10-Q filings; broker research; internal chats/emails; web traffic; social engagement; labor/visa data; patents; lobbying/government-contract data; regulator interactions; CFPB complaints mental-model requirement; short-history caveat; degrees-of-freedom control; cross-source corroboration; fundamental-target validation MEDIUM research/13-multimodal-sources/curious-quant/2020-02-05-vinesh-jha-craft-mining-alternative-data.md
AQR 2019-06-07 AQR provides the core negative-control leak: finance return prediction is small-data, low signal-to-noise, overfit-prone, and ML should complement long-running signals and investment process rather than replace economic theory or human expertise. historical return observations; long-running signals; investment-process datasets; machine-learning feature sets small-data warning; low signal-to-noise warning; overfitting controls; no-advice/no-solicitation boundary MEDIUM sources/06-industry-verticals/aqr-can-machines-learn-finance-raw.md
Situational Awareness / AI-Native Fund Shadow 2026-06-08 A June 2026 profile leaks the public shadow around an AI-native hedge fund: reported $20B-plus AUM, Jane Street backing, a large private Anthropic exposure, market scrutiny of routine filings, and stock-price reactions to disclosed stakes turn fund-flow, backer, private-position, and regulatory-filing data into an AI-infrastructure thesis map. media-reported AUM; named institutional backer information; private Anthropic exposure reporting; 13F and regulatory filings; market reaction to disclosed stakes; senior hiring and inflow reports media-source caveat; reported AUM versus audited AUM separation; private-position claim verification; 13F delay caveat; backer signal not performance proof; market reaction not causality proof LOW-MEDIUM sources/06-industry-verticals/situational-awareness-aum-jane-street-private-positioning-raw.md
ScaleDown Context-Engineering SLM Control 2026-06-06 ScaleDown’s technical docs leak a source-localization control lane: task-specific SLM compression, domain-keyword preservation, AST/BM25 code pruning, semantic retrieval with local embeddings/FAISS, and span extraction with offsets/confidence scores must be evaluated for whether they preserve dates, citations, identifiers, formulas, schema names, and table headers. source-card context; podcast transcript spans; PDF/OCR chunks; code and schema context; domain keywords; local embeddings; FAISS indexes; span offsets exact citation/date preservation; identifier and schema preservation; table/formula/header preservation; offset and confidence audit; BM25 versus semantic baseline; quality-before-token-savings gate; finance-domain accuracy caveat LOW-MEDIUM sources/21-benchmarks/scaledown-context-engineering-slm-raw.md
CurrentAI Investment Agents 2026-06-06 CurrentAI leaks the product shape of investment-team background agents: Excel model updates, scenario analysis, value-chain read-throughs, earnings previews, management-meeting prep, press-release/transcript monitoring, and model-impact analysis. financial operating models; press releases; transcripts; supplier calls; customer commentary; filings; consensus expectations; analyst commentary; news flow; management commentary SOC 2 claim; firm-model boundary; source monitoring caveat; model-integrity checks; vendor-marketing caveat LOW-MEDIUM sources/06-industry-verticals/currentai-proactive-investment-agents-raw.md
Sengil Turkish Gemma Finance Adapter 2026-06-05 Sengil/turkish-gemma-9b-finance-sft leaks a regional Gemma-family adapter lane: Turkish finance instruction following, Turkish local-market terminology, adapter-only loading, Gemma base-license inheritance, synthetic Gemini-generated instruction data, and visible think-tag hygiene all need gates before regional finance tasks are scored. Turkish finance instruction datasets; AlicanKiraz0 Turkish Finance SFT Dataset; Dbmaxwell Turkish finance instruction data; RsGoksel Finansal data; Gemini-generated synthetic instructions; adapter safetensors; Turkish-Gemma base metadata adapter/base revision pin; base Gemma license review; adapter-load smoke test; Turkish fixture scope; synthetic-data caveat; think-tag suppression; no-U.S.-SEC/no-alpha boundary LOW-MEDIUM sources/21-benchmarks/sengil-turkish-gemma-finance-sft-raw.md
Finance Workflow Vendor Platforms 2026-05-31 The finance workflow vendor cluster leaks the investment-team product surface: workflow-bound systems connect to filings, earnings calls, market data, fundamentals, research, CRM, SharePoint, Excel, VDRs, expert calls, proprietary data, MCP/data connectors, and time-series analytics to produce models, memos, decks, diligence artifacts, source-linked financial data, portfolio analyses, scenarios, and trading-signal monitoring. filings; earnings calls; market data; fundamentals; research; CRM; SharePoint; Excel; VDRs; expert calls; proprietary data; MCP/data connectors; time-series analytics source citation; auditability; access control; tenant isolation; human-review requirement; vendor-outcome caveat LOW-MEDIUM sources/06-industry-verticals/finance-workflow-vendor-platforms-raw.md
Fara Browser Source-Discovery Agents 2026-05-31 Fara leaks the ingestion-agent lane for this corpus: browser agents can be evaluated on finding official pages, papers, PDFs, vendor docs, podcast pages, transcripts, and source provenance while preserving publication dates, source authority, stop conditions, and human-in-the-loop controls. official arXiv pages; paper PDFs; vendor documentation; whitepapers; podcast episode pages; transcripts; audio URLs; browser screenshots; web trajectories primary-source versus mirror distinction; URL/date/author preservation; sandboxed browser execution; allow-list controls; stop before login/payment/personal data; human approval for critical points; artifact-specific runtime caveat LOW-MEDIUM sources/21-benchmarks/fara15-browser-source-discovery-raw.md
YorkFr Qwen3 Tiny Sentiment Router 2026-05-30 YorkFr leaks the tiny-model sentiment candidate lane: a 0.6B Qwen3 LoRA can be cheap enough for headline and snippet triage, but the held-out fixture shows label accuracy is not enough when route accuracy, abstention, rationale, and artifact preflight are required. financial headline and statement data; Qwen3-0.6B merged LoRA weights; finance-sentiment-routing fixture; preflight JSON; local MPS run outputs artifact revision pin; license check; tokenizer/runtime check; held-out fixture; label/route/abstention split; negative-evidence preservation LOW-MEDIUM sources/21-benchmarks/qwen3-0-6b-financial-sentiment-yorkfr-raw.md
Qwen3.6 Long-Context Open-Base Runtime 2026-05-30 Qwen3.6-35B-A3B leaks a current open-base control for long finance-source ingestion: 35B total / 3B active parameters, 262K native context, YaRN extension claims, Apache-2.0 metadata, and vLLM/SGLang/Transformers routes make it a candidate for large transcript, filing, and source-ledger replay before any finance specialization claims. HF model weights; Qwen model card; long transcript corpora; SEC filing spans; finance source ledgers; runtime traces; prompt-format records HF revision pin; license check; context-length selection; backend-specific smoke test; cost/latency/memory record; finance held-out replay; no-alpha caveat LOW-MEDIUM sources/21-benchmarks/qwen3-6-35b-a3b-model-card-raw.md
Gemma 4 31B Open-Base Runtime Control 2026-05-30 Gemma 4 31B leaks the open-base counterweight for finance replay: an Apache-2.0 30.7B dense model with 256K context and documented Transformers, vLLM, SGLang, Docker Model Runner, and quantization routes can test whether a strong general open model beats finance-specialist adapters on transcript, filing, spreadsheet, and source-card tasks. HF safetensors; Gemma 4 model card; long transcript packs; filing spans; financial tables; source cards; runtime traces HF revision pin; Apache-2.0 license check; effective-context selection; backend-specific smoke test; cost/latency/memory record; held-out finance replay; no-alpha/no-finance-specialist caveat LOW-MEDIUM sources/21-benchmarks/gemma-4-31b-model-card-raw.md
Fino1-8B Finance Reasoning Control 2026-05-30 Fino1-8B leaks the larger open finance-reasoning control lane: a Qwen3-8B finance checkpoint trained with SFT and GRPO over FinCoT-style reasoning paths from FinQA, TATQA, DocMath-Eval, Econ-Logic, BizBench-QA, and DocFinQA can anchor table, document-math, and reasoning replay if lineage and runtime revisions stay pinned. FinCoT; FinQA; TATQA; DocMath-Eval; Econ-Logic; BizBench-QA; DocFinQA; Open FinLLM Reasoning Leaderboard; HF safetensors Qwen3 lineage verification; HF revision pin; dataset revision pin; leaderboard snapshot date; held-out StateBench replay; backend-specific runtime check; stale-lineage caveat LOW-MEDIUM sources/21-benchmarks/fino1-8b-finance-reasoning-model-card-raw.md
AssetOpsBench Tool-Knowledge Internalization 2026-05-30 AssetOpsBench leaks a tool-catalog training gate relevant to finance agents: fixed schemas, tools, policies, and execution traces can be internalized into small models with QLoRA, but only after plan/schema/tool-call structure, prompt-overhead savings, and catastrophic forgetting are measured. AssetOpsBench tool catalogs; tool descriptions; question-to-plan mappings; execution-style traces; 1,700 fine-tuning examples; Gemma 4 E4B runs; Qwen3-4B runs prompted baseline comparison; AT-F1; LLM-judge planning score caveat; input-length reduction measurement; memory/speed comparison; general-benchmark regression check; stable-tool-catalog prerequisite LOW-MEDIUM sources/21-benchmarks/assetopsbench-tool-knowledge-qlora-raw.md
Amsi-fin-o1.5 MLX Finance VLM 2026-05-30 Amsi-fin-o1.5 leaks a multimodal finance-runtime lane: a Qwen3.5-VL-derived finance VLM is converted to MLX mxfp8 with chart, document OCR, screenshot, options-pricing, technical-analysis, portfolio-management, long-context, and tool-call claims, but each claim needs artifact-specific visual fixtures and runtime smoke tests before promotion. financial charts; financial documents; screenshots; OCR text; options-pricing prompts; technical-analysis images; portfolio-management prompts; MLX conversion metadata; upstream HF model card MLX artifact separation; upstream-vs-conversion revision pin; chart/table/document fixture replay; visual citation check; memory requirement check; backend-specific smoke test; options/trading claim refusal LOW-MEDIUM sources/21-benchmarks/amsi-fin-o1-5-mlx-finance-vlm-raw.md
TraceAlchemy Gemma 4 E4B Finance Instruction 2026-05-29 TraceAlchemy leaks a Gemma-family filing/table reasoning lane: financial statement reasoning, SEC-style table and excerpt extraction, revenue/margin/growth/ratio calculations, unit and scale conversion, sign/direction checks, multi-step table reasoning, and BF16-vs-GGUF artifact separation all need held-out replay before promotion. FinanceBench; TAT-QA; ConvFinQA; FinanceReasoning; Finance-Instruct-500k; synthetic SEC extraction examples; financial statement tables; GGUF metadata BF16-vs-GGUF separation; base/model license review; held-out table fixture; numeric reconciliation; sign/direction check; validation-loss-not-accuracy caveat; MLX/Neuron absence caveat LOW-MEDIUM sources/21-benchmarks/tracealchemy-gemma-4-e4b-finance-model-card-raw.md
Third Point / Loeb AI Adoption 2026-05-29 Dan Loeb’s podcast excerpts leak Third Point’s internal AI adoption pattern: Claude is framed as an employee-autonomy amplifier, native computer scientists coach teams on specific projects, adoption varies from query use to overnight agent runs, and token-heavy workflows are visible enough to become an operating signal. podcast excerpts; Claude sessions; employee AI usage patterns; overnight agent runs; token-consumption behavior; AI project coaching artifacts; investment postmortem comments podcast/media-source caveat; tool adoption not alpha proof; Claude usage not autonomous capital allocation; token intensity not productivity proof; computer-scientist coaching not model disclosure; no current-holdings inference LOW-MEDIUM sources/06-industry-verticals/third-point-loeb-ai-adoption-raw.md
QuantConnect Research Pipeline 2026-05-29 QuantConnect documents a public quant-research operating system: ideas, research notebooks, validation, backtests, paper trading, live monitoring, assistant teams, MCP access, project files, logs, datasets, Object Store writes, and notification channels. Jupyter notebooks; backtest logs; project files; datasets; Object Store artifacts; paper-trading telemetry; live deployment logs tool permissioning; handoff compression; live-action guardrails; audit traces; paper-vs-backtest comparison LOW-MEDIUM sources/06-industry-verticals/quantconnect-assistant-research-pipeline-raw.md
DianJin-R1 Chinese Finance Reasoning Control 2026-05-29 DianJin-R1 leaks a Chinese finance/compliance reasoning lane: Qwen2.5-based 7B/32B models use CFLUE, FinQA, and proprietary Chinese compliance-check data with SFT and GRPO, format rewards, accuracy rewards, and explicit think/answer tags for jurisdiction-specific finance replay. CFLUE; FinQA; Chinese Compliance Check corpus; DianJin-R1-Data; Qwen2.5 base models; reasoning traces; HF model-card metadata MIT license check; artifact revision pin; dataset revision pin; language/jurisdiction scope; think/answer format parser; runtime backend smoke test; no-U.S.-SEC/no-alpha boundary LOW-MEDIUM sources/21-benchmarks/dianjin-r1-chinese-finance-reasoning-raw.md
Local Podcast ASR Runtime Watchlist 2026-05-28 The local ASR ledger leaks a direct ingestion bottleneck: podcast and YouTube mining should compare WhisperX-MLX, Parakeet MLX, Qwen3-ASR, forced alignment, diarization, SRT/VTT/JSON outputs, word timestamps, long-audio chunking, and Apple Silicon memory/RTFx before quote-level finance evidence is promoted. finance podcast audio; YouTube audio; transcript segments; SRT/VTT/JSON outputs; word timestamps; speaker diarization labels; ASR leaderboard CSV; Apple Silicon runtime traces same-episode ASR comparison; WER and quote-error sampling; timestamp-quality check; diarization terms/token gate; chunk overlap audit; memory/latency logging; human quote review before promotion LOW-MEDIUM sources/21-benchmarks/local-podcast-transcription-mlx-models-2026-raw.md
LiquidAI LFM2.5 Portable Source-Ingestion Runtime 2026-05-28 LiquidAI LFM2.5-8B-A1B leaks a portable source-ingestion runtime lane: finance corpus agents need separate HF, GGUF, MLX, ONNX, vLLM, and SGLang scorecards for transcript cleanup, source-card extraction, tool routing, local RAG/source-pack QA, and browser-agent experiments rather than treating a tool-calling runtime as a finance specialist. podcast transcripts; YouTube transcripts; source cards; raw PDFs/OCR; tool-call schemas; local RAG source packs; runtime latency traces; HF/GGUF/MLX/ONNX artifacts runtime artifact separation; license review; HF/GGUF/MLX/ONNX separate scorecards; held-out source-ingestion fixture; tool-call schema replay; latency and memory logging; finance-specialist comparison LOW-MEDIUM sources/21-benchmarks/liquidai-lfm2-5-8b-a1b-runtime-model-card-raw.md
Financial Data Connector Ecosystem 2026-05-28 The connector ecosystem leaks the emerging control plane for finance agents: skills, connectors, subagents, Excel/PowerPoint/Word/Outlook add-ins, long-running sessions, per-tool permissions, managed credential vaults, audit logs, and licensed data access across FactSet, S&P Capital IQ, MSCI, PitchBook, Morningstar, LSEG, Daloopa, Moody’s, D&B, Guidepoint, Third Bridge, and Intralinks. licensed financial datasets; internal warehouses; repositories; CRMs; FactSet; S&P Capital IQ; MSCI; PitchBook; Morningstar; LSEG; Daloopa; Moody’s; Dun & Bradstreet; Guidepoint; Third Bridge; SS&C Intralinks per-tool permissions; managed credential vault; audit logs; tool-call provenance; human approval before client/filing/action use; vendor-survey caveat LOW-MEDIUM sources/06-industry-verticals/financial-data-ai-connector-ecosystem-2026-raw.md
Alternative-Data AI Buy-Side Source Ledger 2026-05-28 The alternative-data source ledger leaks the buy-side ingestion frontier: surveys and event agendas converge on AI over messy alternative data, vendor licensing, point-in-time mapping, company/KPI mapping, data evaluation difficulty, provenance, explainability, compliance, MNPI/PII, model-training rights, and agenda-level movement from alpha to agents. alternative data; web-scraped data; transactional data; employment data; company/KPI mappings; point-in-time histories; fundamental data; vendor datasets; job postings; market transcripts survey/source-tier caveat; MNPI/PII review; licensing and model-training restrictions; point-in-time history check; data-evaluation gate; provenance/explainability requirement LOW-MEDIUM sources/06-industry-verticals/alternative-data-ai-buy-side-2026-raw.md
Sprocket GEX Options LoRA 2026-05-27 Sprocket GEX leaks a narrow derivatives schema lane: options/Gamma Exposure agents can classify dealer-positioning, stock-pinning, 0DTE hedging, persistent gamma, transitional, and low-conviction regimes only after schema, ambiguous-input, no-pattern, license, and adapter-runtime gates pass. SPY options/GEX rows; QQQ options/GEX rows; 21.7M claimed historical options rows; 2,027 SFT rows; 32-case schema eval; LoRA adapter files; chat template license review; base-model revision review; adapter revision pin; schema/label smoke eval caveat; ambiguous/no-pattern cases; final-answer-only output hygiene; no trading-action refusal LOW-MEDIUM sources/21-benchmarks/sprocket-gex-options-lora-raw.md
FundaPod Fundamental Research Agent Pods 2026-05-27 FundaPod leaks the institutional fundamental-research agent architecture: persona-distilled value-investor and macro-strategist agents conduct independent evidence gathering under a shared provenance contract, then surface disagreements to a human PM through knowledge-graph memory instead of collapsing into autonomous trading signals. public investor materials; investment memos; verifiable source evidence; tickers; analyst personas; themes; knowledge-graph memory; persona-based memo comparisons fundamental-research versus trading-signal boundary; shared provenance contract; claim-level source grounding; persona independence check; post-hoc disagreement review; human PM decision boundary; no deployment or alpha inference LOW-MEDIUM sources/21-benchmarks/fundapod-fundamental-research-agent-pods-raw.md
Mihenk Qwen3.6 Turkish BIST Finance Model 2026-05-25 Mihenk leaks a regional Turkish/BIST finance-runtime lane: a Qwen3.6-35B-A3B MoE fine-tune and separate community GGUF conversion expose Turkish/English finance, BIST, crypto, financial-analysis, and risk-management scope, but regional fixtures, exact artifact selection, and quantization-specific replay are required before use. Turkish finance text; BIST market language; crypto finance prompts; risk-management prompts; HF safetensors metadata; community GGUF metadata; Qwen3.6-35B-A3B base metadata HF revision pin; GGUF revision pin; base-vs-finetune separation; exact GGUF file selection; Turkish/English regional held-out fixture; prompt-template check; quantization behavior caveat; no U.S. SEC transfer LOW-MEDIUM sources/21-benchmarks/mihenk-qwen36-turkish-finance-raw.md
Spreadsheet-RL Finance Workbook Agent Training 2026-05-22 Spreadsheet-RL leaks the workbook-agent training loop: collect start-goal spreadsheet pairs from public forums, build oracle spreadsheets with coding agents, run agents in real Microsoft Excel with filesystem-isolated rollouts, route operations through spreadsheet-native tools, and reward final workbooks by recalculated cell matches. public ExcelForum attachments; discussion thread solutions; initial-final spreadsheet pairs; Domain-Spreadsheet finance templates; CPA/CFA/FRM concepts; investment banking tasks; asset management tasks; Excel recalculation outputs rule-based oracle filtering; Excel error rejection; formula computability check; filesystem-isolated workspace; serialized write operations; inspect-modify-verify workflow; Excel recalculation reward; KL-regularized GRPO training LOW-MEDIUM research/21-benchmarks/raw/spreadsheet-rl-2605.22642.chandra.txt
Situational Awareness / AI Infrastructure 13F Options 2026-05-18 Situational Awareness leaks an AI-infrastructure thesis through delayed 13F/options reporting: the former OpenAI researcher’s fund held large put-option positions on Nvidia, the VanEck Semiconductor ETF, Broadcom, Oracle, AMD, TSMC, ASML, and Intel, while holding bullish call options on selected infrastructure or energy-adjacent names such as SanDisk, CoreWeave, CleanSpark, and Bloom Energy. 13F filing; option position disclosures; AI chipmaker universe; semiconductor ETF exposure; AI infrastructure equities; media-parsed position tables; quarter-end filing snapshots 13F delay and partial-coverage caveat; puts/calls versus direct ownership separation; no delta-adjusted exposure inference without strikes and expiries; position-notional caveat; post-filing-exit uncertainty; media-source caveat; no alpha proof LOW-MEDIUM sources/06-industry-verticals/situational-awareness-ai-infrastructure-13f-options-raw.md
Point72 / Turion AI-Focused Strategy 2026-05-14 Point72’s media trail leaks AI as a separate strategy sleeve, not only internal tooling: Turion is reported as a rare standalone AI-focused equities strategy with a named PM, material assets, and visible sensitivity to the AI hardware/capex cycle alongside Point72’s broader multi-strategy expansion. media-reported fund structure; reported Turion assets; PM assignment; Point72 executive-committee memo context; AI hardware and semiconductor exposure reports; performance snippets; hedge-fund sector-overweight data media-source caveat; performance snippet not audited alpha; strategy sleeve not holdings disclosure; AI hardware exposure not model disclosure; reporting-date freshness check; semiconductor cyclicality caveat LOW-MEDIUM sources/06-industry-verticals/point72-turion-ai-focused-strategy-raw.md
Hedgineer / Discretionary Asset Managers 2026-05-09 A vendor/practitioner stack leak: tenant-local hedge-fund AI deployments combine agent management, telemetry/usage mining, MCP/data-access layers, skill libraries, and earnings-prep workflows over fundamentals, consensus, transcripts, filings, and firm model templates. fundamentals; consensus data; earnings transcripts; filings; Excel models; Bloomberg/S&P-style datasets; session traces; data catalogs data-residency boundary; OTEL-style logging; data-quality checks; failure triage; connector entitlement checks LOW-MEDIUM research/13-multimodal-sources/quant-financial-engineering/2026-05-09-hedgineer-hedge-funds-ai.md
Tudor Investment Corp. / AI Boom Regime Risk 2026-05-07 Paul Tudor Jones leaks a discretionary macro control frame for the AI trade: participate through a basket while treating AI as a 1999-like technology regime that may still have one to two years and roughly 40% upside, but whose valuation, market-cap-to-GDP, IPO-supply, and crowding risks can create a severe correction. AI equity basket exposures; earnings and multiple data; market-cap-to-GDP ratios; AI IPO and megacap listing supply; historical technology-cycle analogies; CNBC and Invest Like the Best interview excerpts; ChatGPT and Claude Code adoption milestones media-source caveat; analogy not deterministic forecast; basket exposure not holdings disclosure; market-cap-to-GDP stress check; earnings/multiple support check; IPO-supply monitoring; no current-book inference LOW-MEDIUM sources/06-industry-verticals/tudor-ai-boom-regime-risk-signal-raw.md
FinSenti-Qwen3.5 Structured Sentiment Model 2026-05-06 FinSenti-Qwen3.5 leaks a parser-sensitive sentiment lane: short financial headlines, earnings snippets, and market commentary are trained into a Qwen3.5 specialist with a strict <reasoning>...</reasoning><answer>...</answer> contract, GRPO format rewards, vLLM/SGLang hints, and separate GGUF metadata that must be scored apart from the base artifact. FinSenti-Dataset; financial headlines; earnings snippets; market commentary; balanced SFT sentiment samples; held-out validation/test splits; GGUF metadata HF revision pin; strict XML-like parser check; sentiment correctness score; format compliance score; reasoning-quality caveat; GGUF-vs-safetensors separation; no portfolio-action gate LOW-MEDIUM sources/21-benchmarks/finsenti-qwen3-5-model-card-raw.md
Walleye Capital / Claude Code Firmwide 2026-05-05 Walleye leaks a firmwide coder-agent adoption boundary: a named CEO/CIO quote says 100% of employees at a 400-person hedge fund use Claude Code, framing the tool as an AI-first operating layer for technical and non-technical staff rather than as an autonomous investment engine. employee work artifacts; Claude Code sessions; financial-services agent templates; skills; connectors; subagents; tool-call audit logs; credential-vault permissions; internal repositories and data systems vendor-hosted quote caveat; per-tool permissions; managed credential vault; Claude Console audit log; human review and approval before action; adoption-versus-productivity separation; no autonomous capital allocation LOW-MEDIUM sources/06-industry-verticals/walleye-capital-claude-code-firmwide-raw.md
Citadel / Claude for Excel Coverage Models 2026-05-05 Citadel leaks a spreadsheet-native analyst workflow: a named CTO quote says Claude for Excel is used by investment professionals for coverage models, signal/noise separation, and more rigorous pressure-testing, making Excel models and scenario review a concrete finance-agent surface. Excel coverage models; financial model formulas; scenario assumptions; signal/noise review notes; pressure-test cases; FactSet/S&P Capital IQ/MSCI/PitchBook-style connector data; internal warehouses and repositories; tool-call traces vendor-hosted quote caveat; formula and assumption inspection; source-grounding requirement; connector entitlement boundary; Claude Console audit log; human review and approval before action; no autonomous capital allocation LOW-MEDIUM sources/06-industry-verticals/citadel-claude-excel-coverage-models-raw.md
Sand Grove / AI Deal Document Speed 2026-05 FT reporting leaks a concrete event-driven research workflow: Sand Grove Capital Management uses Claude, Microsoft Copilot, and ChatGPT to analyze long deal documents around complex corporate events such as M&A, compressing document-review cycle time while preserving human decision authority and hallucination/privacy caveats. merger and acquisition documents; corporate-event filings; deal terms and conditions; lengthy transaction documents; assistant outputs from Claude, Microsoft Copilot, and ChatGPT; AIMA hedge-fund AI survey context public-snippet caveat; subscription-gated source caveat; source citation requirement; human investment-decision boundary; hallucination review; data-privacy review; tool-name not alpha proof LOW-MEDIUM sources/06-industry-verticals/sand-grove-ai-deal-document-speed-raw.md
NVIDIA Quant Finance Blueprint 2026-05 NVIDIA’s blueprints leak the closed-loop quant research architecture: Signal Agent generates structured formula JSON, Code Agent emits executable Python, Eval Agent computes Rank IC/p-values over OHLCV data, and GPU optimization accelerates scenario generation and Mean-CVaR portfolio construction. OHLCV price-volume data; S&P 500 histories; signal formula JSON; iteration feedback; IC metrics; scenario returns; CVaR optimization inputs Rank IC threshold; p-value gate; best-effort status; iteration history; out-of-sample caveat; human review before production; risk/diversification constraints LOW-MEDIUM research/06-industry-verticals/nvidia-quantitative-signal-discovery-agent-2026.md
QAnchor Chinese Financial Reranker 2026-04-29 QAnchor leaks a specialist retrieval-control lane: Chinese annual-report QA needs same-document reranking over hybrid/RRF candidates, preserved Qwen3 pair formatting, weak-supervision caveats, score traces, and separation from answer generation. Chinese A-share annual reports; financial filing chunks; hybrid RRF retrieval candidates; weakly supervised train queries; gold eval queries; LoRA adapter metadata; merged sequence-classification model MRR@10; NDCG@10; P@10; 50-query gold eval caveat; qwen3_template preservation; query/candidate/score/rank trace; same-document scope boundary; model-card metric reproduction gate LOW-MEDIUM sources/21-benchmarks/qanchor-qwen3-financial-reranker-raw.md
Viking Global / VikingGPT 2026-04-23 Viking leaks a named internal GenAI adoption pattern: VikingGPT helps staff discuss trade ideas, speed routine operations, and absorb employee questions at reported high frequency, but the closer the workflow gets to an investment decision, the more the stated boundary shifts back to human judgment. employee chatbot questions; trade-idea discussion prompts; routine analyst workflow artifacts; internal usage telemetry; operations queries; investment-decision context human-decision boundary near investment action; usage telemetry provenance; source-grounding requirement; hallucination review; trade-idea versus recommendation separation; profile/reporting caveat; no portfolio-action authority LOW-MEDIUM sources/06-industry-verticals/viking-global-vikinggpt-trade-idea-workflow-raw.md
Hedge-Fund Analyst Role Compression 2026-04-23 Recruiter/media reporting leaks the labor-stack change around hedge-fund AI adoption: some junior analyst/research roles are being pulled or compressed, one AI-equipped analyst is framed as replacing multiple conventional analyst seats, and new AI reliability, data-engineering, and solutions-architecture roles are emerging to operate fund AI systems. recruiter demand signals; withdrawn role requisitions; junior analyst hiring plans; AI-enabled analyst workflow assumptions; AI reliability role descriptions; data engineering and solutions-architecture hiring signals recruiter-source caveat; role compression not measured productivity; avoid double-counting Balyasny/Walleye adoption; job-post and HR-ledger verification needed; new role title not mature control function; no alpha or model-quality inference LOW-MEDIUM sources/06-industry-verticals/hedge-fund-ai-analyst-role-compression-raw.md
Mad Lab Qwen3 1.7B Sentiment GGUF 2026-04-16 Mad Lab leaks the local GGUF sentiment control lane: earnings reports, SEC filings, and financial news can be routed through a 1.7B Qwen3 Q8_0 artifact only after exact GGUF selection, fixture replay, label-vs-route split, synthetic-label leakage checks, and backend-specific smoke tests. earnings reports; SEC filings; financial news; Qwen3-1.7B Q8_0 GGUF file; GGUF metadata; sentiment labels; synthetic or heuristic labels HF API revision pin; GGUF metadata check; label-only fixture lane; route/rationale fixture lane; synthetic-label leakage caveat; backend compatibility caveat; no-Ollama policy LOW-MEDIUM sources/21-benchmarks/mad-lab-qwen3-1-7b-sentiment-gguf-raw.md
Qwen Open Finance R 8B GGUF 2026-04-04 Qwen Open Finance R leaks a local multilingual finance-control lane: a community GGUF conversion of a DragonLLM Qwen finance/economics/business model exposes Qwen3 architecture, 40,960-token context, imatrix IQ4 quantization, and English/French/German tags, but must be scored separately from upstream weights and from every non-GGUF backend. HF model card; GGUF metadata; DragonLLM upstream model; finance/economics/business QA prompts; English finance text; French finance text; German finance text; llama.cpp runtime traces HF revision pin; GGUF metadata check; upstream-vs-conversion separation; multilingual held-out fixture; source-grounding/citation check; backend portability caveat; no-Ollama policy LOW-MEDIUM sources/21-benchmarks/qwen-open-finance-r-8b-gguf-raw.md
PolyBench Prediction-Market Forecasting Agent 2026-04-03 PolyBench leaks a contamination-resistant forecasting harness: live prediction-market agents must combine official resolution rules, timestamp-locked news, Polymarket event metadata, exact CLOB order-book states, spreads, liquidity, confidence thresholds, and simulated order execution before return claims mean anything. Polymarket Gamma API events; binary market metadata; token identifiers; Central Limit Order Book snapshots; bid-ask spreads; live liquidity; Google News streams; official resolution criteria strict snapshot timestamp; resolution-rule primacy; no-future-data prompt baseline; CLOB execution simulation; confidence-weighted return; APY and Sharpe scoring; instruction adherence; pretraining contamination resistance LOW-MEDIUM research/21-benchmarks/raw/polybench-prediction-market-llm-trading-2604.14199.chandra.txt
Hubble DSL-Sandboxed Alpha Factor Mining 2026-04 Hubble leaks the safe alpha-mining loop: an LLM proposes formulas only inside a domain-specific operator language, an AST sandbox rejects unsafe or invalid expressions, dual positive/negative RAG steers away from crowded motifs, a deterministic evaluator scores IC, bucket returns, turnover, coverage, complexity, and HAC significance, and persistent artifacts feed the next round. daily U.S. equity OHLCV panels; S&P 500 universe file; operator registry; positive factor corpus; negative crowded-template corpus; candidate formulas; per-factor diagnostics; prompt and usage metadata; in-sample and held-out OOS panels whitelisted AST nodes; depth and node-count limits; registered operator and variable checks; duplicate detection; RankIC and Pearson IC reporting; HAC t-statistics; turnover and coverage checks; crowded-template penalty; family-concentration penalty; transaction-cost caveat; LLM temporal-leakage caveat LOW-MEDIUM research/21-benchmarks/raw/hubble-agentic-alpha-factor-discovery-2604.09601.chandra.txt
Fino1-4B Finance Reasoning Control 2026-03-31 Fino1-4B leaks a cheap Qwen3 finance-reasoning control: a 4B BF16 Qwen3-family model fine-tuned with LoRA on FinQA-derived reasoning paths can test whether small finance specialists handle table/numerical QA, but the license conflict and synthetic GPT-4o reasoning paths require strict artifact and overfitting gates. FinQA-derived reasoning dataset; GPT-4o-generated reasoning paths; HF safetensors; Qwen3-4B base metadata; financial table QA prompts; runtime snippets; model-card metadata HF revision pin; license conflict review; dataset revision pin; held-out FinQA-style fixture; synthetic-data caveat; backend-specific smoke test; no-compliance/no-alpha boundary LOW-MEDIUM sources/21-benchmarks/fino1-4b-finance-reasoning-model-card-raw.md
Distil LFM2.5 Banking Voice Tool Agent 2026-03-30 The Distil LFM2.5 banking card leaks a tiny voice-agent tool-routing lane: a 354M BF16 LiquidAI/LFM fine-tune over a banking voice-assistant dataset advertises tool/function calling and multi-turn behavior, but promotion needs license review, revision pinning, schema fixtures, parser checks, audio-to-tool handoff tests, and runtime-specific support verification. distil voice assistant banking dataset; banking voice-assistant prompts; tool-call schemas; function-call traces; multi-turn dialogue examples; HF API metadata; safetensors artifact HF revision pin; license review; tool-schema fixture; function-call parser check; multi-turn dialogue replay; backend-specific runtime test; investment/trading scope refusal LOW-MEDIUM sources/21-benchmarks/distil-lfm25-banking-voice-assistant-raw.md
Quant Financial Engineering / Notebook Options Workflow 2026-03-29 Rubnerine leaks a practical analyst-agent workflow: use AI to review earnings and goodwill line items, debate theses, inspect spreadsheets, build Jupyter/yfinance screens with explicit P/E and growth definitions, inspect options volume and deltas down to microsecond fields, propose code patches, and debug discrepancies. earnings releases; goodwill line items; spreadsheet models; yfinance pulls; Yahoo Finance or Bloomberg checks; options quotes; volume; delta fields; Jupyter notebooks trust-but-verify posture; metric definition before filtering; second-source reconciliation; editable parameter check; debug trace requirement; subjective tool-preference caveat LOW-MEDIUM research/13-multimodal-sources/quant-financial-engineering/2026-03-29-hedge-fund-manager-and-ai.md
HarryS64 10-K Financial SLM 2026-03-27 HarryS64’s 10-K SLM leaks a tiny filing-prefilter lane: an 11.5M-parameter GPT-style model trained only on financial-company SEC 10-K filings is useful as a local language-statistics, anomaly, similarity, and chunk-prioritization control, but not as a chat, reasoning, citation, accounting, or alpha model. SEC 10-K filings; financial-company SIC 6000-6411 filings; BPE tokenizer artifacts; PyTorch model.pt; benchmark_results.json; filing-language validation text artifact SHA pin; tokenizer requirement check; model.pt/train.py smoke test; held-out 10-K chunk set; general-text control set; compression-only scoring lane; no user-facing generation caveat LOW-MEDIUM sources/21-benchmarks/harrys64-10k-financial-slm-raw.md
OpenAI / Anthropic Hedge-Fund Quant Talent Pull 2026-03-13 OpenAI and Anthropic are reported hiring named quant, data-science, and recruiting talent from Balyasny, Two Sigma, and Citadel, leaking the research-labor-market value of noisy-data pattern finding, finance-domain knowledge, and hedge-fund recruiting machinery for frontier AI labs. media-reported role moves; LinkedIn-derived employment histories; hedge-fund quant and data-science resumes; AI-lab recruiting-team changes; AI-lab funding and compensation context; recruiter commentary media-source caveat; LinkedIn role-transition verification; talent move not confidential-data transfer; skill-demand signal not strategy disclosure; compensation context not offer proof; no production-model inference LOW-MEDIUM sources/06-industry-verticals/ai-labs-hedge-fund-quant-talent-raw.md
AFIB SuperInvesting Financial QA Benchmark 2026-03 AFIB leaks a production-financial-QA failure taxonomy: institutional research questions need data recency, news integration, valuation logic, analytical depth, hallucination resistance, repeated-query consistency, local-market source grounding, and negative-response analysis. Indian equity market filings; annual reports; SEBI-regulated filings; stock exchange disclosures; Reserve Bank of India material; Ministry of Finance material; financial disclosures; production negative responses factual-accuracy scoring; hallucination-rate tracking; analytical-depth scoring; data-recency scoring; model-consistency checks; unsupported numerical assertion detection; vendor benchmark caveat LOW-MEDIUM sources/21-benchmarks/afib-superinvesting-financial-intelligence-raw.md
Oaktree / Howard Marks AI Investor Judgment Boundary 2026-02-27 Howard Marks’ February 2026 AI commentary leaks a high-level investor-authority boundary: Claude impressed him as a fast, unemotional information processor with hypothesis-generation and judgment-like behavior, but he preserves the human edge for taste, subjective judgment, novelty handling, valuation discipline, and bubble-risk control. Claude tutorial outputs; public Howard Marks memos; AI valuation commentary; AI infrastructure capex debate; active-management displacement analysis; media-reported Oaktree context memo/commentary source caveat; alternative-asset-manager adjacency caveat; capability-versus-valuation separation; human taste and judgment boundary; novel-situation review; no internal-deployment inference; no alpha or position inference LOW-MEDIUM sources/06-industry-verticals/oaktree-howard-marks-ai-investor-judgment-boundary-raw.md
Chinese Financial Report Metric Extraction Adapter 2026-02-23 The Chinese report-metric LoRA leaks a narrow research-report extraction lane: broker or issuer paragraphs can be turned into structured metric candidates only after schema validation, metric taxonomy checks, license review, adapter/base revision pinning, held-out gold spans, and baseline comparison against rules, BM25, NER, and general Qwen prompting. Chinese financial research-report paragraphs; financial metric spans; metric taxonomy labels; LoRA adapter files; Qwen3-14B base metadata; training examples; structured JSON outputs license review; adapter revision pin; base-model revision pin; JSON schema validation; held-out metric-span fixture; rules/BM25/NER baseline; output-hygiene check; small-training-data caveat LOW-MEDIUM sources/21-benchmarks/lmxxf-qwen3-14b-financial-report-metric-lora-raw.md
Kirkoswald / AI Software Oracle Put Options 2026-02-21 Regulatory-filing reporting leaks an AI-software disruption trade map: several funds entered 2026 with large Salesforce, Workday, or Oracle exposure, while Kirkoswald’s Greg Coffey reportedly held close to $400 million of Oracle put options, turning delayed filings into a short-side AI-obsolescence thesis clue. 13F and regulatory filings; media-parsed position tables; Salesforce and Workday equity exposure; Oracle equity and put-option exposure; AI software-obsolescence market narrative; quarter-end filing snapshots filing-lag caveat; media-source caveat; puts versus direct short separation; unknown strike and expiry caveat; post-filing exit uncertainty; no current-position inference; no alpha proof LOW-MEDIUM sources/06-industry-verticals/kirkoswald-ai-software-oracle-put-options-raw.md
Lone Pine / AI Adoption Cost-Takeout Thesis 2026-02-13 Lone Pine’s co-CIO frames AI as an early platform shift whose investable evidence chain runs from improving models and scarce capacity to CEO-reported agent productivity and future CFO cost-takeout claims at large incumbent companies, the ‘revenge of the dinosaurs’ thesis. Goldman Sachs Exchanges podcast excerpts; model-quality progress indicators; hyperscaler and inference capacity data; CEO productivity anecdotes; coding automation and agent-workflow examples; earnings-call CFO cost-takeout claims; financial-statement cost lines podcast/media-source caveat; adoption anecdote not realized ROI; capacity demand not alpha proof; earnings-call confirmation requirement; financial-statement cost-line check; no current-holdings inference LOW-MEDIUM sources/06-industry-verticals/lone-pine-ai-revenge-of-dinosaurs-raw.md
Qwen3 14B Bookkeeper GGUF 2026-01-12 The Qwen3 14B Bookkeeper GGUF card leaks a local accounting-control lane: transaction categorization, journal-entry suggestions, reconciliation-style reasoning, professional-accounting exam data, synthetic financial data, and GGUF export belong in an artifact-separated bookkeeping harness before any finance workflow promotion. professional accounting questions; finance-alpaca data; synthetic financial data; OpenMathReasoning-mini; FineTome-100k; GGUF metadata; Qwen3-14B base metadata HF revision pin; GGUF metadata check; base-vs-conversion separation; license review; held-out bookkeeping fixture; synthetic-data leakage caveat; GAAP/tax/audit refusal gate LOW-MEDIUM sources/21-benchmarks/qwen3-14b-bookkeeper-gguf-raw.md
MARS Risk-Aware Multi-Agent Portfolio Controller 2026 MARS leaks the portfolio-agent risk-control shape: instead of one monolithic RL trader, use heterogeneous safety-critic agents with distinct risk tolerances, a meta-adaptive controller that weights them by market state, and a final risk-management overlay that validates executed actions. portfolio holdings; cash balance; technical indicators; market state vectors; agent Q-values; agent safety-critic estimates; replay buffers; training/validation/testing market regimes Safety-Critic risk threshold; risk-aversion penalty; rolling volatility penalty; rolling max-drawdown penalty; transaction-cost penalty; rule-based action validation; train/validation/test regime split; monolithic-agent baseline LOW-MEDIUM research/21-benchmarks/raw/mars-risk-aware-multi-agent-portfolio-management-aaai-2026.chandra.txt
Alpha-R1 Context-Aware Alpha Screening 2025-12-23 Alpha-R1 leaks a regime-aware factor activation stack: backtest each alpha into quantitative performance vectors, convert factors into semantic profiles, synthesize daily market state from atomic price/news units, then use a GRPO-trained reasoning model to activate sparse factor sets under market-feedback rewards. historical market data; financial news; social media; technical indicators; factor backtest outputs; semantic factor profiles; daily market state descriptions; CSI 300 and CSI 1000 asset pools global historical memory; factor decay profile; market-performance reward; reasoning-quality judge caveat; KL regularization; action-validity and sparsity penalties; TopN/holding-day sensitivity; VWAP and transaction-cost caveat LOW-MEDIUM research/21-benchmarks/raw/alpha-r1-alpha-screening-llm-reasoning-2512.23515.chandra.txt
Coatue / Public Research Leakage Channel 2025-12-22 Coatue leaks a public research-distribution channel: Business Insider reports that the normally private Tiger Cub is now highly visible through podcasts, conference panels, CNBC appearances, and daily public research notes called C:/Takes, while a new access product counts OpenAI and Stripe as significant positions. C:/Takes public research notes; podcast appearances; conference-panel remarks; CNBC appearances; public fund materials; AI and technology thesis vocabulary; public position-theme references media-source caveat; public-note curation caveat; position-theme versus sizing separation; no internal-model inference; no alpha attribution; topic-cadence tracking; cross-source timestamp preservation LOW-MEDIUM sources/06-industry-verticals/coatue-public-research-leakage-channel-raw.md
Citadel / Internal Chatbot Fundamental Equity 2025-12-03 Citadel leaks a fundamental-equity research intake layer: Business Insider reports that an internal AI chatbot rolled out to fundamental equity investors earlier in 2025 helps locate hidden details in public filings, summarize sell-side research, and track executive keyword mentions, while the CTO-level boundary says PMs should not offload human investment judgment to AI. public filings; sell-side research; executive commentary and transcripts; keyword mention trackers; fundamental-equity research prompts; stock-picking team usage telemetry media-source caveat; source provenance and citation requirement; public-filing timestamp and as-of controls; human PM judgment boundary; no investment-decision offload; adoption-versus-performance separation; no alpha proof LOW-MEDIUM sources/06-industry-verticals/citadel-internal-chatbot-fundamental-equity-raw.md
Balyasny / BAMChatGPT Analyst Workflow 2025-11-28 Balyasny leaks a named internal-chatbot adoption layer: Business Insider reports that BAMChatGPT is among AI tools used by roughly 80% of staff and that the firm built an AI bot to take on senior-analyst grunt work for investment teams, backed by senior data-science hiring. analyst work products; investment-team prompts; internal chatbot sessions; AI tool usage telemetry; senior analyst task queues; central data-science rollout artifacts media-source caveat; named-tool provenance; adoption-versus-productivity separation; permissioned internal-data access; source citation and grounding checks; human investment-judgment boundary; no alpha proof LOW-MEDIUM sources/06-industry-verticals/balyasny-bamchatgpt-analyst-grunt-work-raw.md
Davidson Kempner / AI Capex Prisoner’s Dilemma 2025-11-03 Davidson Kempner’s CIO frames Big Tech AI capex as a prisoner’s-dilemma-like forced-spending race whose effects spill into broad portfolios because mega-cap tech dominates equity indices, making AI-return disappointment an index-level ‘AI wobble’ risk rather than only a single-stock thesis. Big Tech AI capex plans; mega-cap index weights; earnings and cash-flow data; AI return-on-capex evidence; market valuation expectations; historical technology bubble analogies; Goldman Sachs Exchanges podcast excerpts podcast/media-source caveat; risk discussion not position disclosure; historical analogy not forecast proof; index-concentration check; return-on-capex evidence requirement; no alpha or current-holdings inference LOW-MEDIUM sources/06-industry-verticals/davidson-kempner-ai-capex-prisoners-dilemma-raw.md
Finance Embeddings Gemma 300M v2 2025-09-26 Finance Embeddings Gemma 300M v2 leaks a retrieval-specialist lane: financial RAG should score finance-term recognition, document similarity, source-card localization, 512-token chunking, mean-pooling/L2-normalized embeddings, and 2.89M-sample finance fine-tuning against BM25 and hybrid baselines before any research-agent context is trusted. financial text samples; SEC filing chunks; investor-report paragraphs; podcast/source-card metadata; finance wiki pages; embedding vectors; retrieval gold sets base-license review; artifact revision pin; BM25 baseline; hybrid retrieval comparison; source/page precision; FinMTEB-style held-out set; latency and cost measurement; no dense-only assumption LOW-MEDIUM sources/21-benchmarks/finance-embeddings-gemma-300m-v2-raw.md
Trading-R1 Financial Trading Reasoning Model 2025-09-11 Trading-R1 leaks the public trading-reasoning training recipe: build daily ticker contexts from technical data, fundamentals, news, sentiment, insider signals, and macro factors; teach structured thesis form first; then reward evidence-grounded claims and volatility-adjusted five-tier decisions. Tauric-TR1-DB; technical market data; company fundamentals; Google News scraped articles; public sentiment; insider sentiment; macroeconomic indicators; 14 major tickers across 18 months input signal-to-noise review; source and quote grounding; structured thesis format reward; volatility-adjusted label construction; market-outcome reward caveat; multi-horizon decision check; reward-hacking caveat; paper-snapshot reproduction requirement LOW-MEDIUM research/21-benchmarks/raw/trading-r1-financial-trading-llm-reasoning-2509.11420.chandra.txt
Kuvera Qwen3 Personal Finance Advisor Model 2025-09 Kuvera leaks a behaviorally grounded personal-finance advisor lane: public Reddit-style personal-finance questions, psychological context, suitability/refusal behavior, calculation checks, and blind LLM-jury evaluation need revision-pinned Qwen3 scoring before any advisor workflow is trusted. public Reddit personal-finance questions; Kuvera-PersonalFinance-V2.1 dataset; behavioral-finance reasoning chains; advisor suitability prompts; calculation fixtures; blind LLM-jury evaluations HF revision pin; dataset/paper separation; license review; held-out suitability fixture; calculation check; refusal/escalation gate; human-review requirement; runtime-specific smoke test LOW-MEDIUM sources/21-benchmarks/kuvera-qwen3-personal-finance-model-raw.md
Balyasny Applied AI / Macro Research 2025-07-08 The profile leaks a granular Balyasny Applied AI workflow: an early-career researcher was building an LLM to forecast market events, with usefulness feedback from macro research leadership and a fail-fast culture that treats unhelpful experiments as disposable rather than turning every model output into a trade thesis. market event definitions; macro research feedback; intern project artifacts; Applied AI experiment logs; macro trading context; expert usefulness labels timestamp boundary for market events; expert macro-research review; fail-fast experiment termination; separation of event forecast from trade recommendation; profile-source caveat; no performance attribution LOW-MEDIUM sources/06-industry-verticals/balyasny-applied-ai-market-events-llm-raw.md
Agentar-Fin-R1 / Finova Finance-Agent Evaluation 2025-07 Agentar-Fin-R1 leaks a finance-agent evaluation lane rather than a deployable model lane: the paper describes task labels, trustworthy knowledge engineering, multi-agent trustworthy data synthesis, validation governance, difficulty-aware optimization, dynamic attribution, and Finova agent checks for reasoning, compliance, tool-call accuracy, and entity recognition, while exposing no verified runnable artifact. financial task labels; trustworthy knowledge resources; synthetic finance reasoning data; validation-governance records; Finova agent tasks; tool-call traces; entity-recognition labels; safety/compliance cases paper-only artifact boundary; model/dataset/Space absence check; model-card/license/weights verification before runtime; tool-call accuracy metric; complex reasoning accuracy metric; safety/compliance detection; entity F1 LOW-MEDIUM sources/21-benchmarks/agentar-fin-r1-qwen3-finance-reasoning-raw.md
Rebellion Research / LLM Backtest Triage 2025-06-28 Fleiss leaks a pre-backtest research triage layer: instead of running a historical backtest for every natural-language strategy query, researchers can use an LLM to frame probability measures around asset price moves, prioritize higher-likelihood mean-reversion candidates, and broaden rigid stock screens with qualitative filters such as earnings-call tone plus improving margins. strategy hypothesis text; asset price histories; moving-average deviations; mean-reversion motifs; probability-measure prompts; earnings-call language; financial ratio screens; profit-margin trends; researcher query logs strict LLM-triage versus backtest boundary; point-in-time price and transcript data; model cutoff check; calibration and false-positive tracking; deterministic historical backtest after triage; transaction-cost and capacity review; profile-source caveat LOW-MEDIUM sources/06-industry-verticals/rebellion-research-llm-backtest-triage-raw.md
WorldQuant / Milken Multimodal Data Expansion 2025-05-07 WorldQuant leaks a multimodal ingestion frontier: its deputy CIO said AI lets the firm expand data brought into models by restructuring data from images and audio, while also warning that black-box AI can generate noise rather than signal and should not replace human judgment. images; audio; restructured multimodal data; model input features; conference-panel practitioner claims; document/news/filing/data context from adjacent panel remarks media-source caveat; modality-specific OCR/ASR error checks; timestamp and source-provenance preservation; noise-versus-signal review; black-box skepticism; human judgment boundary; no performance-proof inference LOW-MEDIUM sources/06-industry-verticals/worldquant-milken-multimodal-ai-data-expansion-raw.md
Freestone Grove / Analyst Edge Control 2025-05-07 Freestone Grove leaks the fundamental-equity over-automation control: its cofounder says the firm spends more time thinking about how people use AI tools than building new AI capabilities, teaches people to keep paying attention, and avoids killing the edge investors get from doing the work themselves. analyst research work products; investment theses; tool-use patterns; human attention checkpoints; manual research artifacts; independent-view formation notes media-source caveat; human-attention check; manual-work preservation gate; independent-view review; over-automation risk control; no architecture inference; no alpha proof LOW-MEDIUM sources/06-industry-verticals/freestone-grove-ai-analyst-edge-control-raw.md
Fin-R1 Finance Reasoning Baseline 2025-03-16 Fin-R1 leaks an older but useful finance-reasoning baseline: a Qwen2.5-lineage 7B/8B model trained with SFT and GRPO over FinCorpus, Ant_Finance, FinPEE, FinCUGE, FinanceIQ, Finance-Instruct, FinQA, TFNS, ConvFinQA, and FinanceQT can anchor premise, numerical reasoning, and abstention tests against newer Qwen3/Gemma/LiquidAI candidates. FinCorpus; Ant_Finance; FinPEE; FinCUGE; FinanceIQ; Finance-Instruct-500K; FinQA; TFNS; ConvFinQA; FinanceQT artifact revision pin; license and tokenizer check; structured think/answer format check; held-out StateBench replay; backend-specific runtime test; reported-score caveat; no-trading/no-advice boundary LOW-MEDIUM sources/21-benchmarks/fin-r1-finance-reasoning-model-raw.md
High-Flyer / DeepSeek Compute Spillover 2025-01-28 High-Flyer leaks the quant-fund-to-AI-lab infrastructure pathway: a machine-learning quant platform reportedly accumulated GPU clusters, engineering talent, and model-building habits before redirecting those assets into DeepSeek, making compute provenance and organizational spillover part of the research-machine signal rather than a trading-alpha claim. quant trading model artifacts; stock-trading data; GPU cluster inventory; accelerator procurement history; AI-lab training workloads; reported A100/H800 counts; model-release documentation; market-impact news traces reported-GPU-count uncertainty; trading-infrastructure versus AI-lab boundary; source provenance check; no-alpha inference gate; market-impact causality caveat; model-training-cost caveat; separate portfolio-use evidence requirement LOW-MEDIUM sources/06-industry-verticals/high-flyer-deepseek-quant-compute-spillover-raw.md
Regional Qwen3 Finance SLM Watchlist 2026-05-31 The regional Qwen3 watchlist leaks the jurisdiction/task coverage gap in finance replay: India budget/policy QA, Thai/English finance language, Qwen3 sentiment GRPO, and a misnamed FinanceQA adapter each require exact metadata, lineage, language, and runtime checks before they can be used as controls. India Budget 2026 instructions; ThaiLLM/Qwen3 merge metadata; FinGPT sentiment train; FinanceQA adapter data; HF API metadata; GGUF sibling metadata; language tags exact revision preflight; language and jurisdiction scope; base-model lineage check; adapter/merged-artifact distinction; backend-specific runtime check; task-specific suite match; no-Ollama/no-alpha boundary LOW sources/21-benchmarks/regional-qwen3-finance-slm-watchlist-raw.md
Dev9124 Qwen3 Finance Instruction Watchlist 2026-05-30 Dev9124/qwen3-finance-model leaks a practical low-cost Qwen3 finance-instruction lane: a 4B Unsloth/TRL artifact trained on Finance-Instruct-500K-style data advertises Transformers, TGI, vLLM, SGLang, and endpoint compatibility, making it useful as a broad finance-language source-synthesis control only after preflight. Finance-Instruct-500K; Qwen3-4B base metadata; Unsloth training traces; TRL training stack; HF safetensors; model-card runtime snippets; preflight record HF revision pin; training-data license review; runtime-specific smoke test; held-out finance replay; source-citation quality check; no-independent-benchmark caveat LOW sources/21-benchmarks/qwen3-finance-model-dev9124-raw.md

Extraction Rules

Promote only clues that reveal at least one concrete data source, workflow, validation control, model/agent role, or human authority boundary.

Reject generic AI optimism, host speculation, marketing without workflow detail, and any claim that implies alpha without audited evidence.

Next Ingestion Queue

  • Man Group / Man Institute / AHL Explains official pages and videos
  • Two Sigma official articles, ACM CAIS references, and conference talks
  • Bridgewater AIA Labs pages, videos, and job posts
  • Schonfeld FE AI Lab and Millennium / Cubist / Point72 / Citadel EQR official pages
  • Tower Research Capital, G-Research, XTX, QRT, Winton, and Renaissance tech/careers pages
  • Risknet Quantcast, Hedge Fund Huddle, Top Traders Unplugged, Odd Lots, Flirting with Models, and AIMA/CAIA/CFA panels
  • Kaggle Jane Street / Optiver / Two Sigma / G-Research competition artifacts and discussion timestamps

Generated by scripts/hedge_fund_leakage_loop.py.