State of AI · Public evidence reconciled August 12, 2026
What quantitative managers are choosing to disclose about AI.
This is an evidence map, not a league table. It separates named model artifacts, embeddings and retrieval, agent/tool surfaces, research workflows, deployment claims, and hiring intent. A missing public signal is not evidence of a missing private capability.
01 · Correction pass
Several artifacts were already in the repository but missing from the public map.
The correction pass reconciled the public page with the newer firm-source synthesis, benchmark notes, and technical-paper ledger.
Numerai Predictive LLM
Numerai’s NumerCon recap describes an 8-billion-parameter model trained on more than one million articles to generate features for the Faith dataset.
Firm-published source · weights and attribution not publicBalyasny BAM embeddings
Balyasny-affiliated researchers describe finance embeddings fine-tuned on 14.3M query-passage pairs and a retrieval/RAG service.
Published paper · not a general finance LLMCFM financial NER
CFM reports a task-specific named-entity recognition fine-tune with a measured F1 change in its case study.
Firm case study · extraction, not investment reasoningAQR-affiliated language models
Academic work reports point-in-time finance-language-model experiments with chronological training controls.
Affiliation is not proof of AQR deploymentThe public model inventory now distinguishes firm-published claims, affiliated research, firm-associated code, hiring intent, and general benchmark models. See the audit and BAM embedding synthesis.
02 · Evidence lanes
Different public signals answer different questions.
Rows describe what a source supports and the boundary that remains. They are not scores.
What changed?
Look for a named model, parameter count, pretraining, fine-tuning, or released model-generated feature set.
Example: Numerai Predictive LLM; CFM financial NER.What is retrieved?
Look for corpus construction, query-passage training, hard negatives, retrieval evaluation, and RAG deployment context.
Example: Balyasny BAM embeddings.What can the system do?
Look for research tools, MCP, code execution, memory, evaluation, permissions, and human-approval boundaries.
Examples: Bridgewater PAT; Numerai Skills and MCP.Where is it used?
Separate research assistance, investor-facing tools, shadow mode, production workflow, and live investment authority.
Public adoption claims remain attributed claims.03 · Public evidence register
The landscape is a source map, not a maturity ladder.
A filled cell records a public signal in that lane. It does not imply that another firm lacks a private capability.
| Firm / case | Named model or weights | Embedding / retrieval | Agent / tools | Workflow / deployment | Hiring / infrastructure |
|---|
04 · Signal density
What the public register actually contains.
These charts count rows in the register above. They describe the reviewed public record; they are not capability scores and do not rank firms.
Rows with a public signal by lane
A row counts once when its matrix cell is populated by direct, affiliated, or hiring/infrastructure evidence.
Evidence type within each lane
The bars separate direct public evidence from affiliated/firm-reported and hiring/infrastructure signals.
05 · Artifact register
Named artifacts should be read with their provenance.
The inventory below is deliberately small. General finance models and embedding benchmarks remain in the benchmark pillar unless a firm ownership link is public.
06 · People and recruiting
Public personnel data reveals ownership signals, not team size.
Named roles, public profiles, and job descriptions are useful for tracing responsibility and intended architecture. They do not prove a filled role, a live system, or investment authority.
Bridgewater · Balyasny · Millennium
Public surfaces name AI leadership, applied-AI programs, internal platforms, evaluation controls, and research-facing tools.
Personnel and vendor/customer claims remain attributed.Acadian · Man AHL · Two Sigma
Public material describes AI or ML within research, text processing, signal development, or workflow tooling.
Model ownership and return attribution are separate questions.G-Research · Point72 · QRT
Roles and public code expose serving, RAG, MCP, fine-tuning, embeddings, evaluation, and secure execution vocabulary.
Role specifications are not deployment logs.HRT · Jane Street · XTX
Public material describes predictive or deep-learning systems and research constraints that should not be relabeled as finance language models.
Trading-model evidence is a distinct category.The useful question
Ask which evidence contract is actually public.
For each manager, separate the model artifact, data and retrieval layer, agent/tool permissions, evaluation gates, human approvals, deployment stage, and any performance claim. The public record is richest where a source names the artifact and its boundary. It is weakest where a role description or marketing phrase is treated as proof of a live investment system.
Maintain a dated source ledger and refresh every model, embedding, personnel, and deployment claim independently. Keep “not found” visibly separate from “does not exist.”
Open the full audit, source boundaries, and verification ledger →