Read the full evidence audit →

State of AI · Public evidence reconciled August 12, 2026

What quantitative managers are choosing to disclose about AI.

This is an evidence map, not a league table. It separates named model artifacts, embeddings and retrieval, agent/tool surfaces, research workflows, deployment claims, and hiring intent. A missing public signal is not evidence of a missing private capability.

6evidence lanes kept separate
8BNumerai Predictive LLM publicly described
BAMBalyasny-affiliated finance embeddings
0complete public model inventories

01 · Correction pass

Several artifacts were already in the repository but missing from the public map.

The correction pass reconciled the public page with the newer firm-source synthesis, benchmark notes, and technical-paper ledger.

Named model

Numerai Predictive LLM

Numerai’s NumerCon recap describes an 8-billion-parameter model trained on more than one million articles to generate features for the Faith dataset.

Firm-published source · weights and attribution not public
Embedding / retrieval

Balyasny BAM embeddings

Balyasny-affiliated researchers describe finance embeddings fine-tuned on 14.3M query-passage pairs and a retrieval/RAG service.

Published paper · not a general finance LLM
Narrow fine-tune

CFM financial NER

CFM reports a task-specific named-entity recognition fine-tune with a measured F1 change in its case study.

Firm case study · extraction, not investment reasoning
Affiliated research

AQR-affiliated language models

Academic work reports point-in-time finance-language-model experiments with chronological training controls.

Affiliation is not proof of AQR deployment

The public model inventory now distinguishes firm-published claims, affiliated research, firm-associated code, hiring intent, and general benchmark models. See the audit and BAM embedding synthesis.

02 · Evidence lanes

Different public signals answer different questions.

Rows describe what a source supports and the boundary that remains. They are not scores.

Model / weights

What changed?

Look for a named model, parameter count, pretraining, fine-tuning, or released model-generated feature set.

Example: Numerai Predictive LLM; CFM financial NER.
Embeddings / retrieval

What is retrieved?

Look for corpus construction, query-passage training, hard negatives, retrieval evaluation, and RAG deployment context.

Example: Balyasny BAM embeddings.
Agent / tools

What can the system do?

Look for research tools, MCP, code execution, memory, evaluation, permissions, and human-approval boundaries.

Examples: Bridgewater PAT; Numerai Skills and MCP.
Workflow / deployment

Where is it used?

Separate research assistance, investor-facing tools, shadow mode, production workflow, and live investment authority.

Public adoption claims remain attributed claims.

03 · Public evidence register

The landscape is a source map, not a maturity ladder.

A filled cell records a public signal in that lane. It does not imply that another firm lacks a private capability.

Firm / case Named model or weights Embedding / retrieval Agent / tools Workflow / deployment Hiring / infrastructure
Direct public evidence Affiliated / firm-reported Hiring or infrastructure signal Not found in reviewed sources

04 · Signal density

What the public register actually contains.

These charts count rows in the register above. They describe the reviewed public record; they are not capability scores and do not rank firms.

Rows with a public signal by lane

A row counts once when its matrix cell is populated by direct, affiliated, or hiring/infrastructure evidence.

Source: the evidence register above; reviewed public artifacts and dated personnel/infrastructure signals.

Evidence type within each lane

The bars separate direct public evidence from affiliated/firm-reported and hiring/infrastructure signals.

Counts are descriptive. “Not found” means not found in the reviewed source set.

05 · Artifact register

Named artifacts should be read with their provenance.

The inventory below is deliberately small. General finance models and embedding benchmarks remain in the benchmark pillar unless a firm ownership link is public.

06 · People and recruiting

Public personnel data reveals ownership signals, not team size.

Named roles, public profiles, and job descriptions are useful for tracing responsibility and intended architecture. They do not prove a filled role, a live system, or investment authority.

Central AI / platform

Bridgewater · Balyasny · Millennium

Public surfaces name AI leadership, applied-AI programs, internal platforms, evaluation controls, and research-facing tools.

Personnel and vendor/customer claims remain attributed.
Investment workflow

Acadian · Man AHL · Two Sigma

Public material describes AI or ML within research, text processing, signal development, or workflow tooling.

Model ownership and return attribution are separate questions.
Runtime / hiring

G-Research · Point72 · QRT

Roles and public code expose serving, RAG, MCP, fine-tuning, embeddings, evaluation, and secure execution vocabulary.

Role specifications are not deployment logs.
Predictive ML

HRT · Jane Street · XTX

Public material describes predictive or deep-learning systems and research constraints that should not be relabeled as finance language models.

Trading-model evidence is a distinct category.

The useful question

Ask which evidence contract is actually public.

For each manager, separate the model artifact, data and retrieval layer, agent/tool permissions, evaluation gates, human approvals, deployment stage, and any performance claim. The public record is richest where a source names the artifact and its boundary. It is weakest where a role description or marketing phrase is treated as proof of a live investment system.

Decision

Maintain a dated source ledger and refresh every model, embedding, personnel, and deployment claim independently. Keep “not found” visibly separate from “does not exist.”

State of AI · Executive Briefings stateofai.pages.dev