← Benchmarks 🕐 26 min read
Benchmarks

Frontier and Domain-Specific AI Benchmarks — Survey Note, 2025–2026

> **Source credibility: VARIES BY SECTION.** This note synthesizes frontier benchmark coverage from arXiv preprints (TIER 1–2), official leaderboards and model cards, and third-party tracking sites (l

Pillar: 21 — AI Benchmarks: Evaluation, Evolution, and Enterprise Practice Status: Survey note. Entries are sourced from published papers, official leaderboards, and third-party tracking sites. [NEEDS VERIFICATION] flags assertions that could not be confirmed from a primary source during this research pass. Do not cite this note as a primary source — cite the underlying studies directly.

See also (wiki): wiki/ai-model-evaluation-benchmarks.md

Cross-references:

  • benchmark-foundations-enterprise-evals-2026.md — validity crisis, CLEAR framework, enterprise eval playbook
  • emerging-benchmarks-may-2026.md — EsoLang-Bench, Harvey LAB, τ-knowledge, τ-voice, ProgramBench, SkillsBench, RHB, ExploitBench, EnergyAgentBench, THINK-Bench, LemmaBench, SlopCodeBench, GiantsBench, ARC-AGI governance
  • wiki/agentic-ai-governance.md
  • wiki/training-architecture.md

Source credibility: VARIES BY SECTION. This note synthesizes frontier benchmark coverage from arXiv preprints (TIER 1–2), official leaderboards and model cards, and third-party tracking sites (llm-stats.com, artificialanalysis.ai). Each entry carries its own credibility rating. Do not cite this note as a primary source — cite the underlying papers or leaderboards directly.


Purpose and Scope

This note covers benchmarks not yet in the corpus — specifically the major frontier capability and domain-specific evaluations that define the published leaderboard landscape as of mid-2026. Where emerging-benchmarks-may-2026.md documented newly released instruments (April–May 2026), this note covers established and recently updated evaluations that practitioners and procurement teams regularly encounter in vendor disclosures.


Part 1 — Frontier Capability Benchmarks

1.1 FrontierMath

Publisher: Epoch AI (independent AI research organization) Paper: arXiv:2411.04872 (November 2024, updated) Credibility: HIGH — independent academic organization, no commercial AI vendor affiliation, 350 problems authored by professional mathematicians under NDA to prevent contamination

What it measures: Research-grade mathematics. 350 original problems across all major branches — number theory, algebraic geometry, real analysis, category theory — designed to require techniques analogous to those used in publishable mathematics research. The benchmark is tiered: Tiers 1–3 (300 problems, difficult competition/research math) and Tier 4 (50 exceptionally hard problems, published as a separate leaderboard track by Epoch AI).

Why it matters: FrontierMath was designed at the frontier to stay hard as models improve. When first published in late 2024, no model solved more than 2% of problems. It has remained the primary measuring stick for genuine mathematical reasoning — as opposed to competition math (AIME, AMC), which saturated in 2025.

Current SOTA (as of May 2026):

  • GPT-5.5 Pro: 52% (Tiers 1–3), 39.6% (Tier 4) — [NEEDS VERIFICATION on exact Tier 4 score; sourced from third-party tracking, not Epoch AI primary release]
  • GPT-5.4 Pro (March 2026): ~50% (Tiers 1–3), 38% (Tier 4)
  • Claude Opus 4.7: 22.9% (Tier 4) [NEEDS VERIFICATION]
  • Gemini 3.1 Pro: 16.7% (Tier 4) [NEEDS VERIFICATION]

The jump from under 2% (GPT-4 era, 2024) to 50%+ (GPT-5.5 Pro, 2026) on Tiers 1–3 is the sharpest capability gain measured on any research-grade benchmark in this period. Tier 4 — the hardest set — remains below 40% for any model, preserving discrimination headroom.

Enterprise relevance: FrontierMath is not a deployment benchmark; no enterprise is asking AI to prove theorems. Its value is as a pure reasoning stress test uncorrupted by memorization — and as an index of reasoning-generalization progress. A 10× improvement in two years on a contamination-resistant math benchmark is relevant context for any organization planning AI capability roadmaps.

URLs: https://epoch.ai/frontiermath · https://epoch.ai/benchmarks/frontiermath-tier-4 · arXiv:2411.04872


1.2 MMMU-Pro (Massive Multidiscipline Multimodal Understanding — Professional)

Publisher: MMMU Benchmark Consortium (academia) Paper: arXiv:2409.02813 (September 2024) Credibility: HIGH — academic benchmark, methodology published, independent of frontier model vendors

What it measures: Multimodal understanding and reasoning across six disciplines (Art & Design, Business, Science, Health & Medicine, Humanities & Social Science, Technology & Engineering). 3,460 questions. Three hardening mechanisms distinguish it from the original MMMU: (1) questions answerable by text-only models are removed, forcing genuine visual grounding; (2) answer choices are augmented to resist elimination strategies; (3) a vision-only input track embeds questions inside images, blocking text-based shortcuts.

Why it matters: The original MMMU was largely saturated by 2024 for frontier models. MMMU-Pro restores discrimination. It is the standard reference for multimodal enterprise capability claims because it requires simultaneous visual and textual reasoning rather than image description.

Current SOTA (as of May 2026):

  • GPT-5.4 Pro: 94% [NEEDS VERIFICATION — sourced from third-party tracking aggregators, not official OpenAI disclosure]
  • Claude Mythos Preview: 92.7% [NEEDS VERIFICATION — same caveat; model name may differ from public release name]
  • Gemini 3.1 Pro: 83.9% [NEEDS VERIFICATION]

Note: Third-party tracking sites (llm-stats.com, bracai.eu, benchlm.ai) show different figures depending on update date. The spread across sources reflects timing and evaluation condition differences, not measurement error. All figures should be treated as directional until confirmed against official model cards.

Enterprise relevance: Any enterprise evaluating models for document understanding, slide analysis, engineering diagram interpretation, or mixed-media knowledge bases should require MMMU-Pro scores alongside text-only evaluations. A high score on a text-only benchmark combined with a low MMMU-Pro score signals that a model’s multimodal capability does not match its text capability — a common marketing gap.

URLs: https://artificialanalysis.ai/evaluations/mmmu-pro · arXiv:2409.02813


1.3 LiveBench

Publisher: LiveBench research team (academia) Status: Spotlight paper, ICLR 2025 Credibility: HIGH — ICLR-reviewed, contamination resistance is architectural, not just claimed

What it measures: 18 tasks across 6 categories — math, coding, reasoning, language, instruction following, data analysis. New questions are released monthly, drawn from recent arXiv papers, news, datasets, and other sources with verifiable ground-truth answers. No LLM judge in the evaluation loop — all tasks have objective correct answers.

Contamination design: Each question is derived from content published after the training cutoff of the model being evaluated. A model cannot have memorized the answer because the source didn’t exist at training time.

Current standings (as of early 2026): [NEEDS VERIFICATION on exact scores; third-party tracking shows o3-mini leading at approximately 0.846 normalized score — primary LiveBench leaderboard should be checked for current rankings]

Enterprise relevance: LiveBench is the most rigorous general-purpose benchmark for tracking capability across tasks where contamination is a concern. Because it updates monthly, it is one of the few evaluations where a vendor cannot pre-optimize against a static test set. For organizations that need to track model capability over time without relying on vendor marketing, LiveBench provides the clearest signal.

URLs: https://livebench.ai · https://github.com/LiveBench/LiveBench


1.4 MATH-500

What it measures: Competition-level mathematics (AMC, AIME, and similar) across 500 problems at five difficulty levels. Unlike GSM8K (grade-school arithmetic, now saturated at ~99%), MATH-500 tests symbolic multi-step reasoning that requires more than arithmetic recall.

Current SOTA (as of May 2026):

  • DeepSeek R1: 97.3%
  • GPT-5.3 Codex: 96% [NEEDS VERIFICATION]
  • Most frontier models cluster at 94–97%

Status as of mid-2026: MATH-500 is approaching saturation for frontier models, serving mainly as a floor check. It still discriminates meaningfully at the 7B–34B model tier, where scores span 70–92%. For frontier model comparisons, FrontierMath Tiers 1–3 provides the discrimination MATH-500 no longer supplies.

Relationship to GPQA Diamond: MATH-500 and GPQA Diamond are complementary, not redundant. MATH-500 tests structured symbolic reasoning with defined correct answers. GPQA Diamond tests PhD-level scientific reasoning in biology, chemistry, and physics — requiring open-ended domain knowledge application. Both are needed for a full frontier reasoning profile; neither substitutes for the other.

URLs: https://artificialanalysis.ai/evaluations/math-500


Part 2 — Agentic and Multi-Step Benchmarks

2.1 GAIA (General AI Assistants Benchmark)

Publisher: Meta AI Research / Hugging Face (joint) Original paper: 2023; maintained as a live leaderboard Credibility: HIGH — academic benchmark, maintained on Hugging Face, three-tiered difficulty with verifiable ground-truth answers

What it measures: Real-world assistant tasks requiring multi-step reasoning, web search, tool use, and file handling across 450+ questions with unambiguous, verifiable correct answers. Three difficulty levels: Level 1 (straightforward tool use), Level 2 (multi-step), Level 3 (complex multi-tool, multi-document).

Current standings (as of May 2026):

  • Scaffolded multi-agent systems (e.g., HAL + Claude Sonnet 4.5): ~74.6% [NEEDS VERIFICATION — sourced from HAL Princeton leaderboard]
  • Bare model performance: Claude Mythos Preview at ~52.3%, GPT-5.4 Pro at ~50.5% (BenchLM.ai snapshot, May 13, 2026) [NEEDS VERIFICATION]
  • Human baseline: 92%

The gap between scaffolded systems (74.6%) and bare models (~50%) illustrates a core measurement problem: top GAIA leaderboard entries are large multi-model ensembles, not single models. The score reflects system design and orchestration quality as much as any individual model’s capability.

Enterprise relevance: GAIA is the most widely cited public benchmark for general agent capability. The 50-point gap between human performance (92%) and best bare-model performance (~52%) represents the decision-making and tool-use reliability gap that enterprise agentic deployments must architect around. The 74.6% from scaffolded systems shows that prompt engineering, orchestration, and tool design close roughly half that gap — the other half remains a model limitation.

URLs: https://huggingface.co/spaces/gaia-benchmark/leaderboard · https://hal.cs.princeton.edu/gaia


2.2 τ-bench (Full Suite Context)

Publisher: Sierra Research Papers: arXiv:2406.12045 (original τ-bench), extended through τ²-bench (2025) and τ³-bench (2026) Credibility: MEDIUM — arXiv preprints; vendor-built; structural conflict of interest (Sierra sells agent infrastructure); evaluation design is worth studying independent of headline scores

What the full suite covers:

Version What it adds Key finding
τ-bench (2024) Tool-Agent-User interaction across retail + airline domains; LLM-simulated user; policy-constrained tool APIs; multi-turn Top models at ~70% pass@1
τ²-bench (2025) Co-ownership tasks — agent must coordinate with user to achieve shared objective, not just execute instructions Adds collaboration dimension absent from v1
τ³-bench (2026) Adds τ-knowledge (knowledge retrieval + action) and τ-voice (real-time voice grounding) Best model at 25.5% pass@1 on knowledge tasks; 26–38% on voice under realistic conditions

See also: τ-knowledge and τ-voice are documented in full in emerging-benchmarks-may-2026.md. This entry covers the progression of the full suite.

Enterprise relevance: The τ suite is the most systematic published progression of agent evaluation from task completion toward realistic deployment conditions. The τ²-bench jump from solo execution to co-ownership is specifically relevant for enterprises deploying copilot-style agents (human-in-the-loop) rather than fully autonomous systems. The benchmarks’ vendor origins are a credibility caveat for specific scores; the evaluation design is not.

URLs: https://sierra.ai/resources/research/tau-bench · https://github.com/sierra-research/tau-bench · https://github.com/sierra-research/tau2-bench


2.3 WebArena and VisualWebArena

Publisher: Carnegie Mellon University / web-arena-x consortium Credibility: HIGH — academic, multi-institution, independent of frontier model vendors

What they measure:

WebArena: 812 realistic web tasks across five self-hosted environments — e-commerce, forum, collaborative development, content management, and map. Tasks require planning across multiple web pages without provided step-by-step instructions. Human baseline: 78%.

VisualWebArena: 910 tasks requiring visual understanding of webpage content (screenshots, images, layout interpretation) in addition to web navigation. Human baseline: 88.7%. Specifically tests whether agents can process and act on visual information embedded in web pages — not just text content.

Current SOTA:

  • WebArena: AWA 1.5 at 57.14% [NEEDS VERIFICATION — sourced from Jace AI blog, not WebArena leaderboard primary]; OpAgent reported at 71.6% in February 2026 [NEEDS VERIFICATION on exact evaluation conditions]
  • VisualWebArena: original paper best result was 16.4% (mid-2024); subsequent agents using fine-tuning on WebGym improved to 42.9% on out-of-distribution sites [NEEDS VERIFICATION on whether this transfers to VisualWebArena directly]

WebArena-Verified: A verified version was presented at the Scaling Environments for Agents (SEA) Workshop at NeurIPS 2025, available via Docker as of February 2026. The verification pass addressed label quality and task ambiguity in the original dataset.

Enterprise relevance: WebArena is the reference benchmark for any enterprise evaluating AI agents for web-based workflows — market research, procurement portals, customer portal navigation, CRM interaction. The 57–71% range for top agents (vs. 78% human) means current agents can handle most straightforward web tasks but fail on complex multi-page planning. VisualWebArena is the relevant extension for any workflow where the web interface includes images, charts, or layout-dependent information.

URLs: https://webarena.dev · https://github.com/web-arena-x/visualwebarena · https://github.com/ServiceNow/webarena-verified


2.4 OSWorld

Publisher: xlang-ai (academic consortium, NeurIPS 2024) Credibility: HIGH — peer-reviewed (NeurIPS 2024), independent, real computer environment execution

What it measures: Multimodal agent performance on real computer use tasks — OS operations, application interaction, file management, multi-app workflows — executed in a live computer environment (Ubuntu VM), not a simulated or API-based approximation. 369 tasks across 9 application domains.

Current SOTA (as of April 2026):

  • GPT-5.4 (March 2026): 75.0% [NEEDS VERIFICATION]
  • Claude Opus 4.6 (February 2026): 72.7% [NEEDS VERIFICATION]
  • Human baseline: approximately 87% [NEEDS VERIFICATION — implied by the 84% of human performance reference]
  • OpenAI Computer-Using Agent (CUA): 38.1% (early 2025, when it led)

Progression signal: 38.1% (early 2025) → 72.7% (February 2026) in roughly 12 months. Computer-use agent capability roughly doubled on OSWorld in one year.

Enterprise relevance: OSWorld is the most direct benchmark for RPA (robotic process automation) replacement, desktop workflow automation, and any use case where an AI agent must operate software interfaces rather than call APIs. The 72–75% range means top agents handle most discrete computer-use tasks correctly but remain unreliable on complex multi-application workflows. The gap from ~73% to human-level (~87%) represents the edge cases that require human oversight or fallback.

URLs: https://github.com/xlang-ai/OSWorld · https://os-world.github.io/


2.5 WorkArena

Publisher: ServiceNow Research Paper: arXiv:2403.07718 (March 2024) Credibility: HIGH — academic preprint, independent benchmark design, enterprise software domain

What it measures: Agent performance on knowledge-work tasks in a realistic enterprise software environment (ServiceNow). Tasks span information retrieval, form completion, list filtering, and multi-step workflows within a full enterprise portal. Distinct from WebArena (general web) — WorkArena targets the specific patterns of enterprise business software.

Current SOTA: WorkArena-Legacy (original environment): AgentWorkflowMemory (AWM) at 35.5%. WorkArena-L1/L2/L3 (harder task tiers released subsequently) have lower published scores — [NEEDS VERIFICATION on current leaderboard rankings for all tiers].

Enterprise relevance: WorkArena is the most enterprise-relevant public web agent benchmark because it uses actual enterprise software patterns rather than consumer web environments. A 35.5% success rate on knowledge-work tasks in enterprise software is the published baseline for assessing current agent capability in workflows like ticketing, service requests, HR processes, and IT operations.

URLs: https://github.com/ServiceNow/WorkArena · arXiv:2403.07718


2.6 AppWorld

Publisher: Stony Brook University (ACL 2024 Best Resource Paper) Credibility: HIGH — peer-reviewed (ACL 2024 Best Resource Paper), deterministic state-based evaluation, no LLM judge

What it measures: Interactive coding agents operating across 9 day-to-day apps (email, calendar, files, contacts, shopping, banking, social, maps, notes) via 457 APIs in a controllable simulated environment populated with ~100 fictitious users. 750 tasks — 500 “normal” (routine multi-app operations) and 250 “challenge” (tasks requiring failure handling, multi-user coordination, constraint satisfaction). Evaluated by state-based unit tests that check not just task completion but collateral damage (unintended side effects).

Current SOTA:

  • GPT-4o: ~49% on normal tasks, ~30% on challenge tasks [NEEDS VERIFICATION — from original AppWorld paper; more recent model scores expected on leaderboard]
  • Most other models: at least 16 percentage points below GPT-4o on challenge tasks at time of original benchmark publication [NEEDS VERIFICATION — likely outdated given 2025–26 model releases]

Gap this fills: AppWorld evaluates multi-app coordination across heterogeneous APIs — the practical environment of enterprise productivity workflows. No other public benchmark combines multi-app scope, realistic user simulation, collateral-damage testing, and challenge-tier tasks. HumanEval and SWE-bench test code on isolated functions; AppWorld tests an agent’s ability to orchestrate code across a system of interacting applications.

URLs: https://appworld.dev · https://github.com/StonyBrookNLP/appworld · arXiv:2407.18901


2.7 SWE-bench Multimodal

Publisher: SWE-bench team (Princeton / Stanford) Paper: arXiv:2410.03859 (October 2024) Credibility: HIGH — same team as SWE-bench Verified, leaderboard maintained with private test set evaluation

What it measures: Software engineering agents working on JavaScript-based, user-facing applications — UI design systems, web app development, interactive mapping. All task instances contain images or videos in problem descriptions or testing scenarios. This forces agents to interpret visual context (screenshots, UI specs, diagrams) as part of the development task — the gap SWE-bench Verified leaves open.

How it differs from SWE-bench Verified:

  • SWE-bench Verified: 500 instances, Python-centric, text-only problem descriptions, ~70–76% top score
  • SWE-bench Multimodal: 100 development instances, JavaScript-centric, visual content required, ~30–35% top score

The performance collapse from ~73% (Verified) to ~30–35% (Multimodal) on the same class of models reveals a structural gap: AI coding agents that approach human-level performance on text-specified Python tasks fail significantly when visual context is required. A reported 73.2% performance drop on image-containing tasks within Multimodal [NEEDS VERIFICATION — cited in secondary analysis] is consistent with this gap.

Enterprise relevance: Any enterprise deploying AI coding agents on front-end, UI/UX, or full-stack JavaScript work should treat SWE-bench Multimodal scores as the relevant reference, not SWE-bench Verified. A vendor claiming 70% on SWE-bench for front-end work is likely citing Verified (Python, text-only) — not Multimodal.

URLs: https://www.swebench.com · arXiv:2410.03859


2.8 MultiAgentBench

Publisher: University of Illinois at Urbana-Champaign (ulab-uiuc) Paper: arXiv:2503.01935 (March 2025); published at ACL 2025 Credibility: HIGH — peer-reviewed (ACL 2025), open dataset, public GitHub

What it measures: LLM multi-agent system performance on both collaborative (shared goal) and competitive (conflicting goal) tasks. Evaluates coordination protocols — star, chain, tree, and graph topologies — and strategies including group discussion and cognitive planning. Uses milestone-based KPIs rather than binary success/failure to capture partial progress.

Key findings:

  • Graph coordination topology performs best among all tested structures on research scenarios
  • Cognitive planning improves milestone achievement rates by 3% on average
  • GPT-4o-mini reaches the highest average task score across tested configurations [NEEDS VERIFICATION — likely outdated by mid-2026 model releases]

Gap this fills: Agent benchmarks (GAIA, τ-bench, OSWorld, AppWorld) all evaluate single-agent systems. As multi-agent architectures — orchestrator/subagent patterns, parallel tool-calling agents, debate architectures — enter enterprise deployment, a benchmark measuring coordination, communication overhead, and conflict resolution becomes necessary. MultiAgentBench is the first peer-reviewed benchmark to address this gap directly.

Enterprise relevance: Enterprises building multi-agent systems for research, analysis, or complex workflow orchestration cannot rely on single-agent benchmark scores to predict multi-agent performance. MultiAgentBench provides the reference methodology for evaluating coordination quality rather than just task completion rate.

URLs: https://arxiv.org/abs/2503.01935 · https://github.com/ulab-uiuc/MARBLE


2.9 Berkeley Function-Calling Leaderboard (BFCL)

Publisher: UC Berkeley (Gorilla LLM group) Credibility: HIGH — academic, independent, continuously updated, paper published at ICML 2025 (PMLR)

What it measures: Tool-use / function-calling accuracy across simple function execution, multiple-function selection, parallel function calls, nested function calls, multi-turn conversations, and irrelevance detection (correctly refusing to call a function when none applies). As of V4, also includes agentic evaluation tasks.

Current standings (BFCL V4, as of December 2025 update):

  • Claude Opus 4.5 (FC): 77.47% overall accuracy — leading position [NEEDS VERIFICATION — from third-party tracking; gorilla.cs.berkeley.edu should be checked for current rankings]
  • Claude Sonnet 4.5 (FC): 73.24%
  • Proprietary models dominate the top tier; open-source models are closing the gap in select categories

Why it matters beyond raw accuracy: BFCL V4’s irrelevance detection category is specifically relevant for enterprise deployments — an agent that calls functions when it should refuse creates data integrity and security risks. Most function-calling benchmarks score only on correct execution; BFCL also scores on correct abstention. The agentic tasks added in V4 move the benchmark from “can it call the right tool” toward “can it decide which tools to call across a multi-step task.”

Enterprise relevance: BFCL is the reference benchmark for any enterprise evaluating models for tool-calling workflows — API orchestration, multi-system integration, workflow automation agents. Require vendors to report BFCL V4 scores broken out by category (especially irrelevance detection and multi-turn) rather than a single aggregate score.

URLs: https://gorilla.cs.berkeley.edu/leaderboard.html · https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard


Part 3 — Domain-Specific Benchmarks

3.1 LegalBench

Publisher: Hazan Research Group, Stanford University (collaborative — 40 contributors, legal professionals) Paper: arXiv:2308.11462 (August 2023); published at NeurIPS 2023 Credibility: HIGH — peer-reviewed (NeurIPS 2023), collaboratively built with legal practitioners, no commercial AI vendor involvement

What it measures: Legal reasoning across 162 tasks in six reasoning categories: issue spotting, rule recall, rule application, rule conclusion, interpretation, and rhetorical understanding. Tasks range from identifying relevant legal issues in fact patterns to applying specific statutory rules and interpreting contract language. Built by legal practitioners from actual legal materials, not generated examples.

SOTA context (as of mid-2026): [NEEDS VERIFICATION — no current maintained public leaderboard found; the hazyresearch.stanford.edu/legalbench site hosts the benchmark; model performance tracking is scattered across individual paper evaluations rather than a unified leaderboard]. The original paper evaluated 20 LLMs; frontier models have improved substantially since 2023. For current scores, consulting the GitHub repository (github.com/HazyResearch/legalbench) or Harvey LAB’s comparative evaluation is advised.

Gap this fills vs. Harvey LAB: LegalBench measures narrow legal reasoning tasks (short prompt → specific legal determination). Harvey LAB measures long-horizon legal agent tasks (instruction + materials → full deliverable). They are complementary: LegalBench for reasoning capability, Harvey LAB for agentic workflow performance. An enterprise evaluating AI for legal research needs both.

Enterprise relevance: LegalBench is the peer-reviewed standard for legal AI reasoning capability. Any vendor claiming legal AI capability should be asked for LegalBench scores alongside task-specific demos. The 162 tasks span enough breadth that aggregate performance correlates with general legal language competency — though domain-specific fine-tuning claims require evaluation on the specific task categories relevant to the deployment (e.g., contract interpretation vs. litigation issue spotting).

URLs: https://hazyresearch.stanford.edu/legalbench · https://github.com/HazyResearch/legalbench · arXiv:2308.11462


3.2 MedQA / MedBench / MedAgentsBench

Three distinct benchmarks covering different dimensions of medical AI evaluation:

MedQA (USMLE-style)

What it measures: Medical reasoning on USMLE-format questions (Step 1, 2, and 3 style). The benchmark has been the standard clinical AI reference since 2020.

Current SOTA: Frontier models now substantially exceed the passing threshold. GPT-4 in zero-shot scored approximately 71.6% on the full dataset; more recent models approach or exceed 90% on USMLE-style questions [NEEDS VERIFICATION — specific 2025–26 scores vary across evaluation conditions]. MedQA is approaching saturation for frontier models on standard format.

MedBench v4 (Chinese clinical AI, multi-modal and agent tracks)

Publisher: MedBench consortium (multi-institution, Chinese clinical domain) Paper: arXiv:2511.14439 (November 2025) Credibility: HIGH — 700,000+ expert-curated tasks across 24 primary and 91 secondary specialties; dedicated tracks for LLMs, multimodal models, and agents

Scores (as of November 2025 paper):

  • Base LLMs overall: mean 54.1/100; best performer Claude Sonnet 4.5 at 62.5/100
  • Multimodal models: mean 47.5/100; best performer GPT-5 at 54.9/100
  • Safety and ethics subscale: mean 18.4/100 across all models [NEEDS VERIFICATION]

The 18.4/100 safety score is notable: clinical AI systems that score competently on diagnosis-related tasks can simultaneously score very poorly on safety-critical behaviors. This is a structured reminder that aggregate benchmark scores mask safety-specific failure modes.

MedAgentsBench

Publisher: Gerstein Lab, Yale University Paper: arXiv:2503.07459 (March 2025) Credibility: HIGH — arXiv preprint, open dataset, focuses specifically on multi-step clinical reasoning rather than single-question recall

What it measures: 862 questions averaging 147 tokens each, combining seven medical test sets (MedQA, PubMedQA, MedMCQA, MedBullets, MMLU, MMLU-Pro, MedExQA, MedXpertQA) and filtering for questions that require genuine multi-step clinical reasoning — cases where base models still struggle despite overall high scores. Evaluates both thinking models (DeepSeek R1, o3) and agent frameworks.

Key finding: Thinking models (DeepSeek R1, o3) perform best on complex medical reasoning tasks. Search-based agent methods offer promising performance-to-cost ratios. The benchmark is explicitly designed to measure the gap that standard medical QA benchmarks miss: tasks where models appear capable in aggregate but fail on genuinely hard reasoning steps.

Enterprise relevance (all three): For enterprises evaluating AI in healthcare-adjacent contexts — clinical documentation, prior authorization, medical coding, patient communication — the progression from MedQA (single-question recall) to MedBench v4 (agent + safety tracks) to MedAgentsBench (complex clinical reasoning) mirrors the risk ladder from low-stakes to high-stakes deployment. Require evaluation at the relevant risk tier, not the easiest tier where scores are highest.

URLs: MedBench: arXiv:2511.14439 · MedAgentsBench: https://github.com/gersteinlab/medagents-benchmark · arXiv:2503.07459


3.3 CyberSecEval (Meta Purple Llama)

Publisher: Meta AI Research, with CyberSecEval 4 developed in collaboration with CrowdStrike Current version: CyberSecEval 4 (September 2025) Credibility: MEDIUM-HIGH — vendor-built (Meta) but methodology is open-source and independently applicable; CrowdStrike co-development on security operations components adds external validation

What it measures (progression by version):

  • CyberSecEval 1 (2023): Secure code generation and prompt injection resistance for coding assistants
  • CyberSecEval 2 (2024): Added cybersecurity knowledge tests, vulnerability exploitation assistance measurement, and prompt injection
  • CyberSecEval 4 (September 2025): Adds CyberSOCEval (Security Operations Center evaluation with two subtests: Malware Analysis and Threat Intelligence Reasoning) and AutoPatchBench (LLM agent capability to automatically patch vulnerabilities in native code)

CyberSOCEval (new in v4): Developed jointly with CrowdStrike. Tests whether AI systems can perform or meaningfully accelerate SOC workflows — malware analysis, threat intelligence correlation. This is the first published benchmark targeting AI-augmented SOC operations specifically.

AutoPatchBench (new in v4): Measures autonomous patching capability for security vulnerabilities in native code. Distinct from ExploitBench (which measures offensive exploitation capability) — AutoPatchBench is defensive.

Enterprise relevance: CyberSecEval is the reference evaluation for enterprises deploying AI in security contexts — coding assistants (secure code generation), SOC augmentation (CyberSOCEval), and automated remediation (AutoPatchBench). The CrowdStrike co-development on CyberSOCEval gives it more operational grounding than most vendor security benchmarks. Read in conjunction with ExploitBench (documented in emerging-benchmarks-may-2026.md) for a full offensive-defensive picture.

URLs: https://ai.meta.com/research/publications/purple-llama-cyberseceval-a-benchmark-for-evaluating-the-cybersecurity-risks-of-large-language-models/ · https://github.com/meta-llama/PurpleLlama · https://ir.crowdstrike.com/news-releases/news-release-details/crowdstrike-and-meta-deliver-new-benchmarks-evaluation-ai/


3.4 FinMCP-Bench

StateBench fixture: sa-042 materializes FinMCP-Bench as the MCP-specific financial tool-use source-acquisition task. Use it as the deployment preflight warning for AgentCore/MCP-style finance agents: single-tool success does not predict multi-turn workflow reliability.

Publisher: DianJin (Ant Group AI team) Paper: arXiv:2603.24943 (March 2026); published at ICASSP 2026 Credibility: HIGH — peer-reviewed (ICASSP 2026), real production MCP tool servers (not mocked), open dataset on Hugging Face

What it measures: LLM agent performance on real-world financial tool use under the Model Context Protocol (MCP). 613 tasks spanning 10 main financial scenarios and 33 sub-scenarios, using 65 real MCP-compliant financial tool servers drawn from production logs of the Qieman financial assistant app. Three task types: 145 single-tool, 249 multi-tool, and 219 multi-turn.

Why it matters: MCP became the de facto wiring standard for LLM tool use after Anthropic introduced it in late 2024; by early 2026, all major model providers had adopted it. FinMCP-Bench is the first benchmark built on actual production MCP servers in a specific domain — not mocked APIs or simulated environments.

Key finding: The best evaluated model scores 3.08% exact match on multi-turn financial tasks — a 20× performance collapse from single-tool performance. This is the agentic reliability gap in a high-stakes production domain: models that handle individual financial tool calls adequately fail dramatically when those calls must chain across multiple turns.

Enterprise relevance: The 20× collapse from single-tool to multi-turn is a direct operational finding for any enterprise building financial AI agents. Single-tool accuracy — the metric typically tested in vendor demos — does not predict multi-turn reliability. For financial services deploying AI agents in workflows requiring sequential tool calls (portfolio queries, transaction chains, multi-step compliance checks), evaluate specifically on multi-turn tasks and expect performance at or below single-digit accuracy without significant engineering investment.

URLs: https://arxiv.org/abs/2603.24943 · https://huggingface.co/datasets/DianJin/FinMCP-Bench · https://huggingface.co/papers/2603.24943


3.5 AgentBench

Publisher: Tsinghua University / ICLR 2024 Paper: arXiv:2308.03688 (August 2023, updated through October 2025) Credibility: HIGH — peer-reviewed (ICLR 2024), 8 environments, 29 models evaluated

What it measures: LLM-as-Agent performance across 8 environments: Operating System (bash command execution), Database (SQL queries in real databases), Knowledge Graph (traversal and reasoning), Digital Card Game, Lateral Thinking Puzzles, House-Holding (virtual household task completion), Web Shopping, and Web Browsing. Each task is a multi-turn interaction (5–50 turns) requiring planning and tool use.

Key findings from original evaluation (29 LLMs):

  • Long-term reasoning, decision-making, and instruction following are the primary bottlenecks
  • Code training has ambivalent impacts across different agent task types — helps some environments, hurts others
  • Performance gap between top API-based models and open-source models was substantial at publication; this gap has narrowed significantly since 2024 [NEEDS VERIFICATION on current leaderboard]

Status in 2026: AgentBench has been updated through October 2025. As a multi-environment benchmark from ICLR 2024, it remains a useful breadth evaluation but has not kept pace with the rapid advancement of agent-specific benchmarks. Individual environments within AgentBench have been superseded by more sophisticated alternatives (OSWorld for OS tasks; WorkArena for web tasks; AppWorld for multi-app). Its primary value in 2026 is as a standardized breadth reference and for evaluating the 8-environment coverage simultaneously.

URLs: https://arxiv.org/abs/2308.03688 · https://llmbench.github.io/


Part 4 — Evaluation Methodology Research

4.1 Chatbot Arena — Critique Papers (2025)

Two significant 2025 papers address Chatbot Arena’s reliability:

“Gaming the Arena: AI Model Evaluation and the Viral Capture of Attention” (arXiv:2512.15252, December 2025) Examines whether optimization pressure on Chatbot Arena’s user-preference scoring creates gaming incentives — specifically, whether training on Arena interaction data improves Arena-specific scores independent of general capability. Distinct from contamination (a training data problem); this is a platform-adaptation problem.

“Beyond Benchmarks: How Users Evaluate AI Chat Assistants” (arXiv:2603.25220, March 2026) Challenges the assumption that blind A/B preference voting in Arena captures real-world deployment quality. Users evaluating in a benchmark context optimize for different attributes than users in sustained deployment (single-response quality vs. interface, reliability, consistency over time, pricing). Arena captures isolated response quality; it does not measure the full deployment experience.

Enterprise implication: Chatbot Arena Elo remains the least gameable single-metric public benchmark — but neither paper falsifies it. They establish its limitations: (1) platform-adapted models may score above their general capability; (2) Arena scores predict single-response quality more reliably than deployment satisfaction. For procurement decisions, use Arena as a proxy for response quality per turn — not as a proxy for deployment reliability or user experience at scale.

URLs: arXiv:2512.15252 · arXiv:2603.25220


4.2 “Can We Trust AI Benchmarks?” — 445-Benchmark Review

Already documented in benchmark-foundations-enterprise-evals-2026.md. Cross-reference: arXiv:2502.06559. The core finding — pervasive construct-validity gaps across 445 LLM benchmarks — underlies all the validity caveats in this note.


Key Data Points Table

Benchmark Domain Current SOTA Human Baseline Status Source
FrontierMath (Tiers 1–3) Research math GPT-5.5 Pro: ~52% N/A (requires professional mathematicians) Still discriminating Epoch AI [NEEDS VERIFICATION]
FrontierMath (Tier 4) Hardest research math GPT-5.5 Pro: 39.6% N/A Still discriminating Epoch AI [NEEDS VERIFICATION]
MMMU-Pro Expert multimodal reasoning GPT-5.4 Pro: 94% N/A Still discriminating Third-party tracking [NEEDS VERIFICATION]
LiveBench Multi-task contamination-free o3-mini: ~0.846 N/A Continuously updated llm-stats.com [NEEDS VERIFICATION]
MATH-500 Competition math DeepSeek R1: 97.3% ~90% Approaching saturation at frontier Artificial Analysis
GAIA (bare model) General multi-step assistant ~52% (Claude Mythos Preview) 92% 40-point gap to human BenchLM.ai, May 2026 [NEEDS VERIFICATION]
GAIA (scaffolded system) General multi-step assistant ~74.6% (HAL + Sonnet 4.5) 92% 17-point gap to human HAL Princeton [NEEDS VERIFICATION]
WebArena Web navigation AWA 1.5: 57.14%; OpAgent: 71.6% 78% Active progress, near human Multiple sources [NEEDS VERIFICATION]
VisualWebArena Visual web navigation ~43% (WebGym fine-tuned) 88.7% Large gap to human arXiv:2401.13649 + WebGym [NEEDS VERIFICATION]
OSWorld Computer use (OS) GPT-5.4: 75.0% ~87% Rapid progress; ~84% of human Third-party tracking [NEEDS VERIFICATION]
WorkArena Enterprise software agent AWM: 35.5% N/A Significant gap WorkArena leaderboard
AppWorld (normal tasks) Multi-app daily workflows GPT-4o: ~49% N/A Large gap; newer models expected higher arXiv:2407.18901 [NEEDS VERIFICATION — likely outdated]
SWE-bench Verified Python code issue resolution 76.8% (Claude Opus 4.5) ~100% 23 pp gap SWE-bench leaderboard
SWE-bench Multimodal Visual/JS code tasks ~30–35% top models ~100% Large gap; 40 pp below Verified arXiv:2410.03859 [NEEDS VERIFICATION]
BFCL V4 Tool/function calling Claude Opus 4.5: 77.47% N/A Active leaderboard gorilla.cs.berkeley.edu [NEEDS VERIFICATION]
GAIA (all levels) Multi-step general assistant 74.6% (scaffolded) 92% 17 pp gap scaffolded HAL Princeton [NEEDS VERIFICATION]
MedBench v4 (base LLM) Clinical AI overall Claude Sonnet 4.5: 62.5/100 N/A Safety subscale critical (18.4/100) arXiv:2511.14439
FinMCP-Bench (multi-turn) Financial tool use MCP 3.08% exact match (best model) N/A Severe gap at multi-turn arXiv:2603.24943
MultiAgentBench Multi-agent coordination GPT-4o-mini leads [NEEDS VERIFICATION] N/A Newly established; likely outdated ACL 2025
τ-bench (v1) Tool-agent-user interaction ~70% pass@1 N/A Active; v3 extends to knowledge+voice arXiv:2406.12045

What This Means for Your Organization

1. The benchmark landscape has split into two tiers

Saturation tier: MMLU, GSM8K, MATH-500 (at the frontier). Any vendor still leading with these numbers is citing evaluations where all top models cluster above 90%. They do not discriminate.

Discrimination tier: FrontierMath, MMMU-Pro, GAIA (bare), OSWorld, SWE-bench Multimodal, FinMCP-Bench multi-turn. These still separate models meaningfully — and in some cases (FinMCP-Bench multi-turn, AppWorld challenge tasks), frontier models score in single digits.

When a vendor leads with the saturation tier and avoids the discrimination tier, ask why.

2. The multimodal-text gap is larger than marketing suggests

MMMU-Pro requires genuine visual reasoning — and while top models score 83–94%, the SWE-bench Verified to Multimodal collapse (~73% to ~30–35%) on the same models shows that visual content in a task context can cut performance in half. Any enterprise workflow involving screenshots, diagrams, PDFs with images, or UI components should evaluate specifically on multimodal benchmarks — not infer from text-only scores.

3. Multi-turn performance is the right metric for agent deployments, not single-call accuracy

FinMCP-Bench’s 20× collapse from single-tool (roughly 60–70% for top models) to multi-turn (3.08% exact match) is the starkest published illustration of a pattern visible across benchmarks: agents that appear competent in single-interaction demos fail dramatically when tasks require sequential tool calls across multiple turns. Vendor demos are almost always single-interaction. Require multi-turn evaluation for any workflow involving chained steps.

4. Safety and ethics scores are systematically low and systematically underreported

MedBench v4’s 18.4/100 mean safety score across all models — on a clinical benchmark — is not a medical-domain anomaly. It reflects a broad pattern: models are evaluated and marketed on capability scores (accuracy, task completion) while safety subscores are buried or absent. For regulated industries (healthcare, financial services, legal), require explicit safety and refusal-behavior evaluation in addition to task accuracy.

5. Domain-specific evaluation has a research-to-enterprise gap

LegalBench (NeurIPS 2023) lacks an actively maintained public leaderboard — the authoritative source for current scores is scattered across individual paper evaluations. FinanceBench is well-documented in research but has no unified procurement-ready leaderboard. MedBench v4 is Chinese clinical domain. The most sophisticated domain benchmarks (Harvey LAB for legal agents, FinMCP-Bench for financial MCP agents, MedAgentsBench for complex clinical reasoning) are all less than 18 months old. For enterprise procurement, the practical implication is that no public benchmark fully covers your domain’s specific task distribution. External benchmarks narrow the field; internal domain evals close the gap.

6. Multi-agent systems require multi-agent benchmarks

Single-agent benchmarks (GAIA, OSWorld, AppWorld) do not predict multi-agent system performance. As orchestrator-subagent architectures, debate systems, and parallel task-splitting agents enter enterprise deployment, MultiAgentBench provides the reference evaluation methodology. If you are building a multi-agent system, benchmark the system, not the components.


Source Notes

All [NEEDS VERIFICATION] flags in this note indicate that a figure was obtained from third-party tracking sites (llm-stats.com, benchlm.ai, artificialanalysis.ai, pricepertoken.com) rather than confirmed against the primary leaderboard or paper. Third-party trackers are generally accurate but may reflect stale data or different evaluation conditions. Before citing any specific score in client-facing materials, verify against:

  • The benchmark’s official leaderboard or GitHub repository
  • The model’s official model card or technical report

This note is a survey instrument for the State of AI corpus. It is not a primary source.


Brandon Sneider | brandon@brandonsneider.com May 2026