Technical Workbench
Benchmarks
25 documents
Artificial Analysis — Inference Benchmarks and Independent Evaluation Landscape 2026
> **Source credibility: MEDIUM-HIGH. TIER 1.**
May 2026
Benchmark Evolution and Enterprise Evals — 2026
> **Source credibility: LOW-MEDIUM. TIER 1–2.**
May 2026
Benchmark Evolution and Enterprise Evaluation Practice — 2026
> **Source credibility: HIGH (Stanford HAI AI Index 2026 — independent academic) / MEDIUM-HIGH (peer-reviewed preprints). TIER 1.**
May 2026
Corporate Evals, Trace Mining, and Automatic Skill Improvement — 2026
The cutting edge in enterprise evals is no longer "which model scored highest." It is **which workflow environment produces trustworthy evidence, which failures become reusable improvements, and which
March 2026
Data-Agent and DataOps Evaluation Landscape — 2026
This synthesis is backed by primary source cards under `sources/21-benchmarks/`
March 2026
Decentralized GitLab Eval Federation — 2026
> **Source status:** This is an architecture pattern grounded in GitLab's documented CI/CD runner, artifacts, downstream-pipeline, and CI/CD component capabilities. It is not a GitLab product claim.
June 2026
Domain YAML Generation Optimization 2026
For domain-specific YAML generation, the first optimization target should not
June 2026
Emerging AI Benchmarks — Inventory, April–May 2026
> **Source credibility: MIXED** — Reference inventory of benchmarks announced or discussed April–May 2026. Individual entries carry their own credibility rating (HIGH/MEDIUM/LOW).
May 2026
Enterprise AI Evals Field Guide — 2026
> **Ingestion status:** The Awesome Evals ingestion materializes 318 cited rows, 318 golden-quality exemplar rows, 269 raw mirrors, 143 mirrored BenchFlow notes, and 83 Chandra-extracted PDF rows, inc
June 2026
Enterprise Eval for Fine-Tuned SLMs and RAG Pipelines
> **Source credibility:** GitHub repo data is TIER 1 (primary source, observed directly). Peer-reviewed papers are TIER 1. Vendor documentation is TIER 2.
May 2026
Enterprise Evaluation for Knowledge Bases and Coding Agents — 2026
> **Source status:** This report ingests Andrei Lopatenko's `LLMEvaluation` compendium as a discovery index and taxonomy source. The compendium itself is a curated map, not a measurement result.
June 2026
EnterpriseClawBench: Enterprise Agent Evals from Real Workplace Sessions
Hugging Face paper page: https://huggingface.co/papers/2606.23654
June 2026
Eval Framework Schema Crosswalk — 2026
> **Source status:** This is a schema and operating-model crosswalk, not a claim that each framework exposes a stable public JSON schema for all internals.
June 2026
Finance AI Benchmarks 2025–2026: What the Evidence Actually Says About AI in Financial Work
This synthesis is backed by the source cards and mirrored artifacts listed in
May 2026
Finance-Domain Models and Benchmarks: What We Were Missing
> **Source credibility: MEDIUM. TIER 1-2.**
May 2026
Finance Kaggle Competition Artifacts: Missing Medium for Quant Eval Design
> **Source credibility: MEDIUM. TIER 2-3.**
May 2026
Finance Measured Runtime Evidence Ledger
> **Document role:** separate measured model/runtime evidence from static
June 2026
Frontier and Domain-Specific AI Benchmarks — Survey Note, 2025–2026
> **Source credibility: VARIES BY SECTION.** This note synthesizes frontier benchmark coverage from arXiv preprints (TIER 1–2), official leaderboards and model cards, and third-party tracking sites (l
May 2026
Frontier Lab Evaluation Research and SDK Primitive Crosswalk — 2026
> **Source status:** This report synthesizes primary lab documentation, SDK references, model-card benchmark disclosures, and eval papers retrieved on June 21, 2026.
June 2026
Local Model Benchmarks: A CIO's Hardware and Quality Reference Guide
> Hardware throughput benchmarks (llama.cpp GitHub discussions, MLX Apple Silicon benchmark paper arxiv:2510.18921, singhajit.com measured results, llmcheck.net standardized tests) = **HIGH** — all re
May 2026
Local Podcast Transcription MLX Models (May 2026)
task-relevant, and includes finance-like `earnings21` / `earnings22` subsets,
March 2026
StateBench Finance Replay Model Candidate Registry
This registry turns the finance/quant model landscape into executable
May 2026
Trace-to-Training-Pair Conversion — Layer 1 of the Enterprise SLM Pipeline
> **Source credibility:** GitHub repos and official SDK documentation are TIER 1 (primary source, observed directly). Vendor product claims without independent validation are TIER 2.
May 2026
Unsolved Hard Problems in RAG and Agentic AI — May 2026
Cross-references within pillar: rag-retrieval-leaderboard-landscape-2026 · surprising-slm-bench-toppers-2026 · enterprise-eval-rag-finetuned-slm-2026 · enterprise-abstention-whitepaper-2026
May 2026
VLM, Computer Use, and Web Agent Benchmarks — 2026
> **Source credibility: VARIES BY SECTION.** Covers academic benchmarks (TIER 1 — arXiv + peer review), vendor-published benchmarks (TIER 2), and practitioner signals (TIER 3).
May 2026