Benchmarks

34 documents

Artificial Analysis — Inference Benchmarks and Independent Evaluation Landscape 2026
Source credibility: MEDIUM-HIGH. TIER 1.
May 2026
Automatic Term Extraction for Enterprise RAG: From Chat Vocabulary to a Governed Search Index
and current data-curation documentation retrieved and rechecked 2026-08-01.
August 2026
BAM Embeddings (Greenback): Domain-Specific Financial Text Retrieval from Balyasny
Anderson et al. (Balyasny Asset Management), "Greenback Bears and Fiscal Hawks: Finance is a Jungle and Text Embeddings Must Adapt," EMNLP 2024 Industry Track. arXiv:2411.07142.
March 2026
Benchmark Evolution and Enterprise Evals — 2026
Source credibility: LOW-MEDIUM. TIER 1–2.
May 2026
Benchmark Evolution and Enterprise Evaluation Practice — 2026
Source credibility: HIGH (Stanford HAI AI Index 2026 — independent academic) / MEDIUM-HIGH (peer-reviewed preprints). TIER 1.
May 2026
Corporate Evals, Trace Mining, and Automatic Skill Improvement — 2026
The cutting edge in enterprise evals is no longer "which model scored highest." It is which workflow environment produces trustworthy evidence, which failures become reusable improvements, and which
March 2026
Data-Agent and DataOps Evaluation Landscape — 2026
This synthesis is backed by primary source cards under sources/21-benchmarks/
March 2026
Decentralized GitLab Eval Federation — 2026
Source status: This is an architecture pattern grounded in GitLab's documented CI/CD runner, artifacts, downstream-pipeline, and CI/CD component capabilities. It is not a GitLab product claim.
June 2026
Domain YAML Generation Optimization 2026
For domain-specific YAML generation, the first optimization target should not
June 2026
Emerging AI Benchmarks — Inventory, April–May 2026
Source credibility: MIXED — Reference inventory of benchmarks announced or discussed April–May 2026. Individual entries carry their own credibility rating (HIGH/MEDIUM/LOW).
May 2026
Enterprise AI Evals Field Guide — 2026
Ingestion status: The Awesome Evals ingestion materializes 318 cited rows, 318 golden-quality exemplar rows, 269 raw mirrors, 143 mirrored BenchFlow notes, and 83 Chandra-extracted PDF rows, inc
June 2026
Enterprise AI Evals, Explained: Test the Work, Not the Model
The enterprise evaluation stack: source, retrieval, response, outcome, and monitoring
August 2026
Enterprise Eval for Fine-Tuned SLMs and RAG Pipelines
Source credibility: GitHub repo data is TIER 1 (primary source, observed directly). Peer-reviewed papers are TIER 1. Vendor documentation is TIER 2.
May 2026
Enterprise Evaluation for Knowledge Bases and Coding Agents — 2026
Source status: This report ingests Andrei Lopatenko's LLMEvaluation compendium as a discovery index and taxonomy source. The compendium itself is a curated map, not a measurement result.
June 2026
EnterpriseClawBench: Enterprise Agent Evals from Real Workplace Sessions
Hugging Face paper page: https://huggingface.co/papers/2606.23654
June 2026
Competing Eval-Authoring Frameworks
Pydantic Evals is the typed contract layer.
March 2026
Eval Framework Schema Crosswalk — 2026
Source status: This is a schema and operating-model crosswalk, not a claim that each framework exposes a stable public JSON schema for all internals.
June 2026
Eval Frameworks for GitLab Containers and Agent Harnesses
These frameworks were previously present at uneven depth.
March 2026
Finance AI Benchmarks 2025–2026: What the Evidence Actually Says About AI in Financial Work
This synthesis is backed by the source cards and mirrored artifacts listed in
May 2026
Finance-Domain Models and Benchmarks: What We Were Missing
Source credibility: MEDIUM. TIER 1-2.
May 2026
Finance Kaggle Competition Artifacts: Missing Medium for Quant Eval Design
Source credibility: MEDIUM. TIER 2-3.
May 2026
Finance Measured Runtime Evidence Ledger
Document role: separate measured model/runtime evidence from static
June 2026
Financial RAG: Hybrid+Rerank Achieves Recall@5 of 0.816; BM25 Outperforms Dense Retrieval on 23k Financial Queries
Akarso, Karaman & Mierbach (Technische Hochschule Ingolstadt / TU Eindhoven / Raduate).
March 2026
Frontier and Domain-Specific AI Benchmarks — Survey Note, 2025–2026
Source credibility: VARIES BY SECTION. This note synthesizes frontier benchmark coverage from arXiv preprints (TIER 1–2), official leaderboards and model cards, and third-party tracking sites (l
May 2026
Frontier Lab Evaluation Research and SDK Primitive Crosswalk — 2026
Source status: This report synthesizes primary lab documentation, SDK references, model-card benchmark disclosures, and eval papers retrieved on June 21, 2026.
June 2026
Evals Skills, Harness Engineering, and Agent-Operable Evaluation
Source status: primary repository and primary OpenAI engineering article reviewed 2026-07-25; discovery began from an authenticated browser capture on X.
March 2026
LangChain Eval Engineering Skill on Harbor
LangChain's July 22, 2026 eval-engineering skill turns repository context and selected production traces into reviewed, executable Harbor tasks.
March 2026
Local Model Benchmarks: A CIO's Hardware and Quality Reference Guide
Hardware throughput benchmarks (llama.cpp GitHub discussions, MLX Apple Silicon benchmark paper arxiv:2510.18921, singhajit.com measured results, llmcheck.net standardized tests) = HIGH — all re
May 2026
Local Podcast Transcription MLX Models (May 2026)
task-relevant, and includes finance-like earnings21 / earnings22 subsets,
March 2026
Semantica-AGI in the State of AI Corpus
Semantica-AGI belongs in the corpus as a context-graph, ontology, provenance, temporal-state, reasoning, and decision-lineage layer.
August 2026
StateBench Finance Replay Model Candidate Registry
This registry turns the finance/quant model landscape into executable
May 2026
Trace-to-Training-Pair Conversion — Layer 1 of the Enterprise SLM Pipeline
Source credibility: GitHub repos and official SDK documentation are TIER 1 (primary source, observed directly). Vendor product claims without independent validation are TIER 2.
May 2026
Unsolved Hard Problems in RAG and Agentic AI — May 2026
Cross-references within pillar: rag-retrieval-leaderboard-landscape-2026 · surprising-slm-bench-toppers-2026 · enterprise-eval-rag-finetuned-slm-2026 · enterprise-abstention-whitepaper-2026
May 2026
VLM, Computer Use, and Web Agent Benchmarks — 2026
Source credibility: VARIES BY SECTION. Covers academic benchmarks (TIER 1 — arXiv + peer review), vendor-published benchmarks (TIER 2), and practitioner signals (TIER 3).
May 2026