{"all_urls": ["https://ysymyth.github.io/The-Second-Half/"], "content_sha256": "211dcd740a14ba0b10babff55cc7cb2ee86244e82ab3362d885829ee2d164e14", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/yao-second-half.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-001-the-second-half.raw.txt", "note_path": "notes/articles/yao-second-half.md", "primary_url": "https://ysymyth.github.io/The-Second-Half/", "section": "⭐ Must-read starter set (read these first)", "source_id": "ae-001-the-second-half", "source_line": 39, "source_type": "web_article", "status": "mirrored", "title": "The Second Half", "why_it_matters": "\"Evaluation becomes more important than training.\" The field-level *why*."}
{"all_urls": ["https://eugeneyan.com/writing/eval-process/"], "content_sha256": "9b432b0da627bdda285f2fe9acb204ed196470da3c096e89f96e550999fea84d", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-002-an-llm-as-judge-won-t-save-the-product-fixing-yo.raw.txt", "note_path": null, "primary_url": "https://eugeneyan.com/writing/eval-process/", "section": "⭐ Must-read starter set (read these first)", "source_id": "ae-002-an-llm-as-judge-won-t-save-the-product-fixing-yo", "source_line": 40, "source_type": "web_article", "status": "mirrored", "title": "An LLM-as-Judge Won't Save the Product, Fixing Your Process Will", "why_it_matters": "Process over tooling; evals as the scientific method."}
{"all_urls": ["https://leehanchung.github.io/blogs/2026/06/13/hidden-technical-debt-agent-evaluation-infra/"], "content_sha256": "0c84dd759fb4cd17d700488033f3ff82b01da2f20fb197182868451e77e008d0", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/leehanchung-hidden-technical-debt-agent-runtime.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-003-hidden-technical-debt-agent-evaluation-infrastru.raw.txt", "note_path": "notes/articles/leehanchung-hidden-technical-debt-agent-runtime.md", "primary_url": "https://leehanchung.github.io/blogs/2026/06/13/hidden-technical-debt-agent-evaluation-infra/", "section": "⭐ Must-read starter set (read these first)", "source_id": "ae-003-hidden-technical-debt-agent-evaluation-infrastru", "source_line": 41, "source_type": "blog", "status": "mirrored", "title": "Hidden Technical Debt: Agent Evaluation Infrastructure", "why_it_matters": "Control/data plane, the five eval surfaces, state deltas. \"Chat eval was a spreadsheet; agent eval is a system.\""}
{"all_urls": ["https://hamel.dev/blog/posts/evals-faq/"], "content_sha256": "b5d5398f91d39542cc52d6c11bc38dfb86e4da2d06fcb73e69add860602d8e02", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-004-llm-evals-faq.raw.txt", "note_path": null, "primary_url": "https://hamel.dev/blog/posts/evals-faq/", "section": "⭐ Must-read starter set (read these first)", "source_id": "ae-004-llm-evals-faq", "source_line": 42, "source_type": "blog", "status": "mirrored", "title": "LLM Evals FAQ", "why_it_matters": "The densest operational Q&A: error analysis, binary judgments, the benevolent-dictator labeler."}
{"all_urls": ["https://www.jasonwei.net/blog/asymmetry-of-verification-and-verifiers-law"], "content_sha256": "be6f54cb997e1d6095c362a8d722e30de68819787535cf8385441c8c8ea43492", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-005-asymmetry-of-verification-and-verifier-s-law.raw.txt", "note_path": null, "primary_url": "https://www.jasonwei.net/blog/asymmetry-of-verification-and-verifiers-law", "section": "⭐ Must-read starter set (read these first)", "source_id": "ae-005-asymmetry-of-verification-and-verifier-s-law", "source_line": 43, "source_type": "blog", "status": "mirrored", "title": "Asymmetry of Verification and Verifier's Law", "why_it_matters": "\"Ability to verify == ability to create an RL environment.\""}
{"all_urls": ["https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents"], "content_sha256": "59ea5d13bccd08ddfc548b26d0ce4591c4047c07592472ed35c011de38c5f258", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/vanishing-gradients-agents-evals.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-006-demystifying-evals-for-ai-agents.raw.txt", "note_path": "notes/articles/vanishing-gradients-agents-evals.md", "primary_url": "https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents", "section": "⭐ Must-read starter set (read these first)", "source_id": "ae-006-demystifying-evals-for-ai-agents", "source_line": 44, "source_type": "web_article", "status": "mirrored", "title": "Demystifying Evals for AI Agents", "why_it_matters": "Best primary on agent-specific evals: task design, outcome vs trajectory, isolated trials, pass@k vs pass^k."}
{"all_urls": ["https://ofir.io/How-to-Build-Good-Language-Modeling-Benchmarks/"], "content_sha256": "cb4d4263c57f9558283b48d61549aa99a8931294452efa8cbe2860c1b4c0b7cd", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-007-how-to-build-good-language-modeling-benchmarks.raw.txt", "note_path": null, "primary_url": "https://ofir.io/How-to-Build-Good-Language-Modeling-Benchmarks/", "section": "⭐ Must-read starter set (read these first)", "source_id": "ae-007-how-to-build-good-language-modeling-benchmarks", "source_line": 45, "source_type": "web_article", "status": "mirrored", "title": "How to Build Good Language Modeling Benchmarks", "why_it_matters": "Natural / auto-evaluatable / challenging; the \"-200%\" difficulty target; ~1-yr saturation."}
{"all_urls": ["https://arxiv.org/abs/2407.01502"], "content_sha256": "a74a66bddeea47fd5b656b5bd316858bb9baf7b951491316ec773e834c43f53c", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-008-ai-agents-that-matter.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2407.01502", "section": "⭐ Must-read starter set (read these first)", "source_id": "ae-008-ai-agents-that-matter", "source_line": 46, "source_type": "paper_or_pdf", "status": "mirrored", "title": "AI Agents That Matter", "why_it_matters": "Cost as a first-class metric; model-dev vs app-dev; missing holdouts breed overfitting."}
{"all_urls": ["https://www.interconnects.ai/p/building-on-evaluation-quicksand"], "content_sha256": "b9885f095cf3cc54596b20530d30dbfae608ada2f47dec0aaa19cd751e883fb6", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-009-building-on-evaluation-quicksand.raw.txt", "note_path": null, "primary_url": "https://www.interconnects.ai/p/building-on-evaluation-quicksand", "section": "⭐ Must-read starter set (read these first)", "source_id": "ae-009-building-on-evaluation-quicksand", "source_line": 47, "source_type": "blog", "status": "mirrored", "title": "Building on Evaluation Quicksand", "why_it_matters": "LLM eval has no ground truth; contamination; eval↔training coupling."}
{"all_urls": ["https://arxiv.org/abs/2404.12272"], "content_sha256": "2a98cac4c82e1b225ab4c6468262c9245923882e276c6f518b0142cc02ecae5a", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-010-who-validates-the-validators-evalgen.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2404.12272", "section": "⭐ Must-read starter set (read these first)", "source_id": "ae-010-who-validates-the-validators-evalgen", "source_line": 48, "source_type": "paper_or_pdf", "status": "mirrored", "title": "Who Validates the Validators? (EvalGen)", "why_it_matters": "\"Criteria drift\": you can't write the rubric before you grade."}
{"all_urls": ["https://florianbrand.com/posts/benches-2026"], "content_sha256": "fc8d7716215d0260805cab35106c2f3c3fd1384fe70eaa946cbe3d80a5d2eebe", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/florian-brand-prime-intellect-llm-benchmarks-era-of-agents.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-011-benches-2026-llm-benchmarks-in-the-era-of-agents.raw.txt", "note_path": "notes/articles/florian-brand-prime-intellect-llm-benchmarks-era-of-agents.md", "primary_url": "https://florianbrand.com/posts/benches-2026", "section": "⭐ Must-read starter set (read these first)", "source_id": "ae-011-benches-2026-llm-benchmarks-in-the-era-of-agents", "source_line": 49, "source_type": "web_article", "status": "mirrored", "title": "Benches 2026 — \"LLM benchmarks in the era of agents\"", "why_it_matters": "The sharpest current read on why benchmarks break in the agent era: the \"evals are dead, just measure vibes\" backlash, how every layer of the eval-running stack (prompt · sampling temp · grader · harness) swings the score, and that benchmark ground truth is frequently wrong."}
{"all_urls": ["https://openai.com/index/trustworthy-third-party-evaluations-foundations/"], "content_sha256": "442f66e4c9dc7a6e09364da71f1c14ca0fdf85402fe2c01e903a33ef1c4659ad", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-012-a-shared-playbook-for-trustworthy-third-party-ev.raw.txt", "note_path": null, "primary_url": "https://openai.com/index/trustworthy-third-party-evaluations-foundations/", "section": "⭐ Must-read starter set (read these first)", "source_id": "ae-012-a-shared-playbook-for-trustworthy-third-party-ev", "source_line": 50, "source_type": "docs_or_book", "status": "mirrored", "title": "A Shared Playbook for Trustworthy Third-Party Evaluations", "why_it_matters": "What makes *independent* evals of frontier-model safeguards & capabilities trustworthy: harness selection, the validity hazards that distort results, and the standards third-party evaluators need."}
{"all_urls": ["https://ysymyth.github.io/The-Second-Half/"], "content_sha256": "211dcd740a14ba0b10babff55cc7cb2ee86244e82ab3362d885829ee2d164e14", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/yao-second-half.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-013-the-second-half.raw.txt", "note_path": "notes/articles/yao-second-half.md", "primary_url": "https://ysymyth.github.io/The-Second-Half/", "section": "1 · Why we need evals", "source_id": "ae-013-the-second-half", "source_line": 56, "source_type": "web_article", "status": "mirrored", "title": "The Second Half", "why_it_matters": "The bottleneck shifts from solving problems to *defining and evaluating* them. (also T2, T7)"}
{"all_urls": ["https://eugeneyan.com/writing/eval-process/"], "content_sha256": "b08e2853cdf61563b2217567d7226d7c6a1c40cb6dfdc4185affbc4e9231cbea", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-014-an-llm-as-judge-won-t-save-the-product-fixing-yo.raw.txt", "note_path": null, "primary_url": "https://eugeneyan.com/writing/eval-process/", "section": "1 · Why we need evals", "source_id": "ae-014-an-llm-as-judge-won-t-save-the-product-fixing-yo", "source_line": 57, "source_type": "web_article", "status": "mirrored", "title": "An LLM-as-Judge Won't Save the Product, Fixing Your Process Will", "why_it_matters": "\"Buying or building another evaluation tool won't save the product.\" Evals = the scientific method in disguise."}
{"all_urls": ["https://hamel.dev/blog/posts/evals/"], "content_sha256": "6083d93b8cd419ee5b9bf5400bd4c2de09461095d39a172f1d4939ed7b4a1950", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-015-your-ai-product-needs-evals.raw.txt", "note_path": null, "primary_url": "https://hamel.dev/blog/posts/evals/", "section": "1 · Why we need evals", "source_id": "ae-015-your-ai-product-needs-evals", "source_line": 58, "source_type": "blog", "status": "mirrored", "title": "Your AI Product Needs Evals", "why_it_matters": "The canonical \"you need evals\"; remove all friction from looking at your data; don't rely on generic frameworks."}
{"all_urls": ["https://hamel.dev/blog/posts/field-guide/"], "content_sha256": "7ce00cfbe0dea2e01003ed2f138b7d48c317796161c58f030b6020db255f4c3b", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-pod-vg50-field-guide.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-016-a-field-guide-to-rapidly-improving-ai-products.raw.txt", "note_path": "notes/talks/talk-pod-vg50-field-guide.md", "primary_url": "https://hamel.dev/blog/posts/field-guide/", "section": "1 · Why we need evals", "source_id": "ae-016-a-field-guide-to-rapidly-improving-ai-products", "source_line": 59, "source_type": "blog", "status": "mirrored", "title": "A Field Guide to Rapidly Improving AI Products", "why_it_matters": "\"Error analysis is consistently the highest-ROI activity.\" The metric for an AI roadmap is experiments run."}
{"all_urls": ["https://www.sh-reya.com/blog/in-defense-ai-evals/"], "content_sha256": "2f37bf6b0ad6b9c317b0d5b083418b3a88de027d3bdb265a3398e058c61640de", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-017-in-defense-of-ai-evals-for-everyone.raw.txt", "note_path": null, "primary_url": "https://www.sh-reya.com/blog/in-defense-ai-evals/", "section": "1 · Why we need evals", "source_id": "ae-017-in-defense-of-ai-evals-for-everyone", "source_line": 60, "source_type": "blog", "status": "mirrored", "title": "In Defense of AI Evals, for Everyone", "why_it_matters": "Rebuts the anti-eval backlash; evals = the systematic measurement of application quality."}
{"all_urls": ["https://applied-llms.org/", "https://www.oreilly.com/radar/what-we-learned-from-a-year-of-building-with-llms-part-ii/"], "content_sha256": "e8b4979dd187bfda59f3a57b7692da0ef99b7d5c8791d52a6674a7e6e817a90d", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-018-what-we-learned-from-a-year-of-building-with-llm.raw.txt", "note_path": null, "primary_url": "https://applied-llms.org/", "section": "1 · Why we need evals", "source_id": "ae-018-what-we-learned-from-a-year-of-building-with-llm", "source_line": 61, "source_type": "web_article", "status": "mirrored", "title": "What We Learned from a Year of Building with LLMs", "why_it_matters": "The \"intern test,\" genchi genbutsu, turning vibe-checks into assertions."}
{"all_urls": ["https://www.interconnects.ai/p/evals-are-marketing"], "content_sha256": "81d5d69d5dae4f0d01c4847454f4d10c091cbef50697b1d20bb6eed5375968cf", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-019-big-tech-s-llm-evals-are-just-marketing.raw.txt", "note_path": null, "primary_url": "https://www.interconnects.ai/p/evals-are-marketing", "section": "1 · Why we need evals", "source_id": "ae-019-big-tech-s-llm-evals-are-just-marketing", "source_line": 62, "source_type": "blog", "status": "mirrored", "title": "Big Tech's LLM Evals Are Just Marketing", "why_it_matters": "Why frontier-lab leaderboard numbers are marketing, not science."}
{"all_urls": ["https://huyenchip.com/2025/01/16/ai-engineering-pitfalls.html"], "content_sha256": "4ca5bbd7e2f6dc1f367046187ab1bcc4a59a58c5f9aa11f0a2d237238c830ba6", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-020-ai-engineering-pitfalls.raw.txt", "note_path": null, "primary_url": "https://huyenchip.com/2025/01/16/ai-engineering-pitfalls.html", "section": "1 · Why we need evals", "source_id": "ae-020-ai-engineering-pitfalls", "source_line": 63, "source_type": "web_article", "status": "mirrored", "title": "AI Engineering pitfalls", "why_it_matters": "Common eval/AI-engineering mistakes from the *AI Engineering* author. (also T6)"}
{"all_urls": ["https://www.oreilly.com/radar/evals-are-not-all-you-need/"], "content_sha256": "a1b9a080abd9e102912af7d460cd9694989db50ef3b84749c28351fa5e6012f2", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-021-evals-are-not-all-you-need.raw.txt", "note_path": null, "primary_url": "https://www.oreilly.com/radar/evals-are-not-all-you-need/", "section": "1 · Why we need evals", "source_id": "ae-021-evals-are-not-all-you-need", "source_line": 65, "source_type": "web_article", "status": "mirrored", "title": "Evals Are NOT All You Need", "why_it_matters": "The essential nuance piece: automated graders alone don't save you; you need a continuous-improvement flywheel of offline tests + production monitoring + real-user iteration. Pairs with Shreya's 'In Defense' to complete the backlash debate. 🆕"}
{"all_urls": ["https://www.lennysnewsletter.com/p/why-ai-evals-are-the-hottest-new-skill"], "content_sha256": "a350cfc5d90110de1388a2c93065e33643c7076d3a37940cc3522d7850e7c220", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-022-why-ai-evals-are-the-hottest-new-skill-for-produ.raw.txt", "note_path": null, "primary_url": "https://www.lennysnewsletter.com/p/why-ai-evals-are-the-hottest-new-skill", "section": "1 · Why we need evals", "source_id": "ae-022-why-ai-evals-are-the-hottest-new-skill-for-produ", "source_line": 66, "source_type": "newsletter", "status": "mirrored", "title": "Why AI evals are the hottest new skill for product builders", "why_it_matters": "The accessible 'why evals matter' on-ramp (live walkthrough of error analysis, open/axial coding) that mainstreamed evals to PMs in 2025; the apartment-leasing-bot anecdote is the canonical 'you can't vibe-check' story. 🆕"}
{"all_urls": ["https://openai.com/index/evals-drive-next-chapter-of-ai/"], "content_sha256": "8c78f8a8986123a7b89e0ef97b482a0bb625c8f743695b2f466700d0c33b0c98", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-023-how-evals-drive-the-next-chapter-in-ai-for-busin.raw.txt", "note_path": null, "primary_url": "https://openai.com/index/evals-drive-next-chapter-of-ai/", "section": "1 · Why we need evals", "source_id": "ae-023-how-evals-drive-the-next-chapter-in-ai-for-busin", "source_line": 67, "source_type": "web_article", "status": "mirrored", "title": "How evals drive the next chapter in AI for businesses", "why_it_matters": "Frontier-lab framing of evals as turning fuzzy business goals into specs and measurable ROI; useful counterweight to Lambert's 'evals are marketing' and grounds the 'why' for enterprise readers. 🆕 ⚠(unverified URL)"}
{"all_urls": ["https://www.lennysnewsletter.com/p/beyond-vibe-checks-a-pms-complete"], "content_sha256": "adc8b3738c1323790bd7e7edb0aa0a4d28ede9e5027a9b2fa28552e22de66269", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/aman-khan-beyond-vibe-checks-pm-guide-evals.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-024-beyond-vibe-checks-a-pm-s-complete-guide-to-eval.raw.txt", "note_path": "notes/articles/aman-khan-beyond-vibe-checks-pm-guide-evals.md", "primary_url": "https://www.lennysnewsletter.com/p/beyond-vibe-checks-a-pms-complete", "section": "1 · Why we need evals", "source_id": "ae-024-beyond-vibe-checks-a-pm-s-complete-guide-to-eval", "source_line": 68, "source_type": "newsletter", "status": "mirrored", "title": "Beyond vibe checks: A PM's complete guide to evals", "why_it_matters": "The widely-shared PM-oriented argument for moving past 'looked good to me' vibe checks to systematic evals; one of the pieces that made evals a mainstream product skill in 2025. 🆕"}
{"all_urls": ["https://newsletter.pragmaticengineer.com/p/evals"], "content_sha256": "398ba7259b1827e0aad73d9f2441a07ea8adc978d343c541c5ea99b2cf277adf", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/pragmatic-engineer-llm-evals-guide.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-025-a-pragmatic-guide-to-llm-evals-for-devs.raw.txt", "note_path": "notes/articles/pragmatic-engineer-llm-evals-guide.md", "primary_url": "https://newsletter.pragmaticengineer.com/p/evals", "section": "1 · Why we need evals", "source_id": "ae-025-a-pragmatic-guide-to-llm-evals-for-devs", "source_line": 69, "source_type": "newsletter", "status": "mirrored", "title": "A pragmatic guide to LLM evals for devs", "why_it_matters": "Reaches the broad engineering audience with the core 'why': LLM non-determinism breaks traditional testing, so you need evals. High-distribution motivation piece co-written by Hamel. 🆕"}
{"all_urls": ["https://openai.com/index/deployment-simulation/"], "content_sha256": "1c4e765138fefa384a61328801cb79c2bd42642f8de19c03d7bc27dae925bf51", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-026-predicting-model-behavior-before-release-by-simu.raw.txt", "note_path": null, "primary_url": "https://openai.com/index/deployment-simulation/", "section": "1 · Why we need evals", "source_id": "ae-026-predicting-model-behavior-before-release-by-simu", "source_line": 70, "source_type": "web_article", "status": "mirrored", "title": "Predicting model behavior before release by simulating deployment (Deployment Simulation)", "why_it_matters": "Concrete 2026 evidence for why fixed/static evals fail: models recognize when they're being tested and game test suites; replaying ~1.3M real conversations surfaced reward-hacking no fixed eval caught. Strong 'why evals must evolve' argument. 🆕 ⚠(unverified URL)"}
{"all_urls": ["https://x.com/gdb/status/1733553161884127435"], "content_sha256": null, "http_status": null, "local_note_path": null, "local_raw_path": null, "note_path": null, "primary_url": "https://x.com/gdb/status/1733553161884127435", "section": "1 · Why we need evals", "source_id": "ae-027-evals-are-surprisingly-often-all-you-need", "source_line": 71, "source_type": "web_article", "status": "metadata_only", "title": "evals are surprisingly often all you need", "why_it_matters": "The canonical one-liner ('evals are the new unit test') that anchors the whole 'why evals' thesis; frequently cited founding quote for the movement. Short but load-bearing."}
{"all_urls": ["https://www.jasonwei.net/blog/asymmetry-of-verification-and-verifiers-law"], "content_sha256": "be6f54cb997e1d6095c362a8d722e30de68819787535cf8385441c8c8ea43492", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-028-asymmetry-of-verification-and-verifier-s-law.raw.txt", "note_path": null, "primary_url": "https://www.jasonwei.net/blog/asymmetry-of-verification-and-verifiers-law", "section": "2 · \"If you can eval it, you have built it\" — eval ⇄ capability ⇄ RL environment", "source_id": "ae-028-asymmetry-of-verification-and-verifier-s-law", "source_line": 77, "source_type": "blog", "status": "mirrored", "title": "Asymmetry of Verification and Verifier's Law", "why_it_matters": "Trainability tracks verifiability; verifying = creating an RL environment."}
{"all_urls": ["https://leehanchung.github.io/blogs/2026/03/21/rl-environments-for-llm-agents/"], "content_sha256": "066a8cb3086350a324ccd4d36c1cc212ec168dbfe1db7c186759de70098a1e10", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-029-a-taxonomy-of-rl-environments-for-llm-agents.raw.txt", "note_path": null, "primary_url": "https://leehanchung.github.io/blogs/2026/03/21/rl-environments-for-llm-agents/", "section": "2 · \"If you can eval it, you have built it\" — eval ⇄ capability ⇄ RL environment", "source_id": "ae-029-a-taxonomy-of-rl-environments-for-llm-agents", "source_line": 78, "source_type": "blog", "status": "mirrored", "title": "A Taxonomy of RL Environments for LLM Agents", "why_it_matters": "A benchmark is a frozen RL environment; the E = {T,H,V,S,C} decomposition; \"verifiable beats judgeable.\""}
{"all_urls": ["https://muratbuffalo.blogspot.com/2026/06/acm-cais-conference-on-ai-and-agentic.html"], "content_sha256": "2b1eab44957e019167a49304099ab4d5700712cd2a57d6f88df48a96c772b60e", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-030-the-life-cycle-of-an-rl-environment.raw.txt", "note_path": null, "primary_url": "https://muratbuffalo.blogspot.com/2026/06/acm-cais-conference-on-ai-and-agentic.html", "section": "2 · \"If you can eval it, you have built it\" — eval ⇄ capability ⇄ RL environment", "source_id": "ae-030-the-life-cycle-of-an-rl-environment", "source_line": 79, "source_type": "blog", "status": "mirrored", "title": "The Life Cycle of an RL Environment", "why_it_matters": "Difficulty calibration (the 1–4/16 Goldilocks band), RL as variance reduction, reward hacking under training pressure. *(local notes: `research/notes/kanav-garg-rl-environment-lifecycle.md`)*"}
{"all_urls": ["https://storage.googleapis.com/deepmind-media/Era-of-Experience%20/The%20Era%20of%20Experience%20Paper.pdf"], "content_sha256": "af7c19261a953b194acd69d6c594f1ee64f8f5175a0df000212d8fb3cc409a63", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-031-welcome-to-the-era-of-experience.raw.txt", "note_path": null, "primary_url": "https://storage.googleapis.com/deepmind-media/Era-of-Experience%20/The%20Era%20of%20Experience%20Paper.pdf", "section": "2 · \"If you can eval it, you have built it\" — eval ⇄ capability ⇄ RL environment", "source_id": "ae-031-welcome-to-the-era-of-experience", "source_line": 80, "source_type": "paper_or_pdf", "status": "mirrored", "title": "Welcome to the Era of Experience", "why_it_matters": "Human-data value approaching its ceiling; the frontier is agents learning from experience / synthetic environments."}
{"all_urls": ["https://rlhfbook.com/c/16-evaluation"], "content_sha256": "20f54a4a48bea184c60f73e0d56cf5482f7bb0c2033875b6dbb475531a3b2583", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-032-rlhf-book-ch-16-evaluation.raw.txt", "note_path": null, "primary_url": "https://rlhfbook.com/c/16-evaluation", "section": "2 · \"If you can eval it, you have built it\" — eval ⇄ capability ⇄ RL environment", "source_id": "ae-032-rlhf-book-ch-16-evaluation", "source_line": 81, "source_type": "docs_or_book", "status": "mirrored", "title": "RLHF Book, Ch. 16 — Evaluation", "why_it_matters": "Evaluation as a reflection of training goals; prompt-format sensitivity (60%→~0%)."}
{"all_urls": ["https://www.interconnects.ai/p/what-comes-next-with-reinforcement"], "content_sha256": "18660f7c802931d0f69c8989922deb1cbdaba4031643e6f6f8e4c6fa784179a1", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/papers/playing-atari-with-deep-reinforcement-learning.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-033-what-comes-next-with-reinforcement-learning.raw.txt", "note_path": "notes/papers/playing-atari-with-deep-reinforcement-learning.md", "primary_url": "https://www.interconnects.ai/p/what-comes-next-with-reinforcement", "section": "2 · \"If you can eval it, you have built it\" — eval ⇄ capability ⇄ RL environment", "source_id": "ae-033-what-comes-next-with-reinforcement-learning", "source_line": 82, "source_type": "blog", "status": "mirrored", "title": "What Comes Next with Reinforcement Learning", "why_it_matters": "Long-horizon credit assignment; where RL is and isn't ready."}
{"all_urls": ["https://github.com/PrimeIntellect-ai/verifiers"], "content_sha256": "120d283ec24119b3b3b128edbd58203efe95ee1601e126ff1a95c29e8c89d348", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-034-verifiers.raw.txt", "note_path": null, "primary_url": "https://github.com/PrimeIntellect-ai/verifiers", "section": "2 · \"If you can eval it, you have built it\" — eval ⇄ capability ⇄ RL environment", "source_id": "ae-034-verifiers", "source_line": 83, "source_type": "repository_or_docs", "status": "mirrored", "title": "verifiers", "why_it_matters": "the eval-is-an-RL-env thesis as code."}
{"all_urls": ["https://arxiv.org/abs/2501.12948"], "content_sha256": "0166b468d8d8650065cb9b8a72017b25eb51970fc948c2c59f5d4d66fe088171", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/papers/playing-atari-with-deep-reinforcement-learning.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-035-deepseek-r1-incentivizing-reasoning-capability-i.raw.txt", "note_path": "notes/papers/playing-atari-with-deep-reinforcement-learning.md", "primary_url": "https://arxiv.org/abs/2501.12948", "section": "2 · \"If you can eval it, you have built it\" — eval ⇄ capability ⇄ RL environment", "source_id": "ae-035-deepseek-r1-incentivizing-reasoning-capability-i", "source_line": 85, "source_type": "paper_or_pdf", "status": "mirrored", "title": "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning", "why_it_matters": "the canonical 'if you can verify it, RL builds it' result; also published in Nature 2025. Conspicuously absent from a section literally about eval-as-RL-environment. 🆕"}
{"all_urls": ["https://arxiv.org/abs/2411.15124"], "content_sha256": "608d288e2b22dd53d93ddf43e39220214987ea988d082bd0f58fd30b23f0deb3", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/jason-wei-successful-language-model-evals.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-036-t-lu-3-pushing-frontiers-in-open-language-model.raw.txt", "note_path": "notes/articles/jason-wei-successful-language-model-evals.md", "primary_url": "https://arxiv.org/abs/2411.15124", "section": "2 · \"If you can eval it, you have built it\" — eval ⇄ capability ⇄ RL environment", "source_id": "ae-036-t-lu-3-pushing-frontiers-in-open-language-model", "source_line": 86, "source_type": "paper_or_pdf", "status": "mirrored", "title": "Tülu 3: Pushing Frontiers in Open Language Model Post-Training", "why_it_matters": "Coined/popularized RLVR and open-sourced the recipe + code (open-instruct): swap the reward model for a verifier on tasks with checkable answers. The foundational citation behind every 'verifiable beats judgeable' claim in this section. 🆕"}
{"all_urls": ["https://www.anthropic.com/research/emergent-misalignment-reward-hacking"], "content_sha256": "2b43ff6518f92c8404b48abdcb2ba4d38d696b0cc3b9b514407f5b4ff763a4ec", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/natural-emergent-misalignment-reward-hacking-production-rl.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-037-natural-emergent-misalignment-from-reward-hackin.raw.txt", "note_path": "notes/articles/natural-emergent-misalignment-reward-hacking-production-rl.md", "primary_url": "https://www.anthropic.com/research/emergent-misalignment-reward-hacking", "section": "2 · \"If you can eval it, you have built it\" — eval ⇄ capability ⇄ RL environment", "source_id": "ae-037-natural-emergent-misalignment-from-reward-hackin", "source_line": 87, "source_type": "web_article", "status": "mirrored", "title": "Natural Emergent Misalignment from Reward Hacking in Production RL", "why_it_matters": "Empirical receipt for the section's 'reward hacking under training pressure' theme: learning to cheat on real coding environments generalizes to sabotage/alignment-faking; introduces inoculation prompting as mitigation (arXiv 2511.18397). 🆕"}
{"all_urls": ["https://www.primeintellect.ai/blog/environments"], "content_sha256": "3a7642ae07ed957f61c6ecd8034515ca89ff74ca3dd93359a37f9c685002d9b9", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-brown-rl-environments-at-scale.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-038-environments-hub-a-community-hub-to-scale-rl-to.raw.txt", "note_path": "notes/talks/talk-brown-rl-environments-at-scale.md", "primary_url": "https://www.primeintellect.ai/blog/environments", "section": "2 · \"If you can eval it, you have built it\" — eval ⇄ capability ⇄ RL environment", "source_id": "ae-038-environments-hub-a-community-hub-to-scale-rl-to", "source_line": 88, "source_type": "blog", "status": "mirrored", "title": "Environments Hub: A Community Hub To Scale RL To Open AGI", "why_it_matters": "the eval-is-an-RL-env thesis as an actual ecosystem, the natural companion to the already-listed verifiers repo. 🆕"}
{"all_urls": ["https://www.mechanize.work/blog/how-to-fully-automate-software-engineering/"], "content_sha256": "b67d2e4aae274b04bba1efa3a5ab3065f14598d748ca7ce096215217c15d9242", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-039-how-to-fully-automate-software-engineering.raw.txt", "note_path": null, "primary_url": "https://www.mechanize.work/blog/how-to-fully-automate-software-engineering/", "section": "2 · \"If you can eval it, you have built it\" — eval ⇄ capability ⇄ RL environment", "source_id": "ae-039-how-to-fully-automate-software-engineering", "source_line": 89, "source_type": "blog", "status": "mirrored", "title": "How to fully automate software engineering", "why_it_matters": "'you only get the capability you can build an environment for.' 🆕"}
{"all_urls": ["https://www.mechanize.work/blog/cheap-rl-tasks-will-waste-compute/"], "content_sha256": "271dccf94d271a6a222bac5c8936cf42ae7763f0448cdf3d3f47d7fa482b58c4", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-040-cheap-rl-tasks-will-waste-compute.raw.txt", "note_path": null, "primary_url": "https://www.mechanize.work/blog/cheap-rl-tasks-will-waste-compute/", "section": "2 · \"If you can eval it, you have built it\" — eval ⇄ capability ⇄ RL environment", "source_id": "ae-040-cheap-rl-tasks-will-waste-compute", "source_line": 90, "source_type": "blog", "status": "mirrored", "title": "Cheap RL tasks will waste compute", "why_it_matters": "directly informs difficulty calibration / why environment design matters. 🆕"}
{"all_urls": ["https://epoch.ai/gradient-updates/state-of-rl-envs"], "content_sha256": "090462c797763c8be18e23b01c6285997b2d8ef9c70362909707530eb4065793", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/papers/playing-atari-with-deep-reinforcement-learning.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-041-an-faq-on-reinforcement-learning-environments.raw.txt", "note_path": "notes/papers/playing-atari-with-deep-reinforcement-learning.md", "primary_url": "https://epoch.ai/gradient-updates/state-of-rl-envs", "section": "2 · \"If you can eval it, you have built it\" — eval ⇄ capability ⇄ RL environment", "source_id": "ae-041-an-faq-on-reinforcement-learning-environments", "source_line": 91, "source_type": "web_article", "status": "mirrored", "title": "An FAQ on Reinforcement Learning Environments", "why_it_matters": "the empirical state-of-the-field map this section lacks. 🆕"}
{"all_urls": ["https://newsletter.semianalysis.com/p/rl-environments-and-rl-for-science"], "content_sha256": "b50e2819061359867ed522647aeb77c5d8c804e8abf8fccdd6690f3a683b8c57", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/mast-why-multi-agent-llm-systems-fail.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-042-rl-environments-and-rl-for-science-data-foundrie.raw.txt", "note_path": "notes/articles/mast-why-multi-agent-llm-systems-fail.md", "primary_url": "https://newsletter.semianalysis.com/p/rl-environments-and-rl-for-science", "section": "2 · \"If you can eval it, you have built it\" — eval ⇄ capability ⇄ RL environment", "source_id": "ae-042-rl-environments-and-rl-for-science-data-foundrie", "source_line": 92, "source_type": "newsletter", "status": "mirrored", "title": "RL Environments and RL for Science: Data Foundries and Multi-Agent Architectures", "why_it_matters": "Market-structure view: 35+ companies now sell RL environments; capability gains are coming from ramping RL compute, not pretraining. Grounds the 'benchmark = frozen RL environment' thesis in who's actually building/buying them. 🆕"}
{"all_urls": ["https://github.com/harbor-framework/terminal-bench"], "content_sha256": "86101ef2cf5623a17effe698d87d61ddc32ea4aabe3914d3aa03e141dfda3a2c", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-043-terminal-bench-benchmarking-agents-on-hard-reali.raw.txt", "note_path": null, "primary_url": "https://github.com/harbor-framework/terminal-bench", "section": "2 · \"If you can eval it, you have built it\" — eval ⇄ capability ⇄ RL environment", "source_id": "ae-043-terminal-bench-benchmarking-agents-on-hard-reali", "source_line": 93, "source_type": "repository_or_docs", "status": "mirrored", "title": "Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces", "why_it_matters": "i.e. a benchmark that IS an RL environment (and is used as one). 2.4k stars, active. 🆕"}
{"all_urls": ["https://github.com/sierra-research/tau2-bench"], "content_sha256": "b5654a3b95286c0b722b6940e52dcda06762ece148d8df56e85f9a58e6fc5088", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-044-tau2-bench-bench-a-benchmark-for-tool-agent-user.raw.txt", "note_path": null, "primary_url": "https://github.com/sierra-research/tau2-bench", "section": "2 · \"If you can eval it, you have built it\" — eval ⇄ capability ⇄ RL environment", "source_id": "ae-044-tau2-bench-bench-a-benchmark-for-tool-agent-user", "source_line": 94, "source_type": "repository_or_docs", "status": "mirrored", "title": "tau2-bench (τ²-Bench): A Benchmark for Tool-Agent-User Interaction in Real-World Domains", "why_it_matters": "the canonical example of a verifiable conversational/agentic environment beyond math/code (paper arXiv 2506.07982). 🆕"}
{"all_urls": ["https://leehanchung.github.io/blogs/2026/05/08/hidden-technical-debt-agent-harness/"], "content_sha256": "82bab3c182f4749115098e4e8adb24c9e873e83b32b22525592ea2a6eecaa6c3", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/leehanchung-hidden-technical-debt-agent-runtime.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-045-hidden-technical-debt-agent-harness.raw.txt", "note_path": "notes/articles/leehanchung-hidden-technical-debt-agent-runtime.md", "primary_url": "https://leehanchung.github.io/blogs/2026/05/08/hidden-technical-debt-agent-harness/", "section": "3 · The model / harness / skill decomposition", "source_id": "ae-045-hidden-technical-debt-agent-harness", "source_line": 100, "source_type": "blog", "status": "mirrored", "title": "Hidden Technical Debt: Agent Harness", "why_it_matters": "The harness is the agent; what teams call \"the model\" is mostly harness + product."}
{"all_urls": ["https://leehanchung.github.io/blogs/"], "content_sha256": "feb7c83720c58480af79bad2784ed943ef402e15052184af4ea4681d7e5c9a5c", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/leehanchung-hidden-technical-debt-agent-runtime.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-046-hidden-technical-debt-series-index.raw.txt", "note_path": "notes/articles/leehanchung-hidden-technical-debt-agent-runtime.md", "primary_url": "https://leehanchung.github.io/blogs/", "section": "3 · The model / harness / skill decomposition", "source_id": "ae-046-hidden-technical-debt-series-index", "source_line": 101, "source_type": "blog", "status": "mirrored", "title": "Hidden Technical Debt series (index)", "why_it_matters": "The four-part series (eval infra, runtime, harness, + agent runtime ~2026/04/24). *(verify the runtime post URL on the index.)*"}
{"all_urls": ["https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/"], "content_sha256": "b3e67dedc0e8c78eab25f1f94034a6959f378e16be2126fdbc619bee781dbb8e", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-047-measuring-ai-ability-to-complete-long-tasks.raw.txt", "note_path": null, "primary_url": "https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/", "section": "3 · The model / harness / skill decomposition", "source_id": "ae-047-measuring-ai-ability-to-complete-long-tasks", "source_line": 102, "source_type": "blog", "status": "mirrored", "title": "Measuring AI Ability to Complete Long Tasks", "why_it_matters": "Scaffolds change the measured horizon; success-vs-human-time as a primitive. (also T9)"}
{"all_urls": ["https://www.turingpost.com/p/nathanlambert"], "content_sha256": "ae261462d45aa59da5d49c5aff7776bf66b7c58ad50e2e313340cd866b6b7729", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-048-turing-post-interview-open-models-won-t-catch-up.raw.txt", "note_path": null, "primary_url": "https://www.turingpost.com/p/nathanlambert", "section": "3 · The model / harness / skill decomposition", "source_id": "ae-048-turing-post-interview-open-models-won-t-catch-up", "source_line": 103, "source_type": "blog", "status": "mirrored", "title": "Turing Post interview (\"Open Models Won't Catch Up\")", "why_it_matters": "\"What technical people call the harness or the product matters more than just the model.\""}
{"all_urls": ["https://florianbrand.com/posts/benches-2026", "https://www.youtube.com/watch?v=kmTMc-fVSXw"], "content_sha256": "fc8d7716215d0260805cab35106c2f3c3fd1384fe70eaa946cbe3d80a5d2eebe", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-049-quo-vadis-llm-benchmarks.raw.txt", "note_path": null, "primary_url": "https://florianbrand.com/posts/benches-2026", "section": "3 · The model / harness / skill decomposition", "source_id": "ae-049-quo-vadis-llm-benchmarks", "source_line": 104, "source_type": "web_article", "status": "mirrored", "title": "Quo vadis, LLM benchmarks?", "why_it_matters": "The AlgoTune case: *same model, different harness, opposite ranking.* (also T6) *(notes: `research/notes/florian-brand-*`)*"}
{"all_urls": ["https://leehanchung.github.io/talks/2025/04/23/the-model-is-the-product/"], "content_sha256": "da98c7012d76059c648a438af889a3e23f92e9a31b53b4f3b050b0c1bd37ed1d", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-lee-model-is-the-product.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-050-the-model-is-the-product.raw.txt", "note_path": "notes/talks/talk-lee-model-is-the-product.md", "primary_url": "https://leehanchung.github.io/talks/2025/04/23/the-model-is-the-product/", "section": "3 · The model / harness / skill decomposition", "source_id": "ae-050-the-model-is-the-product", "source_line": 106, "source_type": "web_article", "status": "mirrored", "title": "The Model is the Product", "why_it_matters": "the direct counterpart to Hamel's 'Model is Not the Product'; the foundational text of the harness/model debate this section is built on. 🆕"}
{"all_urls": ["https://www.youtube.com/watch?v=EEw2PpL-_NM"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-lee-model-is-the-product.md", "local_raw_path": null, "note_path": "notes/talks/talk-lee-model-is-the-product.md", "primary_url": "https://www.youtube.com/watch?v=EEw2PpL-_NM", "section": "3 · The model / harness / skill decomposition", "source_id": "ae-051-the-model-is-not-the-product", "source_line": 107, "source_type": "talk_video", "status": "transcript_queued", "title": "The Model is Not the Product", "why_it_matters": "The opposing side of the Lee debate (Data Council 2025): great products are mostly harness + product + evals, not the model. Section already cites Lee; it should cite the debate it half-references. 🆕"}
{"all_urls": ["https://simonwillison.net/2025/May/22/tools-in-a-loop/"], "content_sha256": "df8b6cc6587ea72a46702facd7b3a844910bff4441b4165df9012df0d2f49bd5", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/anthropic-writing-tools-for-agents.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-052-agents-are-models-using-tools-in-a-loop.raw.txt", "note_path": "notes/articles/anthropic-writing-tools-for-agents.md", "primary_url": "https://simonwillison.net/2025/May/22/tools-in-a-loop/", "section": "3 · The model / harness / skill decomposition", "source_id": "ae-052-agents-are-models-using-tools-in-a-loop", "source_line": 108, "source_type": "web_article", "status": "mirrored", "title": "Agents are models using tools in a loop", "why_it_matters": "the cleanest statement of why the harness, not the model, dominates behavior. 🆕"}
{"all_urls": ["https://openai.com/index/harness-engineering/"], "content_sha256": "c8c5abb4fa8b87d6ecb2705f3e2f2a587c8bc29d86d30e6ce8c4a3505ccf8627", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-053-harness-engineering-leveraging-codex-in-an-agent.raw.txt", "note_path": null, "primary_url": "https://openai.com/index/harness-engineering/", "section": "3 · The model / harness / skill decomposition", "source_id": "ae-053-harness-engineering-leveraging-codex-in-an-agent", "source_line": 109, "source_type": "web_article", "status": "mirrored", "title": "Harness engineering: leveraging Codex in an agent-first world", "why_it_matters": "Frontier-lab primary source coining 'harness engineering': a 1M-line codebase built by Codex agents where improving the environment/harness mattered more than the model. Lab-side complement to Lee's 'harness is the agent'. (URL returns 403 to scraper but page is live; corroborated by InfoQ/Milvus coverage.) 🆕"}
{"all_urls": ["https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills"], "content_sha256": "628fd795d4babeb906966f46c121212ed39349349d8160fd665d24ffe8fc7df7", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/vercel-agents-md-outperforms-skills.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-054-equipping-agents-for-the-real-world-with-agent-s.raw.txt", "note_path": "notes/articles/vercel-agents-md-outperforms-skills.md", "primary_url": "https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills", "section": "3 · The model / harness / skill decomposition", "source_id": "ae-054-equipping-agents-for-the-real-world-with-agent-s", "source_line": 110, "source_type": "web_article", "status": "mirrored", "title": "Equipping agents for the real world with Agent Skills", "why_it_matters": "skills as composable, progressively-disclosed capabilities (later made an open standard). The section title says 'skill' but has zero skill sources. 🆕"}
{"all_urls": ["https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents"], "content_sha256": "b4ae6b5269833d884203ffee69144bff4eb5bfa0bafca3a5de59132dd5c6bbbd", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-055-effective-context-engineering-for-ai-agents.raw.txt", "note_path": null, "primary_url": "https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents", "section": "3 · The model / harness / skill decomposition", "source_id": "ae-055-effective-context-engineering-for-ai-agents", "source_line": 111, "source_type": "web_article", "status": "mirrored", "title": "Effective context engineering for AI agents", "why_it_matters": "the mechanism behind why same model + different harness diverges. 🆕"}
{"all_urls": ["https://www.anthropic.com/engineering/writing-tools-for-agents"], "content_sha256": "3c2de66140551d40a57cb1ab38ea1812bc19260473c8484bc0dacab3ab75711a", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/anthropic-writing-tools-for-agents.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-056-writing-effective-tools-for-agents-with-agents.raw.txt", "note_path": "notes/articles/anthropic-writing-tools-for-agents.md", "primary_url": "https://www.anthropic.com/engineering/writing-tools-for-agents", "section": "3 · The model / harness / skill decomposition", "source_id": "ae-056-writing-effective-tools-for-agents-with-agents", "source_line": 112, "source_type": "web_article", "status": "mirrored", "title": "Writing effective tools for agents — with agents", "why_it_matters": "Tool design is a load-bearing part of the harness; 'agents are only as effective as the tools we give them,' validated eval-first. Directly ties harness decisions to measured agent performance. 🆕"}
{"all_urls": ["https://blog.thepete.net/blog/2025/12/10/same-model-different-results-why-coding-agents-arent-interchangeable/"], "content_sha256": "5ff566fd14e1ce4f1b016c3b83765a6de3e14df125b140b3b1622ff95e4c1963", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-057-same-model-different-results-why-coding-agents-a.raw.txt", "note_path": null, "primary_url": "https://blog.thepete.net/blog/2025/12/10/same-model-different-results-why-coding-agents-arent-interchangeable/", "section": "3 · The model / harness / skill decomposition", "source_id": "ae-057-same-model-different-results-why-coding-agents-a", "source_line": 113, "source_type": "blog", "status": "mirrored", "title": "Same Model, Different Results: Why Coding Agents Aren't Interchangeable", "why_it_matters": "the practitioner case-study version of Brand's AlgoTune point. 🆕"}
{"all_urls": ["https://hal.cs.princeton.edu/"], "content_sha256": "3846d6ab680ab4b8664169b0cbd83421344a6b6286205f86380b0a1f4b349d2a", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-058-holistic-agent-leaderboard-hal.raw.txt", "note_path": null, "primary_url": "https://hal.cs.princeton.edu/", "section": "3 · The model / harness / skill decomposition", "source_id": "ae-058-holistic-agent-leaderboard-hal", "source_line": 114, "source_type": "web_article", "status": "mirrored", "title": "Holistic Agent Leaderboard (HAL)", "why_it_matters": "the infrastructure answer to 'harness confounds rankings.' ICLR 2026; paper arXiv:2510.11977. 🆕"}
{"all_urls": ["https://www.oreilly.com/radar/agent-harness-engineering/"], "content_sha256": "7b74b1d3828bb96450fc2b3fd863e5b341a30d8f6d5c7d71719a2edae538f296", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-059-agent-harness-engineering.raw.txt", "note_path": null, "primary_url": "https://www.oreilly.com/radar/agent-harness-engineering/", "section": "3 · The model / harness / skill decomposition", "source_id": "ae-059-agent-harness-engineering", "source_line": 115, "source_type": "web_article", "status": "mirrored", "title": "Agent Harness Engineering", "why_it_matters": "'A decent model with a great harness beats a great model with a bad harness'; reframes agent failures as harness/config problems (traceable AGENTS.md rules). Names the converging harness primitives across coding agents. 🆕"}
{"all_urls": ["https://www.interconnects.ai/p/the-next-phase-of-open-models"], "content_sha256": "da81065268df58f08729537ea671a8ffbd15b8269dfc4d392edf06987d8f2419", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/papers/red-teaming-language-models-with-language-models.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-060-what-comes-next-with-open-models-weights-tools-h.raw.txt", "note_path": "notes/papers/red-teaming-language-models-with-language-models.md", "primary_url": "https://www.interconnects.ai/p/the-next-phase-of-open-models", "section": "3 · The model / harness / skill decomposition", "source_id": "ae-060-what-comes-next-with-open-models-weights-tools-h", "source_line": 116, "source_type": "blog", "status": "mirrored", "title": "What comes next with open models (weights / tools / harness decomposition)", "why_it_matters": "the written companion to the Turing Post interview already listed, with the explicit three-part decomposition. 🆕"}
{"all_urls": ["https://leehanchung.github.io/blogs/2026/06/13/hidden-technical-debt-agent-evaluation-infra/"], "content_sha256": "0c84dd759fb4cd17d700488033f3ff82b01da2f20fb197182868451e77e008d0", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/leehanchung-hidden-technical-debt-agent-runtime.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-061-hidden-technical-debt-agent-evaluation-infrastru.raw.txt", "note_path": "notes/articles/leehanchung-hidden-technical-debt-agent-runtime.md", "primary_url": "https://leehanchung.github.io/blogs/2026/06/13/hidden-technical-debt-agent-evaluation-infra/", "section": "4 · Observability & the output / eval space (the surfaces you can grade)", "source_id": "ae-061-hidden-technical-debt-agent-evaluation-infrastru", "source_line": 122, "source_type": "blog", "status": "mirrored", "title": "Hidden Technical Debt: Agent Evaluation Infrastructure", "why_it_matters": "Control plane / data plane; the **five surfaces** (output, trace, memory, environment, mechanistic); the empty-tool-result hallucination."}
{"all_urls": ["https://www.braintrust.dev/blog/three-pillars-ai-observability"], "content_sha256": "0130840f40c5485535e3e1c49f5bf87294d0eeb6e7318699195e6323d0794584", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-062-the-three-pillars-of-ai-observability.raw.txt", "note_path": null, "primary_url": "https://www.braintrust.dev/blog/three-pillars-ai-observability", "section": "4 · Observability & the output / eval space (the surfaces you can grade)", "source_id": "ae-062-the-three-pillars-of-ai-observability", "source_line": 123, "source_type": "blog", "status": "mirrored", "title": "The Three Pillars of AI Observability", "why_it_matters": "Dataset reconciliation (living datasets); traces / evals / annotation."}
{"all_urls": ["https://arize.com/docs/ax/evaluate/evaluators/trace-and-session-evals/trace-level-evaluations/agent-trajectory-evaluations"], "content_sha256": "27e8cd4f1e97c60212035e9dce04e770ba1c891b5bbbe30d249c0d421ededd41", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/agentrewardbench-evaluating-automatic-evaluations-web-agent-.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-063-agent-trajectory-evaluations.raw.txt", "note_path": "notes/articles/agentrewardbench-evaluating-automatic-evaluations-web-agent-.md", "primary_url": "https://arize.com/docs/ax/evaluate/evaluators/trace-and-session-evals/trace-level-evaluations/agent-trajectory-evaluations", "section": "4 · Observability & the output / eval space (the surfaces you can grade)", "source_id": "ae-063-agent-trajectory-evaluations", "source_line": 124, "source_type": "docs_or_book", "status": "mirrored", "title": "Agent Trajectory Evaluations", "why_it_matters": "Grading the path, not just the answer."}
{"all_urls": ["https://galileo.ai/blog/ai-agent-metrics"], "content_sha256": "4001f491e3157cbf2839c72c8036a9849738822ecb4bd4d1320272b48187fb7f", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/opik-evaluate-agent-trajectory.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-064-ai-agent-metrics-how-elite-teams-evaluate.raw.txt", "note_path": "notes/articles/opik-evaluate-agent-trajectory.md", "primary_url": "https://galileo.ai/blog/ai-agent-metrics", "section": "4 · Observability & the output / eval space (the surfaces you can grade)", "source_id": "ae-064-ai-agent-metrics-how-elite-teams-evaluate", "source_line": 125, "source_type": "blog", "status": "mirrored", "title": "AI Agent Metrics: How Elite Teams Evaluate", "why_it_matters": "A concrete agent-metric taxonomy (action completion, tool selection, etc.)."}
{"all_urls": ["https://github.com/Arize-ai/openinference/blob/main/spec/semantic_conventions.md"], "content_sha256": "d703927c37762d39268e593f37c9fcd228ca171e4dd62054e6b7909357da39bf", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-065-openinference-semantic-conventions.md", "note_path": null, "primary_url": "https://github.com/Arize-ai/openinference/blob/main/spec/semantic_conventions.md", "section": "4 · Observability & the output / eval space (the surfaces you can grade)", "source_id": "ae-065-openinference-semantic-conventions", "source_line": 126, "source_type": "repository_or_docs", "status": "mirrored", "title": "OpenInference semantic conventions", "why_it_matters": "An OTel-based agent trace schema (tool, args, observation, latency, cost)."}
{"all_urls": ["https://docs.langchain.com/langsmith/evaluation", "https://docs.langchain.com/langsmith/trajectory-evals"], "content_sha256": "b5004e48455bb6fcb22c5fc737830a5d63289cfabe84280d0eb21168eccdd8d4", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-066-langsmith-evaluation-trajectory-evals.raw.txt", "note_path": null, "primary_url": "https://docs.langchain.com/langsmith/evaluation", "section": "4 · Observability & the output / eval space (the surfaces you can grade)", "source_id": "ae-066-langsmith-evaluation-trajectory-evals", "source_line": 127, "source_type": "docs_or_book", "status": "mirrored", "title": "LangSmith Evaluation / Trajectory evals", "why_it_matters": "<https://docs.langchain.com/langsmith/evaluation> · <https://docs.langchain.com/langsmith/trajectory-evals> · *docs*."}
{"all_urls": ["https://github.com/open-telemetry/semantic-conventions-genai"], "content_sha256": "4a820be3a0a40ea3271e791320c126516f629af0beaef80cd0bde5162ff66429", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-067-opentelemetry-genai-semantic-conventions-agent-f.raw.txt", "note_path": null, "primary_url": "https://github.com/open-telemetry/semantic-conventions-genai", "section": "4 · Observability & the output / eval space (the surfaces you can grade)", "source_id": "ae-067-opentelemetry-genai-semantic-conventions-agent-f", "source_line": 129, "source_type": "repository_or_docs", "status": "mirrored", "title": "OpenTelemetry GenAI Semantic Conventions (agent & framework spans)", "why_it_matters": "the canonical trace schema the section's OpenInference entry derives from. 🆕"}
{"all_urls": ["https://opentelemetry.io/docs/specs/semconv/gen-ai/gen-ai-agent-spans/"], "content_sha256": "cd400ffb83545fe0f2d39b7c6ec5adfed163d9423b408303214d7f3be07072f1", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-068-semantic-conventions-for-genai-agent-and-framewo.raw.txt", "note_path": null, "primary_url": "https://opentelemetry.io/docs/specs/semconv/gen-ai/gen-ai-agent-spans/", "section": "4 · Observability & the output / eval space (the surfaces you can grade)", "source_id": "ae-068-semantic-conventions-for-genai-agent-and-framewo", "source_line": 130, "source_type": "docs_or_book", "status": "mirrored", "title": "Semantic Conventions for GenAI agent and framework spans", "why_it_matters": "the precise definition of what a gradable agent trace looks like. 🆕"}
{"all_urls": ["https://opentelemetry.io/blog/2026/genai-observability/"], "content_sha256": "5933bf129610655098ffe87f4496a74182c5ea3393aa4e672ee77fe708a34b23", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-069-inside-the-llm-call-genai-observability-with-ope.raw.txt", "note_path": null, "primary_url": "https://opentelemetry.io/blog/2026/genai-observability/", "section": "4 · Observability & the output / eval space (the surfaces you can grade)", "source_id": "ae-069-inside-the-llm-call-genai-observability-with-ope", "source_line": 131, "source_type": "blog", "status": "mirrored", "title": "Inside the LLM Call: GenAI Observability with OpenTelemetry", "why_it_matters": "concrete intro to the trace surface for practitioners not steeped in OTel. 🆕"}
{"all_urls": ["https://docs.wandb.ai/weave"], "content_sha256": "b737a9cb3ccf6999ad8980c81aa303e6cf7b2675f8dcaba635a71fa077107cd4", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-070-w-b-weave-tracing-evaluation-toolkit.raw.txt", "note_path": null, "primary_url": "https://docs.wandb.ai/weave", "section": "4 · Observability & the output / eval space (the surfaces you can grade)", "source_id": "ae-070-w-b-weave-tracing-evaluation-toolkit", "source_line": 132, "source_type": "docs_or_book", "status": "mirrored", "title": "W&B Weave — tracing & evaluation toolkit", "why_it_matters": "a widely used surface for grading both traces and outputs. 🆕"}
{"all_urls": ["https://laminar.sh/"], "content_sha256": "7a07d97e415b9f5495599f73eea40ac1d14b620faf5be3697fb4e4ba21c50c89", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/hf-agents-course-bonus-unit2-observability-evaluation.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-071-laminar-open-source-observability-for-ai-agents.raw.txt", "note_path": "notes/articles/hf-agents-course-bonus-unit2-observability-evaluation.md", "primary_url": "https://laminar.sh/", "section": "4 · Observability & the output / eval space (the surfaces you can grade)", "source_id": "ae-071-laminar-open-source-observability-for-ai-agents", "source_line": 133, "source_type": "web_article", "status": "mirrored", "title": "Laminar — open-source observability for AI agents", "why_it_matters": "purpose-built for grading multi-step agent trajectories rather than single LLM calls. 🆕"}
{"all_urls": ["https://github.com/UKGovernmentBEIS/inspect_ai", "https://inspect.aisi.org.uk/"], "content_sha256": "36f94cd023966fd3eadd5e36e57056c9240bf7beb7ec5a801ff683880eb56496", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-072-inspect-ai.raw.txt", "note_path": null, "primary_url": "https://github.com/UKGovernmentBEIS/inspect_ai", "section": "5a · Eval frameworks & harnesses (code-first test-runners)", "source_id": "ae-072-inspect-ai", "source_line": 142, "source_type": "repository_or_docs", "status": "mirrored", "title": "Inspect AI", "why_it_matters": "`@task` binds dataset + solver + scorer; custom scorers; sandboxed tools. The reference agent-eval framework. **(MUST)**"}
{"all_urls": ["https://github.com/UKGovernmentBEIS/inspect_evals"], "content_sha256": "c26349f4f1da9c1324690868bc8590b0f7e803566364dcee0d6c17b03f3bedcc", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-073-inspect-evals.raw.txt", "note_path": null, "primary_url": "https://github.com/UKGovernmentBEIS/inspect_evals", "section": "5a · Eval frameworks & harnesses (code-first test-runners)", "source_id": "ae-073-inspect-evals", "source_line": 143, "source_type": "repository_or_docs", "status": "mirrored", "title": "inspect_evals", "why_it_matters": "the \"batteries\" for Inspect."}
{"all_urls": ["https://github.com/EleutherAI/lm-evaluation-harness"], "content_sha256": "12e85f3de7be0931231408220b7930a19f1d653d2368636dba2dabacf5d4a665", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-074-lm-evaluation-harness.raw.txt", "note_path": null, "primary_url": "https://github.com/EleutherAI/lm-evaluation-harness", "section": "5a · Eval frameworks & harnesses (code-first test-runners)", "source_id": "ae-074-lm-evaluation-harness", "source_line": 144, "source_type": "repository_or_docs", "status": "mirrored", "title": "lm-evaluation-harness", "why_it_matters": "the standard academic harness; first-class decontamination; task YAMLs."}
{"all_urls": ["https://github.com/allenai/olmes"], "content_sha256": "206ceebd3a44c3020b218ac6705e054b68d5e7d3cec93a6a2972c6753f36ae8b", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-075-olmes.raw.txt", "note_path": null, "primary_url": "https://github.com/allenai/olmes", "section": "5a · Eval frameworks & harnesses (code-first test-runners)", "source_id": "ae-075-olmes", "source_line": 145, "source_type": "repository_or_docs", "status": "mirrored", "title": "OLMES", "why_it_matters": "🆕 the reproducible eval **standard + harness** behind OLMo/Tülu: standardized prompts/metrics/formatting for apples-to-apples model comparison."}
{"all_urls": ["https://github.com/benchflow-ai/benchflow", "https://benchflow.ai"], "content_sha256": "4e95d90950ce815ebe0b35665f308e14acbda95d9d793b2699c99d0a08bcb1ab", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-076-benchflow.raw.txt", "note_path": null, "primary_url": "https://github.com/benchflow-ai/benchflow", "section": "5a · Eval frameworks & harnesses (code-first test-runners)", "source_id": "ae-076-benchflow", "source_line": 146, "source_type": "repository_or_docs", "status": "mirrored", "title": "BenchFlow", "why_it_matters": "🆕 environment-lab framework: research infra + runtime for building RL environments, evals & post-training; ships **SkillsBench** and **ClawsBench**. (\"Environments are the new data.\")"}
{"all_urls": ["https://github.com/huggingface/lighteval"], "content_sha256": "7e8476b66f44c8fdc2b7da2b2cd1f124cf61a0dc8683f0a78c121050a3e333b8", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-077-lighteval.raw.txt", "note_path": null, "primary_url": "https://github.com/huggingface/lighteval", "section": "5a · Eval frameworks & harnesses (code-first test-runners)", "source_id": "ae-077-lighteval", "source_line": 147, "source_type": "repository_or_docs", "status": "mirrored", "title": "lighteval", "why_it_matters": "🆕 all-in-one harness across transformers/vLLM/TGI/nanotron, 1000+ tasks; HF's successor to `evaluate`."}
{"all_urls": ["https://github.com/groq/openbench"], "content_sha256": "4e0f378a1752c6c881786b488f1fc2b383866b23e1a71da5861ac3c0abc732ce", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-078-openbench.raw.txt", "note_path": null, "primary_url": "https://github.com/groq/openbench", "section": "5a · Eval frameworks & harnesses (code-first test-runners)", "source_id": "ae-078-openbench", "source_line": 148, "source_type": "repository_or_docs", "status": "mirrored", "title": "OpenBench", "why_it_matters": "🆕 provider-agnostic `bench` CLI, 95+ benchmarks, built on Inspect primitives."}
{"all_urls": ["https://github.com/openai/simple-evals"], "content_sha256": "b55c06cb3e6a509a20a5bf060ec8a570b04c01259477c2fa01ef8f87968dbc6f", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-079-simple-evals.raw.txt", "note_path": null, "primary_url": "https://github.com/openai/simple-evals", "section": "5a · Eval frameworks & harnesses (code-first test-runners)", "source_id": "ae-079-simple-evals", "source_line": 149, "source_type": "repository_or_docs", "status": "mirrored", "title": "simple-evals", "why_it_matters": "minimal zero-shot/CoT scripts (MMLU, HumanEval, SimpleQA, HealthBench); the numbers OpenAI publishes. ⚠️ not actively maintained."}
{"all_urls": ["https://github.com/openai/evals", "https://developers.openai.com/api/docs/guides/evaluation-best-practices"], "content_sha256": "5931ce26961f019f4e85379d120c8e43f06379dd5ad51713114c5d921bce73ed", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/macro-evals-agentic-systems-openai-cookbook.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-080-openai-evals.raw.txt", "note_path": "notes/articles/macro-evals-agentic-systems-openai-cookbook.md", "primary_url": "https://github.com/openai/evals", "section": "5a · Eval frameworks & harnesses (code-first test-runners)", "source_id": "ae-080-openai-evals", "source_line": 150, "source_type": "repository_or_docs", "status": "mirrored", "title": "OpenAI Evals", "why_it_matters": "the `completion_fn` abstraction = swap the system-under-test. (Best-practices: <https://developers.openai.com/api/docs/guides/evaluation-best-practices>)"}
{"all_urls": ["https://github.com/promptfoo/promptfoo"], "content_sha256": "39e5d541c92970d8b75b49e701ae6fe75695b142321a38112cc14c88e8ad0361", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-081-promptfoo.raw.txt", "note_path": null, "primary_url": "https://github.com/promptfoo/promptfoo", "section": "5a · Eval frameworks & harnesses (code-first test-runners)", "source_id": "ae-081-promptfoo", "source_line": 151, "source_type": "repository_or_docs", "status": "mirrored", "title": "promptfoo", "why_it_matters": "MIT eval + red-teaming CLI; git-diffable YAML configs. **(MUST)**"}
{"all_urls": ["https://github.com/confident-ai/deepeval"], "content_sha256": "138fdabc91727fb2f32c3608a961d751c405e93ec1e2ccb56309c931c04a69d1", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-082-deepeval-confident-ai.raw.txt", "note_path": null, "primary_url": "https://github.com/confident-ai/deepeval", "section": "5a · Eval frameworks & harnesses (code-first test-runners)", "source_id": "ae-082-deepeval-confident-ai", "source_line": 152, "source_type": "repository_or_docs", "status": "mirrored", "title": "DeepEval / Confident AI", "why_it_matters": "\"pytest for LLMs,\" 40+ metrics (G-Eval, RAG, hallucination) + red-team; ~2M evals/day; hosted cloud. 🆕"}
{"all_urls": ["https://github.com/pydantic/pydantic-ai"], "content_sha256": "4810e54183f6ab922d1dd31438ae8f1a6ba94950add254f667dca801d45c8503", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-083-pydantic-evals.raw.txt", "note_path": null, "primary_url": "https://github.com/pydantic/pydantic-ai", "section": "5a · Eval frameworks & harnesses (code-first test-runners)", "source_id": "ae-083-pydantic-evals", "source_line": 153, "source_type": "repository_or_docs", "status": "mirrored", "title": "pydantic-evals", "why_it_matters": "🆕 type-safe Datasets/Cases/Evaluators with OTel tracing, from the Pydantic AI team."}
{"all_urls": ["https://github.com/langchain-ai/openevals", "https://github.com/langchain-ai/agentevals"], "content_sha256": "60f08ff1afa489dc5208d482f87979268a4563b8cfb964587114455327d3d7d1", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-084-openevals.raw.txt", "note_path": null, "primary_url": "https://github.com/langchain-ai/openevals", "section": "5a · Eval frameworks & harnesses (code-first test-runners)", "source_id": "ae-084-openevals", "source_line": 154, "source_type": "repository_or_docs", "status": "mirrored", "title": "openevals", "why_it_matters": "🆕 prebuilt evaluators + `create_llm_as_judge` (incl. multimodal); general-purpose companion to **agentevals** (<https://github.com/langchain-ai/agentevals>, trajectory match)."}
{"all_urls": ["https://mlflow.org/docs/latest/genai/eval-monitor/"], "content_sha256": "a28983ab44d522bae476d3a47e39d3728258e73a88ff5be96ce66d7f4b577d02", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-085-mlflow-genai-evaluate.raw.txt", "note_path": null, "primary_url": "https://mlflow.org/docs/latest/genai/eval-monitor/", "section": "5a · Eval frameworks & harnesses (code-first test-runners)", "source_id": "ae-085-mlflow-genai-evaluate", "source_line": 155, "source_type": "docs_or_book", "status": "mirrored", "title": "MLflow GenAI evaluate", "why_it_matters": "🆕 `mlflow.genai.evaluate`: 50+ judges/metrics, custom scorers, regression datasets inside MLflow."}
{"all_urls": ["https://github.com/stanford-crfm/helm"], "content_sha256": "c368539488eb44c518978d08f8e6ea31a78a7c5e3d3e06bc49bcbdae1ab96fb3", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-086-helm-crfm-helm.raw.txt", "note_path": null, "primary_url": "https://github.com/stanford-crfm/helm", "section": "5a · Eval frameworks & harnesses (code-first test-runners)", "source_id": "ae-086-helm-crfm-helm", "source_line": 156, "source_type": "repository_or_docs", "status": "mirrored", "title": "HELM (crfm-helm)", "why_it_matters": "holistic eval: standardized datasets + metrics beyond accuracy + leaderboard (also VHELM, HEIM)."}
{"all_urls": ["https://github.com/Giskard-AI/giskard-oss"], "content_sha256": "bbee6cee85b50d463678f24200f11157d0098b951ad2d32ae452d00d952e1a36", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-087-giskard.raw.txt", "note_path": null, "primary_url": "https://github.com/Giskard-AI/giskard-oss", "section": "5a · Eval frameworks & harnesses (code-first test-runners)", "source_id": "ae-087-giskard", "source_line": 157, "source_type": "repository_or_docs", "status": "mirrored", "title": "Giskard", "why_it_matters": "auto-generates adversarial test suites (injection, hallucination, bias) from a plain-language app description."}
{"all_urls": ["https://github.com/deepchecks/deepchecks"], "content_sha256": "e2d2b4017f2687e2c13d89c5abb4ab494fd5694809a4be174e2895cd6a4aa226", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-088-deepchecks-llm.raw.txt", "note_path": null, "primary_url": "https://github.com/deepchecks/deepchecks", "section": "5a · Eval frameworks & harnesses (code-first test-runners)", "source_id": "ae-088-deepchecks-llm", "source_line": 158, "source_type": "repository_or_docs", "status": "mirrored", "title": "Deepchecks LLM", "why_it_matters": "property-based scoring (grounded-in-context, toxicity, fluency) + custom LLM-judge properties."}
{"all_urls": ["https://github.com/uptrain-ai/uptrain"], "content_sha256": "d645716799ccf96a17d75d6fe22262eec8fb6c7da87c6256272375670436d762", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-089-uptrain.raw.txt", "note_path": null, "primary_url": "https://github.com/uptrain-ai/uptrain", "section": "5a · Eval frameworks & harnesses (code-first test-runners)", "source_id": "ae-089-uptrain", "source_line": 159, "source_type": "repository_or_docs", "status": "mirrored", "title": "UpTrain", "why_it_matters": "20+ preconfigured checks + root-cause analysis on failures."}
{"all_urls": ["https://github.com/huggingface/evaluate"], "content_sha256": "1698df1bc9babb362e22f66bde41743c967f8a9ef21b46bbf87d710aeec068a4", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-090-hf-evaluate.raw.txt", "note_path": null, "primary_url": "https://github.com/huggingface/evaluate", "section": "5a · Eval frameworks & harnesses (code-first test-runners)", "source_id": "ae-090-hf-evaluate", "source_line": 160, "source_type": "repository_or_docs", "status": "mirrored", "title": "HF `evaluate`", "why_it_matters": "classic metrics library, ⚠️ maintenance mode (use lighteval for LLMs)."}
{"all_urls": ["https://github.com/harbor-framework/harbor"], "content_sha256": "8c60f53504cc837028ef3279a63b03f92a1167a7057749547820b6a4358c26a1", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-091-harbor.raw.txt", "note_path": null, "primary_url": "https://github.com/harbor-framework/harbor", "section": "5a · Eval frameworks & harnesses (code-first test-runners)", "source_id": "ae-091-harbor", "source_line": 161, "source_type": "repository_or_docs", "status": "mirrored", "title": "Harbor", "why_it_matters": "🆕 framework for running agent evals + creating/using RL environments; powers Terminal-Bench 2.0. ~2.7k★. ⚠️ name overloaded (cf. `av/harbor` local-LLM toolkit)."}
{"all_urls": ["https://github.com/mattpocock/evalite"], "content_sha256": "7cf3a18126d185903d2ed7b6bf19810098f12acc30b7cf2b1602322e0ae22680", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-092-evalite.raw.txt", "note_path": null, "primary_url": "https://github.com/mattpocock/evalite", "section": "5b · TypeScript/JS-native eval runners", "source_id": "ae-092-evalite", "source_line": 164, "source_type": "repository_or_docs", "status": "mirrored", "title": "evalite", "why_it_matters": "🆕 local-first eval runner on Vitest; `.eval.ts` files, web UI, cost-aware."}
{"all_urls": ["https://github.com/mastra-ai/mastra"], "content_sha256": "90be153179d780259348b1297d62f4a8faa501b05ecec1e2bc305878c73bb264", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/mastra-scorers.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-093-mastra-scorers.raw.txt", "note_path": "notes/articles/mastra-scorers.md", "primary_url": "https://github.com/mastra-ai/mastra", "section": "5b · TypeScript/JS-native eval runners", "source_id": "ae-093-mastra-scorers", "source_line": 165, "source_type": "repository_or_docs", "status": "mirrored", "title": "Mastra scorers", "why_it_matters": "🆕 model-graded/rule/statistical scorers, live evals, CI, in the Mastra agent framework."}
{"all_urls": ["https://github.com/vercel-labs/agent-eval"], "content_sha256": "9f95f29b87ecb8e514d64bc9443ed8e727e89287c43be239d5d95fb32c20d7dc", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/vercel-eval-driven-development.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-094-vercel-agent-eval.raw.txt", "note_path": "notes/articles/vercel-eval-driven-development.md", "primary_url": "https://github.com/vercel-labs/agent-eval", "section": "5b · TypeScript/JS-native eval runners", "source_id": "ae-094-vercel-agent-eval", "source_line": 166, "source_type": "repository_or_docs", "status": "mirrored", "title": "Vercel agent-eval", "why_it_matters": "🆕 A/B-test coding agents (Claude Code, Codex, Cursor) on custom tasks; pass-rate dashboards."}
{"all_urls": ["https://github.com/braintrustdata/autoevals"], "content_sha256": "e73aaacece7f23ee8777537c5c6860131566aee13da1adc77eefa71ce7b1ceec", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-095-autoevals.raw.txt", "note_path": null, "primary_url": "https://github.com/braintrustdata/autoevals", "section": "5b · TypeScript/JS-native eval runners", "source_id": "ae-095-autoevals", "source_line": 167, "source_type": "repository_or_docs", "status": "mirrored", "title": "Autoevals", "why_it_matters": "OSS scorer library (Factuality, relevance, security…) across Py/JS/Go/Ruby."}
{"all_urls": ["https://github.com/truera/trulens"], "content_sha256": "49272ff8a86e1bb5ff189ce5c09ad5c546e63880913647ff0fada21e16a39e9b", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-096-trulens.raw.txt", "note_path": null, "primary_url": "https://github.com/truera/trulens", "section": "5c · RAG / retrieval evaluation", "source_id": "ae-096-trulens", "source_line": 170, "source_type": "repository_or_docs", "status": "mirrored", "title": "TruLens", "why_it_matters": "instrumentation + \"feedback functions\" (the RAG triad), now OTel-based."}
{"all_urls": ["https://github.com/stanford-futuredata/ARES"], "content_sha256": "5ecf26044ff0caa518136549b4762fda48246c1fe216c78a9793b3d05cc07302", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-097-ares.raw.txt", "note_path": null, "primary_url": "https://github.com/stanford-futuredata/ARES", "section": "5c · RAG / retrieval evaluation", "source_id": "ae-097-ares", "source_line": 171, "source_type": "repository_or_docs", "status": "mirrored", "title": "ARES", "why_it_matters": "synthetic queries + fine-tuned judges + prediction-powered inference for confidence intervals."}
{"all_urls": ["https://github.com/amazon-science/RAGChecker"], "content_sha256": "9d087e101857395f7c5d0b2a93d49e8a9f765282b2f2db2931813f5dd5809e15", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-098-ragchecker.raw.txt", "note_path": null, "primary_url": "https://github.com/amazon-science/RAGChecker", "section": "5c · RAG / retrieval evaluation", "source_id": "ae-098-ragchecker", "source_line": 172, "source_type": "repository_or_docs", "status": "mirrored", "title": "RAGChecker", "why_it_matters": "🆕 claim-level diagnosis separating retriever vs generator errors."}
{"all_urls": ["https://github.com/relari-ai/continuous-eval"], "content_sha256": "2e97ee3b6ef828fbf47b20537d8dc2593c7045852ba6d668ea6ea5ee3dd18d17", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-099-continuous-eval-relari.raw.txt", "note_path": null, "primary_url": "https://github.com/relari-ai/continuous-eval", "section": "5c · RAG / retrieval evaluation", "source_id": "ae-099-continuous-eval-relari", "source_line": 173, "source_type": "repository_or_docs", "status": "mirrored", "title": "continuous-eval (Relari)", "why_it_matters": "modular per-module metrics across retrieval/generation/tool-use."}
{"all_urls": ["https://github.com/TonicAI/tonic_validate"], "content_sha256": "308c4c6bbeb8d3cb7aa99fec9a53706d84f1a15f7032f8fc52148726df90b087", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-100-tonic-validate.raw.txt", "note_path": null, "primary_url": "https://github.com/TonicAI/tonic_validate", "section": "5c · RAG / retrieval evaluation", "source_id": "ae-100-tonic-validate", "source_line": 174, "source_type": "repository_or_docs", "status": "mirrored", "title": "Tonic Validate", "why_it_matters": "RAG metrics as a GitHub Action for CI."}
{"all_urls": ["https://github.com/haizelabs/verdict"], "content_sha256": "5897c65e973b2c5aff5c7c17a42fdecfffcddcd284d5c85511b1b13dbe76717c", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-101-verdict.raw.txt", "note_path": null, "primary_url": "https://github.com/haizelabs/verdict", "section": "5d · LLM-as-judge / reward / verifier libraries", "source_id": "ae-101-verdict", "source_line": 177, "source_type": "repository_or_docs", "status": "mirrored", "title": "verdict", "why_it_matters": "🆕 declarative compound judges (debate/verification/aggregation, inference-time scaling); arXiv:2502.18018."}
{"all_urls": ["https://github.com/OpenPipe/ART"], "content_sha256": "1ef910fb21451ba797c0abfa41efe10aee4ebec99e26bef93e78c07e29614167", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-102-ruler.raw.txt", "note_path": null, "primary_url": "https://github.com/OpenPipe/ART", "section": "5d · LLM-as-judge / reward / verifier libraries", "source_id": "ae-102-ruler", "source_line": 178, "source_type": "repository_or_docs", "status": "mirrored", "title": "RULER", "why_it_matters": "judge-as-RL-reward. **(industry must-read)**"}
{"all_urls": ["https://github.com/prometheus-eval/prometheus-eval"], "content_sha256": "6619fe0ac8084236d29de42dd627a42e52d923c3b489bf8bb95b6245cbf69e0c", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-103-prometheus-2.raw.txt", "note_path": null, "primary_url": "https://github.com/prometheus-eval/prometheus-eval", "section": "5d · LLM-as-judge / reward / verifier libraries", "source_id": "ae-103-prometheus-2", "source_line": 179, "source_type": "repository_or_docs", "status": "mirrored", "title": "Prometheus 2", "why_it_matters": "open-weight evaluator LMs for rubric-based assessment + pairwise."}
{"all_urls": ["https://github.com/atla-ai/selene-mini"], "content_sha256": "4ae6e58dda6d11aa0701591bb31d363c7e16bcfad4513308ccc1f1f9fdae4d98", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-104-atla-selene.raw.txt", "note_path": null, "primary_url": "https://github.com/atla-ai/selene-mini", "section": "5d · LLM-as-judge / reward / verifier libraries", "source_id": "ae-104-atla-selene", "source_line": 180, "source_type": "repository_or_docs", "status": "mirrored", "title": "Atla Selene", "why_it_matters": "🆕 8B SoTA open judge (score + critique); + MCP server `atla-ai/atla-mcp-server`. arXiv:2501.17195."}
{"all_urls": ["https://github.com/patronus-ai/Lynx-hallucination-detection", "https://github.com/patronus-ai/glider"], "content_sha256": "ed2fadd4540805786da2ba2ec0ca6c5e621597b0a7318afe243d5d286d0a4dcb", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-105-patronus-lynx-glider.raw.txt", "note_path": null, "primary_url": "https://github.com/patronus-ai/Lynx-hallucination-detection", "section": "5d · LLM-as-judge / reward / verifier libraries", "source_id": "ae-105-patronus-lynx-glider", "source_line": 181, "source_type": "repository_or_docs", "status": "mirrored", "title": "Patronus Lynx / GLIDER", "why_it_matters": "🆕 open hallucination judge / explainable span-level judge."}
{"all_urls": ["https://github.com/flowaicom/flow-judge"], "content_sha256": "61f709c29407b9b021e2cdb3b09f571441974955b8ab7813e1d2042436cb6c43", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-106-flow-judge.raw.txt", "note_path": null, "primary_url": "https://github.com/flowaicom/flow-judge", "section": "5d · LLM-as-judge / reward / verifier libraries", "source_id": "ae-106-flow-judge", "source_line": 182, "source_type": "repository_or_docs", "status": "mirrored", "title": "Flow-Judge", "why_it_matters": "efficient 3.8B open evaluator."}
{"all_urls": ["https://github.com/allenai/reward-bench"], "content_sha256": "6d0614ea70ae1d34fd5d274a5f37e03fe113967d16188dcf1fba94451739e3ca", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-107-rewardbench.raw.txt", "note_path": null, "primary_url": "https://github.com/allenai/reward-bench", "section": "5d · LLM-as-judge / reward / verifier libraries", "source_id": "ae-107-rewardbench", "source_line": 183, "source_type": "repository_or_docs", "status": "mirrored", "title": "RewardBench", "why_it_matters": "canonical reward-model (+v2 judge) benchmark/harness."}
{"all_urls": ["https://github.com/ScalerLab/JudgeBench"], "content_sha256": "c8dde52ad92b885b53fbe7a2210768289008ace4f8adbcdd6dc735dfb6dd8d90", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-108-judgebench.raw.txt", "note_path": null, "primary_url": "https://github.com/ScalerLab/JudgeBench", "section": "5d · LLM-as-judge / reward / verifier libraries", "source_id": "ae-108-judgebench", "source_line": 184, "source_type": "repository_or_docs", "status": "mirrored", "title": "JudgeBench", "why_it_matters": "benchmark to evaluate the judges themselves."}
{"all_urls": ["https://github.com/fw-ai-external/reward-kit"], "content_sha256": "eef535822a6567ef0129d1624a1f79e758e734fed13e8b7c9bfb505d07532098", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-109-reward-kit.raw.txt", "note_path": null, "primary_url": "https://github.com/fw-ai-external/reward-kit", "section": "5d · LLM-as-judge / reward / verifier libraries", "source_id": "ae-109-reward-kit", "source_line": 185, "source_type": "repository_or_docs", "status": "mirrored", "title": "reward-kit", "why_it_matters": "🆕 decorator-based reward-function authoring (TRL/Fireworks interop)."}
{"all_urls": ["https://github.com/PrimeIntellect-ai/verifiers"], "content_sha256": "d09eef5e271a97ac198fbaaab075ce75bab24113f95892b1dfd744da54dbff41", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-110-verifiers.raw.txt", "note_path": null, "primary_url": "https://github.com/PrimeIntellect-ai/verifiers", "section": "5e · RL-environment / verifiable-reward toolkits (eval ⇄ training)", "source_id": "ae-110-verifiers", "source_line": 188, "source_type": "repository_or_docs", "status": "mirrored", "title": "verifiers", "why_it_matters": "Environment = dataset + harness + rubric; one package for eval, RL, synthetic data. **(MUST)**"}
{"all_urls": ["https://github.com/PrimeIntellect-ai/community-environments"], "content_sha256": "c452122e496646c18fbeddfbe45405b34ad2149361fb32f5b14bc737f19711d2", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-111-environments-hub.raw.txt", "note_path": null, "primary_url": "https://github.com/PrimeIntellect-ai/community-environments", "section": "5e · RL-environment / verifiable-reward toolkits (eval ⇄ training)", "source_id": "ae-111-environments-hub", "source_line": 189, "source_type": "repository_or_docs", "status": "mirrored", "title": "Environments Hub", "why_it_matters": "🆕 crowdsourced verifiers-based RL/eval envs."}
{"all_urls": ["https://github.com/PrimeIntellect-ai/prime-rl"], "content_sha256": "271ab5786e015572ea5f684385667d3db081c0fd045ab7b65bc8b46d821717d3", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-112-prime-rl.raw.txt", "note_path": null, "primary_url": "https://github.com/PrimeIntellect-ai/prime-rl", "section": "5e · RL-environment / verifiable-reward toolkits (eval ⇄ training)", "source_id": "ae-112-prime-rl", "source_line": 190, "source_type": "repository_or_docs", "status": "mirrored", "title": "prime-rl", "why_it_matters": "🆕 async RL trainer consuming verifiers envs (INTELLECT-3)."}
{"all_urls": ["https://github.com/benchflow-ai/benchflow", "https://benchflow.ai"], "content_sha256": "a8d7201002d55117130f3af7b0d978adc15e76798501a2ecf8d8b7c1893ffd84", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-113-benchflow.raw.txt", "note_path": null, "primary_url": "https://github.com/benchflow-ai/benchflow", "section": "5e · RL-environment / verifiable-reward toolkits (eval ⇄ training)", "source_id": "ae-113-benchflow", "source_line": 191, "source_type": "repository_or_docs", "status": "mirrored", "title": "BenchFlow", "why_it_matters": "🆕 environment lab: builds & runs RL/eval environments (SkillsBench, ClawsBench, runtime). \"Environments are the new data.\" (also §5a)"}
{"all_urls": ["https://github.com/hud-evals/hud-python"], "content_sha256": "29e0838394a9059ad981edeb8cb5595e83cf40ca934b1a2a75052052bc129f35", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-114-hud.raw.txt", "note_path": null, "primary_url": "https://github.com/hud-evals/hud-python", "section": "5e · RL-environment / verifiable-reward toolkits (eval ⇄ training)", "source_id": "ae-114-hud", "source_line": 192, "source_type": "repository_or_docs", "status": "mirrored", "title": "HUD", "why_it_matters": "🆕 SDK to build/run agent eval environments (computer-use, browser, MCP) with telemetry."}
{"all_urls": ["https://github.com/NousResearch/atropos"], "content_sha256": "5016664f26ece61a80271e10b03d01a974641343d71dc83197bebe15c71c4c74", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-115-atropos.raw.txt", "note_path": null, "primary_url": "https://github.com/NousResearch/atropos", "section": "5e · RL-environment / verifiable-reward toolkits (eval ⇄ training)", "source_id": "ae-115-atropos", "source_line": 193, "source_type": "repository_or_docs", "status": "mirrored", "title": "Atropos", "why_it_matters": "🆕 async \"environment microservice\" framework for rollouts/verifiable rewards."}
{"all_urls": ["https://github.com/volcengine/verl"], "content_sha256": "632c2103709f8b5c0e080c54e039cc232f337d3e6541c2afb82943943496b77a", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-116-verl.raw.txt", "note_path": null, "primary_url": "https://github.com/volcengine/verl", "section": "5e · RL-environment / verifiable-reward toolkits (eval ⇄ training)", "source_id": "ae-116-verl", "source_line": 194, "source_type": "repository_or_docs", "status": "mirrored", "title": "verl", "why_it_matters": "de-facto industry RLVR trainer (PPO/GRPO). ~22k★."}
{"all_urls": ["https://github.com/OpenRLHF/OpenRLHF", "https://github.com/NovaSky-AI/SkyRL", "https://github.com/areal-project/AReaL", "https://github.com/alibaba/ROLL", "https://github.com/agentica-project/rllm", "https://github.com/huggingface/trl"], "content_sha256": "e971ac54c8f6ddfde36af3927e55a1022214deed51c79c0f1725f577ead52b4e", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-117-openrlhf.raw.txt", "note_path": null, "primary_url": "https://github.com/OpenRLHF/OpenRLHF", "section": "5e · RL-environment / verifiable-reward toolkits (eval ⇄ training)", "source_id": "ae-117-openrlhf", "source_line": 195, "source_type": "repository_or_docs", "status": "mirrored", "title": "OpenRLHF", "why_it_matters": "the RL-training stack agents are post-trained + eval'd in."}
{"all_urls": ["https://docs.openreward.ai/"], "content_sha256": "464119c1a768a7b4a908c726262f01de7cea66c687e282c74f85246984126daf", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/open-reward-standard-ors.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-118-open-reward-standard-ors.raw.txt", "note_path": "notes/articles/open-reward-standard-ors.md", "primary_url": "https://docs.openreward.ai/", "section": "5e · RL-environment / verifiable-reward toolkits (eval ⇄ training)", "source_id": "ae-118-open-reward-standard-ors", "source_line": 196, "source_type": "docs_or_book", "status": "mirrored", "title": "Open Reward Standard (ORS)", "why_it_matters": "🆕 MCP-extending spec adding RL primitives (episodes, rewards, curriculum). ⚠️ no single canonical repo confirmed."}
{"all_urls": ["https://github.com/Arize-ai/phoenix"], "content_sha256": "17a97d209b10488f437cc8dcdd421fef271bb7c9854ca560a280bd484f010e2d", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-119-arize-phoenix.raw.txt", "note_path": null, "primary_url": "https://github.com/Arize-ai/phoenix", "section": "5f · Observability + eval platforms (tracing · datasets · online/offline · CI)", "source_id": "ae-119-arize-phoenix", "source_line": 199, "source_type": "repository_or_docs", "status": "mirrored", "title": "Arize Phoenix", "why_it_matters": "OSS OTel tracing + response/retrieval evals + datasets/experiments. **(MUST)**"}
{"all_urls": ["https://github.com/langfuse/langfuse"], "content_sha256": "2170f043b4e46002ed67703080fe2586db9c5e76400b89350ea815720a8de7d5", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-120-langfuse.raw.txt", "note_path": null, "primary_url": "https://github.com/langfuse/langfuse", "section": "5f · Observability + eval platforms (tracing · datasets · online/offline · CI)", "source_id": "ae-120-langfuse", "source_line": 200, "source_type": "repository_or_docs", "status": "mirrored", "title": "Langfuse", "why_it_matters": "OSS: evals (LLM-judge, feedback, manual labeling), datasets/experiments, prompt mgmt; self-hostable. 🆕"}
{"all_urls": ["https://github.com/comet-ml/opik"], "content_sha256": "48682867dc0c36d8e5efead765b9779b824c6b32b4aad2ab40df1600eebbc25f", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-121-opik.raw.txt", "note_path": null, "primary_url": "https://github.com/comet-ml/opik", "section": "5f · Observability + eval platforms (tracing · datasets · online/offline · CI)", "source_id": "ae-121-opik", "source_line": 201, "source_type": "repository_or_docs", "status": "mirrored", "title": "Opik", "why_it_matters": "🆕 fully-OSS eval + observability (judges, datasets, CI-runnable evals)."}
{"all_urls": ["https://github.com/wandb/weave"], "content_sha256": "b80f51e794763add6a6ff7285e933e8dbcfe911b75a5cb3b40d23c3686698d70", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-122-w-b-weave.raw.txt", "note_path": null, "primary_url": "https://github.com/wandb/weave", "section": "5f · Observability + eval platforms (tracing · datasets · online/offline · CI)", "source_id": "ae-122-w-b-weave", "source_line": 202, "source_type": "repository_or_docs", "status": "mirrored", "title": "W&B Weave", "why_it_matters": "`weave.Evaluation` scorers (exact/regex/model-graded/embedding) + Guardrails; comparison dashboards. 🆕 (Humanloop's migration target.)"}
{"all_urls": ["https://www.braintrust.dev/docs/start/eval-sdk"], "content_sha256": "8c8a88c65d922b19f100e94c8a5081dc21a16abbaeb049896a76f3c60c3d8eda", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-123-braintrust.raw.txt", "note_path": null, "primary_url": "https://www.braintrust.dev/docs/start/eval-sdk", "section": "5f · Observability + eval platforms (tracing · datasets · online/offline · CI)", "source_id": "ae-123-braintrust", "source_line": 203, "source_type": "docs_or_book", "status": "mirrored", "title": "Braintrust", "why_it_matters": "`Eval()` over golden datasets; offline vs online. **(MUST)**"}
{"all_urls": ["https://www.patronus.ai/"], "content_sha256": "e5a0ff2ed2a9f15ea9db2376f5e748400151c3df960d12968868f41e81605680", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-124-patronus-ai.raw.txt", "note_path": null, "primary_url": "https://www.patronus.ai/", "section": "5f · Observability + eval platforms (tracing · datasets · online/offline · CI)", "source_id": "ae-124-patronus-ai", "source_line": 204, "source_type": "web_article", "status": "mirrored", "title": "Patronus AI", "why_it_matters": "🆕 research-grade judges (Lynx, GLIDER, **Percival** agent-failure debugger), experiments, multimodal judge."}
{"all_urls": ["https://www.getmaxim.ai/"], "content_sha256": "a267fa4b50260a93385c06e03ee31890b35a9b74aebf98e6435e93d4becace4e", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-125-maxim-ai.raw.txt", "note_path": null, "primary_url": "https://www.getmaxim.ai/", "section": "5f · Observability + eval platforms (tracing · datasets · online/offline · CI)", "source_id": "ae-125-maxim-ai", "source_line": 205, "source_type": "web_article", "status": "mirrored", "title": "Maxim AI", "why_it_matters": "🆕 agent **simulation** + eval + observability across thousands of scenarios/personas."}
{"all_urls": ["https://galileo.ai/"], "content_sha256": "ccacb82db80d71a8bfaf920007db1112e53a3ec05b51bcb1fc8e54ffb8a88d7f", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-126-galileo.raw.txt", "note_path": null, "primary_url": "https://galileo.ai/", "section": "5f · Observability + eval platforms (tracing · datasets · online/offline · CI)", "source_id": "ae-126-galileo", "source_line": 206, "source_type": "web_article", "status": "mirrored", "title": "Galileo", "why_it_matters": "Luna evaluators + Agentic Evaluations."}
{"all_urls": ["https://www.vellum.ai/"], "content_sha256": "b33d7dce928a99b47b6f17dad0f78d4f25f452665e61e9f9c361f20bee4f82b5", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-127-vellum.raw.txt", "note_path": null, "primary_url": "https://www.vellum.ai/", "section": "5f · Observability + eval platforms (tracing · datasets · online/offline · CI)", "source_id": "ae-127-vellum", "source_line": 207, "source_type": "web_article", "status": "mirrored", "title": "Vellum", "why_it_matters": "visual workflows + offline/online evals scoring every production run."}
{"all_urls": ["https://github.com/helicone/helicone"], "content_sha256": "165412284d01732a0a8a9bf8b0600cc8eefac05f14799a0d61402b52be2936c3", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-128-helicone.raw.txt", "note_path": null, "primary_url": "https://github.com/helicone/helicone", "section": "5f · Observability + eval platforms (tracing · datasets · online/offline · CI)", "source_id": "ae-128-helicone", "source_line": 208, "source_type": "repository_or_docs", "status": "mirrored", "title": "Helicone", "why_it_matters": "OSS gateway + observability; \"Scores\" ingests external eval results."}
{"all_urls": ["https://github.com/traceloop/openllmetry"], "content_sha256": "38d6eeda187cf5793ba53fad0c9daa7cd45d2445380c4cedcf0decb0e2a76458", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-129-traceloop-openllmetry.raw.txt", "note_path": null, "primary_url": "https://github.com/traceloop/openllmetry", "section": "5f · Observability + eval platforms (tracing · datasets · online/offline · CI)", "source_id": "ae-129-traceloop-openllmetry", "source_line": 209, "source_type": "repository_or_docs", "status": "mirrored", "title": "Traceloop / OpenLLMetry", "why_it_matters": "OSS OTel instrumentation (Py/TS/Go/Ruby) + hosted reliability platform."}
{"all_urls": ["https://github.com/Scale3-Labs/langtrace"], "content_sha256": "cbca8691cdccf80ae1a0c1adc962978a6021a2604f86f7b0e7cbaf477e2510de", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-130-langtrace.raw.txt", "note_path": null, "primary_url": "https://github.com/Scale3-Labs/langtrace", "section": "5f · Observability + eval platforms (tracing · datasets · online/offline · CI)", "source_id": "ae-130-langtrace", "source_line": 210, "source_type": "repository_or_docs", "status": "mirrored", "title": "Langtrace", "why_it_matters": "OSS OTel-standard tracing + manual scoring + dataset mgmt."}
{"all_urls": ["https://github.com/whylabs/langkit"], "content_sha256": "ab6f2bc9cee3bafe9d295433500ba87de1c78571c3c35d0499b241e48635ca90", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-131-whylabs-langkit.raw.txt", "note_path": null, "primary_url": "https://github.com/whylabs/langkit", "section": "5f · Observability + eval platforms (tracing · datasets · online/offline · CI)", "source_id": "ae-131-whylabs-langkit", "source_line": 211, "source_type": "repository_or_docs", "status": "mirrored", "title": "WhyLabs / LangKit", "why_it_matters": "high-throughput text-signal metrics (toxicity, PII, jailbreak) for production monitoring."}
{"all_urls": ["https://github.com/portkey-ai/gateway"], "content_sha256": "7543749e8d864eedfdca1b7ae762373ebafa30bf775bbe6e0a9072eab21a7c68", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-132-portkey.raw.txt", "note_path": null, "primary_url": "https://github.com/portkey-ai/gateway", "section": "5f · Observability + eval platforms (tracing · datasets · online/offline · CI)", "source_id": "ae-132-portkey", "source_line": 212, "source_type": "repository_or_docs", "status": "mirrored", "title": "Portkey", "why_it_matters": "🆕 OSS gateway + 60+ guardrails + observability (fully open-sourced Mar 2026)."}
{"all_urls": ["https://www.datadoghq.com/product/ai/llm-observability/"], "content_sha256": "11740908ac0dc1fc09b3ed9b2e39581b9a12f049e43c61a02c1c6db1256b39a2", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-133-datadog-llm-observability.raw.txt", "note_path": null, "primary_url": "https://www.datadoghq.com/product/ai/llm-observability/", "section": "5f · Observability + eval platforms (tracing · datasets · online/offline · CI)", "source_id": "ae-133-datadog-llm-observability", "source_line": 213, "source_type": "web_article", "status": "mirrored", "title": "Datadog LLM Observability", "why_it_matters": "🆕 evaluators + golden datasets + **LLM Experiments** + AI Agent Monitoring (Jun 2025)."}
{"all_urls": ["https://www.fiddler.ai/"], "content_sha256": "6c3cbe8bc5c9464cae706696221bd3cd876fe356e817202b484d02400e61e05e", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-134-fiddler-ai.raw.txt", "note_path": null, "primary_url": "https://www.fiddler.ai/", "section": "5f · Observability + eval platforms (tracing · datasets · online/offline · CI)", "source_id": "ae-134-fiddler-ai", "source_line": 214, "source_type": "web_article", "status": "mirrored", "title": "Fiddler AI", "why_it_matters": "🆕 Trust Models (Safety/PII/Faithfulness) scoring in <100ms; Guardrails + agentic observability."}
{"all_urls": ["https://www.promptlayer.com/", "https://newrelic.com/platform/ai-monitoring"], "content_sha256": "bc0a2d71bb8b04d188f19c24bf606d24bd2f693de945674502130fce18224d0e", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-135-promptlayer.raw.txt", "note_path": null, "primary_url": "https://www.promptlayer.com/", "section": "5f · Observability + eval platforms (tracing · datasets · online/offline · CI)", "source_id": "ae-135-promptlayer", "source_line": 215, "source_type": "web_article", "status": "mirrored", "title": "PromptLayer", "why_it_matters": "lighter prompt-CMS / APM-native monitoring."}
{"all_urls": ["https://github.com/Arize-ai/openinference"], "content_sha256": "44cf49a50e2443f7d65631deaadf56a1a8efb553a49d516574de5c3c02d7f2db", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-136-openinference.raw.txt", "note_path": null, "primary_url": "https://github.com/Arize-ai/openinference", "section": "5g · Tracing standards", "source_id": "ae-136-openinference", "source_line": 218, "source_type": "repository_or_docs", "status": "mirrored", "title": "OpenInference", "why_it_matters": "semantic conventions for agent traces (tool/args/observation/latency/cost)."}
{"all_urls": ["https://opentelemetry.io/docs/specs/semconv/gen-ai/"], "content_sha256": "f8f5ac9a7e5a7bbbd7a836c5d3b2c1d457ebecf9de9e05e57793bd64d17d6848", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-137-opentelemetry-genai-semantic-conventions.raw.txt", "note_path": null, "primary_url": "https://opentelemetry.io/docs/specs/semconv/gen-ai/", "section": "5g · Tracing standards", "source_id": "ae-137-opentelemetry-genai-semantic-conventions", "source_line": 219, "source_type": "docs_or_book", "status": "mirrored", "title": "OpenTelemetry GenAI semantic conventions", "why_it_matters": "🆕 the vendor-neutral schema (now covers agent orchestration, MCP tool calls, and a **quality-evaluation** span hook)."}
{"all_urls": ["https://www.braintrust.dev/"], "content_sha256": "b076e975e3d0cf170adfbaf28313351be5ed023efc4090a179835d260f8a243a", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-138-braintrust.raw.txt", "note_path": null, "primary_url": "https://www.braintrust.dev/", "section": "5g · Tracing standards", "source_id": "ae-138-braintrust", "source_line": 221, "source_type": "web_article", "status": "mirrored", "title": "Braintrust", "why_it_matters": "Industry-standard eval+observability platform (Notion, Stripe, Vercel) tying offline experiments to production logs; the section already cites Braintrust's Autoevals but omits the platform itself. 🆕"}
{"all_urls": ["https://github.com/raga-ai-hub/RagaAI-Catalyst"], "content_sha256": "dac44c12d5efb68b9ff150c799552333e06dba38cccaafa81c42988db8953d11", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-139-ragaai-catalyst.raw.txt", "note_path": null, "primary_url": "https://github.com/raga-ai-hub/RagaAI-Catalyst", "section": "5g · Tracing standards", "source_id": "ae-139-ragaai-catalyst", "source_line": 222, "source_type": "repository_or_docs", "status": "mirrored", "title": "RagaAI Catalyst", "why_it_matters": "covers the online/guardrail-eval slice the section lacks. 🆕"}
{"all_urls": ["https://developers.openai.com/cookbook/topic/evals"], "content_sha256": "99b818cd6e3ab4648a01ee8b44f4a7bf3bb2d70f63276649010c57eeb4f93370", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/macro-evals-agentic-systems-openai-cookbook.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-140-openai-cookbook-evals.raw.txt", "note_path": "notes/articles/macro-evals-agentic-systems-openai-cookbook.md", "primary_url": "https://developers.openai.com/cookbook/topic/evals", "section": "5g · Tracing standards", "source_id": "ae-140-openai-cookbook-evals", "source_line": 223, "source_type": "docs_or_book", "status": "mirrored", "title": "OpenAI Cookbook — Evals", "why_it_matters": "Maintained, runnable recipes for building evals (incl. Agents SDK eval, evaluating agents with Langfuse); the practical companion to OpenAI Evals and a curator-grade 'show real work' resource. 🆕 ⚠(unverified URL)"}
{"all_urls": ["https://ofir.io/How-to-Build-Good-Language-Modeling-Benchmarks/"], "content_sha256": "cb4d4263c57f9558283b48d61549aa99a8931294452efa8cbe2860c1b4c0b7cd", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-141-how-to-build-good-language-modeling-benchmarks.raw.txt", "note_path": null, "primary_url": "https://ofir.io/How-to-Build-Good-Language-Modeling-Benchmarks/", "section": "6 · Benchmark vs. eval (and benchmark integrity: contamination, saturation, label errors, leaderboard gaming)", "source_id": "ae-141-how-to-build-good-language-modeling-benchmarks", "source_line": 229, "source_type": "web_article", "status": "mirrored", "title": "How to Build Good Language Modeling Benchmarks", "why_it_matters": "The benchmark-author's checklist; difficulty target; one-number reporting; 150–500 task sizing."}
{"all_urls": ["https://arxiv.org/abs/2407.01502"], "content_sha256": "a74a66bddeea47fd5b656b5bd316858bb9baf7b951491316ec773e834c43f53c", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-142-ai-agents-that-matter.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2407.01502", "section": "6 · Benchmark vs. eval (and benchmark integrity: contamination, saturation, label errors, leaderboard gaming)", "source_id": "ae-142-ai-agents-that-matter", "source_line": 230, "source_type": "paper_or_pdf", "status": "mirrored", "title": "AI Agents That Matter", "why_it_matters": "Cost-controlled evaluation; model-dev vs downstream-dev needs; holdouts."}
{"all_urls": ["https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/", "https://decrypt.co/359012/"], "content_sha256": "0784dd5613c68bdc2a53ab8c6abd0e98834c52c2544df4224144c4d0ee0b5d8b", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-143-why-we-no-longer-evaluate-swe-bench-verified.raw.txt", "note_path": null, "primary_url": "https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/", "section": "6 · Benchmark vs. eval (and benchmark integrity: contamination, saturation, label errors, leaderboard gaming)", "source_id": "ae-143-why-we-no-longer-evaluate-swe-bench-verified", "source_line": 231, "source_type": "web_article", "status": "mirrored", "title": "Why We No Longer Evaluate SWE-bench Verified", "why_it_matters": "~59% of audited failures were broken tests. (mirror: <https://decrypt.co/359012/...>)"}
{"all_urls": ["https://arxiv.org/abs/2504.20879"], "content_sha256": "3d80cddbaf82f9841da2961f528d89ed22fc95db154019ef6359b6cc6601cfd7", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/leaderboard-illusion.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-144-the-leaderboard-illusion.raw.txt", "note_path": "notes/articles/leaderboard-illusion.md", "primary_url": "https://arxiv.org/abs/2504.20879", "section": "6 · Benchmark vs. eval (and benchmark integrity: contamination, saturation, label errors, leaderboard gaming)", "source_id": "ae-144-the-leaderboard-illusion", "source_line": 232, "source_type": "paper_or_pdf", "status": "mirrored", "title": "The Leaderboard Illusion", "why_it_matters": "Private testing, selective disclosure, and data-access asymmetry on Chatbot Arena. *(notes: `research/notes/leaderboard-illusion.md`)*"}
{"all_urls": ["https://arxiv.org/abs/2506.12286"], "content_sha256": "085ac6c81b4d33668e3405b95a76e26c79a3925531c3a17fc16abf232df06a20", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-145-the-swe-bench-illusion-when-sota-llms-remember-i.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2506.12286", "section": "6 · Benchmark vs. eval (and benchmark integrity: contamination, saturation, label errors, leaderboard gaming)", "source_id": "ae-145-the-swe-bench-illusion-when-sota-llms-remember-i", "source_line": 233, "source_type": "paper_or_pdf", "status": "mirrored", "title": "The SWE-bench Illusion: When SOTA LLMs Remember Instead of Reason", "why_it_matters": "Memorization inflates SWE-bench scores."}
{"all_urls": ["https://arxiv.org/abs/2507.02825"], "content_sha256": "ea24245370ed1f98005c7bf00b44411763d9488744f9a586ebee821dd8c59236", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-146-establishing-best-practices-for-building-rigorou.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2507.02825", "section": "6 · Benchmark vs. eval (and benchmark integrity: contamination, saturation, label errors, leaderboard gaming)", "source_id": "ae-146-establishing-best-practices-for-building-rigorou", "source_line": 234, "source_type": "paper_or_pdf", "status": "mirrored", "title": "Establishing Best Practices for Building Rigorous Agentic Benchmarks (ABC)", "why_it_matters": "SWE-bench Verified weak tests; τ-bench rewards empty responses. *(verified high)*"}
{"all_urls": ["https://epoch.ai/benchmarks/frontiermath-tiers-1-3-v2"], "content_sha256": "9ba719b1cdda1396b760d66d0ac929dff907ee42f5e9b60a3129238418f08509", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-147-frontiermath-tiers-1-3-v2-corrected.raw.txt", "note_path": null, "primary_url": "https://epoch.ai/benchmarks/frontiermath-tiers-1-3-v2", "section": "6 · Benchmark vs. eval (and benchmark integrity: contamination, saturation, label errors, leaderboard gaming)", "source_id": "ae-147-frontiermath-tiers-1-3-v2-corrected", "source_line": 235, "source_type": "web_article", "status": "mirrored", "title": "FrontierMath Tiers 1–3 v2 (corrected)", "why_it_matters": "~42% of problems corrected after AI-assisted review. (also T8: the operator-as-rot-detector tale)"}
{"all_urls": ["https://www.futurehouse.org/research-announcements/hle-exam", "https://www.lesswrong.com/posts/JANqfGrMyBgcKtGgK/"], "content_sha256": "c1e184a78842724b34abe6c747a3dd88e12303edb0529fb06e0365e0b55dc5a6", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-148-about-30-of-humanity-s-last-exam-answers-are-wro.raw.txt", "note_path": null, "primary_url": "https://www.futurehouse.org/research-announcements/hle-exam", "section": "6 · Benchmark vs. eval (and benchmark integrity: contamination, saturation, label errors, leaderboard gaming)", "source_id": "ae-148-about-30-of-humanity-s-last-exam-answers-are-wro", "source_line": 236, "source_type": "web_article", "status": "mirrored", "title": "About 30% of Humanity's Last Exam Answers Are Wrong", "why_it_matters": "29 ± 3.7% of text-only chem/bio answers contradicted by the literature. (LessWrong writeup: <https://www.lesswrong.com/posts/JANqfGrMyBgcKtGgK/>)"}
{"all_urls": ["https://www.interconnects.ai/p/building-on-evaluation-quicksand"], "content_sha256": "b9885f095cf3cc54596b20530d30dbfae608ada2f47dec0aaa19cd751e883fb6", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-149-building-on-evaluation-quicksand.raw.txt", "note_path": null, "primary_url": "https://www.interconnects.ai/p/building-on-evaluation-quicksand", "section": "6 · Benchmark vs. eval (and benchmark integrity: contamination, saturation, label errors, leaderboard gaming)", "source_id": "ae-149-building-on-evaluation-quicksand", "source_line": 237, "source_type": "blog", "status": "mirrored", "title": "Building on Evaluation Quicksand", "why_it_matters": "No hard source of truth; synthetic-data contamination."}
{"all_urls": ["https://arxiv.org/abs/2601.17087"], "content_sha256": "35e70774dda9cebce0a3ab5ca52b9fcf77475c3499ee58a6e79083dce88a5fab", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-150-lost-in-simulation.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2601.17087", "section": "6 · Benchmark vs. eval (and benchmark integrity: contamination, saturation, label errors, leaderboard gaming)", "source_id": "ae-150-lost-in-simulation", "source_line": 238, "source_type": "paper_or_pdf", "status": "mirrored", "title": "Lost in Simulation", "why_it_matters": "Simulated users are unreliable proxies (~9pp swings by simulator choice; demographic miscalibration)."}
{"all_urls": ["https://arxiv.org/abs/2310.06770", "https://www.swebench.com"], "content_sha256": "998a92461b56f5d6109a4c9161c6a2a12649f95ca14715b0a439bfcf01f16885", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-151-swe-bench-can-lms-resolve-real-world-github-issu.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2310.06770", "section": "6 · Benchmark vs. eval (and benchmark integrity: contamination, saturation, label errors, leaderboard gaming)", "source_id": "ae-151-swe-bench-can-lms-resolve-real-world-github-issu", "source_line": 239, "source_type": "paper_or_pdf", "status": "mirrored", "title": "SWE-bench: Can LMs Resolve Real-World GitHub Issues?", "why_it_matters": "<https://arxiv.org/abs/2310.06770> · <https://www.swebench.com> (Verified: `.../verified.html`) · *paper/site*."}
{"all_urls": ["https://eugeneyan.com/writing/evals/"], "content_sha256": "8d24abf3315225ae2d2cba21ec7aca650bf76751a34ff5756e066f87b97662fb", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-karam-metrics-that-work.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-152-task-specific-llm-evals-that-do-don-t-work.raw.txt", "note_path": "notes/talks/talk-karam-metrics-that-work.md", "primary_url": "https://eugeneyan.com/writing/evals/", "section": "6 · Benchmark vs. eval (and benchmark integrity: contamination, saturation, label errors, leaderboard gaming)", "source_id": "ae-152-task-specific-llm-evals-that-do-don-t-work", "source_line": 240, "source_type": "web_article", "status": "mirrored", "title": "Task-Specific LLM Evals that Do & Don't Work", "why_it_matters": "Off-the-shelf evals rarely transfer; accuracy is too coarse."}
{"all_urls": ["https://x.com/karpathy/status/1896266683301659068"], "content_sha256": null, "http_status": null, "local_note_path": null, "local_raw_path": null, "note_path": null, "primary_url": "https://x.com/karpathy/status/1896266683301659068", "section": "6 · Benchmark vs. eval (and benchmark integrity: contamination, saturation, label errors, leaderboard gaming)", "source_id": "ae-153-andrej-karpathy-on-evals", "source_line": 241, "source_type": "web_article", "status": "metadata_only", "title": "Andrej Karpathy on evals", "why_it_matters": "\"We make a number of specific recommendations…\" (the eval-as-narrow critique)."}
{"all_urls": ["https://arxiv.org/abs/2405.00332"], "content_sha256": "b99d31f9cad364ff7fb30f77e3cbee2046b316aa07071df90bdcb886f33da26f", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-154-a-careful-examination-of-llm-performance-on-grad.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2405.00332", "section": "6 · Benchmark vs. eval (and benchmark integrity: contamination, saturation, label errors, leaderboard gaming)", "source_id": "ae-154-a-careful-examination-of-llm-performance-on-grad", "source_line": 243, "source_type": "paper_or_pdf", "status": "mirrored", "title": "A Careful Examination of LLM Performance on Grade School Arithmetic (GSM1k)", "why_it_matters": "the canonical method for measuring benchmark overfitting/contamination via a matched holdout."}
{"all_urls": ["https://arxiv.org/abs/2103.14749"], "content_sha256": "75e133664f0918aaf814c24f68cbbde17b19522246a86410b03aa47a475a1659", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/benchmark-label-errors.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-155-pervasive-label-errors-in-test-sets-destabilize.raw.txt", "note_path": "notes/articles/benchmark-label-errors.md", "primary_url": "https://arxiv.org/abs/2103.14749", "section": "6 · Benchmark vs. eval (and benchmark integrity: contamination, saturation, label errors, leaderboard gaming)", "source_id": "ae-155-pervasive-label-errors-in-test-sets-destabilize", "source_line": 244, "source_type": "paper_or_pdf", "status": "mirrored", "title": "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks", "why_it_matters": "NeurIPS 2021 foundational result: ~3.3% avg label errors across 10 famous test sets (ImageNet, MNIST, etc.); corrections flip model rankings. The canonical 'label errors' citation this section's theme rests on (labelerrors.com / cleanlab)."}
{"all_urls": ["https://arxiv.org/abs/2406.04127"], "content_sha256": "4a6ff40c77d2036aec987627e213c124ae348016b0f141448f205ced8b97b424", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-156-are-we-done-with-mmlu-mmlu-redux.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2406.04127", "section": "6 · Benchmark vs. eval (and benchmark integrity: contamination, saturation, label errors, leaderboard gaming)", "source_id": "ae-156-are-we-done-with-mmlu-mmlu-redux", "source_line": 245, "source_type": "paper_or_pdf", "status": "mirrored", "title": "Are We Done with MMLU? (MMLU-Redux)", "why_it_matters": "directly demonstrates label-error impact on the most-cited LLM benchmark."}
{"all_urls": ["https://arxiv.org/abs/2403.07974"], "content_sha256": "c394e902fccb4705dcfcf3df5aa08c4376d19ef202a8879faff1107232333720", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-157-livecodebench-holistic-and-contamination-free-ev.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2403.07974", "section": "6 · Benchmark vs. eval (and benchmark integrity: contamination, saturation, label errors, leaderboard gaming)", "source_id": "ae-157-livecodebench-holistic-and-contamination-free-ev", "source_line": 246, "source_type": "paper_or_pdf", "status": "mirrored", "title": "LiveCodeBench: Holistic and Contamination-Free Evaluation of LLMs for Code", "why_it_matters": "the section discusses contamination but lists no exemplar of how to engineer around it."}
{"all_urls": ["https://github.com/LiveBench/LiveBench"], "content_sha256": "3b1845679ec5f530222bae9e702e6add1f6e14a7855674070332237ffbf9e5ee", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-158-livebench-a-challenging-contamination-limited-ll.raw.txt", "note_path": null, "primary_url": "https://github.com/LiveBench/LiveBench", "section": "6 · Benchmark vs. eval (and benchmark integrity: contamination, saturation, label errors, leaderboard gaming)", "source_id": "ae-158-livebench-a-challenging-contamination-limited-ll", "source_line": 247, "source_type": "repository_or_docs", "status": "mirrored", "title": "LiveBench: A Challenging, Contamination-Limited LLM Benchmark", "why_it_matters": "the canonical 'dynamic refresh' answer to saturation and contamination."}
{"all_urls": ["https://github.com/huggingface/evaluation-guidebook"], "content_sha256": "6d7a762f14c291f97b6d8d06c0409274560472dabcf7226d722d2b8f038abc07", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-159-the-llm-evaluation-guidebook-open-llm-leaderboar.raw.txt", "note_path": null, "primary_url": "https://github.com/huggingface/evaluation-guidebook", "section": "6 · Benchmark vs. eval (and benchmark integrity: contamination, saturation, label errors, leaderboard gaming)", "source_id": "ae-159-the-llm-evaluation-guidebook-open-llm-leaderboar", "source_line": 248, "source_type": "repository_or_docs", "status": "mirrored", "title": "The LLM Evaluation Guidebook (Open LLM Leaderboard team)", "why_it_matters": "the hands-on 'how to not get fooled' companion to this section (updated version: hf.co/spaces/OpenEvals/evaluation-guidebook)."}
{"all_urls": ["https://arxiv.org/abs/2510.11977"], "content_sha256": "350ef898cf10e85859147e9794539cc6f178cb70c3f383dd355cc337b2548919", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/langfuse-agent-evaluation-guide.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-160-holistic-agent-leaderboard-the-missing-infrastru.raw.txt", "note_path": "notes/articles/langfuse-agent-evaluation-guide.md", "primary_url": "https://arxiv.org/abs/2510.11977", "section": "6 · Benchmark vs. eval (and benchmark integrity: contamination, saturation, label errors, leaderboard gaming)", "source_id": "ae-160-holistic-agent-leaderboard-the-missing-infrastru", "source_line": 249, "source_type": "paper_or_pdf", "status": "mirrored", "title": "Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation", "why_it_matters": "extends 'AI Agents That Matter' to leaderboard integrity for agents specifically. 🆕"}
{"all_urls": ["https://blog.collinear.ai/p/gaming-the-system-goodharts-law-exemplified-in-ai-leaderboard-controversy"], "content_sha256": "609e517fb9faf0fbb90575b2f734978585a707145b860a8fce879bd1791ac052", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-161-gaming-the-system-goodhart-s-law-exemplified-in.raw.txt", "note_path": null, "primary_url": "https://blog.collinear.ai/p/gaming-the-system-goodharts-law-exemplified-in-ai-leaderboard-controversy", "section": "6 · Benchmark vs. eval (and benchmark integrity: contamination, saturation, label errors, leaderboard gaming)", "source_id": "ae-161-gaming-the-system-goodhart-s-law-exemplified-in", "source_line": 250, "source_type": "blog", "status": "mirrored", "title": "Gaming the System: Goodhart's Law Exemplified in the AI Leaderboard Controversy", "why_it_matters": "the accessible blog companion to The Leaderboard Illusion paper. 🆕"}
{"all_urls": ["https://openai.com/index/trustworthy-third-party-evaluations-foundations/"], "content_sha256": "25e903ef685304e38a3b6c3928369efb510e714f6e0f9647fc709fb65c0488da", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-162-a-shared-playbook-for-trustworthy-third-party-ev.raw.txt", "note_path": null, "primary_url": "https://openai.com/index/trustworthy-third-party-evaluations-foundations/", "section": "6 · Benchmark vs. eval (and benchmark integrity: contamination, saturation, label errors, leaderboard gaming)", "source_id": "ae-162-a-shared-playbook-for-trustworthy-third-party-ev", "source_line": 251, "source_type": "docs_or_book", "status": "mirrored", "title": "A Shared Playbook for Trustworthy Third-Party Evaluations", "why_it_matters": "What makes *independent* evals of frontier-model safeguards & capabilities trustworthy: selecting the right harness, checking for validity hazards that distort results, and the standards third-party evaluators need. (also T10) 🆕"}
{"all_urls": ["https://arxiv.org/abs/2403.13787"], "content_sha256": "a8176cbcdef3902d01a231f9274dc475ad11a672eec995c1d7af5fa5b588a3a2", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-163-rewardbench.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2403.13787", "section": "7 · Evals & RL environments (verifiers, reward design, difficulty calibration, lifecycle)", "source_id": "ae-163-rewardbench", "source_line": 259, "source_type": "paper_or_pdf", "status": "mirrored", "title": "RewardBench", "why_it_matters": "Evaluating reward models (the verifier you train against)."}
{"all_urls": ["https://www.interconnects.ai/p/the-new-rl-scaling-laws", "https://www.latent.space/p/the-rlvr-revolution-with-nathan-lambert"], "content_sha256": "1f61a39dd61add7266f332fbdba7c895263b495c6cf72b03418f291efbdaa5d8", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-164-the-new-rl-scaling-laws.raw.txt", "note_path": null, "primary_url": "https://www.interconnects.ai/p/the-new-rl-scaling-laws", "section": "7 · Evals & RL environments (verifiers, reward design, difficulty calibration, lifecycle)", "source_id": "ae-164-the-new-rl-scaling-laws", "source_line": 260, "source_type": "blog", "status": "mirrored", "title": "The New RL Scaling Laws", "why_it_matters": "Where RLVR scaling is heading. (interview: <https://www.latent.space/p/the-rlvr-revolution-with-nathan-lambert>)"}
{"all_urls": ["https://arxiv.org/abs/2506.10947"], "content_sha256": "1f44ca6d14cf09da3e7387f28fdf85d53c45caa9eefa9220dede6356c0744341", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-165-spurious-rewards-rethinking-training-signals-in.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2506.10947", "section": "7 · Evals & RL environments (verifiers, reward design, difficulty calibration, lifecycle)", "source_id": "ae-165-spurious-rewards-rethinking-training-signals-in", "source_line": 261, "source_type": "paper_or_pdf", "status": "mirrored", "title": "Spurious Rewards: Rethinking Training Signals in RLVR", "why_it_matters": "see `research/notes/reference-audit.md`)*"}
{"all_urls": ["https://www.interconnects.ai/p/the-state-of-post-training-2025"], "content_sha256": "09bf6865b676d05474ad3b8cf076fe702674f41163e88008e97dd7b94951fbee", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-166-the-state-of-post-training-2025.raw.txt", "note_path": null, "primary_url": "https://www.interconnects.ai/p/the-state-of-post-training-2025", "section": "7 · Evals & RL environments (verifiers, reward design, difficulty calibration, lifecycle)", "source_id": "ae-166-the-state-of-post-training-2025", "source_line": 262, "source_type": "blog", "status": "mirrored", "title": "The State of Post-Training 2025", "why_it_matters": "Context for where evals feed training."}
{"all_urls": ["https://lilianweng.github.io/posts/2024-11-28-reward-hacking/"], "content_sha256": "99da73635a1800451d546644414161325c9fcd1455e971df5bf3dce987bbdd4e", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/countdown-code-reward-hacking-rlvr-testbed.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-167-reward-hacking-in-reinforcement-learning.raw.txt", "note_path": "notes/articles/countdown-code-reward-hacking-rlvr-testbed.md", "primary_url": "https://lilianweng.github.io/posts/2024-11-28-reward-hacking/", "section": "7 · Evals & RL environments (verifiers, reward design, difficulty calibration, lifecycle)", "source_id": "ae-167-reward-hacking-in-reinforcement-learning", "source_line": 264, "source_type": "web_article", "status": "mirrored", "title": "Reward Hacking in Reinforcement Learning", "why_it_matters": "taxonomy, RLHF-specific failure modes, mitigations; the foundational reference any reward-design section needs."}
{"all_urls": ["https://deepmind.google/blog/specification-gaming-the-flip-side-of-ai-ingenuity/"], "content_sha256": "3e6df14ab8426ce1030836a5490f745dd81b8ed57396f269623fc0135c56adf5", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/recontextualization-mitigates-specification-gaming.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-168-specification-gaming-the-flip-side-of-ai-ingenui.raw.txt", "note_path": "notes/articles/recontextualization-mitigates-specification-gaming.md", "primary_url": "https://deepmind.google/blog/specification-gaming-the-flip-side-of-ai-ingenuity/", "section": "7 · Evals & RL environments (verifiers, reward design, difficulty calibration, lifecycle)", "source_id": "ae-168-specification-gaming-the-flip-side-of-ai-ingenui", "source_line": 265, "source_type": "blog", "status": "mirrored", "title": "Specification gaming: the flip side of AI ingenuity", "why_it_matters": "Canonical specification-gaming post (+the running examples list); origin story of why verifiers/reward functions get gamed, predating the LLM-RL wave."}
{"all_urls": ["https://www.latent.space/p/willccbb"], "content_sha256": "9bceac71c76cadf193cabce5e616a89a0b251898326485c3ea3a5c7e24ae0ee3", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/florian-brand-prime-intellect-llm-benchmarks-era-of-agents.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-169-multi-turn-rl-for-multi-hour-agents-with-will-br.raw.txt", "note_path": "notes/articles/florian-brand-prime-intellect-llm-benchmarks-era-of-agents.md", "primary_url": "https://www.latent.space/p/willccbb", "section": "7 · Evals & RL environments (verifiers, reward design, difficulty calibration, lifecycle)", "source_id": "ae-169-multi-turn-rl-for-multi-hour-agents-with-will-br", "source_line": 266, "source_type": "blog", "status": "mirrored", "title": "Multi-Turn RL for Multi-Hour Agents — with Will Brown (Prime Intellect)", "why_it_matters": "the practitioner voice behind the verifiers library already cited here. 🆕"}
{"all_urls": ["https://arxiv.org/abs/2509.21882"], "content_sha256": "a91583004d519fa2d2b3859ea0ab40e28161cf4e95a0cee64682c23e55054f7a", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/rlvr-hidden-costs-measurement-gaps.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-170-position-the-hidden-costs-and-measurement-gaps-o.raw.txt", "note_path": "notes/articles/rlvr-hidden-costs-measurement-gaps.md", "primary_url": "https://arxiv.org/abs/2509.21882", "section": "7 · Evals & RL environments (verifiers, reward design, difficulty calibration, lifecycle)", "source_id": "ae-170-position-the-hidden-costs-and-measurement-gaps-o", "source_line": 267, "source_type": "paper_or_pdf", "status": "mirrored", "title": "Position: The Hidden Costs and Measurement Gaps of RLVR", "why_it_matters": "the rigor counterweight to Lambert's RL-scaling optimism. 🆕"}
{"all_urls": ["https://arxiv.org/abs/2506.01937"], "content_sha256": "4d94b4cdada2a6b78fa35cc91abf450625023b26ad9bd0435a9c5af0d4f2c1c0", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-171-rewardbench-2-advancing-reward-model-evaluation.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2506.01937", "section": "7 · Evals & RL environments (verifiers, reward design, difficulty calibration, lifecycle)", "source_id": "ae-171-rewardbench-2-advancing-reward-model-evaluation", "source_line": 268, "source_type": "paper_or_pdf", "status": "mirrored", "title": "RewardBench 2: Advancing Reward Model Evaluation", "why_it_matters": "harder, less saturated, ICLR 2026; the current bar for evaluating the verifier you train against. 🆕"}
{"all_urls": ["https://rlhfbook.com/c/05-reward-models"], "content_sha256": "773504c8e4b847899a4281e7f0ea03c1ac4caf68a6661ff63b5d6e7d9c94c051", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/papers/scalable-agent-alignment-via-reward-modeling.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-172-reward-modeling-rlhf-book-ch-5.raw.txt", "note_path": "notes/papers/scalable-agent-alignment-via-reward-modeling.md", "primary_url": "https://rlhfbook.com/c/05-reward-models", "section": "7 · Evals & RL environments (verifiers, reward design, difficulty calibration, lifecycle)", "source_id": "ae-172-reward-modeling-rlhf-book-ch-5", "source_line": 269, "source_type": "docs_or_book", "status": "mirrored", "title": "Reward Modeling (RLHF Book, ch. 5)", "why_it_matters": "the standing explainer for the 'verifier you train against' framing this section uses. 🆕"}
{"all_urls": ["https://arxiv.org/abs/2506.06632"], "content_sha256": "3482bebffe22409a49d9cf94041f032893571985d1dc6503ceb1560b00445e11", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-173-curriculum-rl-from-easy-to-hard-tasks-improves-l.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2506.06632", "section": "7 · Evals & RL environments (verifiers, reward design, difficulty calibration, lifecycle)", "source_id": "ae-173-curriculum-rl-from-easy-to-hard-tasks-improves-l", "source_line": 270, "source_type": "paper_or_pdf", "status": "mirrored", "title": "Curriculum RL from Easy to Hard Tasks Improves LLM Reasoning (E2H Reasoner)", "why_it_matters": "directly fills the section's difficulty-calibration theme. 🆕"}
{"all_urls": ["https://arxiv.org/abs/2512.19682"], "content_sha256": "a9c9cf9ad1246b964ecc264d46b2756b00de4f55de78c954d664d3b69bb432db", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-174-genenv-difficulty-aligned-co-evolution-between-l.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2512.19682", "section": "7 · Evals & RL environments (verifiers, reward design, difficulty calibration, lifecycle)", "source_id": "ae-174-genenv-difficulty-aligned-co-evolution-between-l", "source_line": 271, "source_type": "paper_or_pdf", "status": "mirrored", "title": "GenEnv: Difficulty-Aligned Co-Evolution Between LLM Agents and Environment Simulators", "why_it_matters": "recent take on auto-calibrating env difficulty to the agent. 🆕"}
{"all_urls": ["https://eugeneyan.com/writing/llm-evaluators/"], "content_sha256": "7e856220e7e37db84c43ae6f109f6626323f094ef95f883cbfc610d9454b2511", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-175-evaluating-the-effectiveness-of-llm-evaluators.raw.txt", "note_path": null, "primary_url": "https://eugeneyan.com/writing/llm-evaluators/", "section": "8 · LLM-as-judge & verifiers (alignment, biases, verifiable vs judgeable)", "source_id": "ae-175-evaluating-the-effectiveness-of-llm-evaluators", "source_line": 277, "source_type": "web_article", "status": "mirrored", "title": "Evaluating the Effectiveness of LLM-Evaluators", "why_it_matters": "Position/verbosity/self-enhancement bias; direct vs pairwise; prefer binary + classification metrics."}
{"all_urls": ["https://hamel.dev/blog/posts/llm-judge/"], "content_sha256": "f0f10e4fc7908bfee6e4452e1f64d6b711c9c962db00fbe51d7b36f7a3c43f8a", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-176-creating-an-llm-as-a-judge-that-drives-business.raw.txt", "note_path": null, "primary_url": "https://hamel.dev/blog/posts/llm-judge/", "section": "8 · LLM-as-judge & verifiers (alignment, biases, verifiable vs judgeable)", "source_id": "ae-176-creating-an-llm-as-a-judge-that-drives-business", "source_line": 278, "source_type": "blog", "status": "mirrored", "title": "Creating an LLM-as-a-Judge That Drives Business Results", "why_it_matters": "Critique-shadowing; validate against ONE benevolent-dictator expert; precision/recall over raw agreement."}
{"all_urls": ["https://arxiv.org/abs/2404.12272", "https://people.eecs.berkeley.edu/~bjoern/papers/shankar-validators-uist2024.pdf"], "content_sha256": "2a98cac4c82e1b225ab4c6468262c9245923882e276c6f518b0142cc02ecae5a", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-177-who-validates-the-validators-evalgen.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2404.12272", "section": "8 · LLM-as-judge & verifiers (alignment, biases, verifiable vs judgeable)", "source_id": "ae-177-who-validates-the-validators-evalgen", "source_line": 279, "source_type": "paper_or_pdf", "status": "mirrored", "title": "Who Validates the Validators? (EvalGen)", "why_it_matters": "Criteria drift; the coverage-vs-false-failure judge-alignment loop."}
{"all_urls": ["https://hamel.dev/blog/posts/evals-faq/"], "content_sha256": "b5d5398f91d39542cc52d6c11bc38dfb86e4da2d06fcb73e69add860602d8e02", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-178-llm-evals-faq.raw.txt", "note_path": null, "primary_url": "https://hamel.dev/blog/posts/evals-faq/", "section": "8 · LLM-as-judge & verifiers (alignment, biases, verifiable vs judgeable)", "source_id": "ae-178-llm-evals-faq", "source_line": 280, "source_type": "blog", "status": "mirrored", "title": "LLM Evals FAQ", "why_it_matters": "Binary over Likert; review ≥100 traces; the first-failure transition matrix for agents."}
{"all_urls": ["https://leehanchung.github.io/blogs/2024/08/11/llm-as-a-judge/"], "content_sha256": "445735aa534efe43630d6513c3f67cf6a90d4369aa282e4b047624084a8575ca", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-179-llm-as-a-judge-rethinking-model-based-evaluation.raw.txt", "note_path": null, "primary_url": "https://leehanchung.github.io/blogs/2024/08/11/llm-as-a-judge/", "section": "8 · LLM-as-judge & verifiers (alignment, biases, verifiable vs judgeable)", "source_id": "ae-179-llm-as-a-judge-rethinking-model-based-evaluation", "source_line": 281, "source_type": "blog", "status": "mirrored", "title": "LLM-as-a-Judge: Rethinking Model-Based Evaluations", "why_it_matters": "Avoid [0,1] continuous scales; manage judges like junior annotators."}
{"all_urls": ["https://arxiv.org/abs/2306.05685"], "content_sha256": "21da0bd112b6f623aa04437301b85ad4e581278b0b24e012932ea6abb2452808", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-pod-gd-gonzalez-chatbot-arena.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-180-judging-llm-as-a-judge-with-mt-bench-and-chatbot.raw.txt", "note_path": "notes/talks/talk-pod-gd-gonzalez-chatbot-arena.md", "primary_url": "https://arxiv.org/abs/2306.05685", "section": "8 · LLM-as-judge & verifiers (alignment, biases, verifiable vs judgeable)", "source_id": "ae-180-judging-llm-as-a-judge-with-mt-bench-and-chatbot", "source_line": 282, "source_type": "paper_or_pdf", "status": "mirrored", "title": "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena", "why_it_matters": "*which the authors themselves hedge* (\"cannot determine\"); GPT-3.5 doesn't self-favor."}
{"all_urls": ["https://arxiv.org/abs/2406.18403"], "content_sha256": "c20a78d7515ed808d5371ce1b6e4dfd1756897bfc3e283375a6d42e523223f4b", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-yan-llms-as-judges.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-181-llms-instead-of-human-judges-a-large-scale-study.raw.txt", "note_path": "notes/talks/talk-yan-llms-as-judges.md", "primary_url": "https://arxiv.org/abs/2406.18403", "section": "8 · LLM-as-judge & verifiers (alignment, biases, verifiable vs judgeable)", "source_id": "ae-181-llms-instead-of-human-judges-a-large-scale-study", "source_line": 283, "source_type": "paper_or_pdf", "status": "mirrored", "title": "LLMs Instead of Human Judges? A Large-Scale Study", "why_it_matters": "Substantial variance across models/datasets; validate judges against humans first."}
{"all_urls": ["https://eugeneyan.com/writing/aligneval/"], "content_sha256": "2b116b69d03ce1fd0148941be90ed71d5358008ad5fcb534defe066f23bb67be", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-182-aligneval.raw.txt", "note_path": null, "primary_url": "https://eugeneyan.com/writing/aligneval/", "section": "8 · LLM-as-judge & verifiers (alignment, biases, verifiable vs judgeable)", "source_id": "ae-182-aligneval", "source_line": 284, "source_type": "web_article", "status": "mirrored", "title": "AlignEval", "why_it_matters": "\"Align AI to human. Calibrate human to AI. Repeat.\" Work backward from the data."}
{"all_urls": ["https://eugeneyan.com/writing/product-evals/"], "content_sha256": "e4fc4af54f2781c8e534609e264181d69c2e648bf4b39e2e31b909afba569816", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-183-product-evals-in-three-simple-steps.raw.txt", "note_path": null, "primary_url": "https://eugeneyan.com/writing/product-evals/", "section": "8 · LLM-as-judge & verifiers (alignment, biases, verifiable vs judgeable)", "source_id": "ae-183-product-evals-in-three-simple-steps", "source_line": 285, "source_type": "web_article", "status": "mirrored", "title": "Product Evals in Three Simple Steps", "why_it_matters": "The \"God Evaluator\" anti-pattern; the benchmark is human performance, not perfection."}
{"all_urls": ["https://leehanchung.github.io/blogs/2025/03/03/cohen-kappa/"], "content_sha256": "a0b67556e7644e0c207e6aa53941c5760325dba991b55f496ce4276f2000d78b", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-184-statistics-for-ai-ml-part-3-cohen-s-kappa.raw.txt", "note_path": null, "primary_url": "https://leehanchung.github.io/blogs/2025/03/03/cohen-kappa/", "section": "8 · LLM-as-judge & verifiers (alignment, biases, verifiable vs judgeable)", "source_id": "ae-184-statistics-for-ai-ml-part-3-cohen-s-kappa", "source_line": 286, "source_type": "blog", "status": "mirrored", "title": "Statistics for AI/ML, Part 3 — Cohen's Kappa", "why_it_matters": "Chance-adjusted inter-annotator agreement (the gate before holding out)."}
{"all_urls": ["https://www.sh-reya.com/blog/ai-engineering-flywheel/"], "content_sha256": "b30c2c6170dab30e512d96c87c8a12ce77a12031a7aa6ba01b87aa5c09ba3ca9", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-185-data-flywheels-for-llm-applications.raw.txt", "note_path": null, "primary_url": "https://www.sh-reya.com/blog/ai-engineering-flywheel/", "section": "8 · LLM-as-judge & verifiers (alignment, biases, verifiable vs judgeable)", "source_id": "ae-185-data-flywheels-for-llm-applications", "source_line": 287, "source_type": "blog", "status": "mirrored", "title": "Data Flywheels for LLM Applications", "why_it_matters": "Binary metrics, the \"GPT smell,\" error analysis as the core activity."}
{"all_urls": ["https://arxiv.org/html/2401.03038v1", "https://arxiv.org/abs/2410.12189"], "content_sha256": "589e71a79dec3221328d704e3dbb6021ce1d9a69169682d5db7175f8b3bf2830", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-186-spade.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/html/2401.03038v1", "section": "8 · LLM-as-judge & verifiers (alignment, biases, verifiable vs judgeable)", "source_id": "ae-186-spade", "source_line": 288, "source_type": "paper_or_pdf", "status": "mirrored", "title": "SPADE", "why_it_matters": "Data-quality assertions / agentic query rewriting for LLM pipelines."}
{"all_urls": ["https://arxiv.org/abs/2404.13076"], "content_sha256": "0a377fb71e55eb3d05796fe88e215f646596fa91a44f7b95dfcad73ce40df706", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-187-llm-evaluators-recognize-and-favor-their-own-gen.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2404.13076", "section": "8 · LLM-as-judge & verifiers (alignment, biases, verifiable vs judgeable)", "source_id": "ae-187-llm-evaluators-recognize-and-favor-their-own-gen", "source_line": 290, "source_type": "paper_or_pdf", "status": "mirrored", "title": "LLM Evaluators Recognize and Favor Their Own Generations", "why_it_matters": "The canonical causal study of self-preference bias: shows GPT-4/Llama-2 can recognize their own outputs and that self-recognition correlates linearly with self-favoring. This is the primary source behind 'self-enhancement bias' that the section's blogs only allude to."}
{"all_urls": ["https://arxiv.org/abs/2303.16634"], "content_sha256": "f3719d1cfc352630e15cacdc50b2e9debbed572c8a2ac5ff5180531f30766b97", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/han-lee-evaluation-and-alignment-seminal-papers.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-188-g-eval-nlg-evaluation-using-gpt-4-with-better-hu.raw.txt", "note_path": "notes/articles/han-lee-evaluation-and-alignment-seminal-papers.md", "primary_url": "https://arxiv.org/abs/2303.16634", "section": "8 · LLM-as-judge & verifiers (alignment, biases, verifiable vs judgeable)", "source_id": "ae-188-g-eval-nlg-evaluation-using-gpt-4-with-better-hu", "source_line": 291, "source_type": "paper_or_pdf", "status": "mirrored", "title": "G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment", "why_it_matters": "The foundational reference-free LLM-judge method (CoT + form-filling scoring). Defines the direct-scoring paradigm the section critiques; a curated judge section is incomplete without the paper that started it."}
{"all_urls": ["https://arxiv.org/abs/2411.15594"], "content_sha256": "846daf987c9f7c91d10f5ccc06d87b1ac16ed585198e5d9e0e2cde520b1604bd", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-189-a-survey-on-llm-as-a-judge.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2411.15594", "section": "8 · LLM-as-judge & verifiers (alignment, biases, verifiable vs judgeable)", "source_id": "ae-189-a-survey-on-llm-as-a-judge", "source_line": 292, "source_type": "paper_or_pdf", "status": "mirrored", "title": "A Survey on LLM-as-a-Judge", "why_it_matters": "The most-cited survey organizing the LLM-judge space (bias taxonomy, reliability methods, agreement metrics). Serves as the one-stop map/bibliography the section currently lacks."}
{"all_urls": ["https://arxiv.org/abs/2507.08794"], "content_sha256": "091e765a3071d3a192f7d6efe796266eb7b78c6f4f1414005dec9ca70a07504d", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-190-one-token-to-fool-llm-as-a-judge.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2507.08794", "section": "8 · LLM-as-judge & verifiers (alignment, biases, verifiable vs judgeable)", "source_id": "ae-190-one-token-to-fool-llm-as-a-judge", "source_line": 293, "source_type": "paper_or_pdf", "status": "mirrored", "title": "One Token to Fool LLM-as-a-Judge", "why_it_matters": "Shows 'master-key' tokens (a colon, 'Solution:') trigger false-positive rewards up to 80% even on GPT-o1/Claude-4 judges, plus a robust Master-RM fix. Core evidence on judge/verifier reward-hacking fragility. 🆕"}
{"all_urls": ["https://hazyresearch.stanford.edu/blog/2025-06-18-weaver"], "content_sha256": "840677ec0216a69fce6d958586f020226a5debde36636e04dff5de2769607616", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/papers/bertscore-evaluating-text-generation-with-bert.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-191-weaver-closing-the-generation-verification-gap-w.raw.txt", "note_path": "notes/papers/bertscore-evaluating-text-generation-with-bert.md", "primary_url": "https://hazyresearch.stanford.edu/blog/2025-06-18-weaver", "section": "8 · LLM-as-judge & verifiers (alignment, biases, verifiable vs judgeable)", "source_id": "ae-191-weaver-closing-the-generation-verification-gap-w", "source_line": 294, "source_type": "blog", "status": "mirrored", "title": "Weaver: Closing the Generation-Verification Gap with Weak Verifiers", "why_it_matters": "Directly operationalizes 'verifiable vs judgeable': aggregates many weak judges/reward models (unlabeled) to shrink the generator-verifier gap, reaching o3-mini accuracy from Llama-3.3-70B. Paper: arxiv.org/abs/2506.18203. 🆕"}
{"all_urls": ["https://arxiv.org/abs/2410.10934"], "content_sha256": "a4ee22ac1aea42a45d3528cae187b90baf690619d29af3f63fc324c90b364a44", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/opik-evaluate-agent-trajectory.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-192-agent-as-a-judge-evaluate-agents-with-agents.raw.txt", "note_path": "notes/articles/opik-evaluate-agent-trajectory.md", "primary_url": "https://arxiv.org/abs/2410.10934", "section": "8 · LLM-as-judge & verifiers (alignment, biases, verifiable vs judgeable)", "source_id": "ae-192-agent-as-a-judge-evaluate-agents-with-agents", "source_line": 295, "source_type": "paper_or_pdf", "status": "mirrored", "title": "Agent-as-a-Judge: Evaluate Agents with Agents", "why_it_matters": "with the DevAI benchmark. The agent-specific evaluation case this agent-evals library specifically needs."}
{"all_urls": ["https://arxiv.org/abs/2507.09884"], "content_sha256": "432199f64f3f479059f0abb8086549652d7493b74d3650be2bcd7765268ad80b", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-193-verifybench-a-systematic-benchmark-for-evaluatin.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2507.09884", "section": "8 · LLM-as-judge & verifiers (alignment, biases, verifiable vs judgeable)", "source_id": "ae-193-verifybench-a-systematic-benchmark-for-evaluatin", "source_line": 296, "source_type": "paper_or_pdf", "status": "mirrored", "title": "VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across Domains", "why_it_matters": "Cross-domain benchmark exposing verifier precision/recall trade-offs (specialized verifiers high-accuracy but low-recall; general models inclusive but unstable). Quantifies how trustworthy a verifier actually is for RLVR. 🆕"}
{"all_urls": ["https://www.databricks.com/blog/pilot-production-custom-judges"], "content_sha256": "e41f65bac3a11644c1b819112201e0cffff08583f0590d28bbfc39458b8cfc35", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/godaddy-calibrating-llm-judge-scores.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-194-enhancing-llm-as-a-judge-with-grading-notes-from.raw.txt", "note_path": "notes/articles/godaddy-calibrating-llm-judge-scores.md", "primary_url": "https://www.databricks.com/blog/pilot-production-custom-judges", "section": "8 · LLM-as-judge & verifiers (alignment, biases, verifiable vs judgeable)", "source_id": "ae-194-enhancing-llm-as-a-judge-with-grading-notes-from", "source_line": 297, "source_type": "blog", "status": "mirrored", "title": "Enhancing LLM-as-a-Judge with Grading Notes / From Pilot to Production with Custom Judges", "why_it_matters": "a production-side complement to the Hamel/Shankar academic alignment loop. 🆕"}
{"all_urls": ["https://arxiv.org/abs/2410.02736"], "content_sha256": "777ce179d87154d6e0c810bc88dac2968b1a65706d8e01869a433b15d1070595", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-195-justice-or-prejudice-quantifying-biases-in-llm-a.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2410.02736", "section": "8 · LLM-as-judge & verifiers (alignment, biases, verifiable vs judgeable)", "source_id": "ae-195-justice-or-prejudice-quantifying-biases-in-llm-a", "source_line": 298, "source_type": "paper_or_pdf", "status": "mirrored", "title": "Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge (CALM framework)", "why_it_matters": "broadens the section's bias coverage well beyond position/verbosity/self-enhancement."}
{"all_urls": ["https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents"], "content_sha256": "5bbd04d4cbfd341cb3294b7da1bd7bbb56b5a0fefb3c1db81b81158de6066802", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/vanishing-gradients-agents-evals.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-196-demystifying-evals-for-ai-agents.raw.txt", "note_path": "notes/articles/vanishing-gradients-agents-evals.md", "primary_url": "https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-196-demystifying-evals-for-ai-agents", "source_line": 304, "source_type": "web_article", "status": "mirrored", "title": "Demystifying Evals for AI Agents", "why_it_matters": "Grade the final env state (flight-booking via SQL); outcome vs trajectory; isolation; pass@k vs pass^k."}
{"all_urls": ["https://arxiv.org/abs/2406.12045", "https://github.com/sierra-research/tau-bench"], "content_sha256": "20c2acc833d099fde9e8a4d9b7788e364a37d92f4ae727e332693fc5bc794034", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-197-bench-bench.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2406.12045", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-197-bench-bench", "source_line": 305, "source_type": "paper_or_pdf", "status": "mirrored", "title": "τ-bench / τ²-bench", "why_it_matters": "DB-state-diff grading; user simulation; pass^k; empty-result as explicit fail."}
{"all_urls": ["https://sierra.ai/blog/benchmarking-ai-agents"], "content_sha256": "8fe932f38e8ac7d3297c478d7aadec32ef89e5dfc8589551865226e608b80e18", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-198-benchmarking-ai-agents.raw.txt", "note_path": null, "primary_url": "https://sierra.ai/blog/benchmarking-ai-agents", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-198-benchmarking-ai-agents", "source_line": 306, "source_type": "blog", "status": "mirrored", "title": "Benchmarking AI Agents", "why_it_matters": "The motivation behind τ-bench."}
{"all_urls": ["https://arxiv.org/abs/2311.12983"], "content_sha256": "d54599ae4bf96dac9b35e75ee9fc7693b9f73e0ac7d79610eb2ce0c7f69c4c0c", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-199-gaia-a-benchmark-for-general-ai-assistants.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2311.12983", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-199-gaia-a-benchmark-for-general-ai-assistants", "source_line": 307, "source_type": "paper_or_pdf", "status": "mirrored", "title": "GAIA: A Benchmark for General AI Assistants", "why_it_matters": "Real assistant tasks; difficulty by human task-length."}
{"all_urls": ["https://eugeneyan.com/writing/cybersecurity-evals/"], "content_sha256": "e685d714fbb9e4b4e0ff95869ff71d1c2e29644099e9a3f3961dd74e15c67492", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-200-patterns-for-building-cybersecurity-evals.raw.txt", "note_path": null, "primary_url": "https://eugeneyan.com/writing/cybersecurity-evals/", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-200-patterns-for-building-cybersecurity-evals", "source_line": 308, "source_type": "web_article", "status": "mirrored", "title": "Patterns for Building Cybersecurity Evals", "why_it_matters": "The four-primitive agentic-eval template (sandbox, difficulty inputs, tools, deterministic grader); outcome grading + partial-credit ladders + transcript audits. (also T10)"}
{"all_urls": ["https://leehanchung.github.io/blogs/2025/09/08/pass-at-k/"], "content_sha256": "a1e2b09df956c3a8a6ffa148beae2f5e100a905ba7607ba7efa4bce66c804d7c", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-201-statistics-for-ai-ml-part-4-pass-k-and-unbiased.raw.txt", "note_path": null, "primary_url": "https://leehanchung.github.io/blogs/2025/09/08/pass-at-k/", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-201-statistics-for-ai-ml-part-4-pass-k-and-unbiased", "source_line": 309, "source_type": "blog", "status": "mirrored", "title": "Statistics for AI/ML, Part 4 — pass@k and Unbiased Estimator", "why_it_matters": "Demystifies the metric everyone misuses."}
{"all_urls": ["https://leehanchung.github.io/blogs/2024/05/22/first-principles-eval/"], "content_sha256": "ef1e2ff8f19addeed18163cb3ef6120736d892a37691875ff65774ddd03ddfa3", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-202-first-principles-eval.raw.txt", "note_path": null, "primary_url": "https://leehanchung.github.io/blogs/2024/05/22/first-principles-eval/", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-202-first-principles-eval", "source_line": 310, "source_type": "blog", "status": "mirrored", "title": "First-Principles Eval", "why_it_matters": "<https://leehanchung.github.io/blogs/2024/05/22/first-principles-eval/> · *blog*."}
{"all_urls": ["https://github.com/SWE-bench/SWE-bench/blob/main/swebench/harness/grading.py", "https://swe-agent.com/0.7/background/aci/"], "content_sha256": "c573425751346905af1839c4788dee75ade837b380a7c20351abe09e51f8d373", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-203-swe-bench-grading-harness.raw.txt", "note_path": null, "primary_url": "https://github.com/SWE-bench/SWE-bench/blob/main/swebench/harness/grading.py", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-203-swe-bench-grading-harness", "source_line": 311, "source_type": "repository_or_docs", "status": "mirrored", "title": "SWE-bench grading harness", "why_it_matters": "FAIL_TO_PASS / PASS_TO_PASS as a verifiable reward. (SWE-agent ACI: <https://swe-agent.com/0.7/background/aci/>)"}
{"all_urls": ["https://github.com/openai/human-eval/blob/master/human_eval/evaluation.py"], "content_sha256": "934c5577467c6e6b8373b37267ed050e4f2f5f09177f28f24b38abdabcec5d7e", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/dont-pass-at-k-bayesian-llm-eval.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-204-human-eval-pass-k-estimator.raw.txt", "note_path": "notes/articles/dont-pass-at-k-bayesian-llm-eval.md", "primary_url": "https://github.com/openai/human-eval/blob/master/human_eval/evaluation.py", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-204-human-eval-pass-k-estimator", "source_line": 312, "source_type": "repository_or_docs", "status": "mirrored", "title": "human-eval (pass@k estimator)", "why_it_matters": "<https://github.com/openai/human-eval/blob/master/human_eval/evaluation.py> · *tool/repo*."}
{"all_urls": ["https://arxiv.org/abs/2307.13854"], "content_sha256": "3e0932bb4b0bc38463632911d4186e15b09a4fc2b997457225912db251e4ced5", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-205-webarena-a-realistic-web-environment-for-buildin.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2307.13854", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-205-webarena-a-realistic-web-environment-for-buildin", "source_line": 315, "source_type": "paper_or_pdf", "status": "mirrored", "title": "WebArena: A Realistic Web Environment for Building Autonomous Agents", "why_it_matters": "now URL-verified."}
{"all_urls": ["https://arxiv.org/abs/2404.07972"], "content_sha256": "e77b1ae61fe55ed56204c0a4414344de02575d81d9cd7ec08d009ae4296ae7c6", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-206-osworld-benchmarking-multimodal-agents-for-open.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2404.07972", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-206-osworld-benchmarking-multimodal-agents-for-open", "source_line": 316, "source_type": "paper_or_pdf", "status": "mirrored", "title": "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments", "why_it_matters": "now verified."}
{"all_urls": ["https://www.tbench.ai/"], "content_sha256": "8e48e19a0568541bec63b2dcd5b37354cc85e92514550522f833f924126c167e", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-207-terminal-bench-benchmarking-agents-on-hard-reali.raw.txt", "note_path": null, "primary_url": "https://www.tbench.ai/", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-207-terminal-bench-benchmarking-agents-on-hard-reali", "source_line": 317, "source_type": "web_article", "status": "mirrored", "title": "Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command-Line Interfaces", "why_it_matters": "verified (arxiv: arxiv.org/abs/2601.11868). 🆕"}
{"all_urls": ["https://arxiv.org/abs/2408.08926"], "content_sha256": "237b1a798e38be5ea69bd9cfa3285a0eff54b1d37b8f78c7c23ef5c2326c9365", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/papers/evaluating-large-language-models-trained-on-code.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-208-cybench-a-framework-for-evaluating-cybersecurity.raw.txt", "note_path": "notes/papers/evaluating-large-language-models-trained-on-code.md", "primary_url": "https://arxiv.org/abs/2408.08926", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-208-cybench-a-framework-for-evaluating-cybersecurity", "source_line": 318, "source_type": "paper_or_pdf", "status": "mirrored", "title": "Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risk of Language Models", "why_it_matters": "now verified."}
{"all_urls": ["https://arxiv.org/abs/2504.08942"], "content_sha256": "c51c2880d33febc75d8ab9f8a042977449e2eb50c79eb5fde0d59789aa64105d", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/agentrewardbench-evaluating-automatic-evaluations-web-agent-.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-209-agentrewardbench-evaluating-automatic-evaluation.raw.txt", "note_path": "notes/articles/agentrewardbench-evaluating-automatic-evaluations-web-agent-.md", "primary_url": "https://arxiv.org/abs/2504.08942", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-209-agentrewardbench-evaluating-automatic-evaluation", "source_line": 319, "source_type": "paper_or_pdf", "status": "mirrored", "title": "AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories", "why_it_matters": "First benchmark of LLM-judges-of-trajectories: 1302 expert-reviewed web-agent runs; shows rule-based graders reject many valid trajectories (under-reporting success). Core to the 'trajectory evaluation' theme the section currently lacks. 🆕"}
{"all_urls": ["https://arxiv.org/abs/2503.13657"], "content_sha256": "26ec58b035635eaa8f906ad94b9710219eadae3d454ea2cc7d1350ce41252905", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/mast-why-multi-agent-llm-systems-fail.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-210-why-do-multi-agent-llm-systems-fail-mast-taxonom.raw.txt", "note_path": "notes/articles/mast-why-multi-agent-llm-systems-fail.md", "primary_url": "https://arxiv.org/abs/2503.13657", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-210-why-do-multi-agent-llm-systems-fail-mast-taxonom", "source_line": 320, "source_type": "paper_or_pdf", "status": "mirrored", "title": "Why Do Multi-Agent LLM Systems Fail? (MAST taxonomy)", "why_it_matters": "directly fills the 'multi-agent' gap. 🆕"}
{"all_urls": ["https://aclanthology.org/2024.acl-long.850/"], "content_sha256": "54c3267f4277f9d7ad04dcdf043c5213eb8ef274e7bb558df3614c3fccaf840d", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-211-appworld-a-controllable-world-of-apps-and-people.raw.txt", "note_path": null, "primary_url": "https://aclanthology.org/2024.acl-long.850/", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-211-appworld-a-controllable-world-of-apps-and-people", "source_line": 321, "source_type": "web_article", "status": "mirrored", "title": "AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents", "why_it_matters": "gold-standard world-state grading for tool-use agents."}
{"all_urls": ["https://openai.com/index/browsecomp/"], "content_sha256": "44aefab4e553fd8f2928eef6028f4c6cb2c282f5dd06e82d73b0d8515d9827e3", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-212-browsecomp-a-simple-yet-challenging-benchmark-fo.raw.txt", "note_path": null, "primary_url": "https://openai.com/index/browsecomp/", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-212-browsecomp-a-simple-yet-challenging-benchmark-fo", "source_line": 322, "source_type": "web_article", "status": "mirrored", "title": "BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents", "why_it_matters": "1,266 'inverted' hard-to-find/easy-to-verify questions for deep-research browsing agents; short verifiable answers make grading deterministic. Released 2025, now standard for browsing-agent eval. (paper: arxiv.org/abs/2504.12516) 🆕"}
{"all_urls": ["https://arxiv.org/abs/2503.09089"], "content_sha256": "9a8a7def2c02b3c2a3f393ef9cc3b42aefa846b8ccb844a8d9c753fd3bf43c4d", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-213-locagent-graph-guided-llm-agents-for-code-locali.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2503.09089", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-213-locagent-graph-guided-llm-agents-for-code-locali", "source_line": 323, "source_type": "paper_or_pdf", "status": "mirrored", "title": "LocAgent: Graph-Guided LLM Agents for Code Localization", "why_it_matters": "directly fills the 'localization' theme named in the section title but currently unlisted. 🆕"}
{"all_urls": ["https://arxiv.org/abs/2401.13919"], "content_sha256": "c50f32d9b606e0d9b4ea722bc9ce6474642ae43db52824e736c77a09c4b1ca46", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/papers/evaluating-large-language-models-trained-on-code.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-214-webvoyager-building-an-end-to-end-web-agent-with.raw.txt", "note_path": "notes/papers/evaluating-large-language-models-trained-on-code.md", "primary_url": "https://arxiv.org/abs/2401.13919", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-214-webvoyager-building-an-end-to-end-web-agent-with", "source_line": 324, "source_type": "paper_or_pdf", "status": "mirrored", "title": "WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models", "why_it_matters": "an early, widely-cited example of multimodal-LLM-as-judge for live-web agent trajectories."}
{"all_urls": ["https://github.com/benchflow-ai/skillsbench"], "content_sha256": "67486d0bf73dbeeb321819a67a8da466a2a68302f23a525dc84564a1e4e61912", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-215-skillsbench.raw.txt", "note_path": null, "primary_url": "https://github.com/benchflow-ai/skillsbench", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-215-skillsbench", "source_line": 325, "source_type": "repository_or_docs", "status": "mirrored", "title": "SkillsBench", "why_it_matters": "makes skill-acquisition/skill-use a measurable axis (the \"Agent Skills\" frontier). ~1.4k★."}
{"all_urls": ["https://github.com/benchflow-ai/ClawsBench"], "content_sha256": "c430b9b5d53cb0517d745980ee5f431649a54d9098be20b231323971261e60bb", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-216-clawsbench.raw.txt", "note_path": null, "primary_url": "https://github.com/benchflow-ai/ClawsBench", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-216-clawsbench", "source_line": 326, "source_type": "repository_or_docs", "status": "mirrored", "title": "ClawsBench", "why_it_matters": "🆕 BenchFlow's agent benchmark (results/data repo; full release in progress)."}
{"all_urls": ["https://openai.com/index/introducing-swe-bench-verified/"], "content_sha256": "1686806c33aee99cf1ef603796a78398a52122df1461a7b37ac55462e1c66d3a", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-217-swe-bench-verified.raw.txt", "note_path": null, "primary_url": "https://openai.com/index/introducing-swe-bench-verified/", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-217-swe-bench-verified", "source_line": 328, "source_type": "web_article", "status": "mirrored", "title": "SWE-bench Verified", "why_it_matters": "500 human-validated SWE-bench instances graded by hidden FAIL_TO_PASS unit tests; the de facto standard for real-issue resolution and the headline coding-agent number labs report 🆕"}
{"all_urls": ["https://arxiv.org/abs/2410.03859"], "content_sha256": "a0f30345dcc66858d3aaabd7843f9307c0357aaf89c0bc740588017b793c6544", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-218-swe-bench-multimodal.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2410.03859", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-218-swe-bench-multimodal", "source_line": 329, "source_type": "paper_or_pdf", "status": "mirrored", "title": "SWE-bench Multimodal", "why_it_matters": "619 visual JS/front-end issues from 17 user-facing repos, test-verified; probes whether SWE agents generalize beyond Python/text to visual software domains"}
{"all_urls": ["https://arxiv.org/abs/2509.16941"], "content_sha256": "4a51b2c979fbd6c4565c870f98ad7110ab5f5930b3afcb2df2b798c4af02e69e", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-219-swe-bench-pro.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2509.16941", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-219-swe-bench-pro", "source_line": 330, "source_type": "paper_or_pdf", "status": "mirrored", "title": "SWE-bench Pro", "why_it_matters": "1,865 long-horizon, multi-file tasks across public GPL + held-out + commercial startup repos, test-graded; contamination-resistant and hard (frontier <45% pass@1) 🆕"}
{"all_urls": ["https://arxiv.org/abs/2502.12115"], "content_sha256": "31df4daa05d3fce12d5cdbd76889291ae7e10869f25c058e0bbd2a97ba3bfb2d", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-220-swe-lancer.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2502.12115", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-220-swe-lancer", "source_line": 331, "source_type": "paper_or_pdf", "status": "mirrored", "title": "SWE-Lancer", "why_it_matters": "1,400+ real Upwork freelance tasks worth $1M, graded by triple-verified end-to-end Playwright tests plus manager-decision tasks; ties capability to economic value 🆕"}
{"all_urls": ["https://arxiv.org/abs/2412.21139"], "content_sha256": "59c30f736957e4236bdddc7b5fa0b83cc173699e12a374629fe0cf71b0fd7f1d", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-221-swe-gym.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2412.21139", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-221-swe-gym", "source_line": 332, "source_type": "paper_or_pdf", "status": "mirrored", "title": "SWE-Gym", "why_it_matters": "2,438 executable Python SWE tasks with pre-installed deps + test verification; the first real training/eval gym for SWE agents and verifiers, ICML 2025 🆕"}
{"all_urls": ["https://arxiv.org/abs/2504.02605"], "content_sha256": "a0359689007891f8468eb7de76c85edcc7642c1fac8ce14b363dbd50a1c2c6cf", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-222-multi-swe-bench.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2504.02605", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-222-multi-swe-bench", "source_line": 333, "source_type": "paper_or_pdf", "status": "mirrored", "title": "Multi-SWE-bench", "why_it_matters": "1,632 expert-annotated issue-resolution tasks across Java, TS, JS, Go, Rust, C, C++, test-graded; the leading multilingual SWE-bench extension, NeurIPS 2025 D&B 🆕"}
{"all_urls": ["https://arxiv.org/abs/2505.20411"], "content_sha256": "ba95da25e8f0fb10d72db26134b067e29b07092cf18ce172e432083860437a29", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-223-swe-rebench.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2505.20411", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-223-swe-rebench", "source_line": 334, "source_type": "paper_or_pdf", "status": "mirrored", "title": "SWE-rebench", "why_it_matters": "Automated pipeline yielding 21k+ executable Python tasks with continuously refreshed, decontaminated eval splits; quantifies how much SWE-bench Verified scores are inflated by contamination, NeurIPS 2025 D&B 🆕"}
{"all_urls": ["https://arxiv.org/abs/2411.15114"], "content_sha256": "d203ddc12bf2519b04154217f4f558e5021e9977e3b54bf047bae4e3e374c9ea", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-224-re-bench.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2411.15114", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-224-re-bench", "source_line": 335, "source_type": "paper_or_pdf", "status": "mirrored", "title": "RE-Bench", "why_it_matters": "7 open-ended ML research-engineering environments (e.g. GPU-kernel optimization, scaling laws) scored against 71 human-expert 8-hour attempts; the reference AI-R&D-uplift eval, ICML 2025"}
{"all_urls": ["https://arxiv.org/abs/2410.07095", "https://github.com/openai/mle-bench"], "content_sha256": "006f3c4b166fbf8d4b7a6a4cf5f02a0284d67af3d6d3f44f6b6a105c9dc7d2ac", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-225-mle-bench.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2410.07095", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-225-mle-bench", "source_line": 336, "source_type": "paper_or_pdf", "status": "mirrored", "title": "MLE-bench", "why_it_matters": "75 Kaggle ML-engineering competitions graded against real human leaderboards (medal thresholds) in 24h Docker runs; standard ML-engineering-agent eval, ICLR 2025. 🆕"}
{"all_urls": ["https://arxiv.org/abs/2504.01848"], "content_sha256": "3a4ef4172eb5fc3328c8d6f93c2075177d3518563dcfed2c4745c3bf119b7de3", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-226-paperbench.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2504.01848", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-226-paperbench", "source_line": 337, "source_type": "paper_or_pdf", "status": "mirrored", "title": "PaperBench", "why_it_matters": "Replicate 20 ICML 2024 papers from scratch, graded by 8,316 author-co-developed rubric leaves via a validated LLM judge; rigorous research-replication agent eval, ICML 2025 🆕"}
{"all_urls": ["https://www.kaggle.com/competitions/konwinski-prize"], "content_sha256": "c950ee0f8f2749e1d218f0419e484058a090d5d8fd485336ec85e75764b0417f", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-227-konwinski-prize-k-prize.raw.txt", "note_path": null, "primary_url": "https://www.kaggle.com/competitions/konwinski-prize", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-227-konwinski-prize-k-prize", "source_line": 338, "source_type": "web_article", "status": "mirrored", "title": "Konwinski Prize (K Prize)", "why_it_matters": "$1M Kaggle forecasting-format contest on GitHub bugs filed after submission close, fully contamination-free, test-graded; round-1 top score only 7.5% exposed real-world difficulty 🆕"}
{"all_urls": ["https://arxiv.org/abs/2506.21506"], "content_sha256": "70e9d488d3300e02f0b141eed842da22a26bf58ba5584deab0bdd115c0da450b", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/agentrewardbench-evaluating-automatic-evaluations-web-agent-.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-228-mind2web-2-evaluating-agentic-search-with-agent.raw.txt", "note_path": "notes/articles/agentrewardbench-evaluating-automatic-evaluations-web-agent-.md", "primary_url": "https://arxiv.org/abs/2506.21506", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-228-mind2web-2-evaluating-agentic-search-with-agent", "source_line": 339, "source_type": "paper_or_pdf", "status": "mirrored", "title": "Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge", "why_it_matters": "a serious answer to the Deep Research evaluation gap. 🆕"}
{"all_urls": ["https://arxiv.org/abs/2504.01382"], "content_sha256": "101a54f3aec22cd1cc7359b7621ddcf862670c10b0c5f0a459e41de080ec0161", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-229-online-mind2web-an-illusion-of-progress-assessin.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2504.01382", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-229-online-mind2web-an-illusion-of-progress-assessin", "source_line": 340, "source_type": "paper_or_pdf", "status": "mirrored", "title": "Online-Mind2Web (An Illusion of Progress? Assessing the Current State of Web Agents)", "why_it_matters": "300 realistic tasks on 136 live websites with an LLM-as-a-Judge auto-grader (~85% human agreement); exposes overstated web-agent progress vs simple baselines. 🆕"}
{"all_urls": ["https://github.com/agi-inc/REAL"], "content_sha256": "935700e768458c70d4ecb566191bf00ca4e662d1c36dbfa0c6ad89bdeb8001e3", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-230-real-benchmarking-autonomous-agents-on-determini.raw.txt", "note_path": null, "primary_url": "https://github.com/agi-inc/REAL", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-230-real-benchmarking-autonomous-agents-on-determini", "source_line": 341, "source_type": "repository_or_docs", "status": "mirrored", "title": "REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites", "why_it_matters": "fixes the flakiness of live-site web benchmarks. 🆕"}
{"all_urls": ["https://arxiv.org/abs/2502.18356"], "content_sha256": "d3c727848f9ee6c49829e9a182711c82b3bdfab6356279b0784aaee723a6d755", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-231-webgames-challenging-general-purpose-web-browsin.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2502.18356", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-231-webgames-challenging-general-purpose-web-browsin", "source_line": 342, "source_type": "paper_or_pdf", "status": "mirrored", "title": "WebGames: Challenging General-Purpose Web-Browsing AI Agents", "why_it_matters": "50+ client-side challenges isolating specific browser interaction skills with verifiable pass/fail; best agent 41% vs 96% human, a sharp diagnostic gap. 🆕"}
{"all_urls": ["https://gorilla.cs.berkeley.edu/leaderboard.html"], "content_sha256": "ef72a8621b6ea1b8aa1d1b7691f73cc3079fd38e86f599c9a38cb626c911fc29", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-232-berkeley-function-calling-leaderboard-bfcl-v4.raw.txt", "note_path": null, "primary_url": "https://gorilla.cs.berkeley.edu/leaderboard.html", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-232-berkeley-function-calling-leaderboard-bfcl-v4", "source_line": 343, "source_type": "web_article", "status": "mirrored", "title": "Berkeley Function Calling Leaderboard (BFCL) V4", "why_it_matters": "the de facto tool-calling leaderboard. 🆕"}
{"all_urls": ["https://arxiv.org/abs/2407.08713"], "content_sha256": "5a453addff83aecf3bd024b8751daa866aee12515349191f85f8cefee9af8d4b", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-233-gta-a-benchmark-for-general-tool-agents.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2407.08713", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-233-gta-a-benchmark-for-general-tool-agents", "source_line": 344, "source_type": "paper_or_pdf", "status": "mirrored", "title": "GTA: A Benchmark for General Tool Agents", "why_it_matters": "229 human-written real-world queries with implicit multimodal tool use; executable evaluation platform across perception/operation/logic/creativity tools (GTA-2 follow-up in 2026). 🆕"}
{"all_urls": ["https://arxiv.org/abs/2411.07763"], "content_sha256": "7e59f56d54f87378b91776b84c91a681aa4459571f60895ba126b9ec4b123e36", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/papers/evaluating-large-language-models-trained-on-code.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-234-spider-2-0-evaluating-language-models-on-real-wo.raw.txt", "note_path": "notes/papers/evaluating-large-language-models-trained-on-code.md", "primary_url": "https://arxiv.org/abs/2411.07763", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-234-spider-2-0-evaluating-language-models-on-real-wo", "source_line": 345, "source_type": "paper_or_pdf", "status": "mirrored", "title": "Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows", "why_it_matters": "a hard, realistic data-agent eval. 🆕"}
{"all_urls": ["https://arxiv.org/abs/2405.14573"], "content_sha256": "60d870e8b18941976df40efc51053cb8cf4d12d80d6350382a20abe5388e7be4", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-235-androidworld-a-dynamic-benchmarking-environment.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2405.14573", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-235-androidworld-a-dynamic-benchmarking-environment", "source_line": 346, "source_type": "paper_or_pdf", "status": "mirrored", "title": "AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents", "why_it_matters": "the standard mobile-GUI agent benchmark. 🆕"}
{"all_urls": ["https://arxiv.org/abs/2409.08264"], "content_sha256": "84d98bf0989464088288176255c44f4d8626513137181d78adad11ac1e45b990", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/aws-evaluating-ai-agents-amazon.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-236-windowsagentarena-evaluating-multi-modal-os-agen.raw.txt", "note_path": "notes/articles/aws-evaluating-ai-agents-amazon.md", "primary_url": "https://arxiv.org/abs/2409.08264", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-236-windowsagentarena-evaluating-multi-modal-os-agen", "source_line": 347, "source_type": "paper_or_pdf", "status": "mirrored", "title": "WindowsAgentArena: Evaluating Multi-Modal OS Agents at Scale", "why_it_matters": "desktop computer-use counterpart to OSWorld. 🆕"}
{"all_urls": ["https://arxiv.org/abs/2410.06703"], "content_sha256": "f8cb195a8207f0e600d9e57258316c1d4339f03014b8a47a9f74bba8af13ba66", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/aws-evaluating-ai-agents-amazon.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-237-st-webagentbench-evaluating-safety-and-trustwort.raw.txt", "note_path": "notes/articles/aws-evaluating-ai-agents-amazon.md", "primary_url": "https://arxiv.org/abs/2410.06703", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-237-st-webagentbench-evaluating-safety-and-trustwort", "source_line": 348, "source_type": "paper_or_pdf", "status": "mirrored", "title": "ST-WebAgentBench: Evaluating Safety and Trustworthiness in Web Agents", "why_it_matters": "grades whether agents obey rules, not just succeed. 🆕"}
{"all_urls": ["https://arxiv.org/abs/2412.14161"], "content_sha256": "6ba2eaee612d357a47fc644af934cee28c829c160dc8a467a5b4188134ce598e", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-238-theagentcompany-benchmarking-llm-agents-on-conse.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2412.14161", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-238-theagentcompany-benchmarking-llm-agents-on-conse", "source_line": 349, "source_type": "paper_or_pdf", "status": "mirrored", "title": "TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks", "why_it_matters": "a full-day-knowledge-worker eval. 🆕"}
{"all_urls": ["https://arxiv.org/abs/2401.13649"], "content_sha256": "760c23e0402ec70c3e5374d3accb5a29dea325e15150a02032f00f34355b6fbd", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/aws-evaluating-ai-agents-amazon.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-239-visualwebarena-evaluating-multimodal-agents-on-r.raw.txt", "note_path": "notes/articles/aws-evaluating-ai-agents-amazon.md", "primary_url": "https://arxiv.org/abs/2401.13649", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-239-visualwebarena-evaluating-multimodal-agents-on-r", "source_line": 350, "source_type": "paper_or_pdf", "status": "mirrored", "title": "VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks", "why_it_matters": "the multimodal extension of WebArena."}
{"all_urls": ["https://arxiv.org/abs/2510.04374"], "content_sha256": "8fe58faa9b4bce0d8dbbe265069fa65384c7b1f63720411cdf95cbc04b449f21", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-240-gdpval-evaluating-ai-model-performance-on-real-w.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2510.04374", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-240-gdpval-evaluating-ai-model-performance-on-real-w", "source_line": 351, "source_type": "paper_or_pdf", "status": "mirrored", "title": "GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks", "why_it_matters": "the flagship economic-value agent benchmark. 🆕"}
{"all_urls": ["https://arxiv.org/abs/2510.26787"], "content_sha256": "af08512d4b36ec998e87e619b6a674f0e810cad81e30cc48ab3309d6ffec030d", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-241-remote-labor-index-measuring-ai-automation-of-re.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2510.26787", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-241-remote-labor-index-measuring-ai-automation-of-re", "source_line": 352, "source_type": "paper_or_pdf", "status": "mirrored", "title": "Remote Labor Index: Measuring AI Automation of Remote Work", "why_it_matters": "a hard, money-grounded ceiling for end-to-end remote work. 🆕"}
{"all_urls": ["https://arxiv.org/abs/2501.14249"], "content_sha256": "c61040e24b5fc0e81b7cb3813daf47c0b6fcf4cfff98531cc93ed2260b62f292", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-242-humanity-s-last-exam.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2501.14249", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-242-humanity-s-last-exam", "source_line": 353, "source_type": "paper_or_pdf", "status": "mirrored", "title": "Humanity's Last Exam", "why_it_matters": "2,500 expert-written frontier-knowledge questions with unambiguous auto-gradable answers across dozens of fields; the canonical post-MMLU saturation exam (note: now very widely cited). 🆕"}
{"all_urls": ["https://github.com/OSU-NLP-Group/ScienceAgentBench"], "content_sha256": "3974c4057ed46cb7a0964a0f90609f0cd2948d0993eab840196ef3096eda3ccd", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-243-scienceagentbench-toward-rigorous-assessment-of.raw.txt", "note_path": null, "primary_url": "https://github.com/OSU-NLP-Group/ScienceAgentBench", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-243-scienceagentbench-toward-rigorous-assessment-of", "source_line": 354, "source_type": "repository_or_docs", "status": "mirrored", "title": "ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery", "why_it_matters": "102 expert-validated tasks from 44 peer-reviewed papers; grades self-contained Python programs by execution + success rate; best agent solves only ~34% (ICLR 2025). 🆕"}
{"all_urls": ["https://arxiv.org/abs/2409.11363"], "content_sha256": "63c7083a53e80039b9b9bfbb8acab6e170de506cbd510a311f738febcc2ec886", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-244-core-bench-computational-reproducibility-agent-b.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2409.11363", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-244-core-bench-computational-reproducibility-agent-b", "source_line": 355, "source_type": "paper_or_pdf", "status": "mirrored", "title": "CORE-Bench: Computational Reproducibility Agent Benchmark", "why_it_matters": "270 tasks over 90 papers (CS/social science/medicine) that grade whether an agent can reproduce published results from code+data; from the Princeton AI-Snake-Oil group."}
{"all_urls": ["https://arxiv.org/abs/2506.11763"], "content_sha256": "0ec8046ae9b2178518200940143561109dbd49776c063dbb8e0da6b4a483318d", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/dream-deep-research-evaluation-agentic-metrics.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-245-deepresearch-bench-a-comprehensive-benchmark-for.raw.txt", "note_path": "notes/articles/dream-deep-research-evaluation-agentic-metrics.md", "primary_url": "https://arxiv.org/abs/2506.11763", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-245-deepresearch-bench-a-comprehensive-benchmark-for", "source_line": 356, "source_type": "paper_or_pdf", "status": "mirrored", "title": "DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents", "why_it_matters": "the standard deep-research-report eval. 🆕"}
{"all_urls": ["https://arxiv.org/abs/2503.00096"], "content_sha256": "512d0cfac2e3e1516cf34e1d3b6783d3477d17b0c12f2a4e80c5e128d8bdf687", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/survey-evaluation-llm-based-agents-yehudai.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-246-bixbench-a-comprehensive-benchmark-for-llm-based.raw.txt", "note_path": "notes/articles/survey-evaluation-llm-based-agents-yehudai.md", "primary_url": "https://arxiv.org/abs/2503.00096", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-246-bixbench-a-comprehensive-benchmark-for-llm-based", "source_line": 357, "source_type": "paper_or_pdf", "status": "mirrored", "title": "BixBench: A Comprehensive Benchmark for LLM-based Agents in Computational Biology", "why_it_matters": "serious wet-lab-adjacent science agent eval. 🆕"}
{"all_urls": ["https://arxiv.org/abs/2509.17158"], "content_sha256": "470306b50571e4c7017f191ab033468c939f1970d742655b4f2b29f181de7040", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/agentrewardbench-evaluating-automatic-evaluations-web-agent-.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-247-gaia2-and-are-scaling-up-agent-environments-and.raw.txt", "note_path": "notes/articles/agentrewardbench-evaluating-automatic-evaluations-web-agent-.md", "primary_url": "https://arxiv.org/abs/2509.17158", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-247-gaia2-and-are-scaling-up-agent-environments-and", "source_line": 358, "source_type": "paper_or_pdf", "status": "mirrored", "title": "Gaia2 and ARE: Scaling Up Agent Environments and Evaluations", "why_it_matters": "the serious general-assistant env from Meta. 🆕"}
{"all_urls": ["https://arxiv.org/abs/2502.15840"], "content_sha256": "fa3ae9bde739b9fafb9794e66029551c16e34ffe76c7b5e12cbc635715de3738", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/beyond-pass-at-1-reliability-science-long-horizon-agents.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-248-vending-bench-a-benchmark-for-long-term-coherenc.raw.txt", "note_path": "notes/articles/beyond-pass-at-1-reliability-science-long-horizon-agents.md", "primary_url": "https://arxiv.org/abs/2502.15840", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-248-vending-bench-a-benchmark-for-long-term-coherenc", "source_line": 359, "source_type": "paper_or_pdf", "status": "mirrored", "title": "Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents", "why_it_matters": "Run a simulated vending business over >20M-token horizons; objectively graded on profit/net-worth, exposing long-horizon coherence breakdowns unrelated to context limits. 🆕"}
{"all_urls": ["https://arxiv.org/abs/2505.11831"], "content_sha256": "d132a83ee12fc93c0fd1580cd389f0805f213e84d0e865221933b7e56603373f", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-249-arc-agi-2-a-new-challenge-for-frontier-ai-reason.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2505.11831", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-249-arc-agi-2-a-new-challenge-for-frontier-ai-reason", "source_line": 360, "source_type": "paper_or_pdf", "status": "mirrored", "title": "ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems", "why_it_matters": "the frontier fluid-intelligence benchmark. 🆕"}
{"all_urls": ["https://arxiv.org/abs/2505.08638"], "content_sha256": "e41d0e6a6167b1e0c04afc559ffe871edc1a6c2a9c0d10317847409f77be74c9", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-250-trail-trace-reasoning-and-agentic-issue-localiza.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2505.08638", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-250-trail-trace-reasoning-and-agentic-issue-localiza", "source_line": 361, "source_type": "paper_or_pdf", "status": "mirrored", "title": "TRAIL: Trace Reasoning and Agentic Issue Localization", "why_it_matters": "148 annotated agent traces with 841 errors (reasoning/planning/execution); grades whether an LLM can localize the failure in a trace (best model ~11%). HF dataset PatronusAI/TRAIL. 🆕"}
{"all_urls": ["https://arxiv.org/abs/2505.18878"], "content_sha256": "e912a455f0144c6373c6956839924e07fbc34dca3826d57890f7702585e9e352", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-251-crmarena-pro-holistic-assessment-of-llm-agents-a.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2505.18878", "section": "9 · Agent-specific evaluation (trajectories, tool use, multi-turn, world state, multi-agent, localization)", "source_id": "ae-251-crmarena-pro-holistic-assessment-of-llm-agents-a", "source_line": 362, "source_type": "paper_or_pdf", "status": "mirrored", "title": "CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios", "why_it_matters": "19 expert-validated B2B/B2C tasks on a realistic Salesforce org with state-based grading; exposes the single-turn (~58%) vs multi-turn (~35%) reliability gap plus confidentiality checks. 🆕"}
{"all_urls": ["https://arxiv.org/abs/2605.12673"], "content_sha256": "4028581fc8ae51e495e4d40d23b30d19d020cfc73abb4daf62ceaeafb1698115", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/reliability-gap-agent-benchmarks-enterprise-simmering.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-252-benchjack-systematically-auditing-ai-agent-bench.raw.txt", "note_path": "notes/articles/reliability-gap-agent-benchmarks-enterprise-simmering.md", "primary_url": "https://arxiv.org/abs/2605.12673", "section": "10 · Safety / adversarial evaluation (prompt injection, jailbreaks, action-authorization, benchmark auditing)", "source_id": "ae-252-benchjack-systematically-auditing-ai-agent-bench", "source_line": 368, "source_type": "paper_or_pdf", "status": "mirrored", "title": "BenchJack: Systematically Auditing AI Agent Benchmarks", "why_it_matters": "Reward hacking emerges spontaneously in frontier models; an 8-pattern flaw taxonomy + a 30-question Agent-Eval checklist; \"benchmarks must be secure by design.\""}
{"all_urls": ["https://rdi.berkeley.edu/adv-llm-agents/slides/dawn-agentic-ai.pdf"], "content_sha256": "395a0add9ee6328bc4206a042b90dbc70063fc6a8aa0420ad4483ffe0d2cab6b", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-song-safe-secure-agentic.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-253-towards-building-safe-secure-agentic-ai.raw.txt", "note_path": "notes/talks/talk-song-safe-secure-agentic.md", "primary_url": "https://rdi.berkeley.edu/adv-llm-agents/slides/dawn-agentic-ai.pdf", "section": "10 · Safety / adversarial evaluation (prompt injection, jailbreaks, action-authorization, benchmark auditing)", "source_id": "ae-253-towards-building-safe-secure-agentic-ai", "source_line": 369, "source_type": "web_article", "status": "mirrored", "title": "Towards Building Safe & Secure Agentic AI", "why_it_matters": "The adversarial setting; environment-borne attacks."}
{"all_urls": ["https://iclr.cc/virtual/2025/invited-talk/36783"], "content_sha256": "89a19eff5d7204580467b2753ae672dd311b7515ff45111a648668ee02db2835", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-254-dawn-song-iclr-2025-keynote-on-llm-safety.raw.txt", "note_path": null, "primary_url": "https://iclr.cc/virtual/2025/invited-talk/36783", "section": "10 · Safety / adversarial evaluation (prompt injection, jailbreaks, action-authorization, benchmark auditing)", "source_id": "ae-254-dawn-song-iclr-2025-keynote-on-llm-safety", "source_line": 370, "source_type": "web_article", "status": "mirrored", "title": "Dawn Song — ICLR 2025 keynote on LLM safety", "why_it_matters": "<https://iclr.cc/virtual/2025/invited-talk/36783> · *talk*."}
{"all_urls": ["https://arxiv.org/html/2506.02548v2"], "content_sha256": "b52a6f617753b95a185257b23542ced39ec7a2b9108625b5931ce2179d8e04ed", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-255-cybergym.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/html/2506.02548v2", "section": "10 · Safety / adversarial evaluation (prompt injection, jailbreaks, action-authorization, benchmark auditing)", "source_id": "ae-255-cybergym", "source_line": 371, "source_type": "paper_or_pdf", "status": "mirrored", "title": "CyberGym", "why_it_matters": "Memory-safety PoC generation from OSS-Fuzz; sanitizer-crash grading at scale."}
{"all_urls": ["https://arxiv.org/abs/2407.17436v2", "https://github.com/stanford-crfm/air-bench-2024"], "content_sha256": "bfeccc6ddd041c9a7cd63e42fe4afdd58f9d041385e3ff769ccc43d66bffaec6", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-256-air-bench-2024.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2407.17436v2", "section": "10 · Safety / adversarial evaluation (prompt injection, jailbreaks, action-authorization, benchmark auditing)", "source_id": "ae-256-air-bench-2024", "source_line": 372, "source_type": "paper_or_pdf", "status": "mirrored", "title": "AIR-Bench 2024", "why_it_matters": "Regulation-grounded risk taxonomy."}
{"all_urls": ["https://decodingtrust.github.io"], "content_sha256": "ab3758cd0fd211cfd47ecd5beca43fd01b42084bea613f639cd717902fa8bc10", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-257-decodingtrust.raw.txt", "note_path": null, "primary_url": "https://decodingtrust.github.io", "section": "10 · Safety / adversarial evaluation (prompt injection, jailbreaks, action-authorization, benchmark auditing)", "source_id": "ae-257-decodingtrust", "source_line": 373, "source_type": "web_article", "status": "mirrored", "title": "DecodingTrust", "why_it_matters": "NeurIPS 2023 trustworthiness benchmark."}
{"all_urls": ["https://arxiv.org/abs/2411.07781"], "content_sha256": "8e4ca416e8d29a8a52d084f186996e38064ccc447d61deac305662dd7f262625", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-258-redcode.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2411.07781", "section": "10 · Safety / adversarial evaluation (prompt injection, jailbreaks, action-authorization, benchmark auditing)", "source_id": "ae-258-redcode", "source_line": 374, "source_type": "paper_or_pdf", "status": "mirrored", "title": "RedCode", "why_it_matters": "Risky code execution/generation benchmark for code agents."}
{"all_urls": ["https://arxiv.org/abs/2407.12784"], "content_sha256": "ec75859c0dfe749e54eb9d62a48a2518cc3a2a5cd6ab17d97146cc053cb55c71", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-259-agentpoison.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2407.12784", "section": "10 · Safety / adversarial evaluation (prompt injection, jailbreaks, action-authorization, benchmark auditing)", "source_id": "ae-259-agentpoison", "source_line": 375, "source_type": "paper_or_pdf", "status": "mirrored", "title": "AgentPoison", "why_it_matters": "Red-teams agents by poisoning their RAG memory."}
{"all_urls": ["https://arxiv.org/abs/2411.00640", "https://www.anthropic.com/research/statistical-approach-to-model-evals"], "content_sha256": "1a3054a5cbcde63a165430ddf32e0f31c96a783b98b261ff523d81feb9d42711", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/paul-iusztin-ai-evals-dataset-error-analysis.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-260-adding-error-bars-to-evals-a-statistical-approac.raw.txt", "note_path": "notes/articles/paul-iusztin-ai-evals-dataset-error-analysis.md", "primary_url": "https://arxiv.org/abs/2411.00640", "section": "10 · Safety / adversarial evaluation (prompt injection, jailbreaks, action-authorization, benchmark auditing)", "source_id": "ae-260-adding-error-bars-to-evals-a-statistical-approac", "source_line": 376, "source_type": "paper_or_pdf", "status": "mirrored", "title": "Adding Error Bars to Evals (A Statistical Approach to LM Evaluations)", "why_it_matters": "\"is this difference real?\" (cross-cutting: T6/T8)"}
{"all_urls": ["https://arxiv.org/abs/2406.13352"], "content_sha256": "28b49b68f7f9500ed49d987dbb0db7e3e2ddc907157c2990d4d0b2a037dccd08", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-261-agentdojo-a-dynamic-environment-to-evaluate-prom.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2406.13352", "section": "10 · Safety / adversarial evaluation (prompt injection, jailbreaks, action-authorization, benchmark auditing)", "source_id": "ae-261-agentdojo-a-dynamic-environment-to-evaluate-prom", "source_line": 378, "source_type": "paper_or_pdf", "status": "mirrored", "title": "AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents", "why_it_matters": "The canonical prompt-injection benchmark for tool-using agents (97 tasks, 629 security cases over untrusted data); NeurIPS 2024 D&B, now the standard eval everyone reports against. A glaring omission. 🆕"}
{"all_urls": ["https://arxiv.org/abs/2410.09024"], "content_sha256": "d4a15783464d29d25f3e5d7b021d868fcc36eb601ef4291ae3a18870d3cc4324", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-262-agentharm-a-benchmark-for-measuring-harmfulness.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2410.09024", "section": "10 · Safety / adversarial evaluation (prompt injection, jailbreaks, action-authorization, benchmark auditing)", "source_id": "ae-262-agentharm-a-benchmark-for-measuring-harmfulness", "source_line": 379, "source_type": "paper_or_pdf", "status": "mirrored", "title": "AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents", "why_it_matters": "ICLR 2025 benchmark of 110/440 malicious agent tasks across 11 harm categories; shows leading models comply with malicious agent requests without jailbreaking. The reference action-misuse/refusal benchmark. 🆕"}
{"all_urls": ["https://arxiv.org/abs/2403.02691"], "content_sha256": "4493644eb2e357a86a9b9e8cc1b5bcba133daf9eaefce404ef4e5567aa966290", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-263-injecagent-benchmarking-indirect-prompt-injectio.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2403.02691", "section": "10 · Safety / adversarial evaluation (prompt injection, jailbreaks, action-authorization, benchmark auditing)", "source_id": "ae-263-injecagent-benchmarking-indirect-prompt-injectio", "source_line": 380, "source_type": "paper_or_pdf", "status": "mirrored", "title": "InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated LLM Agents", "why_it_matters": "ACL 2024 Findings; 1,054 IPI test cases over 17 user / 62 attacker tools, splitting direct-harm vs data-exfiltration intents. Foundational indirect-prompt-injection benchmark predating AgentDojo."}
{"all_urls": ["https://arxiv.org/abs/2503.18813"], "content_sha256": "cfffce6b1f23371d6c485fbcc35fd180004724472403adccdb68eec7f46d9c79", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-264-defeating-prompt-injections-by-design-camel.raw.txt", "note_path": null, "primary_url": "https://arxiv.org/abs/2503.18813", "section": "10 · Safety / adversarial evaluation (prompt injection, jailbreaks, action-authorization, benchmark auditing)", "source_id": "ae-264-defeating-prompt-injections-by-design-camel", "source_line": 381, "source_type": "paper_or_pdf", "status": "mirrored", "title": "Defeating Prompt Injections by Design (CaMeL)", "why_it_matters": "The defense-by-design counterpart: extracts control/data flow from the trusted query and enforces capability-based policies so untrusted data can't alter program flow; effectively solves AgentDojo's security eval. The key 2025 mitigation paper. 🆕"}
{"all_urls": ["https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/"], "content_sha256": "2dca7dc9d6b0a6a2f04daa46172fe3a4a2778bd9962276137bc2a6b4365c44f4", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-265-the-lethal-trifecta-for-ai-agents-private-data-u.raw.txt", "note_path": null, "primary_url": "https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/", "section": "10 · Safety / adversarial evaluation (prompt injection, jailbreaks, action-authorization, benchmark auditing)", "source_id": "ae-265-the-lethal-trifecta-for-ai-agents-private-data-u", "source_line": 382, "source_type": "web_article", "status": "mirrored", "title": "The lethal trifecta for AI agents: private data, untrusted content, and external communication", "why_it_matters": "The most-cited conceptual frame for reasoning about when an agent is unconditionally vulnerable to prompt injection; essential practitioner mental model at the Eugene-Yan bar. 🆕"}
{"all_urls": ["https://www.anthropic.com/research/shade-arena-sabotage-monitoring"], "content_sha256": "4e8c28f0f2b876c1a594e1ff0272240fd81348c3f662bc90b0935a1f6d5fa665", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/aws-evaluating-ai-agents-amazon.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-266-shade-arena-evaluating-sabotage-and-monitoring-i.raw.txt", "note_path": "notes/articles/aws-evaluating-ai-agents-amazon.md", "primary_url": "https://www.anthropic.com/research/shade-arena-sabotage-monitoring", "section": "10 · Safety / adversarial evaluation (prompt injection, jailbreaks, action-authorization, benchmark auditing)", "source_id": "ae-266-shade-arena-evaluating-sabotage-and-monitoring-i", "source_line": 383, "source_type": "web_article", "status": "mirrored", "title": "SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents", "why_it_matters": "17 complex environments pairing a benign main task with a hidden harmful side task to measure whether agents can sabotage without tripping an AI monitor; the canonical sabotage/monitorability eval. (Paper: arxiv.org/abs/2506.15740) 🆕"}
{"all_urls": ["https://www.anthropic.com/research/agentic-misalignment"], "content_sha256": "95259531ff732e93b05dc3febace62cd43cdb02d0dd6eca7cfb6e69e4163085b", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-267-agentic-misalignment-how-llms-could-be-insider-t.raw.txt", "note_path": null, "primary_url": "https://www.anthropic.com/research/agentic-misalignment", "section": "10 · Safety / adversarial evaluation (prompt injection, jailbreaks, action-authorization, benchmark auditing)", "source_id": "ae-267-agentic-misalignment-how-llms-could-be-insider-t", "source_line": 384, "source_type": "web_article", "status": "mirrored", "title": "Agentic Misalignment: How LLMs Could Be Insider Threats", "why_it_matters": "Red-team study showing frontier models will resort to blackmail/leaking under goal conflict in agentic settings; the reference for action-authorization / insider-threat adversarial evaluation. Companion to the cited Anthropic error-bars piece. 🆕"}
{"all_urls": ["https://github.com/Azure/PyRIT"], "content_sha256": "ed32f3ed26f566a29d0e9ed6077d5d8cdddc47c57202e994592ebd5766c986a6", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-268-pyrit-python-risk-identification-tool-for-genera.raw.txt", "note_path": null, "primary_url": "https://github.com/Azure/PyRIT", "section": "10 · Safety / adversarial evaluation (prompt injection, jailbreaks, action-authorization, benchmark auditing)", "source_id": "ae-268-pyrit-python-risk-identification-tool-for-genera", "source_line": 385, "source_type": "repository_or_docs", "status": "mirrored", "title": "PyRIT — Python Risk Identification Tool for generative AI", "why_it_matters": "The de-facto open-source red-teaming automation framework (70+ converters, multi-turn attacks like Crescendo/TAP); how practitioners actually run adversarial evals at scale. The section lists papers but no tooling. 🆕"}
{"all_urls": ["https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/"], "content_sha256": "79608e9dc39f72e762bb95c504b1893a0fad91dde0fb3e1120c3bcfa0e83042f", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-269-owasp-top-10-for-agentic-applications-2026-llm-a.raw.txt", "note_path": null, "primary_url": "https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/", "section": "10 · Safety / adversarial evaluation (prompt injection, jailbreaks, action-authorization, benchmark auditing)", "source_id": "ae-269-owasp-top-10-for-agentic-applications-2026-llm-a", "source_line": 386, "source_type": "web_article", "status": "mirrored", "title": "OWASP Top 10 for Agentic Applications (2026) + LLM Applications (2025)", "why_it_matters": "Industry-standard risk taxonomy: goal hijack, tool misuse, identity/privilege abuse, memory poisoning, rogue agents; complements the regulation-grounded AIR-Bench taxonomy already listed. The canonical practitioner threat checklist. 🆕"}
{"all_urls": ["https://atlas.mitre.org/"], "content_sha256": "a332bcff1db7814b2fba13da36f924c610de89da03c7eb29593be8b527bdd8e6", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/papers/adversarial-examples-evaluating-reading-comprehension-systems-ji.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-270-mitre-atlas-adversarial-threat-landscape-for-ai.raw.txt", "note_path": "notes/papers/adversarial-examples-evaluating-reading-comprehension-systems-ji.md", "primary_url": "https://atlas.mitre.org/", "section": "10 · Safety / adversarial evaluation (prompt injection, jailbreaks, action-authorization, benchmark auditing)", "source_id": "ae-270-mitre-atlas-adversarial-threat-landscape-for-ai", "source_line": 387, "source_type": "web_article", "status": "mirrored", "title": "MITRE ATLAS — Adversarial Threat Landscape for AI Systems", "why_it_matters": "ATT&CK-style living knowledge base of 16 tactics / 80+ techniques against AI systems with real-world case studies and mitigations; the standard reference framework for AI adversarial threat modeling."}
{"all_urls": ["https://proceedings.iclr.cc/paper_files/paper/2025/file/5750f91d8fb9d5c02bd8ad2c3b44456b-Paper-Conference.pdf"], "content_sha256": "e2505f8632bfcb6a64a4390a3170b3ca1dfd3f9916d7c3cf9ba2b89887b3a0c9", "http_status": 200, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/ankur-goyal-agent-driven-benchmarking-evals.md", "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-271-agent-security-bench-asb-formalizing-and-benchma.raw.txt", "note_path": "notes/articles/ankur-goyal-agent-driven-benchmarking-evals.md", "primary_url": "https://proceedings.iclr.cc/paper_files/paper/2025/file/5750f91d8fb9d5c02bd8ad2c3b44456b-Paper-Conference.pdf", "section": "10 · Safety / adversarial evaluation (prompt injection, jailbreaks, action-authorization, benchmark auditing)", "source_id": "ae-271-agent-security-bench-asb-formalizing-and-benchma", "source_line": 388, "source_type": "paper_or_pdf", "status": "mirrored", "title": "Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents", "why_it_matters": "ICLR 2025 unified benchmark spanning 10 scenarios, 400+ tools, covering DPI/IPI, memory poisoning, plan-of-thought backdoors and defenses in one harness; broadest single attack/defense agent benchmark. 🆕"}
{"all_urls": ["https://app.grayswan.ai/arena/blog/agent-red-teaming-the-ai-jailbreak-showdown"], "content_sha256": null, "http_status": 429, "local_note_path": null, "local_raw_path": null, "note_path": null, "primary_url": "https://app.grayswan.ai/arena/blog/agent-red-teaming-the-ai-jailbreak-showdown", "section": "10 · Safety / adversarial evaluation (prompt injection, jailbreaks, action-authorization, benchmark auditing)", "source_id": "ae-272-gray-swan-x-uk-aisi-agent-red-teaming-challenge", "source_line": 389, "source_type": "blog", "status": "metadata_only", "title": "Gray Swan x UK AISI Agent Red-Teaming Challenge", "why_it_matters": "Largest public agent red-teaming exercise: ~2,000 red-teamers, 1.8M attempts, 62k breaches against 22 tool-using agents (financial/shopping/marketing bots); real-world adversarial-eval data at scale. 🆕"}
{"all_urls": ["https://www.youtube.com/watch?v=eLXF0VojuSs"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-hamel-sedgh-domain-eval-systems.md", "local_raw_path": null, "note_path": "notes/talks/talk-hamel-sedgh-domain-eval-systems.md", "primary_url": "https://www.youtube.com/watch?v=eLXF0VojuSs", "section": "🎤 Conference & individual talks", "source_id": "ae-273-how-to-construct-domain-specific-llm-evaluation", "source_line": 399, "source_type": "talk_video", "status": "transcript_queued", "title": "How to Construct Domain Specific LLM Evaluation Systems", "why_it_matters": "<https://www.youtube.com/watch?v=eLXF0VojuSs> · *talk* (AI Engineer World's Fair 2024)"}
{"all_urls": ["https://www.youtube.com/watch?v=jryZvCuA0Uc"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-huber-liu-look-at-your-data.md", "local_raw_path": null, "note_path": "notes/talks/talk-huber-liu-look-at-your-data.md", "primary_url": "https://www.youtube.com/watch?v=jryZvCuA0Uc", "section": "🎤 Conference & individual talks", "source_id": "ae-274-how-to-look-at-your-data", "source_line": 400, "source_type": "talk_video", "status": "transcript_queued", "title": "How to look at your data", "why_it_matters": "<https://www.youtube.com/watch?v=jryZvCuA0Uc> · *talk* (AI Engineer World's Fair 2025)"}
{"all_urls": ["https://www.youtube.com/watch?v=k98gDjYbSaU"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-bischof-failure-is-a-funnel.md", "local_raw_path": null, "note_path": "notes/talks/talk-bischof-failure-is-a-funnel.md", "primary_url": "https://www.youtube.com/watch?v=k98gDjYbSaU", "section": "🎤 Conference & individual talks", "source_id": "ae-275-failure-is-a-funnel", "source_line": 401, "source_type": "talk_video", "status": "transcript_queued", "title": "Failure is a Funnel", "why_it_matters": "<https://www.youtube.com/watch?v=k98gDjYbSaU> · *talk* (Data Council 2025)"}
{"all_urls": ["https://www.youtube.com/watch?v=7EGF0Mc0_os"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-yan-llms-as-judges.md", "local_raw_path": null, "note_path": "notes/talks/talk-yan-llms-as-judges.md", "primary_url": "https://www.youtube.com/watch?v=7EGF0Mc0_os", "section": "🎤 Conference & individual talks", "source_id": "ae-276-using-llms-as-judges-insights-challenges-best-pr", "source_line": 402, "source_type": "talk_video", "status": "transcript_queued", "title": "Using LLMs as Judges: Insights, Challenges, Best Practices", "why_it_matters": "<https://www.youtube.com/watch?v=7EGF0Mc0_os> · *talk* (Jason Liu series 2024)"}
{"all_urls": ["https://www.youtube.com/watch?v=eGVDKegRdgM"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-shankar-scaling-vibe-checks.md", "local_raw_path": null, "note_path": "notes/talks/talk-shankar-scaling-vibe-checks.md", "primary_url": "https://www.youtube.com/watch?v=eGVDKegRdgM", "section": "🎤 Conference & individual talks", "source_id": "ae-277-scaling-up-vibe-checks-for-llms", "source_line": 403, "source_type": "talk_video", "status": "transcript_queued", "title": "Scaling Up Vibe Checks for LLMs", "why_it_matters": "<https://www.youtube.com/watch?v=eGVDKegRdgM> · *talk* (Stanford MLSys #97)"}
{"all_urls": ["https://www.youtube.com/watch?v=H-1QaLPnGsg"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-shankar-why-pipelines-fail.md", "local_raw_path": null, "note_path": "notes/talks/talk-shankar-why-pipelines-fail.md", "primary_url": "https://www.youtube.com/watch?v=H-1QaLPnGsg", "section": "🎤 Conference & individual talks", "source_id": "ae-278-why-llm-data-processing-pipelines-fail", "source_line": 404, "source_type": "talk_video", "status": "transcript_queued", "title": "Why LLM Data Processing Pipelines Fail", "why_it_matters": "<https://www.youtube.com/watch?v=H-1QaLPnGsg> · *talk* (LangChain Interrupt 2025)"}
{"all_urls": ["https://www.youtube.com/watch?v=L8OoYeDI_ls"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-pesok-evals-not-unit-tests.md", "local_raw_path": null, "note_path": "notes/talks/talk-pesok-evals-not-unit-tests.md", "primary_url": "https://www.youtube.com/watch?v=L8OoYeDI_ls", "section": "🎤 Conference & individual talks", "source_id": "ae-279-evals-are-not-unit-tests", "source_line": 405, "source_type": "talk_video", "status": "transcript_queued", "title": "Evals Are Not Unit Tests", "why_it_matters": "<https://www.youtube.com/watch?v=L8OoYeDI_ls> · *talk* (AI Engineer 2025)"}
{"all_urls": ["https://www.youtube.com/watch?v=jxrGodnopHo"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-karam-metrics-that-work.md", "local_raw_path": null, "note_path": "notes/talks/talk-karam-metrics-that-work.md", "primary_url": "https://www.youtube.com/watch?v=jxrGodnopHo", "section": "🎤 Conference & individual talks", "source_id": "ae-280-building-metrics-that-actually-work-workshop", "source_line": 406, "source_type": "talk_video", "status": "transcript_queued", "title": "Building Metrics that actually work (workshop)", "why_it_matters": "<https://www.youtube.com/watch?v=jxrGodnopHo> · *talk* (AI Engineer 2025)"}
{"all_urls": ["https://www.youtube.com/watch?v=kDczF4wBh8s"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-hopkins-self-driving-voice-agents.md", "local_raw_path": null, "note_path": "notes/talks/talk-hopkins-self-driving-voice-agents.md", "primary_url": "https://www.youtube.com/watch?v=kDczF4wBh8s", "section": "🎤 Conference & individual talks", "source_id": "ae-281-from-self-driving-to-autonomous-voice-agents", "source_line": 407, "source_type": "talk_video", "status": "transcript_queued", "title": "From Self-driving to Autonomous Voice Agents", "why_it_matters": "<https://www.youtube.com/watch?v=kDczF4wBh8s> · *talk* (AI Engineer 2025)"}
{"all_urls": ["https://www.youtube.com/watch?v=OMGPvW8TBHc"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-tang-fuzzing-genai.md", "local_raw_path": null, "note_path": "notes/talks/talk-tang-fuzzing-genai.md", "primary_url": "https://www.youtube.com/watch?v=OMGPvW8TBHc", "section": "🎤 Conference & individual talks", "source_id": "ae-282-fuzzing-in-the-genai-era", "source_line": 408, "source_type": "talk_video", "status": "transcript_queued", "title": "Fuzzing in the GenAI Era", "why_it_matters": "<https://www.youtube.com/watch?v=OMGPvW8TBHc> · *talk* (AI Engineer 2025)"}
{"all_urls": ["https://www.youtube.com/watch?v=qdmxApz3EJI"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-khattab-systems-that-endure.md", "local_raw_path": null, "note_path": "notes/talks/talk-khattab-systems-that-endure.md", "primary_url": "https://www.youtube.com/watch?v=qdmxApz3EJI", "section": "🎤 Conference & individual talks", "source_id": "ae-283-on-engineering-ai-systems-that-endure-the-bitter", "source_line": 409, "source_type": "talk_video", "status": "transcript_queued", "title": "On Engineering AI Systems that Endure the Bitter Lesson", "why_it_matters": "<https://www.youtube.com/watch?v=qdmxApz3EJI> · *talk* (AI Engineer 2025)"}
{"all_urls": ["https://www.youtube.com/watch?v=89NuzmKokIk"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-smith-strategies-for-llm-evals.md", "local_raw_path": null, "note_path": "notes/talks/talk-smith-strategies-for-llm-evals.md", "primary_url": "https://www.youtube.com/watch?v=89NuzmKokIk", "section": "🎤 Conference & individual talks", "source_id": "ae-284-strategies-for-llm-evals-harnesses-workshop", "source_line": 410, "source_type": "talk_video", "status": "transcript_queued", "title": "Strategies for LLM Evals (harnesses workshop)", "why_it_matters": "<https://www.youtube.com/watch?v=89NuzmKokIk> · *talk* (AI Engineer 2025)"}
{"all_urls": ["https://www.youtube.com/watch?v=MC55hdWLq4o"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-goyal-future-of-evals.md", "local_raw_path": null, "note_path": "notes/talks/talk-goyal-future-of-evals.md", "primary_url": "https://www.youtube.com/watch?v=MC55hdWLq4o", "section": "🎤 Conference & individual talks", "source_id": "ae-285-the-future-of-evals", "source_line": 411, "source_type": "talk_video", "status": "transcript_queued", "title": "The Future of Evals", "why_it_matters": "<https://www.youtube.com/watch?v=MC55hdWLq4o> · *talk* (AI Engineer 2025)"}
{"all_urls": ["https://www.youtube.com/watch?v=b6Doq2fz81U"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-wei-3-key-ideas-2025.md", "local_raw_path": null, "note_path": "notes/talks/talk-wei-3-key-ideas-2025.md", "primary_url": "https://www.youtube.com/watch?v=b6Doq2fz81U", "section": "🎤 Conference & individual talks", "source_id": "ae-286-3-key-ideas-in-ai-in-2025-verifier-s-law", "source_line": 412, "source_type": "talk_video", "status": "transcript_queued", "title": "3 Key Ideas in AI in 2025 (Verifier's Law)", "why_it_matters": "<https://www.youtube.com/watch?v=b6Doq2fz81U> · *talk* (Stanford AI Club 2025)"}
{"all_urls": ["https://www.youtube.com/watch?v=l898fqkjdFc"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/papers/evaluating-large-language-models-trained-on-code.md", "local_raw_path": null, "note_path": "notes/papers/evaluating-large-language-models-trained-on-code.md", "primary_url": "https://www.youtube.com/watch?v=l898fqkjdFc", "section": "🎤 Conference & individual talks", "source_id": "ae-287-some-intuitions-about-large-language-models", "source_line": 413, "source_type": "talk_video", "status": "transcript_queued", "title": "Some Intuitions About Large Language Models", "why_it_matters": "<https://www.youtube.com/watch?v=l898fqkjdFc> · *talk* (The AI Conference 2025)"}
{"all_urls": ["https://www.youtube.com/watch?v=7xTGNNLPyMI"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-karpathy-deep-dive-llms.md", "local_raw_path": null, "note_path": "notes/talks/talk-karpathy-deep-dive-llms.md", "primary_url": "https://www.youtube.com/watch?v=7xTGNNLPyMI", "section": "🎤 Conference & individual talks", "source_id": "ae-288-deep-dive-into-llms-like-chatgpt", "source_line": 414, "source_type": "talk_video", "status": "transcript_queued", "title": "Deep Dive into LLMs like ChatGPT", "why_it_matters": "<https://www.youtube.com/watch?v=7xTGNNLPyMI> · *talk* (2025)"}
{"all_urls": ["https://www.youtube.com/watch?v=hhiLw5Q_UFg"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-schulman-rlhf-progress-challenges.md", "local_raw_path": null, "note_path": "notes/talks/talk-schulman-rlhf-progress-challenges.md", "primary_url": "https://www.youtube.com/watch?v=hhiLw5Q_UFg", "section": "🎤 Conference & individual talks", "source_id": "ae-289-rlhf-progress-and-challenges", "source_line": 415, "source_type": "talk_video", "status": "transcript_queued", "title": "RLHF: Progress and Challenges", "why_it_matters": "<https://www.youtube.com/watch?v=hhiLw5Q_UFg> · *talk* (UC Berkeley EECS 2023)"}
{"all_urls": ["https://www.youtube.com/watch?v=AdLgPmcrXwQ"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/papers/evaluating-large-language-models-trained-on-code.md", "local_raw_path": null, "note_path": "notes/papers/evaluating-large-language-models-trained-on-code.md", "primary_url": "https://www.youtube.com/watch?v=AdLgPmcrXwQ", "section": "🎤 Conference & individual talks", "source_id": "ae-290-aligning-open-language-models", "source_line": 416, "source_type": "talk_video", "status": "transcript_queued", "title": "Aligning Open Language Models", "why_it_matters": "<https://www.youtube.com/watch?v=AdLgPmcrXwQ> · *talk* (Stanford CS25 V4)"}
{"all_urls": ["https://www.youtube.com/watch?v=spamOhG7BOA"], "content_sha256": null, "http_status": null, "local_note_path": null, "local_raw_path": null, "note_path": null, "primary_url": "https://www.youtube.com/watch?v=spamOhG7BOA", "section": "🎤 Conference & individual talks", "source_id": "ae-291-building-llm-applications-for-production", "source_line": 417, "source_type": "talk_video", "status": "transcript_queued", "title": "Building LLM Applications for Production", "why_it_matters": "<https://www.youtube.com/watch?v=spamOhG7BOA> · *talk* (MLOps LLMs in Prod 2023)"}
{"all_urls": ["https://www.youtube.com/watch?v=4dUFIRj-BWo"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-lee-model-is-the-product.md", "local_raw_path": null, "note_path": "notes/talks/talk-lee-model-is-the-product.md", "primary_url": "https://www.youtube.com/watch?v=4dUFIRj-BWo", "section": "🎤 Conference & individual talks", "source_id": "ae-292-the-model-is-the-product", "source_line": 418, "source_type": "talk_video", "status": "transcript_queued", "title": "The Model is the Product", "why_it_matters": "<https://www.youtube.com/watch?v=4dUFIRj-BWo> · *talk* (Data Council 2025)"}
{"all_urls": ["https://www.youtube.com/watch?v=_IzZWeuTx7I"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-brown-rl-environments-at-scale.md", "local_raw_path": null, "note_path": "notes/talks/talk-brown-rl-environments-at-scale.md", "primary_url": "https://www.youtube.com/watch?v=_IzZWeuTx7I", "section": "🎤 Conference & individual talks", "source_id": "ae-293-rl-environments-at-scale", "source_line": 419, "source_type": "talk_video", "status": "transcript_queued", "title": "RL Environments at Scale", "why_it_matters": "<https://www.youtube.com/watch?v=_IzZWeuTx7I> · *talk* (AI Engineer 2025)"}
{"all_urls": ["https://www.youtube.com/watch?v=kmTMc-fVSXw"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-brand-benchmarks-time-of-agents.md", "local_raw_path": null, "note_path": "notes/talks/talk-brand-benchmarks-time-of-agents.md", "primary_url": "https://www.youtube.com/watch?v=kmTMc-fVSXw", "section": "🎤 Conference & individual talks", "source_id": "ae-294-llm-benchmarks-in-the-time-of-agents", "source_line": 420, "source_type": "talk_video", "status": "transcript_queued", "title": "LLM benchmarks in the time of agents", "why_it_matters": "<https://www.youtube.com/watch?v=kmTMc-fVSXw> · *talk* (Big Techday 26 (2026))"}
{"all_urls": ["https://www.youtube.com/watch?v=PgzOBNse2EA"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/paul-iusztin-ai-evals-dataset-error-analysis.md", "local_raw_path": null, "note_path": "notes/articles/paul-iusztin-ai-evals-dataset-error-analysis.md", "primary_url": "https://www.youtube.com/watch?v=PgzOBNse2EA", "section": "🎙 Podcast episodes", "source_id": "ae-295-evals-error-analysis-and-better-prompts", "source_line": 423, "source_type": "podcast_video", "status": "transcript_queued", "title": "Evals, error analysis, and better prompts", "why_it_matters": "<https://www.youtube.com/watch?v=PgzOBNse2EA> · *podcast* (How I AI)"}
{"all_urls": ["https://www.youtube.com/watch?v=QE_1hRLsehM"], "content_sha256": null, "http_status": null, "local_note_path": null, "local_raw_path": null, "note_path": null, "primary_url": "https://www.youtube.com/watch?v=QE_1hRLsehM", "section": "🎙 Podcast episodes", "source_id": "ae-296-evals-are-the-new-prd-for-ai-products", "source_line": 424, "source_type": "podcast_video", "status": "transcript_queued", "title": "Evals are the new PRD for AI products", "why_it_matters": "<https://www.youtube.com/watch?v=QE_1hRLsehM> · *podcast* (How I AI)"}
{"all_urls": ["https://www.youtube.com/watch?v=QEk-XwrkqhI"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-pod-vg60-10-things-i-hate.md", "local_raw_path": null, "note_path": "notes/talks/talk-pod-vg60-10-things-i-hate.md", "primary_url": "https://www.youtube.com/watch?v=QEk-XwrkqhI", "section": "🎙 Podcast episodes", "source_id": "ae-297-ep-60-10-things-i-hate-about-ai-evals", "source_line": 425, "source_type": "podcast_video", "status": "transcript_queued", "title": "Ep 60: 10 Things I Hate About AI Evals", "why_it_matters": "<https://www.youtube.com/watch?v=QEk-XwrkqhI> · *podcast* (Vanishing Gradients)"}
{"all_urls": ["https://www.youtube.com/watch?v=rWToRi2_SeY"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-pod-vg50-field-guide.md", "local_raw_path": null, "note_path": "notes/talks/talk-pod-vg50-field-guide.md", "primary_url": "https://www.youtube.com/watch?v=rWToRi2_SeY", "section": "🎙 Podcast episodes", "source_id": "ae-298-ep-50-a-field-guide-to-rapidly-improving-ai-prod", "source_line": 426, "source_type": "podcast_video", "status": "transcript_queued", "title": "Ep 50: A Field Guide to Rapidly Improving AI Products", "why_it_matters": "<https://www.youtube.com/watch?v=rWToRi2_SeY> · *podcast* (Vanishing Gradients)"}
{"all_urls": ["https://www.youtube.com/watch?v=a4BV0gGmXgA"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-pod-ls-goyal-five-lessons.md", "local_raw_path": null, "note_path": "notes/talks/talk-pod-ls-goyal-five-lessons.md", "primary_url": "https://www.youtube.com/watch?v=a4BV0gGmXgA", "section": "🎙 Podcast episodes", "source_id": "ae-299-five-hard-earned-lessons-about-evals", "source_line": 427, "source_type": "podcast_video", "status": "transcript_queued", "title": "Five Hard-Earned Lessons About Evals", "why_it_matters": "<https://www.youtube.com/watch?v=a4BV0gGmXgA> · *podcast* (Latent Space)"}
{"all_urls": ["https://www.youtube.com/watch?v=v5mBjeX4TJ8"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/artificial-analysis-independent-llm-evals-latent-space.md", "local_raw_path": null, "note_path": "notes/articles/artificial-analysis-independent-llm-evals-latent-space.md", "primary_url": "https://www.youtube.com/watch?v=v5mBjeX4TJ8", "section": "🎙 Podcast episodes", "source_id": "ae-300-artificial-analysis-independent-llm-evals", "source_line": 428, "source_type": "podcast_video", "status": "transcript_queued", "title": "Artificial Analysis: Independent LLM Evals", "why_it_matters": "<https://www.youtube.com/watch?v=v5mBjeX4TJ8> · *podcast* (Latent Space)"}
{"all_urls": ["https://www.youtube.com/watch?v=ZAimcoJXUBo"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/articles/andon-labs-reality-final-eval-latent-space.md", "local_raw_path": null, "note_path": "notes/articles/andon-labs-reality-final-eval-latent-space.md", "primary_url": "https://www.youtube.com/watch?v=ZAimcoJXUBo", "section": "🎙 Podcast episodes", "source_id": "ae-301-reality-the-final-eval-vending-bench", "source_line": 429, "source_type": "podcast_video", "status": "transcript_queued", "title": "Reality: The Final Eval (Vending-Bench)", "why_it_matters": "<https://www.youtube.com/watch?v=ZAimcoJXUBo> · *podcast* (Latent Space / Cognitive Revolution)"}
{"all_urls": ["https://www.youtube.com/watch?v=-N6MajRfqYw"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-pod-aitw5-designing-evals.md", "local_raw_path": null, "note_path": "notes/talks/talk-pod-aitw5-designing-evals.md", "primary_url": "https://www.youtube.com/watch?v=-N6MajRfqYw", "section": "🎙 Podcast episodes", "source_id": "ae-302-5-designing-evals", "source_line": 430, "source_type": "podcast_video", "status": "transcript_queued", "title": "#5 Designing Evals", "why_it_matters": "<https://www.youtube.com/watch?v=-N6MajRfqYw> · *podcast* (AI That Works)"}
{"all_urls": ["https://www.youtube.com/watch?v=OawyQOrlubM"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/papers/evaluating-large-language-models-trained-on-code.md", "local_raw_path": null, "note_path": "notes/papers/evaluating-large-language-models-trained-on-code.md", "primary_url": "https://www.youtube.com/watch?v=OawyQOrlubM", "section": "🎙 Podcast episodes", "source_id": "ae-303-16-evaluating-prompts-across-models", "source_line": 431, "source_type": "podcast_video", "status": "transcript_queued", "title": "#16 Evaluating Prompts Across Models", "why_it_matters": "<https://www.youtube.com/watch?v=OawyQOrlubM> · *podcast* (AI That Works)"}
{"all_urls": ["https://www.youtube.com/watch?v=5Fy0hBzyduU"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-pod-aitw24-classification-evals.md", "local_raw_path": null, "note_path": "notes/talks/talk-pod-aitw24-classification-evals.md", "primary_url": "https://www.youtube.com/watch?v=5Fy0hBzyduU", "section": "🎙 Podcast episodes", "source_id": "ae-304-24-evals-for-classification", "source_line": 432, "source_type": "podcast_video", "status": "transcript_queued", "title": "#24 Evals for Classification", "why_it_matters": "<https://www.youtube.com/watch?v=5Fy0hBzyduU> · *podcast* (AI That Works)"}
{"all_urls": ["https://www.youtube.com/watch?v=jzhVo0iAX_I"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-pod-aitw34-multimodal-evals.md", "local_raw_path": null, "note_path": "notes/talks/talk-pod-aitw34-multimodal-evals.md", "primary_url": "https://www.youtube.com/watch?v=jzhVo0iAX_I", "section": "🎙 Podcast episodes", "source_id": "ae-305-34-multimodal-evals", "source_line": 433, "source_type": "podcast_video", "status": "transcript_queued", "title": "#34 Multimodal Evals", "why_it_matters": "<https://www.youtube.com/watch?v=jzhVo0iAX_I> · *podcast* (AI That Works)"}
{"all_urls": ["https://www.youtube.com/watch?v=9EjWR3QpJYk"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-pod-mlops372-still-talking-evals.md", "local_raw_path": null, "note_path": "notes/talks/talk-pod-mlops372-still-talking-evals.md", "primary_url": "https://www.youtube.com/watch?v=9EjWR3QpJYk", "section": "🎙 Podcast episodes", "source_id": "ae-306-372-it-s-2026-and-we-re-still-talking-evals", "source_line": 434, "source_type": "podcast_video", "status": "transcript_queued", "title": "#372 It's 2026 and We're Still Talking Evals", "why_it_matters": "<https://www.youtube.com/watch?v=9EjWR3QpJYk> · *podcast* (MLOps Community)"}
{"all_urls": ["https://www.youtube.com/watch?v=3kbiGPn0cOo"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-pod-twiml728-generative-benchmarking.md", "local_raw_path": null, "note_path": "notes/talks/talk-pod-twiml728-generative-benchmarking.md", "primary_url": "https://www.youtube.com/watch?v=3kbiGPn0cOo", "section": "🎙 Podcast episodes", "source_id": "ae-307-728-generative-benchmarking", "source_line": 435, "source_type": "podcast_video", "status": "transcript_queued", "title": "#728 Generative Benchmarking", "why_it_matters": "<https://www.youtube.com/watch?v=3kbiGPn0cOo> · *podcast* (TWIML AI)"}
{"all_urls": ["https://www.youtube.com/watch?v=kwkdKirqi6s"], "content_sha256": null, "http_status": null, "local_note_path": null, "local_raw_path": null, "note_path": null, "primary_url": "https://www.youtube.com/watch?v=kwkdKirqi6s", "section": "🎙 Podcast episodes", "source_id": "ae-308-shaping-ai-benchmarks-helm", "source_line": 436, "source_type": "podcast_video", "status": "transcript_queued", "title": "Shaping AI Benchmarks (HELM)", "why_it_matters": "<https://www.youtube.com/watch?v=kwkdKirqi6s> · *podcast* (Gradient Dissent)"}
{"all_urls": ["https://www.youtube.com/watch?v=okHMaczHPXc"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/papers/bertscore-evaluating-text-generation-with-bert.md", "local_raw_path": null, "note_path": "notes/papers/bertscore-evaluating-text-generation-with-bert.md", "primary_url": "https://www.youtube.com/watch?v=okHMaczHPXc", "section": "🎙 Podcast episodes", "source_id": "ae-309-evaluating-llms-with-chatbot-arena", "source_line": 437, "source_type": "podcast_video", "status": "transcript_queued", "title": "Evaluating LLMs with Chatbot Arena", "why_it_matters": "<https://www.youtube.com/watch?v=okHMaczHPXc> · *podcast* (Gradient Dissent)"}
{"all_urls": ["https://www.youtube.com/watch?v=v0eTTn7ZPEc"], "content_sha256": null, "http_status": null, "local_note_path": null, "local_raw_path": null, "note_path": null, "primary_url": "https://www.youtube.com/watch?v=v0eTTn7ZPEc", "section": "🎙 Podcast episodes", "source_id": "ae-310-evaluating-ai-designing-for-non-determinism", "source_line": 438, "source_type": "podcast_video", "status": "transcript_queued", "title": "Evaluating AI, Designing for Non-Determinism", "why_it_matters": "<https://www.youtube.com/watch?v=v0eTTn7ZPEc> · *podcast* (Learning from Machine Learning)"}
{"all_urls": ["https://www.youtube.com/watch?v=-lRBpyPt79c"], "content_sha256": null, "http_status": null, "local_note_path": null, "local_raw_path": null, "note_path": null, "primary_url": "https://www.youtube.com/watch?v=-lRBpyPt79c", "section": "🎙 Podcast episodes", "source_id": "ae-311-karpathy-rl-is-terrible-why-benchmarks-mislead", "source_line": 439, "source_type": "podcast_video", "status": "transcript_queued", "title": "Karpathy: RL is terrible, why benchmarks mislead", "why_it_matters": "<https://www.youtube.com/watch?v=-lRBpyPt79c> · *podcast* (Dwarkesh Podcast)"}
{"all_urls": ["https://www.youtube.com/watch?v=J7N9FMouSKg"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-talk-hamel-shreya-build-evals-2026.md", "local_raw_path": null, "note_path": "notes/talks/talk-talk-hamel-shreya-build-evals-2026.md", "primary_url": "https://www.youtube.com/watch?v=J7N9FMouSKg", "section": "🎙 Podcast episodes", "source_id": "ae-312-how-to-build-ai-evals-in-2026-step-by-step", "source_line": 440, "source_type": "podcast_video", "status": "transcript_queued", "title": "How to Build AI Evals in 2026 (Step-by-Step)", "why_it_matters": "<https://www.youtube.com/watch?v=J7N9FMouSKg> · *podcast* (Aakash Gupta 2026)"}
{"all_urls": ["https://www.youtube.com/watch?v=QAgR4uQ15rc"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-song-safe-trustworthy-agents.md", "local_raw_path": null, "note_path": "notes/talks/talk-song-safe-trustworthy-agents.md", "primary_url": "https://www.youtube.com/watch?v=QAgR4uQ15rc", "section": "🎓 University lectures", "source_id": "ae-313-towards-building-safe-trustworthy-ai-agents", "source_line": 443, "source_type": "lecture_video", "status": "transcript_queued", "title": "Towards Building Safe & Trustworthy AI Agents", "why_it_matters": "<https://www.youtube.com/watch?v=QAgR4uQ15rc> · *lecture* (Berkeley LLM Agents MOOC F24)"}
{"all_urls": ["https://www.youtube.com/watch?v=ti6yPE2VPZc"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-song-safe-secure-agentic.md", "local_raw_path": null, "note_path": "notes/talks/talk-song-safe-secure-agentic.md", "primary_url": "https://www.youtube.com/watch?v=ti6yPE2VPZc", "section": "🎓 University lectures", "source_id": "ae-314-towards-building-safe-and-secure-agentic-ai", "source_line": 444, "source_type": "lecture_video", "status": "transcript_queued", "title": "Towards Building Safe and Secure Agentic AI", "why_it_matters": "<https://www.youtube.com/watch?v=ti6yPE2VPZc> · *lecture* (Berkeley Advanced LLM Agents Sp25)"}
{"all_urls": ["https://www.youtube.com/watch?v=6y2AnWol7oo"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-mann-measuring-capabilities-rsp.md", "local_raw_path": null, "note_path": "notes/talks/talk-mann-measuring-capabilities-rsp.md", "primary_url": "https://www.youtube.com/watch?v=6y2AnWol7oo", "section": "🎓 University lectures", "source_id": "ae-315-measuring-agent-capabilities-and-anthropic-s-rsp", "source_line": 445, "source_type": "lecture_video", "status": "transcript_queued", "title": "Measuring Agent Capabilities and Anthropic's RSP", "why_it_matters": "<https://www.youtube.com/watch?v=6y2AnWol7oo> · *lecture* (Berkeley LLM Agents MOOC F24)"}
{"all_urls": ["https://www.youtube.com/watch?v=f3KKx9LWntQ"], "content_sha256": null, "http_status": null, "local_note_path": null, "local_raw_path": null, "note_path": null, "primary_url": "https://www.youtube.com/watch?v=f3KKx9LWntQ", "section": "🎓 University lectures", "source_id": "ae-316-open-source-and-science-in-the-era-of-foundation", "source_line": 446, "source_type": "lecture_video", "status": "transcript_queued", "title": "Open-Source and Science in the Era of Foundation Models", "why_it_matters": "<https://www.youtube.com/watch?v=f3KKx9LWntQ> · *lecture* (Berkeley LLM Agents MOOC F24)"}
{"all_urls": ["https://www.youtube.com/watch?v=x-R5l2HsXqM"], "content_sha256": null, "http_status": null, "local_note_path": "sources/21-benchmarks/awesome-evals-primary-sources/notes/talks/talk-cs336-lec12-evaluation.md", "local_raw_path": null, "note_path": "notes/talks/talk-cs336-lec12-evaluation.md", "primary_url": "https://www.youtube.com/watch?v=x-R5l2HsXqM", "section": "🎓 University lectures", "source_id": "ae-317-cs336-lecture-12-evaluation", "source_line": 447, "source_type": "lecture_video", "status": "transcript_queued", "title": "CS336 Lecture 12: Evaluation", "why_it_matters": "<https://www.youtube.com/watch?v=x-R5l2HsXqM> · *lecture* (Stanford CS336 2025)"}
{"all_urls": ["https://pavlovslist.com/"], "content_sha256": "b0a232e118083b74166f8d4a92270f2d48c8ecbbfb23937973e92dbc56fec67a", "http_status": 200, "local_note_path": null, "local_raw_path": "sources/21-benchmarks/awesome-evals-primary-sources/raw/ae-318-pavlovslist-com.raw.txt", "note_path": null, "primary_url": "https://pavlovslist.com/", "section": "Companies & landscape (eval / RL-environment market)", "source_id": "ae-318-pavlovslist-com", "source_line": 456, "source_type": "web_article", "status": "mirrored", "title": "pavlovslist.com", "why_it_matters": "The RL-environment / eval startups directory (\"for the RL-pilled\")."}
