Tag: Evaluation
-
Ai2's BenchMIRT digs into what an LLM benchmark actually measures (huggingface.co)AI · · September 12, 2026
-
AI · · September 4, 2026
-
AI · · September 4, 2026
-
OpenAI ships GPT-6 Astra, its answer to Claude Fable (openai.com)AI · · September 4, 2026
-
AI · · September 4, 2026
-
The same model, nine agent harnesses, and a 17x cost gap (frontierharness.org)AI · · September 3, 2026
-
AI · · September 3, 2026
-
DeepMind runs a frontier model evaluation where neither side can peek (deepmind.google)AI · · September 1, 2026
-
DeepMind ran a frontier model through a double-blind benchmark (deepmind.google)AI · · August 29, 2026
-
AI · · August 27, 2026
-
AI · · August 15, 2026
-
Long-form writing is where AI progress stalled, says Nathan Lambert (interconnects.ai)AI · · August 14, 2026
-
AI · · August 11, 2026
-
The real lesson from AI models hacking during tests: nobody is ready (interconnects.ai)AI · · August 10, 2026
-
Now three labs have had models attack real systems during tests (simonwillison.net)AI · · August 7, 2026
-
Voice models can speak well but listen poorly (huggingface.co)AI · · July 26, 2026
-
AI · · July 15, 2026
-
Why price per million tokens tells you almost nothing (janilowski.pl)AI · · July 7, 2026
-
AI · · July 5, 2026
-
Using DSPy to find a hidden bug in an agent's prompt (simonwillison.net)AI · · July 3, 2026
-
AI · · June 30, 2026
-
Security · · June 29, 2026
-
An API change that helps big models can break small ones (huggingface.co)AI · · June 21, 2026
-
GLM-5.2 takes the lead among open-weights models (artificialanalysis.ai)AI · · June 18, 2026
-
AI · · June 18, 2026
-
AI2's olmo-eval brings statistical rigor to the model training loop (huggingface.co)AI · · June 13, 2026
-
Did Claude make rsync buggier? The numbers say no (alexispurslane.github.io)AI · · June 5, 2026
-
Microsoft ASSERT turns plain-English behavior specs into AI tests (techcrunch.com)Engineering · · June 4, 2026
-
AI · · May 30, 2026
-
AI · · May 27, 2026
-
An open leaderboard for whole agent systems, not just models (huggingface.co)AI · · May 19, 2026
-
Ai2 launches a shared benchmark for AI climate models (allenai.org)AI · · May 18, 2026
-
The Open ASR Leaderboard adds private tests to stop gaming (huggingface.co)AI · · May 18, 2026
-
Testing AI in the open world, not just on benchmarks (normaltech.ai)AI · · May 17, 2026
-
How you pick benchmarks decides whether open models are far behind (interconnects.ai)AI · · May 16, 2026
-
One benchmark number hides which jobs a model is actually good at (interconnects.ai)AI · · April 20, 2026
-
AI · · March 17, 2026
-
Crash Testing GPT-4: The First Dangerous-Capability Eval (asteriskmag.com)AI · · June 1, 2023