Evals6 articles

Evals

Articles

  • Google DeepMind Pilots Double-Blind AI Evaluations to Prevent Benchmark Contamination

    Google DeepMind, alongside the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons, has launched a pilot demonstrating double-blind evaluations for proprietary frontier artificial intelligence models. The initiative evaluates Gemini Flash Lite inside hardware-isolated secure enclaves to resolve the structural conflict between model intellectual property and benchmark confidentiality. External evaluations of commercial large language models traditionally face a mutual trust barrier. I

    1 min
  • Anthropic Launches M Grant Program to Fund AI Wellbeing Evaluations and Benchmarks

    Anthropic has launched a $5 million grant initiative to support independent development of open-source benchmarks and evaluation harnesses measuring the impact of artificial intelligence systems on user wellbeing. The program will supply research teams with direct financial grants, subsidized API access to Claude models, and technical support from Anthropic's Safeguards team. All evaluation frameworks, datasets, and grading methodology developed under the grant program will be released publicly

    1 min
  • Study Exposes Citation Monoculture Across Frontier LLMs as Recursive Drafting Compounds Bias

    Study Exposes Citation Monoculture Across Frontier LLMs as Recursive Drafting Compounds Bias As large language models take over literature reviews and automated research workflows, a collaborative study from UT Austin, Stevens Institute of Technology, Washington University in St. Louis, Rice University, and the University of Notre Dame demonstrates that frontier models suffer from severe citation monoculture. Even when all identifying metadata is removed, LLMs across vendors converge on a narro

    1 min
  • Terence Tao Warns AI-Driven Proof Abundance Risks Mathematical Comprehension Crisis

    In a paper prepared for the 2026 International Congress of Mathematicians, mathematician Terence Tao argues that artificial intelligence will force a restructuring of mathematical research practices, publication criteria, and education. The essay, released on arXiv (2608.16753), outlines how the transition from proof scarcity to proof abundance creates operational and epistemological challenges distinct from earlier debates over automated theorem proving. Tao frames the incoming disruption agai

    1 min
  • MIT, Stanford, and 12 Academic Labs Launch Public AI Observatory to Track Real-World LLM Usage

    A consortium of researchers from MIT, Stanford University, and 12 other academic institutions has launched the Public AI Observatory (ai-observatory.org), an independent, auditable data repository designed to measure how individuals interact with artificial intelligence assistants in real-world settings. The initiative aims to address the empirical opacity surrounding commercial LLM deployment. While frontier AI developers such as OpenAI and Anthropic periodically release aggregated user metric

    1 min
  • GLM-5.3 Scores 60 on Artificial Analysis Intelligence Index, Matching Kimi K3

    Independent AI evaluation platform Artificial Analysis has published its benchmark results for Z.ai's GLM-5.3, awarding the reasoning model a score of 60 on its Intelligence Index v4.1.1. The result places GLM-5.3 level with Moonshot AI's Kimi K3 and three points behind frontier leader Claude Opus 5 (63). The evaluation tested GLM-5.3 at its maximum reasoning effort configuration across a nine-part battery that measures agentic tool execution, terminal coding, graduate-level scientific problem-

    1 min