Benchmarks3 articles

Benchmarks

Articles

  • Artificial Analysis Launches Search Index Benchmark for AI Agent Search APIs

    Artificial Analysis has released the Search Index, a benchmark suite designed to evaluate web search APIs for autonomous AI agents across retrieval quality, query latency, and end-to-end task economics. The initial evaluation tests seven dedicated search providers: Parallel, Exa, Firecrawl, You.com, Tavily, Keenable, and Brave. Benchmark Setup and Evaluation Methodology To isolate search API performance from model variance, the evaluation executes all tests with GPT-5.6 Luna inside Stirrup,

    1 min
  • Zhipu's GLM-5.2 Narrows the Gap With Anthropic's Fable 5 on Coding Benchmarks

    Zhipu AI's open-weight GLM-5.2 model has placed second globally on the Code Arena coding benchmark, trailing only Anthropic's Claude Fable 5 and sitting within one point of Claude Opus 4.8 on the hardest agentic tasks. The result, recorded after the model's June 13 release under an MIT license, signals a fast-closing capability gap between Chinese open-weight systems and the leading U.S. frontier models. Benchmark Performance On Code Arena's front-end coding leaderboard, GLM-5.2 places second

    1 min