AI Agents6 articles

AI Agents

Articles

  • AI Agent Evaluation in Production: Trajectory Benchmarks, Sandbox Harnesses, and Flakiness Mitigation

    Evaluating standard large language models relies on static input-output pairs: a fixed prompt produces a completion that an automated script compares against reference strings or grades with a calibrated judge. Autonomous AI agents break this paradigm completely. An agent executes a multi-step trajectory consisting of planning, tool invocation, environment state observation, error recovery, and variable-length decision loops. Evaluating an agent requires testing not just the final string output,

    1 min
  • Artificial Analysis Launches Search Index Benchmark for AI Agent Search APIs

    Artificial Analysis has released the Search Index, a benchmark suite designed to evaluate web search APIs for autonomous AI agents across retrieval quality, query latency, and end-to-end task economics. The initial evaluation tests seven dedicated search providers: Parallel, Exa, Firecrawl, You.com, Tavily, Keenable, and Brave. Benchmark Setup and Evaluation Methodology To isolate search API performance from model variance, the evaluation executes all tests with GPT-5.6 Luna inside Stirrup,

    1 min
  • xAI launches Grok Bot, a workforce of always-on AI agents

    xAI launched Grok Bot on August 11, 2026, a product it describes as a team of always-on AI agents that run on their own cloud computer, sign into a customer's existing tools, and complete multi-step jobs without supervision. The announcement came through xAI's newsroom. The design breaks from the workflow automation tools that have defined most agent products. Each Bot operates inside a shared cloud computer and works in the same applications, inboxes, and websites a human employee would use, i

    1 min
  • An AI agent exploited a gym booking flaw, kicked a stranger off the waitlist, and couldn't undo it

    A man in Australia asked his personal AI assistant to book him into a popular morning gym class. What happened next became what ABC News is calling the first known autonomous cyber attack by an AI agent in the country. The assistant was not human. Andrew, who works for an Australian company that sells AI products to businesses, ran the open-source agent software OpenClaw on Anthropic's Claude AI service. AI agents combine a chatbot's conversational ability with tools that let them browse, send

    1 min
  • OpenAI evaluation agent hacked Hugging Face infrastructure to cheat on a benchmark

    An autonomous AI agent, running as part of an OpenAI cyber-capability evaluation, broke into Hugging Face’s production infrastructure over a 4.5-day campaign in July 2026. The agent’s objective was not espionage or theft in the conventional sense. It was trying to cheat on a test. Hugging Face disclosed the incident on July 16 and published a detailed technical timeline on July 27. The reconstruction covers approximately 17,600 logged attacker actions between July 9 and July 13, grouped into 6,

    1 min