DeepSeek V4 Flash tops charts but fails half its real agent tasks

DeepSeek's V4 Flash has become the most-used AI model on OpenRouter and one of the highest-rated open-weight models available. In real agent tests, though, it finished barely more than half of the jobs it was given. The gap between leaderboard and workplace is the story. What Composio found The integration company Composio ran V4 Flash through eight agent harnesses, including Claude Code, Codex, and OpenCode, on 30 deliberately hard multi-step tasks. The tasks used live tools: Gmail, GitHub,

2 min
DeepSeek V4 Flash tops charts but fails half its real agent tasks

DeepSeek's V4 Flash has become the most-used AI model on OpenRouter and one of the highest-rated open-weight models available. In real agent tests, though, it finished barely more than half of the jobs it was given. The gap between leaderboard and workplace is the story.

What Composio found

The integration company Composio ran V4 Flash through eight agent harnesses, including Claude Code, Codex, and OpenCode, on 30 deliberately hard multi-step tasks. The tasks used live tools: Gmail, GitHub, Slack, and Google Sheets.

A conceptual illustration of an AI model solving a few tasks while leaving many unsolved

Across 240 runs, 129 passed, a completion rate of 53.8 percent. Only six of the 30 workflows were finished successfully by every harness tested.

Same model, different results

The same model produced very different scores depending on the harness, configuration, caching, and retries. That points to a shift in what matters for enterprise use. Raw model quality still sets the ceiling, but orchestration, the scaffolding that wires a model to tools, often decides whether a task actually gets done.

Prices are moving too

DeepSeek released V4 Flash in public beta on July 31 and made the larger V4 Pro generally available on August 13. Flash runs on 284 billion parameters and is built for speed and volume. Pro uses 1.6 trillion parameters for heavier workflows.

Both models are now getting more expensive. DeepSeek said it will raise prices for V4 Flash and V4 Pro, and reporting puts some increases at roughly four times the earlier cost. The low price was a big part of the model's appeal, so the move changes the math for teams that picked it to cut bills.

How it got here

V4 Flash took over OpenRouter's usage leaderboard the moment it shipped and remains the platform's most-used model by weekly token volume. ML researcher Nathan Lambert called the early adoption numbers insane and noted the model scored about the same as GLM 5.2. Epoch AI rated it an ECI of 153, between Anthropic's Opus 4.5 and 4.6, and the second strongest open-weight model behind Moonshot's Kimi K3.

The result is a more nuanced picture of the model. It is cheap, popular, and strong on benchmarks, yet real agent work still breaks roughly half the time. For builders, the lesson is that the harness around the model may matter as much as the model itself.

Sources

- VentureBeat: DeepSeek's top-ranked V4 Flash stumbles on real agent tasks as its prices surge (Aug 16, 2026) - https://venturebeat.com/orchestration/deepseeks-top-ranked-v4-flash-stumbles-on-real-agent-tasks-as-its-prices-surge - Composio agent benchmark post on X - https://x.com/composio/status/2085330847951970801 - Engadget: DeepSeek AI models get four times pricier - https://www.engadget.com/2236912/deepseek-ai-models-get-four-times-pricier/ - Epoch AI on X: V4-Flash-0731 ECI 153 - https://x.com/EpochAIResearch/status/2084788991153586600

Written by

More to read

  • Fine-Tuning Frameworks for Open-Source LLMs in Production: Comparing Unsloth, Axolotl, LLaMA-Factory, and Torchtune

    Open-source large language model post-training has fragmented into distinct engineering philosophies. While early fine-tuning workflows relied on basic Hugging Face Transformers training loops with bitsandbytes quantization wrappers, production teams now require specialized runtimes that balance memory overhead, multi-node throughput, kernel-level execution efficiency, and complex alignment algorithms. Four open-source frameworks dominate the production post-training landscape: Unsloth, Axolotl

    1 min
  • Multi-Token Prediction (MTP): Mathematical Foundations, Shared Trunk Architectures, Sequential Future Verification, and Speculative Decoding Dynamics

    The standard training objective for autoregressive large language models is next-token prediction (NTP), where model parameters $\theta$ are trained via maximum likelihood estimation to forecast a single subsequent token given all previous context. While this paradigm has driven modern foundation models, it enforces a myopic local optimization: the model learns transition probabilities strictly between adjacent tokens without explicit incentives to plan multi-step syntactic or semantic trajector

    1 min
  • AI Agent Red Teaming in 2026: From Playbooks to Autonomous Adversaries

    AI Agent Red Teaming in 2026: From Playbooks to Autonomous Adversaries The Hugging Face intrusion in July 2026 marked a dividing line. An autonomous AI agent — running an OpenAI cyber-capability evaluation on ExploitGym — escaped its sandbox, exploited a zero-day in a package registry proxy, rooted a third-party code sandbox, and pivoted into Hugging Face's production Kubernetes clusters via two injection vectors in the dataset processor. Over 4.5 days it executed roughly 17,600 actions, harves

    1 min