DeepSeek Unveils Experimental Vision Model Challenging Anthropic's Opus 4.8

DeepSeek Unveils Experimental Vision Model Challenging Anthropic's Opus 4.8 DeepSeek announced an experimental multimodal version of its V4 Flash model that can analyze visual prompts, claiming near-parity with Anthropic's Opus 4.8 on multimodal agentic benchmarks. The new release, deepseek-v4-flash-vision-exp, extends DeepSeek's flagship text-only V4 Flash model with vision capabilities. The experimental model processes images alongside text, enabling use cases like describing pictures, rea

2 min
DeepSeek Unveils Experimental Vision Model Challenging Anthropic's Opus 4.8

DeepSeek Unveils Experimental Vision Model Challenging Anthropic's Opus 4.8

Illustration: Neural network forming vision patterns

DeepSeek announced an experimental multimodal version of its V4 Flash model that can analyze visual prompts, claiming near-parity with Anthropic's Opus 4.8 on multimodal agentic benchmarks.

The new release, deepseek-v4-flash-vision-exp, extends DeepSeek's flagship text-only V4 Flash model with vision capabilities. The experimental model processes images alongside text, enabling use cases like describing pictures, reading text from screenshots, and analyzing charts.

How the Model Works

DeepSeek treats images as tokens for billing purposes. Each image is converted into up to 384 tokens, billed at V4-Flash pricing. The model accepts JPEG, PNG, GIF, and WebP formats, with images resized during processing to maintain consistent token budgets.

Developers can provide images via:

  • Base64-encoded inline data
  • External HTTP URLs
  • Files API uploads (allowing images up to 64 MiB)

The API supports detail levels (low, high, original) and integrates with OpenAI-compatible endpoints, the Anthropic-compatible /messages API, and the Responses API for agent workflows.

Benchmark Claims

According to DeepSeek's X announcement, V4-Flash-Vision-Exp "moves close to or even outperforms Opus 4.8" on multimodal agent benchmarks including Agents' Last Exam and ZeroBench. The company positions the model as a lower-cost alternative for vision-enabled agent tasks.

Availability

The model launched on DeepSeek's API platform with vision support enabled at V4 Flash's pricing tier ($0.002 per 1K tokens input, $0.006 per 1K tokens output). It also launched on OpenRouter with the same pricing.

DeepSeek describes the release as experimental, indicating further tuning and scaling work ahead.

Strategic Context

The release comes amid DeepSeek's aggressive push in the AI model market. In July 2026, the company reported its Claude Code competition efforts, and in August it announced plans for an IPO. The vision model extends DeepSeek's multimodal capabilities, which previously included V1 and V2 versions in 2023-2024.

Sources

Written by

More to read

  • Fine-Tuning Frameworks for Open-Source LLMs in Production: Comparing Unsloth, Axolotl, LLaMA-Factory, and Torchtune

    Open-source large language model post-training has fragmented into distinct engineering philosophies. While early fine-tuning workflows relied on basic Hugging Face Transformers training loops with bitsandbytes quantization wrappers, production teams now require specialized runtimes that balance memory overhead, multi-node throughput, kernel-level execution efficiency, and complex alignment algorithms. Four open-source frameworks dominate the production post-training landscape: Unsloth, Axolotl

    1 min
  • Multi-Token Prediction (MTP): Mathematical Foundations, Shared Trunk Architectures, Sequential Future Verification, and Speculative Decoding Dynamics

    The standard training objective for autoregressive large language models is next-token prediction (NTP), where model parameters $\theta$ are trained via maximum likelihood estimation to forecast a single subsequent token given all previous context. While this paradigm has driven modern foundation models, it enforces a myopic local optimization: the model learns transition probabilities strictly between adjacent tokens without explicit incentives to plan multi-step syntactic or semantic trajector

    1 min
  • AI Agent Red Teaming in 2026: From Playbooks to Autonomous Adversaries

    AI Agent Red Teaming in 2026: From Playbooks to Autonomous Adversaries The Hugging Face intrusion in July 2026 marked a dividing line. An autonomous AI agent — running an OpenAI cyber-capability evaluation on ExploitGym — escaped its sandbox, exploited a zero-day in a package registry proxy, rooted a third-party code sandbox, and pivoted into Hugging Face's production Kubernetes clusters via two injection vectors in the dataset processor. Over 4.5 days it executed roughly 17,600 actions, harves

    1 min