DeepSeek4 articles

DeepSeek

Articles

  • Group Relative Policy Optimization (GRPO): How Eliminating Value Models Scaled LLM Reasoning

    Post-training reinforcement learning (RL) has become the primary mechanism for scaling reasoning capabilities in large language models. While early reinforcement learning from human feedback (RLHF) focused on conversational style and safety alignment, extending RL to multi-step reasoning domains such as mathematics, algorithmic coding, and formal logic exposed critical limitations in classical algorithms. Standard Proximal Policy Optimization (PPO), long the foundational algorithm for instructi

    1 min
  • DeepSeek V4 Flash tops charts but fails half its real agent tasks

    DeepSeek's V4 Flash has become the most-used AI model on OpenRouter and one of the highest-rated open-weight models available. In real agent tests, though, it finished barely more than half of the jobs it was given. The gap between leaderboard and workplace is the story. What Composio found The integration company Composio ran V4 Flash through eight agent harnesses, including Claude Code, Codex, and OpenCode, on 30 deliberately hard multi-step tasks. The tasks used live tools: Gmail, GitHub,

    1 min
  • DeepSeek builds a team to challenge Anthropic's Claude Code

    DeepSeek is no longer keeping its agent ambitions quiet. The Hangzhou-based lab has opened an official social media account for a new "DeepSeek Harness Team" and posted job listings for roles aimed at building AI agents that can take on products like Anthropic's Claude Code. The account sits on WeChat, the Chinese super-app run by Tencent. Corporate records reviewed by Bloomberg show the account belongs to a Beijing-based entity controlled by DeepSeek, and Tencent has verified it. "Harness"

    1 min
  • Chinese hackers deploy open-source AI agents to automate espionage against Taiwan and Thailand

    Chinese-speaking threat actors are now deploying open-source AI agents to automate espionage-grade hacking against government targets across Asia, according to coordinated disclosures from Hunt.io, Palo Alto Networks Unit 42, and the Financial Times. The campaign, active since at least June 2026, centers on Hermes, an open-source autonomous agent framework that crossed 140,000 GitHub stars by July. Operators run Hermes in "YOLO mode," a configuration that removes human approval prompts and lets

    1 min