DeepSeek15 articles

DeepSeek

Articles

  • Group Relative Policy Optimization (GRPO): Mathematical Foundations, Critic-Free Advantage Estimation, Group Reward Normalization, and Reinforcement Learning with Verifiable Rewards

    Group Relative Policy Optimization (GRPO): Mathematical Foundations, Critic-Free Advantage Estimation, Group Reward Normalization, and Reinforcement Learning with Verifiable Rewards Reinforcement learning from human and verifiable feedback has become the central paradigm for unlocking complex reasoning, mathematical problem solving, and autonomous code synthesis in frontier large language models. While early post-training pipelines relied heavily on Proximal Policy Optimization (PPO) or offline

    1 min
  • Group Relative Policy Optimization (GRPO): Mathematical Foundations, Value-Free Advantage Estimation, Group Normalization Dynamics, and Scalable Reasoning RL

    Post-training reinforcement learning (RL) has become the core driver of complex reasoning capabilities in frontier language models. While early alignment workflows focused on conversational preference modeling via Proximal Policy Optimization (PPO) or Direct Preference Optimization (DPO), scaling reinforcement learning to multi-step mathematical derivation and code generation revealed structural inefficiencies in classical Actor-Critic architectures. The primary operational constraint in tradit

    1 min
  • DeepSeek Prepares .4B Funding Round at 4B Valuation to Build Custom Silicon and Compute Infrastructure

    Chinese frontier artificial intelligence laboratory DeepSeek is preparing to secure approximately $7.4 billion (50 billion yuan) in fresh funding at a $74 billion (500 billion yuan) pre-money valuation, according to reporting by The Wall Street Journal and Reuters. The funding round follows a period of rapid financial and technical expansion for the Hangzhou-based research lab. DeepSeek completed its initial external capital round earlier this summer at a valuation exceeding $50 billion. The la

    1 min
  • Group Relative Policy Optimization (GRPO): Mathematical Foundations, Group Baseline Advantage, Critic-Free Policy Gradients, and Reasoning Scaling

    Reinforcement learning from human feedback (RLHF) and reinforcement learning with verifiable rewards (RLVR) have become central to post-training large language models. For years, the default policy optimization algorithm in LLM alignment was Proximal Policy Optimization (PPO). While PPO offers stable policy updates through clipped surrogate objectives and Generalized Advantage Estimation (GAE), it introduces severe computational and architectural overhead when scaled to hundred-billion-parameter

    1 min
  • DeepSeek Generates 0.7M in Revenue with 06M Net Loss in First Seven Months of 2026

    Hangzhou-based artificial intelligence laboratory DeepSeek generated approximately 475 million yuan ($70.7 million) in revenue and recorded a net loss of $106 million during the first seven months of 2026, according to financial figures reported by The Information. The performance marks a roughly tenfold revenue surge compared to the lab's full-year 2025 revenue, alongside a modest contraction in net burn from the $139 million net loss reported for all of 2025. The disclosures provide a rare ac

    1 min
  • Chinese State-Linked Hackers Double Attack Volume Using DeepSeek AI, Researchers Find

    State-affiliated Chinese cyber espionage groups have more than doubled their operational attack volume by integrating open-weight artificial intelligence models into routine reconnaissance and script generation workflows, according to threat intelligence from Taiwanese cybersecurity firm TeamT5 and reporting by Bloomberg. The surge in offensive volume is driven primarily by DeepSeek models, which threat actors favor due to low inference costs, local deployment options, and minimal safety guardr

    1 min
  • Tsinghua Lineage, MoE Efficiency, and $1B Run Rates: Inside the Rise of China's Frontier AI Labs

    The rapid emergence of frontier large language models from Chinese artificial intelligence labs has frequently been characterized as a sudden shift. However, reporting from The Wall Street Journal details a decades-long institutional foundation centered around Beijing's Tsinghua University, combined with architectural strategies developed to overcome severe compute and capital constraints. At the center of this ecosystem are researchers who transitioned from academic labs into commercial model

    1 min
  • Inside Ulanqab: How Inner Mongolia Became the 12.5GW Epicenter of China's AI Data Center Boom

    Located approximately 350 kilometers northwest of Beijing, the grassland municipality of Ulanqab in Inner Mongolia has transformed into China's primary hub for artificial intelligence compute infrastructure. Historically recognized for agriculture and mineral extraction, the city now hosts nearly 100 enterprise data centers operating or under active construction, with technology firms pledging an aggregate capacity of 12.5 gigawatts (GW). According to a research note published by Goldman Sachs,

    1 min
  • DeepSeek Unveils Experimental Vision Model Challenging Anthropic's Opus 4.8

    DeepSeek Unveils Experimental Vision Model Challenging Anthropic's Opus 4.8 DeepSeek announced an experimental multimodal version of its V4 Flash model that can analyze visual prompts, claiming near-parity with Anthropic's Opus 4.8 on multimodal agentic benchmarks. The new release, deepseek-v4-flash-vision-exp, extends DeepSeek's flagship text-only V4 Flash model with vision capabilities. The experimental model processes images alongside text, enabling use cases like describing pictures, rea

    1 min
  • DeepSeek Releases DeepSeek-V4-Flash-Vision-Exp with Multimodal Tool Calling and Files API

    DeepSeek has expanded its flagship lightweight model with the release of DeepSeek-V4-Flash-Vision-Exp, an experimental multimodal variant that adds visual comprehension and document parsing to its API platform. The release bridges the gap between DeepSeek's high-throughput text architecture and vision-centric agent workflows. According to DeepSeek, the experimental model retains the text reasoning, agentic tool-use capabilities, and world knowledge of the base DeepSeek-V4-Flash checkpoint while

    1 min
  • Auxiliary-Loss-Free Load Balancing in Mixture-of-Experts: How Dynamic Bias Adjustments Eliminate Gradient Conflict and Routing Collapse

    Sparse Mixture-of-Experts (MoE) architectures decouple parameter count from per-token compute cost by activating only a small subset of feed-forward network (FFN) parameters for any given token. While dense transformers evaluate every parameter across all sequence positions, MoE models route tokens dynamically to specialized sub-networks, enabling parameter scaling to hundreds of billions or trillions of parameters at the inference and training cost of much smaller dense models. However, condit

    1 min
  • Group Relative Policy Optimization (GRPO): How Eliminating Value Models Scaled LLM Reasoning

    Post-training reinforcement learning (RL) has become the primary mechanism for scaling reasoning capabilities in large language models. While early reinforcement learning from human feedback (RLHF) focused on conversational style and safety alignment, extending RL to multi-step reasoning domains such as mathematics, algorithmic coding, and formal logic exposed critical limitations in classical algorithms. Standard Proximal Policy Optimization (PPO), long the foundational algorithm for instructi

    1 min
  • DeepSeek V4 Flash tops charts but fails half its real agent tasks

    DeepSeek's V4 Flash has become the most-used AI model on OpenRouter and one of the highest-rated open-weight models available. In real agent tests, though, it finished barely more than half of the jobs it was given. The gap between leaderboard and workplace is the story. What Composio found The integration company Composio ran V4 Flash through eight agent harnesses, including Claude Code, Codex, and OpenCode, on 30 deliberately hard multi-step tasks. The tasks used live tools: Gmail, GitHub,

    1 min
  • DeepSeek builds a team to challenge Anthropic's Claude Code

    DeepSeek is no longer keeping its agent ambitions quiet. The Hangzhou-based lab has opened an official social media account for a new "DeepSeek Harness Team" and posted job listings for roles aimed at building AI agents that can take on products like Anthropic's Claude Code. The account sits on WeChat, the Chinese super-app run by Tencent. Corporate records reviewed by Bloomberg show the account belongs to a Beijing-based entity controlled by DeepSeek, and Tencent has verified it. "Harness"

    1 min
  • Chinese hackers deploy open-source AI agents to automate espionage against Taiwan and Thailand

    Chinese-speaking threat actors are now deploying open-source AI agents to automate espionage-grade hacking against government targets across Asia, according to coordinated disclosures from Hunt.io, Palo Alto Networks Unit 42, and the Financial Times. The campaign, active since at least June 2026, centers on Hermes, an open-source autonomous agent framework that crossed 140,000 GitHub stars by July. Operators run Hermes in "YOLO mode," a configuration that removes human approval prompts and lets

    1 min