AI Models4 articles

AI Models

Articles

  • ByteDance Trains 10 Trillion-Parameter AI Model to Rival Anthropic's Mythos

    ByteDance is pretraining a large model with up to 10 trillion parameters, a scale the Financial Times reports could put it in the same class as Anthropic's most advanced systems. The model, still in early pretraining, would be more than three times the size of Moonshot AI's Kimi K3, currently the largest Chinese model at 2.8 trillion parameters. Three people familiar with the project told the FT the model is in pretraining, a phase that typically lasts three to six months before full training a

    1 min
  • Meta Enters the AI Coding Wars with Muse Code and Muse Spark 1.2

    Meta has launched Muse Code, a terminal-based AI coding agent now in beta, alongside Muse Spark 1.2, a coding-focused update to its proprietary frontier model family. The two releases put Meta in direct competition with Anthropic's Claude Code, OpenAI's Codex, and the growing field of agentic coding tools that have become the default way many developers ship software. Muse Code is installable on macOS or Linux with a single curl command, though it requires a Meta account and billing details. Un

    1 min
  • Moonshot AI's Kimi K3 tops benchmarks as Chinese models reach the frontier

    A two-year-old Beijing startup has produced a model that sits closer to the American frontier than anything from Alphabet, Meta, or SpaceX. Moonshot AI's Kimi K3, released July 16, is a 2.8-trillion-parameter open-weight system that jumped to first place on Arena.ai's Frontend Code leaderboard within hours of release, scoring 1,679 points against 1,631 for Anthropic's Claude Fable 5. It was the first Chinese model ever to top that board. On the Artificial Analysis Intelligence Index, K3 debuted

    1 min
  • Alibaba launches Qwen 3.8-Max, benchmarks fall short of claims

    Alibaba officially launched Qwen 3.8-Max on August 3, calling it the most powerful model in the Qwen family to date. The 2.4 trillion-parameter model supports a one-million-token context window and is multimodal, capable of processing text, images, video, and documents in a single prompt. The company priced Qwen 3.8-Max at $2 per million input tokens and $6 per million output tokens, with implicit caching at $0.25 per million tokens. That positions it well below the pricing of frontier models f

    1 min