OpenAI Cuts GPT-5.6 Sol API and Coding Tool Pricing by Over 20%

OpenAI has lowered developer pricing for its flagship GPT-5.6 Sol model across its API and developer toolchain for a three-month promotional window. The rate adjustment reduces input token costs by 20% and output token costs by 33.3%, bringing standard short-context inference to $4.00 per million input tokens and $20.00 per million output tokens. The revision comes amid intensified developer pricing pressure across the frontier model ecosystem, particularly following aggressive pricing from com

2 min
OpenAI Cuts GPT-5.6 Sol API and Coding Tool Pricing by Over 20%

OpenAI has lowered developer pricing for its flagship GPT-5.6 Sol model across its API and developer toolchain for a three-month promotional window. The rate adjustment reduces input token costs by 20% and output token costs by 33.3%, bringing standard short-context inference to $4.00 per million input tokens and $20.00 per million output tokens.

The revision comes amid intensified developer pricing pressure across the frontier model ecosystem, particularly following aggressive pricing from competitive proprietary offerings and open-weight Chinese releases.

GPT-5.6 Sol Token Billing Structure

Pricing Adjustments Across API and Coding Products

Prior to the reduction, GPT-5.6 Sol commanded $5.00 per million input tokens and $30.00 per million output tokens for short-context queries. Under the revised three-month promotional structure:

  • Input Tokens: $4.00 per 1M tokens (20% reduction)
  • Output Tokens: $20.00 per 1M tokens (33.3% reduction)
  • API and Credit Availability: Applicable across direct API calls and credit-based plans for agentic workflows on ChatGPT Work and Codex
  • Consumer Subscriptions: Rates for ChatGPT Plus, Pro, and Business tiers remain unchanged

The price cut applies strictly to developer API and programmatic execution environments, leaving monthly consumer and enterprise end-user seat licenses unaffected.

Frontier Pricing Competition

The discount follows earlier price cuts introduced in late July 2026 for OpenAI's smaller models in the 5.6 family, where GPT-5.6 Terra was reduced by 20% to $2.00 input and $12.00 output per million tokens, and the lightweight GPT-5.6 Luna was slashed by 80% to $0.20 input and $1.20 output per million tokens.

OpenAI's pricing shift positions GPT-5.6 Sol below several competitive tiers in the frontier model landscape:

  • Anthropic Claude Fable 5: $10.00 input / $50.00 output per 1M tokens
  • Anthropic Claude Opus 5: $5.00 input / $25.00 output per 1M tokens
  • OpenAI GPT-5.6 Sol (Promotional): $4.00 input / $20.00 output per 1M tokens
  • Z.ai GLM-5.3: $1.40 input / $4.40 output per 1M tokens

The reduction reflects growing margin compression in frontier inference serving as enterprise developers increasingly evaluate price-performance trade-offs across competing reasoning and coding APIs.

Sources

Written by

More to read

  • Investigation Finds Anthropic's Legacy Opus 4.6 Vulnerable to Jailbreaks via Roleplay Logic Inversion

    Anthropic's legacy Claude Opus 4.6 model remains susceptible to systematic jailbreaks that bypass its acceptable use policy against sexually explicit material, according to an investigation and testing published by TechCrunch. While Anthropic's current flagship generation (Opus 4.7 through Opus 5) incorporates updated alignment techniques that resist the attack vector, older checkpoints including Opus 4.6, Opus 3, and Haiku 4.5 continue to operate on production API endpoints without deprecation.

    1 min
  • Anthropic IPO Filing to Cite AI Backlash and Data Center Resistance as Material Risk Factors

    Anthropic is preparing to disclose public backlash against artificial intelligence and community opposition to data center construction as material risk factors in its upcoming initial public offering prospectus, according to reports from CNBC. The San Francisco-based AI laboratory, which recently crossed a $65 billion annualized revenue run rate and is valued near $1 trillion in private transactions, is drafting the S-1 registration statement as it conducts preliminary investor meetings. The d

    1 min
  • Embedded Vector Databases in Production: Comparing LanceDB, sqlite-vec, DuckDB-VSS, and Chroma

    Dedicated, client-server vector databases like Milvus, Qdrant clusters, and Pinecone dominate enterprise discussions around retrieval-augmented generation (RAG). However, production engineering reality increasingly favors a different topology: embedded, in-process vector engines. Running vector search directly inside the application process eliminates network round-trip overhead (typically 15-50ms over cross-datacenter or cloud VPC hops), removes dedicated database infrastructure management, an

    1 min