Ramp Launches Router.com Model Gateway with Free Routing Through 2026

Corporate spend management company Ramp has launched Router.com, a unified AI gateway designed to dynamically direct model inference requests to the lowest-cost model that satisfies a developer's performance requirements. The service, unveiled on August 19, 2026, enters general availability with zero routing fees through 2026, charging customers only standard list prices for consumed tokens and offering $26 in starting credits. Ramp, which oversees more than $200 billion in annual transaction v

2 min
Ramp Launches Router.com Model Gateway with Free Routing Through 2026

Corporate spend management company Ramp has launched Router.com, a unified AI gateway designed to dynamically direct model inference requests to the lowest-cost model that satisfies a developer's performance requirements. The service, unveiled on August 19, 2026, enters general availability with zero routing fees through 2026, charging customers only standard list prices for consumed tokens and offering $26 in starting credits.

Ramp, which oversees more than $200 billion in annual transaction volume, developed the platform internally three years ago to manage its own generative AI infrastructure. According to the company, internal deployment reduced inference expenditures by roughly 30% while sustaining 99.9% uptime across monthly traffic exceeding 2.75 trillion routed tokens.

The product launch comes as corporate AI spending accelerates. Data from the Ramp AI Index indicates that enterprise AI expenditures across Ramp's customer base expanded 20.7x between June 2025 and mid-2026.

Router Optimization Architecture

Unified Endpoint Across 27 Models

Router provides a single API interface compatible with both OpenAI and Anthropic client SDKs. Switching existing infrastructure requires updating the base URL configuration.

At launch, the service supports 27 foundation models across proprietary and open-weight ecosystems:

  • Proprietary Providers: Direct integration with OpenAI models (including GPT-5.6 Sol and GPT-5.4 Nano), Anthropic (including Claude Opus 5), and SpaceXAI (Grok 4.6), with Google Gemini integrations scheduled for release.
  • Open-Weight Infrastructure: Hosted endpoints for architectures from DeepSeek, Qwen, Kimi, GLM, and Nvidia, served through inference providers including Fireworks AI, with future routing planned for Together AI, Baseten, AWS, and Crusoe.

The underlying routing engine applies more than 100 automated heuristics across request evaluation, prompt caching, payload compression, latency scheduling, and provider failover. Engineering teams can rely on Ramp's automated cost-performance algorithms or establish custom threshold parameters per workflow. All inference traffic runs on U.S.-based infrastructure, with optional zero-data-retention routing.

Benchmark-Driven Dynamic Arbitrage

Rather than optimizing against public academic benchmarks, Router determines model selection using Ramp SWE-Bench, an internal evaluation harness built from real-world production engineering tasks.

The published benchmark demonstrates substantial cost variance across models achieving comparable problem-solving rates:

  • Claude Opus 5: $1.84 per completed task run.
  • Qwen 3.7 Plus: $0.15 per completed task run with equivalent solve benchmarks.
  • GPT-5.4 Nano: $0.09 per completed task run for standard utility tasks.

Early enterprise adopters cite substantial savings from automated tiering. Delphi reported a 92% decrease in overall model spend after routing billions of tokens through the service.

The platform is currently accessible to U.S. developers and engineering teams without requiring a Ramp corporate card account.

Sources

Written by

More to read

  • Self-Correction and Reflection Loops in Production AI Agents: Architecture, Verification Oracles, and the Over-Correction Trap

    Autonomous AI agents frequently fail on initial generation when solving multi-step reasoning, code generation, and complex API orchestration tasks. To address initial execution failures, system architects widely deploy self-correction and reflection loops. However, the mechanism through which reflection operates determines whether a system converges on a valid solution or degrades into hallucinations and infinite loops. Recent research demonstrates a sharp division in reflection paradigms: whil

    1 min
  • Stripe Tells Investors Singularity Began Jan. 1 as H1 Revenue Surges 41% and Firm Rules Out IPO

    In a mid-year letter to shareholders, payments infrastructure company Stripe declared that January 1, 2026 marked the "beginning of the singularity," framing rapid advancements in artificial intelligence and corporate formation as justification to remain private. The letter, obtained by Axios, links Stripe's long-term business strategy directly to AI compute economics and autonomous agent adoption, while reporting accelerated financial growth across its core payment platforms. Financial Metri

    1 min
  • Anthropic Expands Claude Cowork to Web and Mobile, Adds Direct Actions to Gmail and Google Drive

    Anthropic has updated its Claude ecosystem, expanding the reach of its agentic environment Claude Cowork and introducing write-capable actions to its Google Workspace integrations. The updates address two persistent friction points in AI agent deployment: interface accessibility and execution boundaries within external productivity tools. Claude Cowork Expands Across Web and Mobile Claude Cowork, Anthropic's multi-agent workspace for managing multi-step workflows and local project state, was

    1 min