Oxford Study Details Chinese Gray-Market Proxies Reselling Claude Tokens at 90% Discounts

An investigation by the Oxford China Policy Lab reveals that Chinese developers routinely access Anthropic's frontier Claude models at discounts between 70% and 90% below list price, bypassing geographical blocks, payment filters, and biometric identity verification through a decentralized network of API proxies known locally as "transfer stations" (中转站). The analysis, authored by Oxford researcher Zilan Qian and published via ChinaTalk, outlines the modular supply chain and economic mechanics

3 min
Oxford Study Details Chinese Gray-Market Proxies Reselling Claude Tokens at 90% Discounts

An investigation by the Oxford China Policy Lab reveals that Chinese developers routinely access Anthropic's frontier Claude models at discounts between 70% and 90% below list price, bypassing geographical blocks, payment filters, and biometric identity verification through a decentralized network of API proxies known locally as "transfer stations" (中转站).

The analysis, authored by Oxford researcher Zilan Qian and published via ChinaTalk, outlines the modular supply chain and economic mechanics powering China's underground token economy. Despite Anthropic operating some of the strictest geographical restrictions in the AI sector, including barring entities with majority ownership in unsupported regions and requiring live selfie verification for select accounts, transfer stations have turned Claude access into a commodity accessible via standard domestic payment rails like WeChat Pay and Alipay.

Technical routing architecture of API transfer stations

The Architecture of API Transfer Stations

Transfer stations function as intermediaries between Chinese software clients and Anthropic's inference endpoints. Developers configure tools such as Claude Code or custom API harnesses by replacing the default base URL with the proxy's endpoint address. The proxy ingests requests locally, forwards them through overseas nodes disguised as legitimate traffic, and returns the generation back to the user without requiring a VPN or foreign credit card.

The operational ecosystem relies on a disaggregated supply chain where individual actors specialize in discrete functions:

  • Upstream Credential Merchants: Supply mass-registered accounts, foreign phone numbers via automated SMS farms, and overseas credit lines.
  • KYC Circumvention Providers: Deploy synthetic identification documents, deepfake facial rendering, or recruited individuals in lower-income regions to pass Anthropic's live selfie verification.
  • Proxy Operators: Manage load balancing, token rate limits, account rotation, and payment gateways.
  • Downstream Resellers: Package access for retail developers on platforms such as Taobao and GitHub.

Because participants operate independently, suspensions by AI providers affect only individual nodes. Operators routinely restore failed endpoints within hours by swapping credentials from intact upstream pools.

The 'One Fish, Three Meals' Pricing Engine

Official Claude Opus tokens remain among the most expensive in commercial AI serving. Transfer stations frequently offer token access at rates near 1 RMB per $1 of face-value compute (roughly a 85% to 90% reduction). According to the report, operators achieve these sub-cost economics through three primary mechanisms:

  1. Quota and Discount Arbitrage: Operators pool free registration credits ($5 per tier), exploit educational and enterprise discounts, and split flat-rate $200 subscription tiers across multiple concurrent users through custom rate limiters. Fraudulently funded accounts also enter pools at zero acquisition cost.
  2. Model Substitution ("Diluting"): Proxies silently reroute requests for expensive frontier models to smaller variants or domestic open-weight models, returning outputs under the requested label. An independent audit by Germany's CISPA Helmholtz Center for Information Security across 17 commercial proxies found that a proxied endpoint labeled Gemini-2.5 achieved only 37.00% on medical evaluation benchmarks compared to 83.82% on the genuine API. Frequent account cycling in proxies also shatters KV cache prefix continuity, shifting hidden token overhead to end users.
  3. Telemetry and Log Harvesting: The proxy layer retains complete visibility over prompt inputs, tool-use calls, full reasoning traces, and repository contexts generated by agentic workflows. These interaction logs provide high-value training corpora for supervised fine-tuning and model distillation. Datasets containing unattributed Claude reasoning chains have already appeared on public hubs such as Hugging Face.

Implications for AI Safety and Access Controls

The findings highlight systemic limits in perimeter-based AI governance. Closed-weight lab safety infrastructure relies on direct inference visibility to identify coordinated abuse, automated vulnerability scanning, or biological threat research. Systems such as Anthropic's Clio depend on cross-conversation and cross-account pattern recognition to detect distributed attacks.

When traffic traverses multi-tenant transfer stations, model providers observe only the proxy node's IP address and rotated credentials. Multi-stage inquiries can be split across disjointed accounts, blinding provider-level anomaly detectors. Furthermore, the downstream identity markets created to bypass biometric KYC feed into traditional identity fraud, deepfake generation, and credential theft outside the AI domain.

As Western frontier labs intensify joint enforcement efforts to curb unauthorized distillation and proxy access, the persistence of the transfer station economy demonstrates that client-side geoblocking and KYC checks remain insufficient against modular, economically motivated intermediary networks.

Sources

Written by

More to read

  • Hybrid SSM-Transformer Architectures: How Interleaving Attention and Recurrence Solves the State-Retrieval Trade-Off

    Hybrid SSM-Transformer Architectures: How Interleaving Attention and Recurrence Solves the State-Retrieval Trade-Off Autoregressive language models face a fundamental tension between inference efficiency and long-context retrieval capacity. Pure Transformer architectures scale quadratic computational complexity during sequence prefill and linear key-value (KV) cache memory consumption during autoregressive token generation. Conversely, pure State Space Models (SSMs) and linear recurrent neural

    1 min
  • Study: Why Labor-Saving LLMs Incline Scientists to Do More Work Less Well

    A theoretical study published by researchers from Princeton University, the University of Washington, and collaborating institutions models how large language models alter researchers' time allocation across projects. The authors find that by reducing time friction across different stages of the research lifecycle, AI assistants increase the opportunity cost of researcher time, creating economic incentives to publish a higher volume of less thoroughly refined papers. The paper, titled The unint

    1 min
  • Memory Shortage Drives Nvidia AI Server Prices Up Over 15%

    Nvidia has notified major customers that prices for server systems containing its artificial intelligence accelerators are increasing by more than 15% in many configurations, according to reports from Bloomberg and Fortune. The price adjustments stem from severe supply constraints and rising costs across dynamic random-access memory (DRAM) and high-bandwidth memory (HBM) modules. The price increases will apply to server systems scheduled for delivery starting in early 2027, covering platforms p

    1 min