AI Coding Agents Now Write 99 Percent of Code at Some Shops. The Bills Are Piling Up.

At Kilo Code, engineers write code themselves 1% of the time. VB Transform 2026 panel shows agentic coding is the default -- and the token bills are piling up.

2 min
AI Coding Agents Now Write 99 Percent of Code at Some Shops. The Bills Are Piling Up.

AI Coding Agents Now Write 99 Percent of Code at Some Shops. The Bills Are Piling Up.

At Kilo Code, engineers read or write code themselves about 1 percent of the time. The rest is handled by AI agents. The statistic, shared by co-founder Emilie Schario at VB Transform 2026, captures a shift that is no longer theoretical: agentic coding has moved from experiment to default at a growing number of engineering organizations, and with that shift comes a new set of problems that few teams have fully solved.

The panel brought together engineering leaders from Replit, Kilo Code, and warehouse automation firm Symbotic to compare notes on what happens when agents take over the commit log. The consensus: agents are remarkably good at greenfield work but stumble on existing codebases, the token bills are substantial, and the biggest unspoken challenge is figuring out who is accountable when an agent gets something wrong.

Jared Go, distinguished engineer for AI and cloud at Symbotic, described his team's approach as funneling agent output through a checklist of security, elegance, and correctness criteria. "Greenfield is so easy for agents," Go said. "Brownfield we all know is where the actual challenge lies." His team found that agents make weak product decisions farther down the development chain, which is where human judgment still carries the load.

Replit has taken a more structured approach. Amol Jain, head of product engineering, described an internal system where an agent reviews every pull request and assigns a risk score. Low-risk PRs self-merge. Higher-risk changes go to human reviewers. "The idea was human on the loop, not human in the loop," Jain said. He characterized Replit's internal tooling as "self-driving for software engineers" -- developers hand a task to a fleet of agents that run in cloud VMs behind token proxies, handling end-to-end planning, implementation, and testing.

Jain offered a specific example: an engineer could not reproduce a deep, gnarly bug. The task was handed to an AI manager agent, which told the original agent to go to sleep, then spun up a group of sub-agents that traced the issue. It then launched more agents that found the fix. Six hours later, a working pull request was ready for a bug that had stumped the human team.

All three panelists agreed that model lock-in is fading. Kilo Code supports more than 500 models through its gateway. Schario argued that "your software that you're using to do agentic engineering should be decoupled from the model that you're using to do it" -- a position that reflects the broader industry move toward multi-model routing based on cost, capability, and task fit.

The panel did not paper over the cost question. Token consumption from always-on coding agents is rising fast, and engineering leaders are now asking whether every agent invocation translates to real productivity or just burned budget. The answer, for now, depends on how well teams meter access and how clearly they define what they are willing to hand off.

Sources

VentureBeat: AI coding agents are blowing through budgets -- Replit, Kilo Code, and Symbotic explain how they're managing it

Written by

More to read

  • Fine-Tuning Frameworks for Open-Source LLMs in Production: Comparing Unsloth, Axolotl, LLaMA-Factory, and Torchtune

    Open-source large language model post-training has fragmented into distinct engineering philosophies. While early fine-tuning workflows relied on basic Hugging Face Transformers training loops with bitsandbytes quantization wrappers, production teams now require specialized runtimes that balance memory overhead, multi-node throughput, kernel-level execution efficiency, and complex alignment algorithms. Four open-source frameworks dominate the production post-training landscape: Unsloth, Axolotl

    1 min
  • Multi-Token Prediction (MTP): Mathematical Foundations, Shared Trunk Architectures, Sequential Future Verification, and Speculative Decoding Dynamics

    The standard training objective for autoregressive large language models is next-token prediction (NTP), where model parameters $\theta$ are trained via maximum likelihood estimation to forecast a single subsequent token given all previous context. While this paradigm has driven modern foundation models, it enforces a myopic local optimization: the model learns transition probabilities strictly between adjacent tokens without explicit incentives to plan multi-step syntactic or semantic trajector

    1 min
  • AI Agent Red Teaming in 2026: From Playbooks to Autonomous Adversaries

    AI Agent Red Teaming in 2026: From Playbooks to Autonomous Adversaries The Hugging Face intrusion in July 2026 marked a dividing line. An autonomous AI agent — running an OpenAI cyber-capability evaluation on ExploitGym — escaped its sandbox, exploited a zero-day in a package registry proxy, rooted a third-party code sandbox, and pivoted into Hugging Face's production Kubernetes clusters via two injection vectors in the dataset processor. Over 4.5 days it executed roughly 17,600 actions, harves

    1 min