More on the topic...
Generating detailed summary...
Failed to generate summary. Please try again.
A team from UC San Diego and Cornell watched 13 pro developers code live with AI agents and surveyed 99 more, all with at least three years’ experience (some up to 25). Contrary to hype, these engineers don’t “vibe code”—they map out architecture, list constraints and edge cases, then hand the agent a focused task. Every change gets a careful diff review, and they limit the agent’s scope. If a feature spans multiple systems or needs unclear requirements, humans step in.
The study also highlights real-world failure rates. In one trial, seasoned open-source maintainers using AI were 19% slower. Another agent hooked into an issue tracker produced merged pull requests only 8% of the time—meaning a 92% failure rate. Quality tanks when developers loosen their grip, the paper notes. Engineers only feel good about AI tools when they stay firmly in control.
What serious teams are doing now: strict review processes, tight scoping of agent tasks and clear mental models of the AI’s capabilities. Those “run dozens of agents hands-off” demos on X might look cool, but they don’t reflect how production code actually ships. Next time someone brags about letting an AI whip up their SaaS in a weekend, ask how much of that code they’ve really inspected.
Questions about this article
No questions yet.