More on the topic...
Generating detailed summary...
Failed to generate summary. Please try again.
Investors who despair that every AI startup is just “thin wrappers” around giant models miss how much of real engineering still resists measurement. Early coding benchmarks gave us Devin in 2024—13 percent task success—and by late 2025 agents hit high eighties and joined Goldman Sachs and the U.S. Army. Yet, MIT’s Mert Demirer and coauthors found that although code output rose 180 percent, shipped code climbed only 30 percent. Tests catch simple correctness, but decades-old systems with hidden dependencies and fragile deployment pipelines need human judgment. Real reliability comes from years of live traffic, not a green checkmark on a leaderboard.
Automation isn’t just a smarter model. It’s the tech, the product design, the workflows, and the organization all moving in sync. Companies hire CEOs as much for people skills as for analysis because rolling out an AI rebuild takes years, not quarters. Tasks that become easily verifiable quickly drop into commodity: open and distilled models race on price. At the same time, labs fold their own tooling—retrieval, routing, specialized reasoning—into their core weights, chasing the “absorption frontier.” The winner there is whoever can own custom workflows on private data, not whoever built the most generic benchmark-beating model.
That private correctness—the part you can’t check without access to a company’s systems—and the walls around it define the real AI moat. You can’t train on a bank’s production environment or sign off on liability just by tweaking parameters. You need security reviews, integration contracts, and doctors’ daily trust. Firms that win this corner build translation layers: they integrate models, hand them the right tools, and train users to embed them in daily routines. Maintenance and domain-specialized engineering keep that advantage alive, long after any public benchmark has been absorbed.
Questions about this article
No questions yet.