The reliability gap, not raw capability, is what holds agents back
In an ICML 2026 keynote, Arvind Narayanan argues that the interesting number in AI right now is the one that has not moved. Over the past 24 months, he says, capability rose sharply while reliability, measured across consistency, robustness, calibration, and operational safety, went up only five to ten percentage points. That gap is why fully automated agents keep disappointing while agents that work alongside a person do better. He frames AI as a normal technology whose real effect comes from a slow adaptation stage, closer to how electricity forced factories to be redesigned than to a drop-in replacement for workers, and he claims that stage has barely started.
His software example is concrete. Writing code is only about a third of the job, the "execute" layer; the "decide" layer of understanding requirements and the "deliver" layer of integration and maintenance are not compressed by AI and are arguably growing as the middle shrinks. He points out that software engineering employment kept rising through earlier productivity waves, and draws the familiar parallels to bank tellers after ATMs and to radiologists. He also separates four things people lump together as advanced AI: recursive self-improvement, humanlike intelligence, economically transformative AI, and superintelligence. His prediction is that human effort shifts from building toward evaluation, judgment, and steering, and that good evaluations become a company's real intellectual property. The full talk is here.
Why it matters
If you are planning hiring or automation around coding agents, the binding constraint is reliability and the decide and deliver work around the code, not the model's raw ability to produce it.