The creator of Caffe, the open source deep learning framework he built as a graduate student, and a veteran of Google, Meta, and Alibaba says the next AI bottleneck is not capability.
In a two-hour retrospective out this week, AI scientist 贾扬清 (Jia Yangqing), a veteran of Google, Meta, and Alibaba, recalls the moment, early in his career, when the industry declared AI dead. "At that time the whole industry declared it dead," he says of the 2005–2012 academic winter. "Everyone thought AI was a historical concept." He uses the memory to argue that the field's next bottleneck isn't capability. It's trust.
The episode, a 10th-anniversary crossover on the What's Next 科技早知道 and 声东击西 podcasts hosted by Xu Tao, drops the same week Jia confirmed the launch of Intent Lab, his second new company, after Lepton AI was sold to Nvidia. He frames the new venture as a direct response to the gap he keeps naming in the conversation. A single agent can be smart, he says, but a group of agents is not yet a team. Near the start of the conversation he calls himself essentially a humanities student at heart. The aside turns out to be load-bearing, since the people problem, not the model problem, is the bottleneck he has been circling for twenty years.
Jia's Caffe framework, the open-source deep-learning tool he wrote as a graduate student, became a standard tool for early-2010s computer-vision research and a building block for the image-recognition models that defined the 2013–2022 industry era. The four-era chapter spine of the episode is a wave-meter: 2005–2012, when deep learning was still an academic curiosity; 2013–2022, the industry era that put that work inside Google Brain, Facebook AI Research, and Alibaba's DAMO Academy; 2022–2024, the first-venture stretch capped by Lepton AI's sale to Nvidia; and 2024 onward, which Jia casts as a redefinition phase. Jia's argument is not that capability has stalled. He argues that the conversation about what AI can do has outrun the conversation about what it should do. Around the 50-minute mark he poses the question plainly: can this technology be widely deployed in all fields? The answer, he suggests, is no, not until the trust and accountability layer catches up.
"From completing tasks to gaining trust" is the chapter title at the two-hour mark, and Jia treats it as the load-bearing question for the next phase. The earlier chapters cover agents in the technical sense, where a single model can plan, call tools, and execute. The late chapter is about something messier, which is what an agent system owes a user, an employer, a regulator, a counterparty in a transaction. The mechanism he names is coordination under uncertainty, not raw model performance, and the difference is not academic. A single agent that fails a step can be retried. A team of agents that fails a handoff is a liability.
Jia is also direct about why the next bottleneck is unlikely to be solved inside a large lab. Large companies, he argues, optimize around the capability ceiling because that is what their review cycles, headcount plans, and release cadences reward. The team-of-agents problem, by contrast, is a product and policy problem as much as a research one, which is why he has structured Intent Lab around a small, deliberately mixed team rather than another scaled research org. He is critical, on the record, of the assumption that long-term AI bets inside big companies translate automatically into shipped products. In his telling, big labs optimize for what they can measure, and the next bottleneck is precisely what does not show up in their dashboards.
The episode lands on a specific question, not a sweeping one. Jia asks what kind of people a company needs when AI can already do more of the work, a question that re-frames hiring, evaluation, and trust in one breath. The bet inside Intent Lab is that the answer is small teams that own a complete loop, which means model, product, customer, and accountability in the same room. Whether that structure survives contact with the same enterprise buyers who funded the previous cycle is the watch item. The show notes are the source-of-record in this packet; the full audio is the place to verify the chapter claims and the on-the-record quotes that anchor them.