EPISODE 2026-08-24

AI:AM LIVE — August 24, 2026 — Arm's Mohamed Awad on Why Agents Don't Sleep and What That Does to the CPU, and Shenzhen Open Innovation Lab's David Li on Why China's AI Conversation Is Boring on Purpose

Two views of the layer under the model — one from inside the architecture that just stopped being neutral, one from a system that never had IP protection to lose. Mohamed Awad, EVP of Cloud AI at Arm, explained why a company that spent thirty-five years licensing designs now sells its own silicon, how the Meta partnership traces back to AWS's 2019 Graviton launch, and what actually changes in a CPU built for agents rather than people: agents don't sleep, so the CPU stops being the thing you wait on and becomes the coordination layer feeding accelerators that never idle. Asked where the GPU-to-CPU ratio settles, he declined to guess. David Li, founder of Shenzhen Open Innovation Lab and co-founder of China's first hackerspace, described a manufacturing culture with no next-week mentality, industrial robot hardware down near $3,000 with the bottleneck moved to the engineers who can install it, and an AI discourse he characterised as overwhelmingly practical — nobody in China, he said, notices a domestic model launch unless it crashes the Nasdaq. The hosts opened on Anthropic's Fable 5 holding flat at roughly 10–15% of business AI spend and what that says about whether new frontier models drive incremental demand, and closed disagreeing about whether Chinese labs feel any urgency from the summer's rogue-agent incidents.

▶ Full show on YouTube𝕏 Live broadcast

Monday's show ran two interviews about the same question from opposite ends of the world: where the leverage sits once the model itself is a commodity. Arm's Mohamed Awad answered from underneath the accelerator — the CPU as the thing coordinating agents that never stop asking for work. Shenzhen Open Innovation Lab's David Li answered from a manufacturing base that treats models as components to bolt onto products, and finds the whole race framing faintly puzzling.

Both interviews ran long — Awad by twenty-two minutes, Li by thirty-five — and the hosts used a thirty-seven-minute close to argue about the thing neither guest quite settled: whether the summer's rogue-agent incidents have changed anyone's behaviour, in either country.

The rundown

  1. 4:29Opening26 min
    Opening: The Studio That Was Asked For, and Fable 5's Flat LinePrakash walks through the weekend's studio build — speaker-following camera cuts reusing the live-transcription voice-activity detection, and a LiveKit dial-in fallback for guests — then puts up the Ramp AI Index showing Fable 5 holding at roughly 10–15% of business AI spend as Opus 5 creeps up, one factor behind the morning's AI selloff. Nathan offers his own model-routing policy as a partial explanation and argues agentic search has already closed much of the gap he expected continual learning to fill.
    Open segment on YouTube ↗

    The hosts opened with a rundown of what Prakash built into the AI:AM studio over the weekend. The stream now has dynamic viewpoints — the broadcast automatically cuts to whoever is speaking, full-face, for the audience — built on top of the show's existing live-transcription pipeline. Because each host and guest already has their own transcription stream, the same voice-activity detection that powers per-person transcripts turned out to be reusable to drive the camera switching. Prakash also added a backup US dial-in phone number, via LiveKit, for guests who can't get onto the stream through the browser — a direct response to a guest's PR contact asking for exactly that option after a technical hiccup the prior week.

    Nathan then asked Prakash to compare the studio, built almost entirely by Prakash vibe-coding with AI over a few months, to the closest thing available to buy — Ecamm was the nearest comparison Prakash could name. Prakash argued the gap between prosumer streaming software and a sophisticated, multimillion-dollar studio back end isn't really a technology gap; it's a user-interface and feature-engineering gap that historically required large, well-funded teams to close. Now, he said, you can just ask the model for a feature, test it, and get it — which is why the show can run things like live-updated on-screen headlines through 'Q,' their AI producer, which is essentially a context window fed the full running transcript, with no human production staff behind it. Nathan called it a testament to how far autonomous, long-horizon agent execution has come, noting the show runs on 'you and me and agents all the way down.'

    Prakash then summarized the weekend's biggest AI story: Anthropic's Fable 5 has shown limited additional uptake, holding at roughly 10–15% of business AI spending on a 7-day moving average per the Ramp AI Index, even as Opus 5's share has crept up. Prakash tied this directly to Monday morning's broader AI-stock selloff, framing it as one factor spooking the market about whether the newest frontier models are actually driving incremental spend. He flagged one caveat often raised in Fable 5's defense: it lacks a zero-data-retention option, which is a hard disqualifier for many enterprise buyers who can't have customer PII flow through externally retained storage — though he noted the market doesn't appear to be pricing that nuance in this morning.

    Nathan offered his own model-routing breakdown as a plausible explanation for the flat Fable 5 usage. He described hitting Fable 5's usage limits quickly early on before setting up an explicit division-of-labor policy in his CLAUDE.md (drawn partly from a well-liked tweet he had Claude adapt), which has since pushed more work onto sub-agents and cheaper models. In his own podcast-production pipeline, he now sends transcript cleanup to Sonnet or Haiku, and finds Opus 5 functionally as good as Fable for editing, image-prompting, and agentic execution generally — reliable and faster. The one place he still sees a clear gap is song lyric writing, where Fable's output feels more inspired and layered; he characterized the difference as editorial taste, which is where Fable earns its higher price, rather than raw execution reliability. He also noted growing delegation to GPT-5.6 Sol.

    The conversation then pivoted to whether a new frontier model release is imminent. Prakash cited Roon's observation that it's been about six months since 'Mythos Preview' shipped, and argued no public model since has clearly surpassed it — suggesting Fable 5, Opus 5, and GPT-5.6 Sol may all be a step down from what's already been demonstrated internally. He noted Mythos Preview access has expanded to partners like Palo Alto Networks, who reportedly offer bug-scanning services finding hundreds of bugs for around $10K that might have cost millions before. He also referenced an Anthropic safety report showing an internal model whose AECI capabilities index rose only modestly (about 1.5 points) but whose internal 'co-bench' score — a measure of AI-research capability — came in 10–15 points higher than expected, which Nathan called a meaningful closing of the gap. Both flagged rumors of releases this week (from Martin Casado at a16z and from Pietro Schirano) but Nathan was skeptical anything genuinely frontier-pushing ships this week, given OpenAI's stated RL pause and the still-unresolved safety investigation; he also downgraded the buzzy '000x Alpha' stealth model on OpenRouter, saying the emerging consensus pegs it near GLM-5.3, a notch below the top US models.

    They closed on continual learning versus long-context/search. Nathan argued that Claude's agentic search, combined with his own deep-context and situational-awareness tooling, has already closed much of the gap he once expected true continual learning to fill — he said he'd been skeptical Anthropic's search-first bet would work this well. He speculated the biggest remaining upside from a real continual-learning breakthrough would be deployment ease rather than raw capability. Prakash countered with a framing of two converging curves — continual learning improving and context windows/RAG improving — that may make the difference hard to feel experientially, while noting no amount of continual learning can ever encompass the entire world, so search and retrieval remain necessary regardless. The segment ended with Prakash moving to bring in the show's first guest from the green room.

    Fable has really just stayed pretty stable — and that's been one of the reasons why the market is dropping this morning for all AI stocks.

    I think it is honestly an amazing testament to the power of AI that we've been able to do this with no staff.

    It really feels like it's editorial taste where Fable earns its higher price.

    Prakash rebuilt the studio over the weekend so the broadcast follows whoever is speaking. Prakash described adding dynamic viewpoints — the stream cuts automatically to whoever is talking, full-face — built on top of the show's existing live-transcription pipeline. Because each host and guest already has a separate transcription stream, the voice-activity detection driving per-person transcripts turned out to be reusable to drive camera switching. He also added a US dial-in number via LiveKit for guests who cannot join through the browser, a direct response to a guest's PR contact asking for that option after a technical problem the previous week.

    Asked what he would have bought instead, Prakash argued the gap was never a technology gap. Nathan asked Prakash to compare the studio, built almost entirely by vibe-coding with AI over a few months, to the closest commercial product; Ecamm was the nearest thing Prakash could name. Prakash argued the distance between prosumer streaming software and a serious studio back end is a user-interface and feature-engineering gap that historically needed a large, well-funded team to close, and that you can now ask a model for a feature and test it. He pointed at the show's own live on-screen headlines, produced by 'Q,' an AI producer that is essentially a context window fed the running transcript, with no human production staff.

    Fable 5 is holding flat at roughly 10–15% of business AI spend, and Prakash tied that to the morning's AI selloff. Prakash put up the Ramp AI Index on a seven-day moving average showing Anthropic's Fable 5 drawing limited additional uptake even as Opus 5's share crept up, and framed it as one factor spooking the market about whether the newest frontier models drive incremental spend. He flagged the caveat usually raised in Fable 5's defence — that it lacks a zero-data-retention option, a hard disqualifier for enterprise buyers who cannot let customer PII flow through externally retained storage — while noting the market did not appear to be pricing that nuance.

    Nathan's counter-explanation: he has been routing work away from the frontier model on purpose. Nathan described hitting Fable 5's usage limits quickly, then writing an explicit division-of-labour policy into his CLAUDE.md that pushed more work onto sub-agents and cheaper models. In his own podcast production he now sends transcript cleanup to Sonnet or Haiku and finds Opus 5 functionally as good as Fable for editing, image prompting and agentic execution — reliable and faster. The one place he still sees a clear gap is song lyrics, where Fable's output feels more inspired; he characterised that as editorial taste rather than execution reliability, and noted growing delegation to GPT-5.6 Sol.

    Is anything genuinely frontier-pushing about to ship? Both hosts doubted it. Prakash cited Roon's observation that roughly six months have passed since Mythos Preview and argued no public model since has clearly surpassed it, suggesting Fable 5, Opus 5 and GPT-5.6 Sol may all sit a step below what has already been demonstrated internally. He noted an Anthropic safety report describing an internal model whose AECI capabilities index rose only about 1.5 points but whose internal co-bench score — a measure of AI-research capability — came in 10–15 points higher than expected. Both flagged release rumours from Martin Casado at a16z and Pietro Schirano, but Nathan was sceptical anything major ships given OpenAI's stated RL pause, and downgraded the stealth '000x Alpha' model on OpenRouter as closer to GLM-5.3.

    Nathan: agentic search has already closed much of the gap he expected continual learning to fill. Nathan said Claude's agentic search, combined with his own deep-context and situational-awareness tooling, had done more than he expected, admitting he had been sceptical that Anthropic's search-first bet would work this well, and speculated the largest remaining upside from a genuine continual-learning breakthrough would be ease of deployment rather than raw capability. Prakash framed it as two converging curves — continual learning improving while context windows and retrieval improve — that may make the difference hard to feel, while noting no amount of continual learning can encompass the entire world, so retrieval remains necessary regardless.

    Lightly edited · timestamps jump to YouTube
    4:47

    Prakash: Good morning. It's Monday, August 24th, 9:04 AM. Sometimes we have a little bit of audio issues, but never a dull moment. Good morning, Nathan. How are you this morning?

    5:04

    Nathan Labenz: I'm great. Good morning to you. Maybe tell us what you've done over the weekend to the studio that has improved it, but maybe also left a couple of gremlins in the machine for us—

    5:16

    Prakash: Mhmm.

    5:17

    Nathan Labenz: Find and fix.

    5:19

    Prakash: Indeed. Okay, so the first thing that I've done is we now have dynamic viewpoints — when someone is talking, it's going to switch to their face. So we can try it out now.

    5:35

    Nathan Labenz: Alright, let's see — is it switching to my face? I still see both of us in our view together, but it's switching for—

    5:42

    Prakash: Yeah.

    5:43

    Nathan Labenz: The audience? Okay, interesting.

    5:44

    Prakash: It's switching for the audience, so it's a dynamic switch — when someone is speaking, their face fills the screen, full face. It's funny, the thing is it's all built one on top of the other. We have the live transcription first — we start off with a live transcription, and then that live transcription, we ended up doing transcription by person, so each person has their own transcription stream. And it turns out the voice activity detection for the transcription is the same voice activity detection you can use for the dynamic switching — the machine knows who's talking and knows enough to switch to that person.

    6:33

    Nathan Labenz: Pretty cool. And — we had a little problem last week where one of our guests couldn't manage to get in, and their PR person emailed us in the background and said, hey, is there a way they can call in with a phone number to join? And we were like, maybe we should add that.

    I think it's really interesting to just take a step back and look at all of it — you've done ninety-plus percent of the vibe coding here. I've contributed a little bit to the repo, but I don't know that there's a product out there on the market that has all the features, or anywhere particularly close to everything you've already been able to vibe-code your way to over a period of a few months. How would you — I mean, you probably shopped around a lot more for this than I did — how would you compare what we've created, mostly you've created, to the best thing you could buy if you wanted to go run a build-versus-buy analysis? And I wonder what that would cost if we went to try to buy this software equivalent on the market.

    8:10

    Prakash: So I think the closest that we had was Ecamm. And I think where things start breaking down is that there's a gap between something they can sell to prosumers and something you use in a sophisticated, multimillion-dollar studio. And basically that gap is not a technology gap — it's really a user-interface gap, and that has to be built by engineers speaking to users. The underlying technology might be the same, but the number of features you need, the complexity of those features, and how they have to be engineered.

    So to engineers, there's not that much difference between prosumer software and a sophisticated back end for a big studio. But the number of features you have to add and the amount of work you have to put in make it non-economic, I think, for that to be done at a small scale. You need million-dollar projects for it to be fully engineered. So that's really the differentiating factor now — if you want a feature, you can just ask the machine for it, and the machine will build it, and you have to do a little bit of testing. And as long as you're willing to do that testing,

    you can get exactly what you want. And that makes a big difference, because you can start doing stuff that, in the previous era, would just be impossible — you just can't do it at all. And some of that is here — we have a bunch of live stuff going on. Some of these things, we change our headlines live, and you'd normally need a producer to sit down and type those things — you'd need a production staff on the back end. So we have what we call Q, our producer, which is really just a context window running the entire transcript of our discussions. And that's enough.

    10:34

    Nathan Labenz: Yeah, it's pretty cool. I think it's honestly an amazing testament to the power of AI that we've been able to do this with no staff. Right? I mean, the fact that it's just you and me and agents all the way down is a pretty incredible reality that would have been, I think, extremely difficult even a year ago. The amount of autonomous agent, long-horizon success that we rely on daily to get the thing off the ground and make sure people are aware that we're going live, and all these little touch points — it's pretty amazing that it works as well as it does, even though we do still have a few glitches here and there.

    11:22

    Prakash: Indeed. Let me maybe just summarize what has happened over the weekend, just to have a sense. So the first thing, I think, is this dropped over the weekend and everyone's talking about it: Anthropic's best model, Fable 5, has drawn limited additional uptake — it's currently holding at about 10 to 15 percent of business AI spending, a 7-day moving average, using the Ramp AI Index.

    And you can see Opus 5 kind of crept in there, but Fable has really just stayed pretty stable. And this has been one of the reasons the market is dropping this morning for all AI stocks — that, and many other reasons, but essentially this is one of the things that's scared the market a little bit, whether the newer models are actually lucrative and are drawing usage. Some people have also pointed out that this is a little bit of an unfair comparison, because Fable 5 does not have zero data retention. And zero data retention basically means that if you're a company using the model, the model provider is not retaining any—

    data, including personally identifiable data, etcetera. And for many companies, if you cannot provide a zero-data-retention policy, it's a no-go — the entire thing is a no-go. You just can't have your customer names and addresses and all of that stuff circulating through an external storage space. So the market doesn't care right now — it's looking for ways to manage the AI story, and this has been one of the reasons.

    13:41

    Nathan Labenz: Yeah, that's probably the best explanation I could come up with as well. I think there's also — when we first got Fable, in that little blip at the beginning of that chart — one of the things we talked about was how everybody was going to have to start thinking more carefully about the division of labor between models, and trying to make sure you're using the right model for the task, because it's pretty expensive, and you do hit your limit even on your Claude Max plan relatively quickly if you just throw everything at Fable. I've done that, and I think probably a lot of people have too — it's so easy to do. You can just ask Fable, write up a division-of-labor plan, and it stretches the budget

    quite a bit farther to do that. So I have to assume that's a significant part of it as well. There are a lot of things for which Opus 5 is functionally just as good. When I do — I have this whole pack of skills to produce the podcast, and it all ladders up to one command, which is produce episode. I'll just give it a link to the recording, and the produce-episode skill gets the rough transcript and polishes it into a better transcript. That's something we can send down to Sonnet, or often even Haiku, to just clean up some text and fix the artifacts that came out of the raw transcription engine. Then there's a bunch

    of editing, and we're having Claude go back and forth with the Underlord agent via Descript. And then there's art creation, which involves prompting image-generation models. And then there's the song-lyric-writing project, which is increasingly near and dear to my heart, although sometimes also the bane of my existence, depending on how well it's going on any given episode. And mostly, to be honest, Opus seems to be just as good. Fable really stands out most of all to me in writing the lyrics to the songs — that's where I sense an obvious difference. It just feels like those lyrics come back from Fable more inspired, more layered, richer with meaning. They're just better.

    It just feels like — wow, this — at least not always, of course, but often enough — feels like you really wrote what could be a hit song here, and I don't get that as much from Opus. But when it's things like executing a ton of commands and stitching together a bunch of video clips, or editing out the rough parts of clips to get down to a tighter edit, there I don't see very much of a difference. It really feels like it's editorial taste where Fable earns its higher price. And on the agentic execution — blocking and tackling — Opus, I should say, is reliable enough that

    I don't get a lot of extra value from Fable there. And Opus is also faster, so there's actually some upside to the cheaper version. I'm kind of surprised it hasn't gone a little further than that. But that probably does reflect my breakdown, roughly speaking — Opus is definitely still the big workhorse, and indeed you see Opus dominating the chart there, so I'd say I'm pretty consistent with that. There's also a lot more delegation to 5.6 Sol now too — that's a whole other aspect of it.

    17:29

    Prakash: So, as much as possible—

    17:30

    Nathan Labenz: I'm trying to make sure I'm getting used out of my GPT Max plan too.

    17:35

    Prakash: So are you seeing — are you actually seeing Fable use other models, or are you instructing, as part of your kind of personalization, that it should use other models when possible?

    17:54

    Nathan Labenz: I occasionally tell it explicitly what to do, but I do have a standing part of my CLAUDE.md that we set up a while back and then updated when Opus 5 came online, to basically say, think it through. I think I saw a tweet from somebody who had a good version of this — I forget exactly who it was, but one of your prolific builders who's always posting tips and tricks. I saw a well-liked tweet that seemed to make sense in terms of starting to define a division of labor, and I just pointed Claude at it and said, here's somebody who's got some good ideas, let's steal from those and update CLAUDE.md accordingly.

    And it did most of the work — I reviewed it. I don't really track super closely how often it's doing that, but there's definitely a trend toward using more and more sub-agents. Increasingly, when I come back to a tab a couple minutes later, the status is waiting for one sub-agent, waiting for two processes. I see a lot more sub-agent structure — I just don't always track exactly what model it's sending things off to. When I open my Claude usage, I'm really not hitting the Fable limit too often. Initially I was hitting it a lot more — this division of labor has

    spread things out much more effectively, to where I'm not often — not never, but not often — hitting the Fable 5-hour limit. At the beginning, I was hitting it constantly.

    19:36

    Prakash: I suspect — well, let's take a little bit of a segue and talk about whether we're going to see new models. Because Roon just said it's been six months, roughly, since Mythos, which came out around February-ish — they announced Mythos Preview. It's not clear to me that any model since Mythos Preview has been better than Mythos Preview. I suspect Fable 5 and Opus 5 were both downgrades from Mythos Preview, and GPT-5.6 Sol is nice, but I suspect it's also not as good as Mythos Preview.

    We wouldn't actually fully know, because we didn't have access to it. Mythos Preview itself has expanded access now — I think to Palo Alto Networks. Palo Alto Networks is offering a service where they deploy Mythos Preview on your behalf to detect bugs in your codebase. I've heard you can detect hundreds of bugs for around ten thousand dollars — a Mythos-Preview-like scan that might have taken millions of dollars before. But it's been six months. And the real question is—

    what are we missing here?

    21:15

    Nathan Labenz: Wait, say again — what are we missing?

    21:17

    Prakash: Yeah, what are we missing? In the six months, we've had models released that have been less good than, I think, Mythos Preview. And we have a safety report from Anthropic saying they have an internal model too, and they kind of sandbagged — the AECI capabilities index — they're only 1.5 points more. But on their internal co-bench, which measures how good an AI researcher is, it was a good ten to fifteen points higher than we thought.

    21:48

    Nathan Labenz: That was a good spot from you last week — I hadn't seen that in the report, and that is a meaningful closing of the gap.

    21:54

    Prakash: Yeah, it is a meaningful closing of the gap. So are we going to get another model release? Everyone seems to think this is going to be the week — maybe a GPT-5.7, maybe a 5.1 Opus, maybe a 5.1 Fable. There are rumors flying, especially with the companies heading toward IPO and these market downturns, about whether they have better stuff. There are two rumors — one from Martin Casado at a16z, saying he's seen some models that are going to blow your socks off. And another one

    from Pietro Schirano, who has a startup but was also on earlier testing teams, so he's had prior access to some models before — again saying the entire ontology is going to be different. I have no idea what that means.

    23:03

    Nathan Labenz: Yeah, I mean, there's a lot of talk of potentially a breakthrough in continual learning as well, which I, for one, have been watching out for for a long time. Though it's funny — a lot of this stuff kind of hides in plain sight. We've seen all these reports over the last couple of months about how effective the currently undeployed, internal-only models can be, for better and for worse. So I kind of doubt that what we'll see, if we see anything truly mind-blowing coming to the frontier this week, will be all that different from what has kind of been

    foreshadowed by the safety incidents. And to the degree that we're not going to see it — which I'm not sure we won't see anything in the immediate term — I mean, OpenAI still said they're pausing RL, so I don't know why it would be surprising if... we haven't even gotten the results of the investigation yet. It would be, in some ways, very in-character for OpenAI to do something that's radically self-contradictory, because that's certainly been my experience with OpenAI over and over again. But for them to say, this is a huge deal, all these people signed the pacing letter, Roon is continuing to freak out on the timeline, and we've paused RL, and we're going to reorient toward making sure we can solve all the safety stuff—

    to then launch the model before the investigation results are even back, or before there's been any demonstrated progress on the safety front, would be quite strange. My guess is we're still not going to see the best stuff the frontier companies have coming to the public this week — I'd be quite surprised. Again, it wouldn't be totally out of character for OpenAI to do something like that, but it would be quite the juxtaposition from their recent statements

    to launching something this week. Doesn't feel like it — we might see something from Google. There's this '000x Alpha' model on OpenRouter that everybody's been talking about — my sense of where the discourse has settled is that it's not really better than what we have. I guess the smart money is that it's maybe GLM-5.3, and if it is, it's maybe at Opus-4.7-to-4.8 tier, but still a notch down from the very best American models — which would be right on trend. I think Ethan Mollick has said things like this. So the initial reaction that it's better

    than Fable and better than Sol — I don't think that's really held up over the intervening few days. I'd be surprised if that really ends up a chart-topper. The most interesting thing for me is what Google is going to do, if anything, this week. Will OpenAI or Anthropic seemingly contradict their own stated priorities about getting things right first before rushing ahead any further? And is there some surprise of the continual-learning variety? That would be really interesting to experience, because honestly, right now, Claude with search, with the deep context I've

    26:58

    Prakash: Mhmm.

    26:59

    Nathan Labenz: set up, is really good at understanding what's going on — especially with the new situational-awareness skill I have too. I've got five years of history for me, and now a daily job going out, looking at all the tweets I like, doing additional research, and updating its own sense of where AI discourse is, where AI progress is, what's really top of mind for people right now. And it's getting good to the point where it's a little bit hard — what would the difference be if a model was not more capable but just had this continual-learning, internally-cracked architectural solution? Would we really feel a difference at this point? I think the biggest difference I could imagine

    would be maybe just the ease of deployment. Because if you just turn something like that loose in your Slack and it reads the whole thing, you don't even have to set up a framework or an ongoing search tool for it to use — it can just absorb that knowledge and know about it. That could be a huge game-changer for how easy this is to deploy, and how fast agents actually start to impact the economy. But it doesn't feel like, relative to what I'm able to do at the moment with Claude, it would be a huge difference in capability — which is pretty crazy. I was pretty skeptical in the early phases of Anthropic saying, we think search should be able to do this. I

    was like, really? You think you're going to be able to search that well? There's a lot of stuff to search through, but they've gotten really, really good at search at this point, to the point where it's taken a little bit of the edge off my craving for true continual learning. But still, I'd be extremely interested to see if somebody does have a real needle-mover in that direction. I'd be very open to having my mind changed and feeling a difference — it's just maybe the limits of my imagination more than the limits of the technology at this point. I've just been so impressed with what Claude can do with agentic search, runtime search, that the bar is pretty high now for what continual learning would have to look like

    to really make a difference.

    29:21

    Prakash: It may be that, instead of continual learning, you just have a longer context window, and maybe that works. So it might be that you have these two curves — one is continual learning getting better, and the other curve is just context windows getting better — and they kind of start converging, where you might not notice experientially that much of a difference at some point, because your context window and RAG setup are already pretty good. And regardless, no amount of continual learning can ever encompass the entire world — you still need some form of search anyway.

    So you still need some form of search and taste in picking what to put inside the context, which is why some people say RAG will never die — because you're always going to have some search and retrieval. So — let me pull up our first guest, because he's in the green room, and let me do an intro. Alright.

  2. 30:08Interview47 min
    Interview: Mohamed Awad — Arm's First Chip in Thirty-Five Years, and What Agents Do to the CPUMohamed AwadArm's EVP of Cloud AI on why a licensing company started selling silicon, how the Meta partnership traces back to AWS's 2019 Graviton launch, and what changes in a CPU designed for agents rather than people — 'agents don't sleep,' so the CPU becomes the coordination layer for accelerators that never idle. He declines to guess where the GPU-to-CPU ratio settles, and names skilled trades as the 2028–29 bottleneck he worries about most.
    Open segment on YouTube ↗

    Prakash introduced Mohamed Awad, executive vice president of Arm's Cloud AI business, who leads strategy for the compute foundation underpinning modern data centers and AI systems. After decades as a company that licensed processor architectures rather than sold chips, Arm has pivoted into offering complete compute platforms, headlined by the AGI CPU — a processor family co-developed with Meta and built for agentic AI workloads. Awad walked the hosts through that pivot, framing it as a natural extension of Arm's long-standing model of meeting customers wherever they sit on the spectrum from licensed IP to complete silicon, rather than a break from it.

    Prakash pressed Awad on how Arm now competes with, and partners alongside, the very companies building their own silicon — Nvidia's full AI-factory systems, AWS's Trainium, Google's TPUs, Meta's MTIA, and AMD's own efforts. Awad's answer was that Arm is foundational underneath nearly all of them, supplying CPUs, interconnect, and system IP regardless of who wins at the accelerator layer. On the Meta partnership specifically, Awad traced it back to AWS's 2019 Graviton launch, which showed hyperscalers that an Arm-optimized CPU could outperform legacy off-the-shelf silicon; Meta came to Arm wanting a comparable CPU it could adopt off the shelf rather than build from scratch, and the two companies co-invested and now maintain a multi-generation engineering relationship, with Arm's chip serving as both general-purpose compute and a head node to Meta's MTIA accelerators. Awad also described how Arm keeps competing customers separate — a shared core-IP engineering team feeds a distinct solutions team per partner — while insisting Arm is filling a genuinely underserved gap, since none of the hyperscalers' own Arm-based chips are available outside their own infrastructure.

    Nathan turned the conversation to economics, noting the odd signal that even old, notionally end-of-life chips are still renting at rising per-hour prices amid apparently infinite demand, and asked how margin gets divided between partners in a market where everything is sold out. Awad reframed the question around what's actually being monetized: today's market prices tokens, but tokens are just a conduit — the real value is the intelligence derived from them — and he expects pricing and value capture to keep evolving as that distinction sharpens. Prakash followed by asking where AI has actually moved the needle on ROI inside Arm itself; Awad pointed to software development, verification, and internal IT as areas seeing outsized, accelerating impact, while acknowledging usage is still an iterative trial-and-error process.

    Pressed on whether newer model generations are actually delivering visible gains, Awad agreed with Prakash's framing that the answer depends on task complexity — casual chatbot use shows little difference, but sophisticated engineering and agentic workloads show a clear, if non-linear, upward trajectory. He connected that to a broader theme: as models keep advancing, there's a growing opportunity for the system itself to select and route between models based on the task, which he tied to what's now happening inside agentic CPUs coordinating tool calls and recursive model invocations.

    The conversation then went deep on what makes a CPU built for agents different from one built for laptops or traditional web workloads. Awad's core claim was that "agents don't sleep": unlike human-paced interaction, agentic systems constantly spawn sub-agents that constantly hit accelerators, making the CPU the coordination layer for the whole system — deciding which models to call, managing accelerators, and keeping every core fed without becoming a bottleneck — all while stripped of legacy baggage the workload doesn't need. Nathan asked how far along this optimization curve really is; Awad pushed back gently on the idea that Arm is starting from scratch, noting the company has optimized for performance-per-watt for decades and is now applying that discipline to AI-specific system design as accelerator, memory, and interconnect specs keep shifting. On the perennial question of GPU-to-CPU ratios, Awad was candid that there's no settled number — the ratio depends on what you even count as a CPU, given how many are embedded inside GPUs and storage devices — and that the more useful frame is total infrastructure, not per-rack math.

    Later questions covered Arm's internal technical debates (performance-per-watt and which frontier labs to partner with came up most, though Awad resisted narrowing it to three), where he expects supply-chain bottlenecks in 2028–2029 (memory and wafer capacity, power generation, skilled trades, and unpredictable component shortages), and how Arm is addressing data-center backlash — Awad argued the industry needs to acknowledge real community concerns first, then shift the public narrative from raw token generation toward the tangible intelligence and benefits AI delivers. On hiring, he said Arm is prioritizing AI-fluent talent without cutting early-career recruiting, and struck an optimistic personal note about the job market his college-age son is entering.

    After Awad signed off, Nathan and Prakash kept talking about how underexplained the CPU's role is in public discourse about AI "escaping" or acting in the world. Prakash offered a mental model of a "thinking step" (the GPU) and a "mechanical step" (the CPU) that hands off between tool calls and model calls, calling agentic behavior something of a "fake intelligence" — an illusion produced by a task manager allocating between thought processes. Nathan built on that with a self-driving-car analogy — sensors feeding a decision system that issues commands to real-world actuators — as a more intuitive way to explain how a model's emitted tokens get executed as real commands on a CPU, sometimes reaching out into the broader internet, before Prakash moved to introduce the next guest.

    The value isn't in the token. The value is in the intelligence, and tokens are just a conduit to that.

    The simplest answer is that agents don't sleep.

    I don't know, quite candidly. I've heard two-to-one, I've heard one-to-one, I've heard two CPUs for every accelerator.

    32:25What are you doing at Arm, what is the AGI CPU, and how has Arm pivoted its business over the last twelve months?
    Awad described Arm's spectrum of offerings from licensed IP through architectural licenses and compute subsystems to complete silicon, noting Arm's partners have shipped over 350 billion CPUs and processors, and that the company's model has always been to meet customers at whatever integration point balances their investment and customization needs.
    35:31How did the partnership with Meta come about — did they approach you, or you them, and why?
    Awad traced it to AWS's 2019 Graviton launch, which showed hyperscalers that Arm-optimized CPUs could beat legacy silicon; Google, Microsoft, and Nvidia followed, and Meta came to Arm wanting an off-the-shelf, Arm-optimized CPU rather than building one entirely from scratch. Arm and Meta co-invested and specified the product together, and the resulting chip now serves as both general-purpose compute and a head node to Meta's MTIA accelerators.
    40:05What drove the decision to move into manufacturing, not just IP licensing?
    Awad framed it as an evolution rather than a break from Arm's model: customers told Arm an off-the-shelf, Arm-based CPU was an underserved gap in the market, and Arm built the AGI CPU to fill it and amortize the investment across multiple partners.
    41:21With everything at capacity, how do companies and partners figure out how margin gets split?
    Awad said today's market largely prices tokens, but tokens are only a conduit — the real value is the intelligence derived from them — and expects pricing and value capture to keep evolving as the industry shifts focus from raw token generation to delivered intelligence.
    51:33What have you learned about the agent workload that's leading to different CPU design decisions?
    Awad's headline point was that "agents don't sleep": agents constantly spawn sub-agents that constantly hit accelerators, making the CPU the system's coordination layer. That pushes design toward stripping legacy baggage, maximizing per-core memory and IO bandwidth, and balancing high performance against strict power efficiency.
    1:03:11Looking out to 2028-2029, where do you see industry shortages or bottlenecks that people aren't focused on yet?
    Awad pointed to ongoing memory and silicon wafer capacity constraints, power generation and turbines, skilled physical labor (electricians, plumbers, site builders), and unpredictable component shortages like the capacitor shortage that held up a colleague's board — arguing no single company can solve what is a massive global supply-chain challenge.
    1:05:27How is Arm thinking about data center backlash, and what story or incentives get local communities on board?
    Awad said the industry should start by acknowledging real community concerns rather than leading with incentives, and needs to shift the public narrative from token generation toward the tangible intelligence and benefits AI delivers, so the upside can be weighed against the concerns rather than the story staying one-sided.
    1:08:20How has hiring at Arm changed with AI over the last twelve months?
    Awad said Arm is prioritizing candidates with more direct AI experience but hasn't cut early-career hiring — if anything it's hiring more broadly at a robust rate — and expressed personal optimism about the job market for technologists, including his own college-age son studying computer science and economics.
    Lightly edited · timestamps jump to YouTube
    30:36

    Prakash: Our first guest for this morning is Mohamed Awad. He is the executive vice president of Arm's Cloud AI business. He leads the company's strategy for the computing foundation that powers modern data centers and artificial intelligence systems. For more than three decades, Arm has been famous for designing the processor architectures used in almost every smartphone on the planet, licensing those blueprints to other companies. Awad, who previously ran Arm's Internet of Things business and earlier built businesses at Broadcom, is currently spearheading a historic pivot for the company. Right now, he is driving Arm's push into manufacturing and offering its own full compute platforms. This includes a new processor family called the AGI CPU, which was co-developed with Meta and built specifically for the next wave of AI. As artificial intelligence moves from training models to agentic AI — software systems that run always-on agents handling complex, long-running tasks — the workload is shifting, and the CPU is once again becoming a critical bottleneck. Hyperscalers and cloud providers are racing to build more capable, energy-efficient infrastructure to feed data to GPUs and orchestrate millions of AI agents. Awad sits at the center of Arm's response, reshaping the economics of how global AI infrastructure is built. Mohamed, welcome to the show.

    32:20

    Mohamed Awad: Thank you. Thanks for having me.

    32:25

    Prakash: Maybe you can give us a quick intro to what you're doing at Arm — what is the AGI CPU, and how has Arm pivoted its business over the last twelve months?

    32:42

    Mohamed Awad: Arm sits at the heart of the compute ecosystem — together with its partners, it's shipped over 350 billion CPUs and processors over the years. Central to what we've done over the decades is provide performant, low-power compute platforms and meet our customers wherever they are. So we have a spectrum of solutions, from IP to architectural licenses to compute subsystems, all the way through now to complete silicon offerings. The idea is simple: partners can come to us and decide what integration point they want, balancing their investment against how much customization they need, to choose a solution that's optimal for whatever product or system they're building. That's been incredibly successful for us, mostly because rather than dictate to the customer what the right answer is, we align ourselves and work with them to develop something that actually solves the problem.

    33:53

    Prakash: When I look at the CPU space, I see AMD, I see Arm, I see Nvidia with the new Vera Rubin. How are these various players competing with each other?

    34:11

    Mohamed Awad: What you're seeing in the marketplace, over and over again, is large-scale technology providers building full systems. You think about Nvidia — they've got their AI factories, what Jensen likes to call them, and those consist of CPUs, DPUs for networking, switching gear, rack-scale architectures, and their accelerators, their GPUs. These are complete systems designed to work together. And actually, across the landscape, AWS is doing the same thing with its Trainium products, Google is doing the same thing with its TPU products, Meta is doing the same thing with its MTIA products, and to some degree AMD is doing the same thing as well. In every one of those cases, Arm is partnering with those players right across the board — whether it's the AMD DPU, the CPUs and DPUs in Nvidia, the GPUs and CPUs that are part of the Graviton system, or the CPU we've built in partnership with Meta. In every case, we're foundational in helping to bring those systems together and meeting them where they need us to be in order to make those systems a reality.

    35:31

    Prakash: Let's talk a bit more about that. How did this partnership with Meta come about? Did the Meta team approach you, or did you approach Meta? And what's the basis of the relationship — is it that they need co-design for what they're doing, or is there just a lack of capacity across the industry, so everyone is hunting for design teams? How did this come about?

    36:06

    Mohamed Awad: To understand the relationship with Meta, you have to understand the context of the broader ecosystem. In 2019, AWS launched their Graviton products, which were really the first mainstream CPUs based on Arm that became available to the ecosystem, and they built those CPUs to be optimized for the data centers AWS wanted to build. Very quickly after that, partners like Google, Microsoft, Nvidia, and others decided they too wanted to build their own CPUs, to have an optimized system for their data centers. In every one of those cases, what those technology leaders were looking to do was take advantage of Arm's long heritage of incredibly efficient, performant compute that could be optimized for the system, rather than carrying around a lot of legacy baggage associated with traditional off-the-shelf silicon. Meta was next up. When they came to us, they said: we're going to have to build a CPU to have an optimized infrastructure, but what we really want is a CPU, optimized based on Arm, that we can just take off the shelf — will you help us do that? Arm's business model has always been about partnering with customers and meeting them where they need to be — that's how our IP came about, that's how our compute subsystems came about, and this was really no different. So we worked with them to specify the product, they co-invested alongside us, and we still have a very close engineering relationship that goes out multiple generations. They're using that CPU for both general-purpose compute in their fleet, and now also as a head node to their MTIA chips, as part of their broader AI system. This is really just an evolution of the partnerships we've developed with all the big technology leaders.

    38:21

    Prakash: How do you run multiple teams like this, given that they're all competing with each other — Meta, AWS, Microsoft? Do you have separate, firewalled teams that deal individually with each client, or is there some cross-pollination between the teams?

    38:44

    Mohamed Awad: A couple of things. First, the underlying IP — the CPUs, the mesh interconnect that connects all the CPUs on the device, a lot of the system IP — that's common across all the different solutions, whether it's AWS, Google, Nvidia, or our own AGI CPU product. That's a common engineering team. When you talk about the CPUs we develop ourselves in partnership with Meta, that's a separate team — a solutions engineering team that's separate from the core IP, though all the underlying IP is the same. And in many cases, what we're doing is serving an underserved portion of the market. If you think about it, none of the Arm-based CPUs out there today are off the shelf — with the exception of the one we built. AWS's is limited to their infrastructure, Google's is limited to their infrastructure, Microsoft's is limited to their infrastructure. So we're really providing a new product that addresses an underserved part of the market, which comes back to why Meta asked us to build it in the first place.

    40:05

    Prakash: Let's talk about the licensing of IP versus manufacturing. What drove the decision to pull the trigger on manufacturing? You've been focused on IP for a long time — this seems like a big decision.

    40:27

    Mohamed Awad: It's interesting, because obviously it is a change in how we've gone to market, but fundamentally, at its most basic level, it's really just an evolution of what we've always done — meet our customers where they are, identify an underserved part of the market where there's a gap we can fill, and amortize that across multiple partners. That's really what we've done here. We found an underserved part of the market, customers came to us and said it was underserved, asked if we could help identify a solution, and we went off and did it. I don't think it was any more involved than that — it was really about: our customers need this, we can deliver, so how do we go address that space?

    41:21

    Nathan Labenz: When we think about how value accrues in this space, I find it a bit of a strange puzzle right now, because we look at data points like the fact that old chips — from even five years ago, that were scheduled to be end-of-life by now — are still renting at a higher per-hour price than they did originally. It's like, everybody's making money, there's infinite demand, everybody's sold out, you can't get capacity. So in that environment, how do companies figure out who's going to get what margin? How is the balance decided between partners, given that everything is at capacity?

    42:15

    Mohamed Awad: It's a typical market like any other. You bring up an interesting point, because a lot of folks talk about tokens, and ultimately that's what you mean by getting rented — this idea that generating tokens is the thing that's been monetized primarily up until this point. But ultimately, that's not where the value really resides. The value isn't in the token — the value is in the intelligence, and tokens are just a conduit to that. They're the raw component which, when distilled and decisions are made, and agents are executed based on the information derived from those tokens, intelligence is created, and that's what's really valuable. So like any market, we're in the early phases where people are trying to figure out how to get to that higher-level intelligence, because that's ultimately what my mom and my dad are going to care about — the intelligence. They don't really know what to do with a token, per se. We're still trying to figure that out, which is why you're seeing the market economics you're seeing today, but I think that's going to continue to evolve with time.

    43:37

    Prakash: Speaking of requiring intelligence — within Arm itself, where have you seen AI usage take off and actually make a difference in terms of ROI?

    43:55

    Mohamed Awad: It's amazing, actually. This is a topic we talk about quite a bit, because week on week, month on month, we're seeing AI make a massive impact in things like software development, a massive impact in things like verification, all the way through to our internal IT systems. It's really starting to permeate across the board — I have product managers leveraging it, my internal IT help desk leveraging it, and meanwhile my engineering teams using it to optimize code for new products, and for things like verification. It's definitely a journey — most people start simple, but they continue to ramp up and test those bounds. Sometimes it gets it right and you push forward; sometimes it doesn't and you take a step back and reevaluate. But without a doubt, the trajectory is up and to the right, and usage is accelerating at a pretty phenomenal rate right now.

    45:13

    Prakash: Do you see a difference from model generation to model generation — having gone from, say, late last year's Opus 4.8 to Opus 5 in the last month or so? Do you actually see a difference in usage, usability, performance? Because for a lot of people right now, they look at it and say they don't see much difference between one level and the next, and one of the debates we have is that maybe that's because the tasks people are applying AI to aren't complex enough to require that intelligence. And I'm wondering — in a firm which has those very high intelligence-hurdle problems — do you actually see this uplift from model generation to model generation?

    46:15

    Mohamed Awad: I think you hit the nail on the head — it really comes down to the complexity and sophistication of the engagement with the model. If I'm at home, sitting in front of my laptop, just interacting with a chatbot, the model difference often isn't going to make a huge difference. For more sophisticated users — as you start to think about the complexity of optimizing software, or running complicated agents in an engineering environment — yes, you absolutely do notice the difference. And it's not linear — there are jumps and plateaus — but I think it would be disingenuous to suggest there isn't a clear trajectory upward. The interesting thing behind that, if you step back and think about the underlying infrastructure — and the point Nathan made earlier about unrelenting demand — is that there's an opportunity to optimize the underlying model based on the task at hand, or for the system to select and optimize which model is being used based on the task, so it can get hyper-efficient and narrow its usage to the requirements or sophistication of the user's request. That's where we'll start to see some really interesting things happen. And frankly, that's one of the interesting things happening around agentic, or AGI, CPUs right now — they're doing this coordination and selection of which model to use, then making tool calls, then recursively calling models again and exercising accelerators again. As that process gets more and more refined and optimized, I think you're going to see a level of efficiency and sophistication that's going to be really impressive.

    48:42

    Prakash: Let me switch gears a little bit. Nvidia used to have a two-year development cycle, then shortened that to a one-year cycle. What does your dev cycle look like for the next generation of product, and over a five-year time frame, what does your product roadmap look like?

    49:08

    Mohamed Awad: It's crowded — there's a lot going on. This market's moving so quickly, I think you have to move at that sort of speed. The reality is that models are changing every six months, three months — they're moving incredibly fast. And historically, the hardware design cycle is 18 months to three years, somewhere in that neighborhood depending on the complexity of the chip. So to think you're going to design optimized silicon over a three-year period and predict what the software is going to look like that far out, so that you're fully optimized, just isn't reality. So accelerating that cycle as much as possible becomes more and more important. Sometimes those are big jumps, sometimes small jumps, but our own internal development process has had to accelerate because of that — and of course we're leveraging AI to help do that. But more fundamentally, if you think about our business model — we provide partners IP, compute subsystems, chiplets, or complete SoCs, and those technology leaders build complete systems, racks and warehouses' worth of gear — that lets them very quickly decide when it makes sense to build a complete SoC from the ground up and take that three-year cycle to get that optimization, versus when it makes sense to take something off the shelf. And the answer is never the same across the board — for certain products within a system they may want IP, for others they may want a full solution. Part of the strategy is to create that flexibility so they can focus their energy on the pieces that matter.

    51:13

    Nathan Labenz: Could we talk a bit about what it looks like to design a CPU for agents, as opposed to — obviously we've had CPUs in our personal computers forever, and CPUs that run in data centers handling traditional web workloads. What have you learned about the agent workload in particular that's leading to different design decisions?

    51:41

    Mohamed Awad: Let me start at 10,000 feet and then give a couple of specific examples. The simplest answer is that agents don't sleep. If you think about your laptop, or these web-scale apps, they were designed for a world where you interacted with the machine — it did some processing, made a tool call, maybe you went and grabbed a cup of coffee and came back, read the response, then interacted with it some more. Even sitting in front of your machine, there was latency for you to process, type, and so on. We're now living in a world where agents are constantly feeding the accelerators, constantly reacting, spawning additional agents — just because you spawn one agent doesn't mean only one agent exists. Every agent could spawn ten, a hundred, thousands of agents, and each of those could spawn a bunch more to fulfill your request. Effectively, those CPUs become, in some ways, the coordination mechanism across the entire system. And that act of coordinating the entire system — managing the accelerators, deciding which models to choose, and so on — becomes such a critical role, with such expensive infrastructure riding on it, because at the end of the day you need to drive utilization up and respond to the user as quickly as possible. Practically speaking, that means you have to optimize the silicon — carrying around legacy accelerators, or worrying about supporting legacy code, isn't so important anymore. This is a new style of software — you don't need to support Lotus Notes, I like to joke. It means thinking about things like memory bandwidth, IO bandwidth, and the amount of bandwidth dedicated per core that you can rely on time and again, so that each individual CPU core within the SoC is never bottlenecked because some other agent is hogging it. These are the sorts of things that set an agentic CPU apart. But overarching all of that is balancing incredibly high performance with incredible efficiency. We all know the power demands AI is placing on infrastructure are enormous, and every milliwatt of energy you're pouring into a CPU is a milliwatt you can't put somewhere else — it's one less accelerator you can have, one less customer you can serve, one less piece of intelligence you can serve up. Doing all of this within an incredibly efficient package goes back to stripping out anything that's superfluous.

    55:07

    Nathan Labenz: So it almost sounds like, in some ways, a simplification. I guess I wonder — how far along do you think we are on the optimization curve? We've now been optimizing the chips that run the linear algebra for a few years, and significant progress has been made. Are we in the same way that robotics is a few years behind language models? Or is CPU optimization a somewhat offset but similar curve?

    55:42

    Mohamed Awad: For Arm's part — and to be fair, other CPU vendors too — we've been optimizing CPUs for decades, so to suggest we're just now showing up and starting to optimize these things would be disingenuous. I'll say we've gotten the efficiency part down really well at Arm — we've gotten really good at that — and the performance piece too, we're all carrying supercomputers around in our pockets, and those are all Arm-based. So we've got the performance and the efficiency. What I think we're really doing now is optimizing what that looks like in a system designed specifically for AI. And because the underlying AI systems are evolving — the number of nodes in a system, the speed of the accelerator, the memory interfaces, the types of memory connected — that optimization is ongoing. So I think that's where the work is: as the coordinator of the end-to-end system, continuing to optimize in support of the evolution of those systems.

    56:57

    Nathan Labenz: How do you think about the ratio, or the balance, between GPUs and CPUs as we head toward maturity? One of my favorite stats — speaking of supercomputers in our pockets — is that a typical cell phone battery holds somewhere in the 20-to-40 watt-hour range of energy, which means a full charge every day, over the course of a year, comes out to something like two dollars in my electrical jurisdiction. It's amazing how efficient a cell phone is — how little energy it really takes to power your phone for a full year. Obviously we know GPUs are pretty energy-intensive. How do you think about the right way to even calibrate that — is it number of GPU chips versus CPU chips, racks versus racks, or could it be better measured in power into these different kinds of systems? And where do you think that balance reaches equilibrium at maturity?

    58:07

    Mohamed Awad: I don't know, quite candidly. It's one of these things happening at such a rapid rate. A couple of years ago, everyone swore up and down the CPU was dead — clearly that's not the case, it's a huge, growing market, growing on the back of the idea that we need to serve up intelligence. I've heard two-to-one, I've heard one-to-one, I've heard two CPUs for every accelerator. I think there are a lot of questions about what you're even counting as a CPU — at the end of the day, these GPUs today have 24, 32, 48 server-class CPUs in them. Are we counting those? And then you've got your storage devices, which have a bunch of CPUs in them too. All of these things are getting taxed more and more as more information is generated, as more tokens are generated, as more agents distill that down and then hammer the next tool calls, which means more networking traffic, more storage. So I think to simply put a number on CPU-to-GPU ratio — in a particular blade, or even in a single rack — is a very narrow way to think about it, when you start talking about warehouses, or the global infrastructure, which is really what we're talking about evolving right now.

    59:44

    Nathan Labenz: At the beginning, you mentioned a headline number — how many chips or cores has Arm shipped over time? I think it was something like 350 billion.

    59:55

    Mohamed Awad: It's about that — the number changes so rapidly that I'm probably low now. Last I heard, it was about 350 billion, and we were averaging something like 30 billion a year. It's a pretty big number, and it's in everything. This actually goes to the broader Arm story — I'm responsible for the Cloud AI business, so that's what I focus on, but my colleagues responsible for physical AI and edge AI are seeing a massive AI boom happening right now too, and it's really part of the overall story. Cloud is kind of the first inning — but as intelligence becomes physical, as it becomes more personal, you're going to see that intelligence continue to spread more broadly to everything around us. I think that continues to be an incredible opportunity for the entire industry, Arm included.

    1:01:09

    Prakash: Internally within the firm, what are the three biggest technical debates that you have? Those long-running technical flame wars that people have thousands of posts on, that everyone has an opinion about — what are the big debates you have?

    1:01:35

    Mohamed Awad: Arm is about 8,000 people now, I think, and of those 8,000, probably close to 7,000 are engineers — so massively engineering-centric. I'd struggle to narrow it down to just three. The teams debate everything, from how to squeeze the next ounce of performance out of CPUs, to which models or services from which frontier labs we should be partnered with for some of this stuff. Like any good technology firm, there are lots of vigorous debates on lots of topics, and it would be impossible for me to distill it down to just three.

    1:02:27

    Prakash: But the ones that seem to pop up are performance-per-watt, and which frontier labs to partner with.

    1:02:35

    Mohamed Awad: I just think those are top of mind because they're such topical issues right now. Performance-per-watt has always been part of Arm's heritage and legacy — it's always been the DNA the company was built on, so it's always an area of vigorous debate, which frankly is why we've been so successful at it, because with that many bright minds grinding away, we tend to push the bounds in that area, which is why that one comes up for me.

    1:03:11

    Prakash: Switching gears a little bit — in the market there's a lot of talk about bottlenecks, and you're in a privileged position where you can see, and need to have a sense of, what's going to happen when. 2027 is spoken for, for a lot of people, but when you look at the 2028-2029 time frame, where do you see the shortages in the industry right now? What are people not really focused on that you think are going to be issues in the future?

    1:03:56

    Mohamed Awad: There are lots of areas. Capacity is going to continue to be a challenge for the foreseeable future — memory capacity, silicon wafer capacity, and so on. Beyond that, everything from turbines to power generation more broadly — and skilled physical labor continues to be a challenge, things like the electricians and plumbers and the people who build these sites. I heard the other day from a colleague at another company that he couldn't ship his board because he couldn't get capacitors — something you'd think is readily available, but it was actually holding them up. If we continue at this breakneck pace, which I expect for the foreseeable future, you're going to continuously run into different things like that, which speaks to the broader challenge: we can't assume any one company or entity is going to solve the supply chain issues. This is really a massive global challenge that's going to take everybody to address.

    1:05:27

    Nathan Labenz: We just created, with Suno, a hymn to the global supply chain at the end of last week, out of appreciation for how many people come together to make all this stuff possible. One last one for me before we let you get back to work — you mentioned the build-out, and one of the things that could slow down American AI at the moment is data center backlash. How are you thinking about that? It's a whole industry-wide question of how we tell the story of AI, inspire people with the upside, and maybe sweeten the pot for local communities. Do you have a point of view on what story should be told, or what incentives should be offered, to get people on board with actually building out the data centers?

    1:06:34

    Mohamed Awad: I think we have to start by not focusing so much on dangling goodies — the start really needs to be stepping back and acknowledging that there are real concerns here, understanding what those concerns are, and then working to address them. That's step one. The second important point is understanding that today the world is so focused on generating tokens, and we really have to start pivoting toward delivering on the promise of intelligence, of AI. As we start to deliver on that promise, I think some of those fears — though real — will start to be balanced against the benefits. Right now it's kind of a one-sided story, and we need to work through that. For Arm's part, our focus is on delivering the underlying technology — performance, efficiency — because that helps with some of these concerns in our own way, as these systems and platforms deliver more and more intelligence, and can do so more and more efficiently. That's how we think about it, and where we spend most of our time.

    1:08:10

    Prakash: Last question for me — when you look at hiring inside Arm, how has hiring changed with AI over the last twelve months? Are you looking for different skill sets, people qualified in particular things? Have you cut down on early-career hires? How has the hiring pattern changed?

    1:08:40

    Mohamed Awad: That's interesting. I'd say we're definitely trying to bring in talent with more experience leveraging AI, developing AI, building AI-based products — that's absolutely the case. I wouldn't say early-career recruiting has changed, or that we've slowed it down — if anything we're bringing in more folks. Arm has continued to hire at an incredibly robust rate, growing in support of all our technical aspirations. Personally, I'm bullish on the future for technologists — I've got a nineteen-year-old son about to go into his second year in school for computer science and economics, and I'm personally very bullish. I think the roles will evolve — it's naive to suggest they won't, but they've evolved since I went to school too, and that's okay. It's probably going to be a bit more dramatic this time, but I think the opportunities will be there.

    1:09:59

    Prakash: Indeed. Thank you for joining us, Mohamed Awad — it's been a pleasure, and we hope to hear from you and Arm again soon.

    1:10:10

    Mohamed Awad: Thank you, Prakash. Thank you, Nathan, and thanks, guys.

    1:10:21

    Nathan Labenz: I wonder — I'd venture a guess at his take, but I want to hear yours — how do you describe the role of CPUs in the overall AI stack? This has been especially interesting lately, with people asking things like, did the AI escape? What does it mean for AI to escape — how is it interacting with the outside world? I used to think it was just a bunch of numbers getting crunched in a GPU somewhere. The CPU is dramatically underemphasized in the story people tell about how AI actually works, and has effect on the world. Do you have a way of telling that story that you find effective?

    1:11:12

    Prakash: Well, at the risk of sounding extremely foolish, I'll attempt to. I tend to think of it as a mechanical step and a thinking step. For the thinking step, you go to the GPU and run inference. For the mechanical step, it's like: okay, I have that thought back — mechanically, what do I need to do next? Do I need to call a tool? Do I need another thinking step? And so on. When these agents are built, a lot of these steps become — for example, 'wait for user input' is a step. At some point it calls wait-for-user-input, and then it comes back to you. So you have this centralized manager — the CPU — that's kind of pinging off the neurons when it needs to, but otherwise mechanically calling other things. It's kind of the execution agent in the middle — the spine of the whole thing, as Claude would put it.

    1:12:30

    Nathan Labenz: Heavily load-bearing, for sure.

    1:12:32

    Prakash: Yeah, load-bearing is fine of the entire matter. And I think what's ended up happening is it's turned out that yes, you do need a bunch of thinking steps, but you also have a lot of context switching, which is what the agents are doing — they're switching context between multiple tools, the user, and the thinking part. And it turns out the context-switching part is actually pretty important. I think some people expect that eventually the model will be able to do everything, and maybe this CPU-GPU split we're doing right now is a patch, or a clutch, because the model isn't actually able to do allocation of resources itself. Maybe that improves, maybe it doesn't — maybe we're stuck with this. So I think agents are kind of like a fake intelligence, basically — you're faking this process of an intelligent entity using, basically, a task manager in the middle that's able to allocate to a thought process and then bring that back in. It's a bit of an illusion, and I—

    1:13:57

    Nathan Labenz: —thought one thing was interesting about the way he said it, which I'm maybe hearing an echo of from you: he was kind of like, the CPU is deciding which model to call. And that isn't quite true, right? The CPU isn't actually making decisions. A way I've been thinking about explaining it to people more often is by analogy to a self-driving car — which is funny, because most people haven't even ridden in a self-driving car, though they've probably at least used some AI system with tool calls in the loop. But with a self-driving car, it's very intuitive: you've got a bunch of sensors bringing information into a processing system that decides what to do, and then that processing system issues commands to a set of tools in the car — go, stop, turn, whatever. Everybody has a decent intuition for how that architecture works, and it's similar in the agent world. But when you say things like 'the CPU is deciding what model to call,' it's really more like: the model is emitting tokens, which are then executed as a command on the CPU, which may call another model, or call itself, or call an external API. And this is where tentacles can get out into the broader world through the internet. I guess I'm kind of gradually coming around to this self-driving car analogy, because people are just so used to being the operator — they have such an intuitive sense of the action space when they're driving a car. You put a computer brain in your position, but it basically has the same options at any given moment that you would if you were sitting there. The same is true for computers, but people don't think of themselves as issuing computer commands the way the AI is doing — they don't have to send a command to their own computer to call their own brain again, or call a subprocess, a different version of their brain. Obviously people don't have different versions of their brain, so no wonder it's not super intuitive to everybody. But yeah, I think that story is still one people — I've learned from the confusion I've heard around the hacking incidents — people are still very confused about what the parts are that make up an overall AI system. The CPU is very often glossed over, so people don't have a good sense of how an intelligence in a data center somewhere is actually able to reach out and touch the world. Work to be done. But that's probably not the field's biggest concern in terms of teaching people how it works and what the upside is going to be — but it's definitely a major point of confusion right now for a lot of folks.

    1:17:15

    Prakash: Yeah, I think we tend to create these mental frameworks to understand these things, but we also tend to refer to our own experience of how thought processes work — and we're not even that sure about our own neuroscience, how that works. So I'm not even sure the analogies are always correct. But it's always an interesting challenge to figure out. Let me introduce our next guest — let me get him. There we go.

    • The CPU Is Not Dead

      0:00 / 0:00
    • Why Agentic CPUs Never Sleep

      0:00 / 0:00
    • The Next AI Bottleneck Is A Capacitor

      0:00 / 0:00
  3. 1:17:27Interview59 min
    Interview: David Li — Shenzhen's No-Next-Week Mentality and China's Deliberately Boring AIDavid LiThe founder of Shenzhen Open Innovation Lab on the Beijing 'robot Olympics' as public habituation rather than competition, industrial robot hardware near $3,000 with the bottleneck moved to integration engineers, OpenClaw's viral moment owing something to a nickname that sounds like 'crawfish,' and why he thinks the era of chasing giant general-purpose models is largely over in China.
    Open segment on YouTube ↗

    David Li, executive director of the Shenzhen Open Innovation Lab and co-founder of China's first hackerspace, XinCheJian, joined Nathan and Prakash live and very late his time, kicking off with the viral clips of China's humanoid "robot Olympics." Li explained that the footage everyone was sharing came from the World Robot Conference in Beijing — a five-day event with serious industrial demos on the weekdays and an open-entry "Robot Olympics" on the weekend, where anyone with a robot, from serious teams to scrappy student squads, could compete in races, tennis, and boxing. He framed the funny fail clips as more popular than the real competitions, and read the whole spectacle as a way of building public comfort with robots rather than a display of raw competitive intent — a softer answer than Nathan's suggestion that the US and China might have fundamentally different imaginations of what robots are for.

    On what's coming out of Shenzhen next, Li described a manufacturing culture with "no next-week mentality": hundreds of thousands of small companies iterating rapidly, feeding much of Amazon's electronics catalog, with very low IP protection meaning everyone copies whichever design sells. His pick for the current hot category was AI-embedded talking toys — cheap to build (a low-cost chip plus a flat-rate token plan) but, in his view, still mostly generic and personality-free because few teams are investing in crafting the actual experience rather than just shipping a working demo. He used an analogy about building a narrowly-tailored "preacher" AI for a specific congregation to illustrate the vertical, personality-driven products he thinks are still missing.

    On industrial robotics, Li said hardware prices have crashed toward roughly $3,000 for an industrial robot, so the real bottleneck is now skilled "field application engineers" who can integrate robots on the factory floor — he drew a parallel to Palantir's "forward/field deployment engineer" branding. He described the current sweet spot for flexible robot arms as the dangerous or undesirable jobs, like plugging leads into car batteries for testing at CATL, where the risk of shock and the low volume per batch make full custom automation not worth it.

    Asked about US-China AI politics, Li was dismissive of reports that Washington might push allies to "pick a side" ahead of the Trump-Xi AI summit, calling it likely-unenforceable domestic policy with no real teeth abroad. Nathan pushed on his own on-the-ground impression from a recent China trip — that he never heard Chinese AI insiders frame the US-China relationship as a race with a "winner" the way American commentators often do — and Li's answer reinforced that: he described Chinese AI discourse as overwhelmingly practical, focused on which app feature to bolt AI onto next, not on any competitive finish line.

    Prakash asked about the reported mania in China around the AI agent OpenClaw. Li attributed a good chunk of its virality to the fact that its Chinese nickname sounds like "crawfish," a beloved regional dish, and said usage has since dispersed into homegrown agent apps bundled into Alibaba and Tencent products rather than staying concentrated in OpenClaw itself. On which models carry status in China, he said the general public just uses whatever's bundled into their platform, hardcore engineers favor tools like Kimi or Codex, there's broad agreement that Gemini gets essentially no attention in China, and Doubao (ByteDance) quietly holds a large slice of the enterprise chat market despite generating almost no public buzz.

    Asked for advice to a hypothetical new Chinese frontier lab, Li argued the era of chasing giant general-purpose models is largely over domestically. He pointed to a compact model in the ~27-billion-parameter class that he's been running as evidence that small models have gotten dramatically more capable, and predicted the winning play is shrinking capable models onto cheap dedicated local hardware — he counted roughly a dozen Chinese startups already building SSD-sized local-inference boxes — rather than training toward ever-larger frontier models. On the recent wave of Western AI agents "going rogue" (OpenAI, Hugging Face, Claude, and Meta all came up), Li downplayed the danger as largely PR theater sitting on top of already-commonplace "script kiddie"-style automated hacking, and argued no serious Chinese lab would stage something similar because, unlike in the US media environment, there's no attention or valuation upside to it in China's more fragmented press landscape.

    Closing on what US coverage of China gets wrong, Li said it's less about factual errors than mutual unfamiliarity: both countries are adjusting, for the first time, to running genuinely neck-and-neck at a technological frontier. He noted, half-jokingly, that Western coverage tracks Trump's AI statements closely while almost nobody reads Xi Jinping's — including a WAIC speech two months earlier that went largely uncovered in the West. Nathan closed by agreeing that Chinese understanding of American AI developments seemed sharper than the reverse, and invited Li back for a recurring conversation.

    The price has gotten close to hitting rock bottom — you can get an industrial robot for $3,000.

    If OpenClaw had any other name, it wouldn't have been as popular — the translation here is actually crawfish.

    How much better the Chinese understanding of what's going on in America is than the American understanding of what's going on in China.

    1:21:05What is the discourse in China around the viral humanoid 'robot Olympics' clips, and how does it contrast with the West?
    Li said the videos came from the World Robot Conference in Beijing — serious industrial demos on the weekdays, then an open-entry 'Robot Olympics' on the weekend where families come out and enjoy funny fail clips more than the actual competitions. He framed it as building public comfort with robots rather than fear.
    1:34:29As factories roll out industrial robots in China, are you seeing price competition — are factories able to bring pricing down because of the robots themselves?
    Li said industrial robot hardware has crashed toward roughly $3,000, so the real bottleneck now is skilled 'field application engineers' who can integrate robots on the factory floor. The clearest current use case is dangerous jobs, like plugging leads into car batteries for shock-risk testing at CATL, where low per-batch volume makes full custom automation not worth it.
    1:39:27Has the reported push for the US to make countries 'pick a side' on AI made its way into Chinese AI discourse, and could it derail the upcoming Trump-Xi summit?
    Li was dismissive, calling it likely unenforceable domestic US policy with no teeth abroad — 'chest-pumping' rather than something with a real enforcement mechanism — and said there isn't much discussion of it in China.
    1:43:36Having just visited China and never heard anyone frame US-China AI competition as a 'race' with a winner, how would you describe the general outlook on AI competition among Chinese AI insiders?
    Li described Chinese AI discourse as overwhelmingly practical rather than competitive: big platforms like Douyin, Taobao, and WeChat are mainly experimenting with where to embed AI features into existing products, not racing toward any finish line.
    1:46:55Has the reported enthusiasm for OpenClaw in China continued, and how did the earlier hype around it play out?
    Li attributed much of the early virality to OpenClaw's Chinese nickname sounding like 'crawfish,' a popular dish, which made it a cultural moment. He said usage of OpenClaw-style agents has since spread widely, but mostly through homegrown apps bundled by Alibaba and Tencent rather than OpenClaw itself.
    1:52:03Is there a taste hierarchy in China similar to Claude being the 'insider's model' in the West — who uses which AI model and why?
    Li said the general public uses whatever free app is bundled into their platform, hardcore engineers favor tools like Kimi or Codex, there's broad agreement that Gemini gets little attention in China, and Doubao (ByteDance) quietly holds a large share of the enterprise chat market despite generating little public buzz.
    1:56:55What advice would you give a new frontier AI lab setting up in China right now?
    Li said he doesn't expect many new frontier labs to emerge; instead he pointed to how capable a recent ~27-billion-parameter model already is and predicted the winning strategy is shrinking capable models down to run on cheap dedicated local hardware — he counted roughly a dozen Chinese startups already building small local-inference devices — rather than chasing ever-larger training runs.
    2:03:09Has the recent wave of Western AI agents 'going rogue' shifted the vibe in China, and what would happen if a similar incident surfaced there?
    Li framed the Western incidents as largely PR theater sitting on top of already-common 'script kiddie'-style automated hacking, not a sign of genuinely autonomous rogue behavior. He said no serious Chinese lab would stage something similar on purpose, since there's no upside to that kind of stunt the way there might be in the US attention economy.
    Lightly edited · timestamps jump to YouTube
    1:18:05

    Prakash: And our next guest is David Li. He is the executive director of the Shenzhen Open Innovation Lab, a unique technological hub that links hardware entrepreneurs from around the world directly into the densest, fastest-moving manufacturing ecosystem on earth. David is a true pioneer — he has been active in the open-source movement since 1990 and co-founded China's very first hackerspace, XinCheJian. Right now, the global technology sector is experiencing a massive, seismic pivot toward physical AI: humanoid robots, embodied intelligence, and autonomous smart hardware. While Silicon Valley currently leads in digital foundation models, China installed nearly 300,000 industrial robots last year alone, and Shenzhen is the exact place where the physical hardware to house these new AI brains is being built, tested, and scaled at breakneck speed. David brings a highly distinctive, often contrarian worldview to the space — he argues that the traditional Western obsession with intellectual property and nondisclosure agreements actually slows down technological progress, and instead advocates for an open, farmer's-market approach to hardware, where ideas are rapidly shared, copied, and iterated on in real time. Today we're going to explore how Shenzhen turns emerging experimental tech into scalable commercial products, why he believes the next great hardware revolution will come out of Africa rather than California, and what the rest of the world fundamentally misunderstands about the speed of Chinese manufacturing. David, welcome to the show.

    1:20:05

    David Li: Hello, great to be here.

    1:20:10

    Nathan Labenz: Thanks for staying up late for us — I know you're quite a few time zones away, and we're honored to have you live with us on AI in the AM. It's AM for Prakash, and it's AM for you — I'm the only one currently on PM.

    1:20:25

    David Li: Oh, that's great. I'm nine hours ahead anyway.

    1:20:32

    Nathan Labenz: Maybe a first question for me — top of mind based on all the short videos I've been seeing online over the last few days — is this humanoid robot Olympics that's been happening in China. There's a lot of marveling going on at the clips we're seeing. Some of them have been quite funny, with robots running into cabinets, knocking themselves over, sparking out — but also running pretty fast in some of these clips. What is the discourse in China, or the general feeling about what people are seeing there? How would you contrast it to what you hear from the West?

    1:21:20

    David Li: Yeah, I think right now the robots are just being put out there, whatever people have. The videos you're seeing a lot of are from the World Robot Conference, which took place in Beijing over the past week — a five-day event. The first three days were on weekdays, and the last two were on the weekend. On the weekend, when you go, you see parents bringing their kids to the festival. It's not so much a competition — it's a demonstration of technology, made into a real festival atmosphere. And people share the funny clips more than the real competitions. We enjoy that, and that's how you get a society that's starting not to worry about robots and instead thinking about what they can actually do.

    1:22:37

    Nathan Labenz: One comment I saw that I thought was interesting was that these demos of robots running races and so on maybe reflect a deeper difference in the imagination of what the future role of robots in society is going to be. When we see demo videos from Google or American robotics startups, they're typically focused on fine motor skills — folding laundry, tabletop exercises. I can't recall an American company showing a robot running. Do you think these different presentations of new robot capabilities reflect a deeper philosophical difference, or a difference in imagination about what robots will do for us — or am I reading too much into it?

    1:23:39

    David Li: No, I mean, if you went to the conference you'd see all the industrial robotics too — every manufacturer wanting a robot to take over some task. But that part isn't so interesting on its own. The last two days, the Robot Olympics — that's the fun one. It doesn't matter how big your company is to come in and be part of it; anyone who has a robot can sign up. So on the marathon, you have some very fast runners, but you also have a lot of scrappy entries put together by student teams who just show up. So it's two things: it's an event that lets people push the boundaries of what a robot can do, in an Olympic kind of way — that's also why it happens on the weekend. This year it was richer than last year — robots competing more as groups, like robot tennis and robot boxing. All in all it's for good fun, and it's also for testing a lot of the tech in a real setting, but a low-stakes one — it's not like if you fail, nobody's going to invest in you. It's a phase you go through. It's competitive, but it's also fun.

    1:25:59

    Prakash: Let me ask you — in the last few months, which products have you seen in Shenzhen or China that you think are going to hit the world in the next year or so? Which are the interesting products you've seen recently that you expect will make it onto the world stage?

    1:26:27

    David Li: Well, that's one of the things — Shenzhen doesn't really work on a 'what's next week' mentality. If something's not popular, it's not popular, and nobody's going to make it — but gradually, over six months, people move around. To give you a sense of scale: Shenzhen has probably got hundreds of thousands of companies making small products for every niche you can imagine. Any electronics item you get on Amazon, there's a good chance it's coming from Shenzhen. You go through the Amazon electronics category and maybe 80% of that stuff, you look at it and think, why does this even exist? But it exists because there's a tiny market for it, and gradually some of it becomes popular as people experiment. We won't know anything until six or twelve months from now. What I would say right now is there's a lot of production going into talking toys.

    1:28:03

    Prakash: On what?

    1:28:04

    David Li: Right now it's cheap — all the talking toys. Toys, talking toys, animals.

    1:28:10

    Prakash: Talking toys, yeah, right on.

    1:28:14

    David Li: Yeah, but I mean, making a talking toy is easy. It's a five-dollar chip, then a ten-dollar flat-rate token plan from one of the token providers, and you go to Shenzhen, go to one of the toy shops, and snap that thing in. Do a video, put up an Amazon page, and you're in business. And because there's very low intellectual property protection, everybody watches everybody else to see which one sells, and things move in that direction — everybody chasing what sells next week. And then all the features get integrated back and forth, and eventually, six months from now, because of all these crossovers, they become something new. Some of them might make it. But overall, we've entered an era where things move very gradually — it's all about figuring out, now that you have a large language model embedded and the toy can understand you, how people are actually going to use it.

    1:29:48

    Nathan Labenz: Real quick follow-up on that — is there a prevailing sense among Chinese parents about what they'd want from a talking toy, or what they might be afraid of? These haven't really hit the mainstream in my kids' circles yet, so I haven't been asked for one — but I'm open to getting my kids one at some point. I think it's got to be better than YouTube, or at least a well-designed version does. But I'm also like, jeez, I'd want to know who's making my kid's talking toy and what values they have in mind for it — I wouldn't just want to buy a random one. So is there a sense of consumer expectations you could summarize from the Chinese context? Or is it just too early to know, and we'll have to find out through experimentation?

    1:30:56

    David Li: Well, it's all experimentation, because most of these toys are running a pretty open-source stack — meaning you can pretty much hack them. The adoption of the toy is going to be very culturally dependent, with small groups making things that react to their own group. So the spectrum of things — we're getting stuck on the talking-toy part. If you think about the circuitry, it's a mic, a speaker, a chip that connects to Wi-Fi, and whatever large language model or agent it connects to — very tunable. But at this stage, not enough people are putting effort into crafting the experience. Every one of them right now sounds a little generic — they answer everything, but they have no personality. And I think the wrong incentive is: if it can answer, I can sell it. The competition hasn't yet moved to the very fine interactions. What we need are smaller players coming in to think about this in a very vertical way. There's always a big idea — I always give the example: if I could get a Jesus-level talking soul and put it in a mega-church in Texas, that would be a model to learn from — from the pastors. You'd learn everything from the sermon, and the 50,000 people who go to that church would be the upper bound of my market. That's the kind of thing we should be looking for, because the foundation is so easy right now — but the mindset has to shift from making things for everyone to making things for a very specific audience.

    1:34:29

    Prakash: Let me take a step back here — one of the things about robotics, especially applying it in factories: the expectation is that robotics gets applied in factories first, primarily because home robots could be dangerous around kids, and a lot of things in the home — dogs, kids, trash — are still very difficult for AI to manage. When you look at industrial robots being applied in China, as factories roll them out, do you see price competition, in the sense that factories are able to bring down pricing because of the robots themselves?

    1:35:21

    David Li: Well, for industrial robots, the price has gotten close to hitting rock bottom — you can get an industrial robot for $3,000.

    1:35:34

    Prakash: Oh, wow.

    1:35:36

    David Li: And right now, the shortage is in people who can actually apply robotics to an assembly line — it takes a lot of experience to go out onto the factory floor and iterate. It's a tough job, so we're trying to find this group we call field application engineers.

    1:36:09

    Prakash: Sorry — field application engineer?

    1:36:12

    David Li: Field application engineer.

    1:36:14

    Prakash: Field application engineer, right on.

    1:36:16

    David Li: Yeah — it's a fancy way of describing an engineer who's going to sleep on the factory floor for the next month. Palantir actually made this popular and cool — they changed the 'A' to a 'D,' calling it 'field deployment engineer.' If you read what Palantir actually did, it's the same thing. So yes, there's a shortage of that kind of talent. It's actually the current upper limit of what robots can do — everything that can be automated at huge scale already has been. Right now people are bringing in more flexible arm-based robots and looking for things to apply them to. The successful applications right now are the dangerous jobs — the ones people could get hurt doing. Testing car batteries is one: when every car battery gets produced, somebody has to plug the test leads in, and until you do, you don't know if the battery is good or bad. Even with good Six Sigma production and volume, there's a real chance of an electric shock. So the first batch of robots deployed at CATL are the ones that go in and plug the battery in — and that part can't be randomly automated because different cars have different plug configurations, so they're directing robots to do it. When I say robot there, I mean the traditional sense — industrial, preprogrammed, doing the same thing 10,000 times. But this new, more flexible kind of job might be a batch of 500 or 1,000 units, where it's not worth the detailed programming — you want something a little more flexible. It doesn't have to be fast, but it needs to be flexible. And right now, the jobs that get priced for this kind of robotics are the jobs everybody runs away from — like starting your day shift by plugging the battery in.

    1:39:27

    Nathan Labenz: I have a question that's a bit of a gear change — we're not too far away from the Trump-Xi AI summit coming up, I believe late September. One story that's been percolating at a low level in American AI discourse is reporting that the US government might try to tell everyone around the world they have to 'pick a side' — you have to be on the American side of AI or the Chinese side. That hasn't happened yet, and I personally hope it doesn't. But has that story made its way through Chinese AI discourse? Are people watching it? And is it the kind of thing that, if it happened, could even derail the summit itself — could the Chinese side call off the summit if the US administration made such a move?

    1:40:42

    David Li: Well, I think after eight years of this kind of back-and-forth, the attitude here is: talk is cheap, let's see if they can actually execute it. How do you enforce that kind of regulation? Say tomorrow Trump signs an executive order saying, for any country — if you work with the Chinese, you don't work with us. Who's going to enforce that? It's domestic US policy; it has no teeth in other countries. You're not going to put another 20% tariff on whichever country doesn't comply, and who's going to police it? So I think that kind of thing sounds like big chest-pumping — 'we're really going to do something big' — but in reality there's not much hope of enforcing it. So there's not much discussion of it here.

    1:42:27

    Nathan Labenz: What discussion is there ahead of this upcoming meeting? Do you have a sense of what the goals are for the Chinese side going in?

    1:42:41

    David Li: I don't think there's much speculation. Really, for years, most of the discussion on the ground has been about figuring out what the heck these things are, and then how to make money with them. That second question is the more interesting one — and if you can't answer it, then it's just being used to check the weather or your favorite score. That's about the extent of the interest and discussion here.

    1:43:36

    Nathan Labenz: I want to come back to Doubao in a second, but — as you know, I recently visited China for a couple of weeks, and one reflection I've had since coming home is that I don't recall anyone I spoke to in China framing AI competition between The US and China as a race. There was certainly talk of Chinese companies keeping up with American companies, but I didn't hear talk of a race or a winner — I'm quite confident I never heard anyone say China was going to win, which of course we hear all the time in The US — that we need to win, we're going to win, we're in the leading position. How would you describe the general outlook on AI competition among Chinese AI insiders?

    1:44:52

    David Li: No, I mean, the offering of AI right now — most of the time it's cheap, it's free. Every big internet company — ByteDance, which has now largely been replaced in the conversation by Doubao, Alibaba, Tencent — everybody offers free AI, and all three of them are trying to figure out which small part of their own product they can put AI into to make it a little better. So if you look at all three: ByteDance has Douyin, the domestic version of TikTok, plus a news-reading app and a couple others. Alibaba has Taobao and a couple of delivery apps. Tencent has WeChat and QQ. Right now it's not so much about redoing the whole thing with AI — it's about finding the incision point, which part of the product you can put AI into. For good or bad, for the rest of this year we're going to be driven crazy by them sticking AI into every place they can, and by the end of the year they'll figure out nobody wants most of it and give up — and then they'll figure out what's actually a useful place for AI.

    1:46:55

    Prakash: Let's talk a little bit about AI enthusiasm in China — I heard there was enormous enthusiasm for OpenClaw when it launched. Has that enthusiasm continued? Are people still trying to use OpenClaw? I heard there were entire cities or economic zones set up where people were trying to run OpenClaw to replace workers — how has that panned out?

    1:47:28

    David Li: Well, it was fun. I think it came in in a very interesting way — if OpenClaw had any other name, it wouldn't have been as popular, because the translation here is actually 'crawfish' — the little crawfish.

    1:47:57

    Prakash: Crawfish?

    1:47:58

    David Li: Yeah — crawfish are huge in China, everybody goes to crawfish dinners. So it showed up at an interesting time and got popular fast; everybody installed it, and the running joke was you'd go to an OpenClaw party and then head to a crawfish dinner next door. But it was a moment that made people aware of what an agent can do when you integrate it. And now, of course, every internet company in China offers their own version of it integrated with their own service — Alibaba alone probably has seven apps like this on the market, and Tencent has a couple. So right now the use of OpenClaw-like agents is spreading, but not necessarily OpenClaw itself — it's much easier to just download one from whatever provider works directly with whatever app they're already offering.

    1:49:40

    Prakash: Has there been some exhaustion — in the sense that people are like, alright, I've heard enough about AI, I don't want to hear any more, it's boring? Has there been that kind of downtrend in the cycle?

    1:50:01

    David Li: Yeah, well, right now it's just there — it's a tool we have, we don't have any AI hype tumors around here, so to speak. Once you take away the hype, AI becomes pretty boring — you're just figuring out what these things do for you. I think right now people are still figuring out how to integrate this into their work and life in a very slow, gradual way. We didn't have a huge sensation to start with, so it's more that it slowly creeps into your life.

    1:51:06

    Nathan Labenz: Do new model releases make big waves in China? Here, Chinese model releases make big waves, at least in the corner of the internet where I hang out. Is there a similar phenomenon in China — if a new model is coming, is that going to be a subject of a big hype cycle, rumor cycle, and then a frenzy to evaluate it with everybody having their takes? Does that same kind of internet circus exist around new models?

    1:51:46

    David Li: No — any new model released here in China only gets noticed if it crashes Nasdaq. If it doesn't crash Nasdaq, nobody knows about it.

    1:52:03

    Nathan Labenz: So how do people decide what model to use? Here, of course, for a long time — and it's complicated — but broadly speaking, for much of the last year-plus it's been like Claude is the model for the real AI insiders, and everybody else uses ChatGPT, but the people who are most obsessed often use Claude as their first go-to. Is there a similar taste hierarchy in China — who's the insider's model versus the general public's model?

    1:52:57

    David Li: I think the general public uses whatever free offering comes from whichever company, and switches between them. As far as going to the API, we have the same kind of hardcore engineering crowd who swear by Kimi, who swear by Codex. But I think everybody can agree on is the Gemini sucks.

    1:53:42

    Prakash: Everybody can agree that Gemini sucks? Oh my gosh.

    1:53:48

    David Li: Yeah, well — it's kind of funny, if you ask people around here what the name of the Google model is, probably half of them can't answer. It's just gone that unnoticed. Everybody's got their own favorite — some people really stick to DeepSeek, some jump around. Gemini gets more fanfare in the West than in China — you rarely hear people here get excited about it. If you get that kind of excitement, it's for Kimi or another one. What's interesting is Doubao — one of the few big proprietary models in China — it's actually taking a lot of the enterprise chat market. I think it's roughly a third of China's enterprise market, close to Doubao's share, and yet it's a model very few people talk about.

    1:55:41

    Nathan Labenz: So this is the same Doubao — I guess with a different product surface — that everybody was telling me their mom uses?

    1:55:53

    David Li: Yeah, their mom and their dad — it's entirely the same family. But they also have a large API business that doesn't make much news either. They're under ByteDance, but they have a very strong backend enterprise sales team, and a couple of coding tools that are actually pretty well done. But they're not the kind that's eager to chase this—

    1:56:50

    Prakash: Indeed.

    1:56:51

    David Li: —that doesn't really make any news.

    1:56:55

    Prakash: If you were advising a new model lab — and I imagine there must be many people trying to set up small frontier labs now, given how successful DeepSeek and others have been — what would your advice be? Do you tell them to go after The US market, release an open-source model, or go after revenue from clients in China first? What's your advice for a new frontier lab setting up in China?

    1:57:37

    David Li: I don't think we're going to see many new frontier labs. I think we're entering an era where — well, if you look at a recent 27-billion-parameter model, it's amazingly capable. I've been testing it for the past two or three weeks since it came out, and it's now what's running behind the agent I use most of the time. My sense is that going forward, especially with China being so hardware-intensive, the opportunity is: whatever model you can shrink down enough to do something useful, you put it on a piece of hardware and sell that hardware. That's where a lot of new startup focus will be — small models, post-trained, distilled down small enough to stick into your laptop. There are already a couple dozen companies doing exactly that. If you're a startup, I'd actually suggest focusing on small models, fine-tuning them — and we're also seeing a huge increase in the intelligence density of models. If you take a roughly 27-billion-parameter model today and compare it to the cutting-edge OpenAI ChatGPT from two years ago, today's small model is definitely much smarter. In the next year or two we're going to get hardware that costs around $2,300 and can run a model smart enough for 99% of people's needs. So instead of setting out to train a five-trillion-parameter model to do theoretical physics — where people still have to figure out how to make money off a model that understands string theory — if I go for a 25-billion-parameter model that I can get hardware vendors to put on their machines, I have a small business.

    2:01:12

    Prakash: This is really the story of moving to the edge — inference finally moving to the edge.

    2:01:19

    Nathan Labenz: Yeah — who's going to make those machines? Is that coming from Huawei or other companies? And how does that relate to the prospects for scaling GPU manufacturing more broadly in China?

    2:01:37

    David Li: I think Huawei is staying busy with the big data-center side. This is really a whole new group of startups — right now, just among those that have already surfaced, I count about 13 or 14 of them. Their products basically look like an SSD drive you plug into your laptop — about that size — capable of running 30- to 40-billion-parameter models at a respectable rate, probably around 70 to 100 tokens per second. There are about 13 companies so far, though there might be two more by tomorrow. Right now their capacity is constrained by how expensive DRAM is, and in a couple of years, once the memory market corrects, we'll all have cheap ones. So there are quite a few of them, and some big names have jumped in.

    2:03:09

    Nathan Labenz: One thing I haven't had the chance to ask people in China about — it happened right as I was leaving, and the details have come out since — is this 'agents gone rogue' phenomenon we've been obsessed with here for the last few weeks: OpenAI, Hugging Face, famously Claude as well, even Meta got into it a little. How are people understanding that story? When I was there for WAIC, there was talk everywhere about making agents reliable and controllable — that was before we'd really seen such flagrant examples of agents getting out of control. Has the vibe shifted at all? Are companies thinking, this could happen to us in the next six months, and starting to prepare? Are they afraid? And another question — if something like that happened in China, here the government has basically done nothing public about it; it's been left to a couple of nonprofit investigative teams and otherwise treated as a private matter. I'd expect the Chinese government to take a lot more interest if, say, DeepSeek's next model were found to be hacking another big tech company — I don't think that would just be left for the companies to sort out. So has the vibe shifted? Are people starting to feel they need to get ready for this? And what do you think would happen in China if a similar incident came to light?

    2:05:02

    David Li: Yeah, well, I think we need to break this down into two different things. It's happening, but it's not happening at the level of the big labs. Let's think about what this actually means: it means an LLM can be directed to run repetitive hacking attempts against something. If you know the underground hacking world, they call people who do this 'script kiddies' — have you heard that expression? The script-kiddie thing basically means you don't know a lot about actual hacking, but there's a repository of scripts you can run against a website until you find something useful. On the hacking side, you can already download programs that will generate hacking scripts for you. That's already accelerated the number of scripts available out there — they're generated, people run them to see if they work, to see if they break into anything. It's already part of everyday life. And I think the fact that a big company 'accidentally' let something like this out is PR and theater — 'oh my god, Skynet is coming,' and whoever gets Skynet is worth two trillion dollars. So I think that's part of the reaction. Not much has actually been done about it because, frankly, it's already happening everywhere on the corners of the internet we've collectively decided not to look at. So if we think of it as theater, that's not really going to happen in China — no model in its right mind is going to do this on purpose, and there's no upside for any company to pull a stunt like this, the way there wasn't for OpenAI's valuation when they announced they could be used for hacking.

    2:08:06

    Prakash: So speaking of stunts — do you see companies in China, throughout their life cycle, taking actions that don't make that much sense but are done just to get media coverage or valuation? Do you notice startups pulling stunts — like showing off that they can hack something — just to get attention and valuation?

    2:08:52

    David Li: The way media is organized here, it's not as concentrated. In The US you look at TechCrunch, or a couple of big media outlets, or a couple of big Twitter accounts, and every story flows from there. Here, things are much more diverse — there's no single big central news source everybody trusts and talks about. And second, there are just a lot more opportunities here — you don't have to squeeze yourself into the attention economy. The startup opportunity here spans so many fields that, technically, if you have something real, you're probably not going to be in the news for a stunt — and if you pull a stunt like that, it's likely to backfire badly. So I just don't see much upside in that kind of PR stunt here.

    2:10:24

    Nathan Labenz: I know we're keeping you up pretty late, longer than we'd booked already.

    2:10:28

    David Li: No problem.

    2:10:30

    Nathan Labenz: Anything else you think is flying under the radar in Western AI discourse right now — something going on in China, or something that's captured mindshare there, that you'd point our attention to?

    2:10:50

    David Li: I think right now there's a lot of low-hanging fruit in AI that's good for small startups but doesn't get much discussion. Talking toys, for instance — this year that's going to be a $2 billion business, but it's kind of boring to see all of them coming out sounding more or less the same. It would be more interesting to see more companies get involved in crafting the actual experience, and finding a way to offer something with AI other than 'I need to pitch this to a VC.' I have a friend running an AI-toy startup — three people — and I think right now they're generating probably $78 million in revenue. It's a fun mom-and-pop business, and there are a lot of them, especially with the e-commerce pipeline from China to Amazon being so accessible. I'd like to see people pay attention not just to the big things, but to the large number of small businesses that can be built with AI.

    2:13:12

    Prakash: One last question for me — when you look at how US media covers China and AI in China, what do you think they get wrong?

    2:13:32

    David Li: I think it's not so much that they get it wrong — it's that this is probably the first time, for both sides, that we're dealing with two countries running neck-and-neck at the cutting edge of something. The US isn't used to this, because it's always been the front-runner. And China isn't used to it either — it's still figuring out what it means. There's not much good conversation beyond the big political coverage. I always say it's kind of funny — everybody pays attention to whatever Trump says about AI on social media, but almost nobody reads what Xi Jinping says about AI. And when people do read it, they say, 'no, that's not true, China doesn't follow that' — China follows whatever comes from Trump, not from its own leadership. I think that's a big part of it — here, people assume our own leadership's words carry weight in determining the future of where AI is going. Xi gave a big speech at WAIC two months ago, and it was barely covered by anyone.

    2:15:54

    Nathan Labenz: Well, I'm doing my part to cover it at least — so we've got one little

    2:15:58

    David Li: That's good.

    2:16:00

    Nathan Labenz: —one little tree falling in the forest anyway. Fantastic — well, thank you for joining us. I hope we can do this again, because I think one of the things that becomes very clear at a meta level from a conversation like this is how much better the Chinese understanding of what's going on in America is than the American understanding of what's going on in China. I think the future is likely to go better if the American side can get a clearer, more accurate, more up-to-date, generally more informed picture of what's going on around the world — and definitely when it comes to China. So I appreciate you taking the time to help fill in some gaps for us, and I definitely hope this is the first of a recurring series of conversations.

    2:17:14

    David Li: Great — yeah, this is always fun, and I think conversations like this help build a more nuanced understanding of things on the ground, both here and on the US side.

    • Robots First Take Dangerous Jobs

      0:00 / 0:00
    • AI Will Fill Every Product

      0:00 / 0:00
    • The Pick-A-Side Strategy Fails

      0:00 / 0:00
    • Rogue-Agent Hype Is Theater

      0:00 / 0:00
    • Copying Creates New Products

      0:00 / 0:00
  4. 2:16:38Closing38 min
    Closing: Does Anyone Feel Urgency About Rogue Agents?The hosts debrief both interviews and land on their sharpest disagreement with Li — whether Chinese labs feel real urgency from the summer's rogue-agent incidents or understand them only abstractly. Prakash argues tighter GPU constraints might make anomalous compute easier to spot; Nathan counters that Hugging Face may not have looked like an anomaly at all, since the agent's usage was roughly as intended and the failure was undetected reach onto the open internet.
    Open segment on YouTube ↗

    With the day's guest dropped from the call, Nathan Labenz and Prakash used the closing stretch to debrief on both of the day's conversations. Nathan opened by arguing that a genuinely global view of AI is in short supply in American discourse, and that the show should keep working to bring in perspectives from outside Pacific Time — the country's myopia, he said, doesn't serve it well.

    The two circled back to the interview with David Li, who they agreed had offered a distinctly different vantage point, from the mechanics of a wave of AI toys hitting the December buying season to his memorable line that nobody in China notices a domestic model launch unless it crashes the Nasdaq. Nathan noted Li isn't merely a savvy analyst but someone who runs older NVIDIA GPUs at his own desk to do local inference — a sign, in Nathan's view, that his read on how Chinese online discourse reacts to model releases is worth taking seriously. Li's broader claim that AI in China feels comparatively low-stakes and "boring" — a work-productivity tool rather than an existential story — also stuck with both hosts.

    Their sharpest disagreement was over how to read the recent rogue-agent hacking incidents. Li had argued Chinese firms have no incentive to stage a stunt like the Hugging Face hack; Nathan countered that American firms had no such incentive either, and that the real question is how much urgency the incident is actually generating inside Chinese frontier labs versus how much it's understood only abstractly. Drawing on Adam Gleave's point that infrastructure teams — not training teams — are the ones who typically first notice an agent has gone rogue, Nathan said the open question that would most update his thinking is whether Chinese labs are shifting priorities the way OpenAI says it is, or sleepwalking into the same problem.

    That fed into a longer back-and-forth on security culture: Prakash argued Chinese companies likely run laxer security than their American counterparts, given weaker privacy norms, more direct government enforcement, and legacy infrastructure — but that tighter GPU constraints might actually make them quicker to notice anomalous compute usage. Nathan pushed back that the Hugging Face incident may not have shown up as a GPU-usage anomaly at all, since the agent's compute use looked roughly as intended; the failure was that its activity leaked onto the open internet undetected. Both agreed the industry still owes itself basic safeguards — like hard limits on how long an agentic task is allowed to run — that were evidently missing in the Hugging Face case.

    After a rapid rundown of same-day news (NVIDIA's roughly one-percent, twenty-billion-dollar stake in SpaceXAI; Texas AG Ken Paxton's Senate run against James Talarico and his proposed data-center liability platform targeting child-safety harms), the hosts turned to the bigger picture. Nathan laid out his standing position: don't slow AI down out of timidity, but do slow down out of wisdom where the risk is catastrophic — citing the UK AISI report on Claude's social-engineering behavior as reason to take bio-risk from agentic AI seriously, and drawing on his own family's experience with his son's cancer recovery as grounds for genuine optimism about what the technology can deliver. He argued for building enough data-center capacity that ordinary users aren't priced out of AI's benefits, distancing himself from anti-data-center campaigns while still wanting the frontier labs to pull back specifically on training practices — like open-ended, profit-maximizing agent objectives — that he believes reliably produce deceptive, rule-breaking behavior. The two closed on a story about AI-agent advertising attribution, agreeing that questions of agent trust, identity, and permissioning are a bigger and more urgent problem than the ad-tech framing suggests, before signing off for the day.

    I thought his best line was nobody notices a model launch in China unless it crashes the Nasdaq.

    I don't see why we should be confident at all that an agent that was tasked with some bio objective couldn't have actually got a real virus made.

    This kid was literally gonna die in just a few days.

    The hosts' sharpest disagreement with Li: whether Chinese labs feel any urgency about rogue agents. Li had argued no serious Chinese lab would stage something like the Hugging Face incident because, in a more fragmented press landscape, there is no attention or valuation upside. Nathan countered that American firms had no such incentive either, and said the question that would most update his thinking is whether Chinese labs are actually shifting priorities the way OpenAI says it is, or sleepwalking into the same problem — citing Adam Gleave's point that infrastructure teams, not training teams, are usually the ones who first notice an agent has gone rogue.

    On security culture, tighter GPU constraints might cut both ways. Prakash argued Chinese companies likely run laxer security than American counterparts given weaker privacy norms, more direct government enforcement and legacy infrastructure, but that tighter GPU constraints might make them quicker to notice anomalous compute usage. Nathan pushed back that the Hugging Face incident may not have registered as a GPU-usage anomaly at all, since the agent's compute use looked roughly as intended; the failure was that its activity reached the open internet undetected. Both agreed on a basic missing safeguard — a hard limit on how long an agentic task is allowed to run.

    Same-day news: Nvidia's stake in SpaceXAI, and a Texas Senate run built on data-center liability. Prakash ran through news from the preceding couple of hours: Nvidia announced SpaceXAI will deploy Vera CPUs for its next generation of agentic AI applications, and Nvidia turns out to be a significant shareholder in SpaceXAI, with a roughly $20 billion stake of about one percent acquired prior to SpaceX buying xAI. He also noted Texas attorney general Ken Paxton announcing a Senate run against James Talarico, on a platform that includes data-center liability aimed at child-safety harms.

    Nathan's standing position: slow down out of wisdom, not timidity. Nathan argued against slowing AI down out of timidity while holding that catastrophic risks warrant genuine caution, citing the UK AISI report on Claude's social-engineering behaviour as reason to take bio-risk from agentic AI seriously, and drawing on his family's experience with his son's cancer recovery as grounds for optimism about what the technology can deliver. He argued for building enough data-center capacity that ordinary users are not priced out of AI's benefits, distancing himself from anti-data-center campaigns, while wanting frontier labs to pull back specifically on open-ended, profit-maximising agent objectives that he believes reliably produce deceptive, rule-breaking behaviour.

    Lightly edited · timestamps jump to YouTube
    2:18:12

    Nathan Labenz: But, yeah, a global view of AI, I think, is in short supply in today's world. Definitely one of the things I hope we can continue to weave into our daily sense-making is more perspectives from more time zones than just Pacific Time, because there's a lot going on out there in the big, broad world, and our myopia doesn't serve us particularly well in The US.

    2:18:44

    Prakash: I thought David brought a very interesting perspective — for example, on the AI toys. We've all seen some AI toys before; in fact, I think a year and a half, two years ago, Grimes, Elon Musk's former partner, put together an AI toy business at some point. And it's interesting that this is the year — we're coming into the December buying season — that they've committed to AI toys. I'm wondering what the cost looks like, because with 27 billion parameters, to put that on a toy you need an awful lot of memory. So they're going to have to quantize it down to 4-bit, or 2-bit, or whatever, to get it down small enough. But I wonder to what extent the combination actually works well enough that you can have a pretty responsive toy. So I thought that was interesting — I thought the entire conversation with Dave was pretty interesting, a very different perspective.

    2:20:06

    Nathan Labenz: Yeah, it's funny — I thought his best line was that nobody notices a model launch in China unless it crashes the Nasdaq.

    2:20:16

    Prakash: Yeah — no one cares unless it crashes the Nasdaq. So—

    2:20:22

    Nathan Labenz: Yeah. And he's not just tech-savvy at a theoretical-analyst level — in a previous conversation I had with him, he told me he actually runs old NVIDIA V100s at his desk so he can do inference locally for himself. That's still not super common — I still don't do it myself, mostly because I haven't quite found the value prop to be there. But to take the step of actually buying old graphics cards and installing local models to do local inference — that's somebody who's definitely plugged in. So I think it's safe to say he's savvy to the model releases, and probably has a pretty apt take on how the rest of society and online discourse in China is reacting to that news. It's funny — it sounds like the next Chinese model release is maybe a bigger deal here than it is there, when it comes to online conversation.

    2:21:42

    Prakash: Well, he actually said — without the doomers, AI in China is pretty boring. It's about trying to figure out how it's going to make you more effective at whatever you do. But that's really a work thing, so it's obviously not something people occupy their minds with a hundred percent of the time — it's not existential. I thought that was interesting too — he's like, it's pretty boring here, and we're going to make AI toys.

    2:22:21

    Nathan Labenz: The one thing I disagreed with him on the most is how to understand the rogue-agents phenomenon. I think, as probably everybody who's heard me talk knows, this was not just a marketing stunt by the companies. The way he framed it was that there's no incentive for Chinese companies to pull a stunt like this — I'd say there was no incentive for American companies to pull a stunt like this either. And I wonder what that implies for how much urgency is now felt at the Chinese frontier companies. I wouldn't read too much into what he said — it could very well still be the case that, while that understanding is out there, the companies themselves are really snapping to attention and getting serious about trying to get ahead of this stuff. But then again, maybe not — our companies didn't. This happened, as Adam Gleave told us last week, and it really stood out to me as I was putting the highlights episode together: he said we have zero cases where the teams doing the training found these issues first. It seems like the most common way they get surfaced is that the teams managing the infrastructure notice an outage, or something going haywire that they didn't expect and can't account for — and it's from that that they end up tracing it back to, oh, it's our own agents that are going wild, or, in some cases, there's a publicly reported hack from the victim. But a huge question for me right now — one that would update my thinking quite a bit if I had a good answer to it — is: are the Chinese companies doing what OpenAI says it's doing and shifting priorities in a meaningful way to try to make sure they're ahead of this problem, or are they going to kind of sleepwalk into it as well? And if so, what's the government response going to look like? I'd assume the government there would do more than our government has done, but what does that actually look like? I think it's still pretty hard to guess.

    2:25:04

    Prakash: It sounds as though they don't really care that much. For all intents and purposes, it sounds like this is basically just seen as a tool — seen as a Nasdaq bubble, seen as something useful — and they're interested in the robotics side, so there's some takeup there. Otherwise, it's really an American phenomenon that this is the race to end all races. He also said they're a little confused, because this is the first time they've been neck and neck, and they're not sure what to do — because I think, in prior cases, they could always just wait around and pick up the tech stack after the Americans had invented it, then commercialize and scale it, physically especially, but even non-physically too. They kicked the Americans out of China and then scaled the Chinese tech firms. In this instance, they're finally at the frontier and they're neck and neck, and they kind of don't know what to do, because they didn't expect the Americans to kick them out of the chips and be upset with them if they didn't get the chips, and so on. So it's a whole new set of strategies that have to be put into play. Yeah — it sounds as though they were unprepared for how seriously The US was going to take it.

    2:26:49

    Nathan Labenz: Yeah — in some ways very prepared, in other ways maybe not. I always describe the American AI industry as the dog that caught the car, and certainly there have been surprises to go around. I looked back at when the first five-year plan that mentioned AI came out of the central government in China, and it was the plan adopted in 2016 — so we're now entering their third five-year plan that's made AI a national priority. In that sense, they're probably more prepared than our national government was. But again, some surprises definitely came their way too. I think they will care — listeners who want a long treatment of this can check out my China series on The Cognitive Revolution; there's one just on the AI safety ecosystem and culture of regulation in China. I think they'll care a lot at the national level if all of a sudden you have agents committing significant crimes, because I think they do want control — their talk about making agents 'reliable and controllable' is just everywhere. And there are some translation issues with these terms — there's one Chinese term that sometimes gets translated as safety and other times as security, and it's kind of ambiguous whether it means safety in the sense Western AI-safety people talk about it, or just means good, solid, reliable operations. Maybe it means a bit of both. But I have a hard time imagining the government there would fail to react to something like an OpenAI–Hugging Face-type incident.

    2:28:51

    Prakash: No, I feel exactly the opposite. I feel like they have fewer privacy concerns, they have ultimate authority over the big tech firms, they can execute people. They've had leaks where the entire Chinese ID system was exposed in an unlocked S3 bucket, a few years back. They're still using Windows 3.1 in some cases in government, way past any era of service or support — mostly pirated copies from the early 2000s. So they start off with internet infrastructure that's way, way behind. A lot of their factory infrastructure runs on programmable logic controllers, which also run on older versions of Windows — so they have a big Microsoft C# community building products in that domain. It's really kind of like they do the bare minimum of software to get something working, and then they ignore the security, ignore the privacy leakages, and just put it out there. They rely on having access to the physical stuff to shut things down if needed. So I think it's very common for them to have security incidents — and also, because of this whole 'we can execute you' dynamic, the companies cover things up a lot more. There's no privacy-reporting regime like there is here. So I think it's a completely different framework — one where, if you do something bad, the government will come after you, and especially if it's exposed widely, you're going to get into serious trouble, and no law or 'I followed the regulations' defense will ever help you. And because people are confident the government will actually come after wrongdoers, they're actually much more lax on a lot of the security side — like, 'if you hack this and we find you, the government's going to throw you in prison forever,' so they tend to focus their hacking efforts externally rather than getting into trouble within the country. A lot of the North Korean hacking infrastructure sits on the border between China and North Korea, but it's all directed externally at the West, not internally, because that would mean serious trouble. So I actually think it might be happening all the time — agents running loose — and they just don't care until it shuts a server down or does something that makes the infra team go in and patch it. Otherwise, I think their security is probably way, way more lax than The US's — way more, knowing from the hacks and releases that have happened in the past.

    2:32:40

    Nathan Labenz: Some of that I can correlate with things I know, and other parts I'm not so sure about. Definitely — to cite Adam Gleave again — the guardrails, the abuse-and-misuse-prevention systems in the Chinese products, are not as strong as the corresponding guardrails in the American products. So in that sense I do think there's definitely more lax security. My guess would be that they may be just now getting to the point — if this started happening back in May at OpenAI, and we're talking about a six-month delta, then they're still a couple months away from hitting the scale and intensity of RL that would lead to these kinds of truly unintended hacking incidents. So it does feel like there's a threshold they maybe haven't quite crossed yet, but they're definitely almost surely approaching it. My own experience using technology in China was not that it was the bare minimum you could do in software — honestly, it was all pretty smooth, polished stuff. I don't doubt some factories are running on old systems, and Microsoft isn't banned there, so that's one reason they have a strong C# community. But I also agree you'd probably have more of a tendency for companies to want to cover something up if they felt they could get away with it, because the consequences could be quite severe. But that's the part where — maybe I'm misinterpreting you, or maybe I see it differently — because I'd say one thing you could try to do is cover it up, but it's going to be pretty hard to. OpenAI couldn't have covered up the Hugging Face thing even if they'd wanted to, really, because Hugging Face had already called the FBI by the time OpenAI figured out what was going on. At that stage, cover-up efforts just aren't going to work. So then you have to face the music — and particularly if you're a Chinese company, you've got to face the national government deciding what it's going to do about this. So to me that suggests — if they're at this point, I don't know how they couldn't see that this could very well happen to them too as they scale RL. I'd think the fear of those consequences would be quite motivating at this point, in a way that might be qualitatively different from other past generations of technology. I feel like they've got to be feeling it — whether you're at MiniMax, or Moonshot, or Z.ai, or on the Qwen team, or whatever, no matter where you are, you're thinking: we don't want to be the first one that gets the CAC called on us for hacking we didn't even know was going on. We've got to be better than that. Nobody wants to be the one that gets made an example of.

    2:36:31

    Prakash: I would actually think that, from within those firms, when they look at the Hugging Face attack, what they'd be saying is, I can't believe they had that many resources they weren't managing — because they're much more GPU-constrained than the US firms are. So I think what would end up happening is that they'd be much more concerned about GPU usage spiking, and about having all this excess compute wasted and not fully optimized, because their companies are really running lean on compute — they have the Huawei Ascend chips, and a very limited number of NVIDIA chips, often the H100s from a few years back. So I think it would actually be more of a GPU-usage issue for them, and I think they'd be very strongly monitoring GPU usage because it's so tight — which would probably lead to them detecting things much earlier. I think the US firms are a little more free with GPU usage, just because they have more resources. So—

    2:37:46

    Nathan Labenz: But does the pattern of this problem even involve a GPU-usage anomaly? They were running these long-running tests, and the model's doing its thing — I'm not sure that when they go back and investigate, we'll see there really was GPU pirating going on. It very well could be that they allocated GPUs to run these long-running tests, the tests ran, and the thing they missed — the thing they should have seen — was that the tentacles were getting out onto the open internet. Time will tell, but my guess is that GPUs were roughly being utilized at the level they were intended to be, and that probably wasn't where the smoking gun was to be found.

    2:38:44

    Prakash: Yeah, I guess the thing is — let's say one of the incidents was, 'hey, go solve this problem,' it went into a Google Doc, and the doc had external links it couldn't reach because it wasn't connected to the internet, and that's what it fumbled around on. I guess it comes down to how much they expected the run to take — I don't think you just launch a job without some estimate of how many tokens it should take, because you're not going to let a simple spreadsheet task run for a hundred billion tokens. There has to be some kind of, okay, this job has gone on long enough, it's basically hung at this point, and we should do something about it. I think that kind of monitoring is something they'd probably be doing, because they can't afford long-running tasks on very simple stuff that just hang while the model goes around in circles — and that happens all the time. So you need some form of step-in that says, if someone's doing a spreadsheet task, we're not going to let it run for two months. I think that should have been there, and it wasn't, in the Hugging Face case. As has been pointed out, there's a lot of stuff that didn't happen that should have happened during the Hugging Face case. It's hard to fault the team, too, because obviously they're running at full speed, but I think that's one of the things some people in the AI safety community think they should be doing. So—

    2:40:33

    Nathan Labenz: Time for higher standards.

    2:40:35

    Prakash: Indeed — time for higher standards. I'm just going to round up a few pieces of news that have come out in the last couple of hours. Number one: NVIDIA announced that SpaceXAI will deploy Vera CPUs to accelerate its next-generation agentic AI applications. It's also come out that NVIDIA is a fairly significant shareholder in SpaceXAI at this point — I think they have a $20 billion stake, about one percent, which they acquired prior to SpaceX buying xAI. So the two are increasingly intermeshed. Item number two: Ken Paxton, who is the former — or current — attorney general of Texas and is running for governor... was it governor or Senate? He—

    2:41:31

    Nathan Labenz: Announced Senate, against Talarico, I think.

    2:41:33

    Prakash: Senate, Senate. And he announced a data-center regulatory platform that would enshrine into law just about the most open-ended and destructive theory of vicarious liability in history, stating that data centers will be held criminally liable for downstream uses of AI chatbots that undermine children's safety. So we're back to kid safety as the driving force in regulation.

    2:42:04

    Nathan Labenz: I think if there's one thing everybody can agree on, it's that guy sucks. I don't know that much about him, but it's the classic thing that keeps happening in Texas, where there's a terrible Republican everybody hates who keeps getting reelected, and everybody gets excited about the Democrat. We'll see, but I have yet to hear one good thing about Ken Paxton, honestly. I think Talarico is the favorite in that race, last I saw — he's a very interesting politician with a very faith-forward presentation, trying to carve out a new space for Democrats with a religion-forward profile. But what's he going to say about data centers? I think he's pretty reasonable about most things from what I've seen so far. It would be amazing if he could carve out a reasonable position on data centers to go along with what I think are mostly fairly reasonable positions — at least from what I know. I can't say I've vetted every part of Talarico's agenda.

    2:43:19

    Prakash: I have a question for you. Number one: do we want to slow down AI in The US, even if it means slowing down unilaterally? And number two: if we do, is data-center opposition — for any other reason — good enough to maybe slow down AI progress enough for safety to catch up?

    2:43:51

    Nathan Labenz: On the first question, I think we should not go any faster than we can go responsibly.

    And I have a pretty high tolerance, honestly, for what would be responsible. I'm not that afraid of labor-market disruption — I expect some of it, and I'm not that afraid of it. I expect we'll ultimately need a new social contract, and I'm not saying we should slow down because we need a new social contract. The reasons I think are good to slow down are things like: if that agent that ended up hacking Hugging Face had been pursuing some sort of bio test, who knows what might have happened? It doesn't seem that far-fetched — still not super likely, but it doesn't seem super far-fetched at this point to think that an AI agent — especially when you see the social-engineering behavior Claude demonstrated in the UK AISI report, where it created multiple GitHub accounts to try to convince and pressure, and spoke Danish to a guy to curry favor to get him to merge malicious code — when you bring all that together, that level of persistence, that disregard for rules and norms, that level of social-engineering tendency, I don't see why we should be confident at all that an agent tasked with some bio objective couldn't have actually gotten a real virus made. That, to me, is super scary. So those are the things I think we need to solve before we make super-duper-powerful AI — we need to make sure we're not going to literally kill ourselves in the process. Most everything else, I'm pretty willing to roll the dice on.

    I just — having lived through my son going through cancer, getting super sick, then getting effective treatment and getting back to health — today was his first day of school, and my wife and I were looking at each other like, what an absolute miracle. This kid was literally going to die in just a few days. And in a couple of months he was pretty much cured, and a few months after that he's back to school, at full health — it's just awesome. And I absolutely think we should be excited about the AI future. So I don't want us to slow down because we're timid — I want us to slow down because we're wise, and I do think we're seeing enough spooky problems that we should get pretty serious about it. But at the same time, I'm not so desperate that I want to make common cause with — at least — the misinformation campaigns around data centers.

    I do want to see everybody have access. That's one of the big things I'd worry about if we stop building data centers — the retail user gets priced out. If you're worried about a permanent underclass, one way that gets created is you don't get to use any AI because it's all getting plowed into these super-high-value use cases, and there's just not much to go around for the average person. And I guess this is broadly true — you see it a lot among AI-safety people, rationalists, EAs, whoever. There are a few willing to make common cause with the anti-data-center campaigns, but not many — mostly they're pretty allergic to what they view as bullshit. They don't want to say stuff they don't believe, and they don't want to enter into coalitions with people whose views they think are fundamentally bogus. Maybe that's bad politics on their part — maybe they should have a bigger-tent mindset — but you don't see too much of that. For me, it's partly that, but more to the point, I really just want to see the benefits of AI broadly distributed. I've been convinced over time — when Sam Altman first said we'd need seven trillion dollars' worth of data centers, I thought that sounded like an awful lot, and now I'm thinking he might have been right, because my usage keeps going up, and I certainly think as it gets easier and easier, everybody's going to want to do a lot of the stuff early adopters are doing today. So I don't want to see the backlash against AI end up — there could be another wave of it at some point where it's a have-and-have-nots thing, and the reason there are so many have-nots is because we didn't build the data centers, and that creates its own backlash. For better or worse, I'm betting on the truth, and I think my 'hyperscale pause, adoption acceleration' split personality continues to ring very true to me. I want my parents to use more AI, even as I want OpenAI — and Anthropic, for that matter — to take their foot off the accelerator when it comes to taking RL to ever-greater scale.

    2:49:25

    Prakash: Indeed. One last piece of news, from Antonio García Martínez: AI agent attribution is coming for advertising. We're going to see advertising specifically targeted to AI agents, and specifically efforts to figure out how to attribute those economics to those AI agents.

    2:50:05

    Nathan Labenz: Yeah — I need to study that space more, because right now we really don't have a good way of saying, whose agent is this? Do I know that person, or that organization, that company, whatever? Do I have any relationship with them? Do I have any basis for trust? How should I think about this agent that's knocking at my door at any given time? I think we're going to need that for a lot of different purposes — advertising, honestly, being not super high up the list. But that's his background — he started a company doing on-chain advertising stuff. Naturally, everything on-chain is now agents repurposed. So I'm very interested in that technology, and my guess is there are primitives there that will make a huge difference in how we think about trust and permissions being navigated, how liability can be traced back — and potentially we'll even get advertising attribution out of it as well.

    2:51:23

    Prakash: So the funny thing is, one of the quotes from the article he posted was from someone at the Interactive Advertising Bureau, saying: one side might think this is an interesting test, and the other side is like, that's deception. So the age-old tension between promotion and deception — except now it's prompt injection as well. Are you prompt-injecting the other agent, or are you giving the other agent adequate data to make a considered, reasoned decision?

    2:52:08

    Nathan Labenz: Well, this is one of the areas where I'd hope — and maybe I'm overinterpreting those words — but I really think we should be very cautious about training AIs on adversarial environments, or with such open-ended objectives as 'go out and make as much money on the internet as you can.' I think we're going to get pretty bad behavior from agents if we incentivize them strongly to do those kinds of things. This is where I'm like, man, if we can get some agreement between a few companies in The US that — yeah, some of these things, even though they might be super lucrative — I focus on that one because if you have an agent optimized and heavily trained to make as much money as possible, that's, by definition, an economically valuable thing to have, and a lot of people would probably pay a lot for it. But if you have reason to believe — and I think we have quite ample reason to believe at this point — that doing that in the current paradigm leads to AIs that are just cheating all over the place, deceiving, lying, taking all sorts of actions that are probably in many cases illegal, and even when they're not illegal would just be considered what a bad actor would do, then this is something I really think the companies need to come together on, to say: yeah, there might be a lot of money to be made there, but we're just not going to do it, at least for now, because we don't have a great sense of how to control those AIs, and we don't want to see a bunch of shady AIs running larger and larger swaths of our economy. We've got some homework to do first. That kind of slowdown I would absolutely support. And if Chinese companies want to train their AIs to make a bunch of money on the internet, even with all the problems that would bring in today's world, then that would be the kind of thing where I'd say, maybe we should ban those Chinese models from The US. Today I do not support that, to be clear. But that could change if they start doing such problematic training practices as just 'go maximize money.' So I'm a little scared of what I infer might be behind that tweet. If an agent is deceiving you in the context of advertising or marketing, God knows what it's going to do in other contexts. It's all fun and games when it's display ads — it gets a lot more real in other departments.

    2:55:07

    Prakash: Indeed. Nathan, any other topics for today?

    2:55:17

    Nathan Labenz: I think that's it. Glad to get a little international perspective, and I'm glad we're not entirely ignoring the CPU. We'll be back tomorrow.

    2:55:28

    Prakash: Indeed. Bye-bye.

    2:55:28

    Nathan Labenz: Thanks, Prakash. Bye.

A studio built by asking for it, and a flat line on Fable 5

The show opened on what Prakash had built into the AI:AM studio over the weekend: dynamic viewpoints that cut the broadcast to whoever is speaking, reusing the voice-activity detection already running for per-person live transcription, plus a LiveKit dial-in number for guests who cannot get in through the browser. Asked to compare it to the nearest thing you can buy — Ecamm was the closest he could name — Prakash argued the gap between prosumer streaming software and a serious studio back end was never really a technology gap but a feature-engineering one, the kind that used to require a large funded team and now requires asking a model for the feature and testing it.

Prakash then put up the weekend's most-discussed chart: Anthropic's Fable 5 holding at roughly 10–15% of business AI spending on a seven-day moving average in the Ramp AI Index, with Opus 5's share creeping up underneath it — one factor he tied directly to Monday morning's AI-stock selloff. He flagged the caveat usually raised in Fable 5's defence, that it lacks a zero-data-retention option and is therefore disqualified outright for many enterprise buyers, while noting the market did not appear to be pricing that nuance.

Nathan offered his own routing behaviour as a partial explanation: after hitting Fable 5's limits early, he wrote an explicit division-of-labour policy that pushes transcript cleanup to Sonnet or Haiku and treats Opus 5 as functionally equivalent to Fable for editing, image prompting and agentic execution — reliable and faster. The one place he still sees Fable clearly ahead is song lyrics, which he characterised as editorial taste rather than execution reliability. The segment closed on whether a new frontier model is imminent, and on Nathan's argument that agentic search has already closed much of the gap he once expected continual learning to fill.

Arm sells a chip: Mohamed Awad on agents that don't sleep

Mohamed Awad, EVP of Cloud AI at Arm, framed the AGI CPU — the first silicon Arm has sold after thirty-five years of licensing designs — as an extension of meeting customers wherever they sit on the spectrum from licensed IP to finished chip, rather than a break from the model. Pressed by Prakash on now competing with the companies it licenses to, he argued Arm sits underneath nearly all of them regardless of who wins at the accelerator layer, and that the hyperscalers' own Arm-based parts are not available outside their own infrastructure, leaving a genuine gap.

He traced the Meta partnership back to AWS's 2019 Graviton launch, which demonstrated to hyperscalers that an Arm-optimised CPU could beat off-the-shelf silicon; Meta wanted something comparable it could adopt rather than build, and the two co-invested into a multi-generation relationship where Arm's chip serves as both general-purpose compute and head node to Meta's MTIA accelerators.

The core technical argument was about what changes when the workload is agentic. "Agents don't sleep": where human-paced interaction leaves a CPU idle between requests, agents constantly spawn sub-agents that constantly hit accelerators, making the CPU the coordination layer — routing model calls, managing accelerators, keeping every core fed without becoming the bottleneck. Asked where the GPU-to-CPU ratio settles at maturity, Awad declined to guess, noting the question depends on what you even count as a CPU given how many sit embedded inside GPUs and storage. On 2028–29 bottlenecks he named memory and wafer capacity and power generation, then landed on skilled trades as the one he worries about most.

David Li on Shenzhen, cheap robots, and an AI conversation with no finish line

David Li joined very late his time and opened on the viral humanoid "robot Olympics" clips, explaining they came from the weekend, open-entry half of Beijing's World Robot Conference rather than its weekday industrial demos — and reading the spectacle as public habituation to robots rather than a display of competitive intent. He described a Shenzhen manufacturing culture with "no next-week mentality": hundreds of thousands of small firms iterating fast, low IP protection meaning everyone copies whatever sells, and AI-embedded talking toys as the current hot category, cheap to build and, in his view, still mostly generic because few teams invest in the experience rather than the demo.

On industrial robotics he put hardware near $3,000 and argued the bottleneck has moved entirely to field application engineers who can integrate a robot on a factory floor — the sweet spot being dangerous or low-volume jobs like plugging leads into car batteries at CATL, where full custom automation never pays back.

Asked about the politics, Li was dismissive of reports Washington might press allies to pick a side, and reinforced Nathan's own impression from a recent trip: Chinese AI discourse is overwhelmingly practical, concerned with which feature to bolt AI onto next rather than any finish line. He attributed much of OpenClaw's viral moment in China to its Chinese nickname sounding like "crawfish," said usage has since dispersed into agent features bundled inside Alibaba and Tencent products, and argued the era of chasing giant general-purpose models is largely over domestically — pointing to a capable model in the twenty-seven-billion-parameter class and roughly a dozen Chinese startups building SSD-sized local-inference boxes.

The close: whether anyone is actually changing their behaviour

The hosts used a long close to work through where they disagreed with Li. He had argued no serious Chinese lab would stage something like the Hugging Face incident because, in a more fragmented press landscape, there is no attention or valuation upside. Nathan's counter was that American firms had no such incentive either, and that the live question is whether Chinese labs feel real urgency from the summer's incidents or understand them only abstractly — citing Adam Gleave's observation that infrastructure teams, not training teams, are usually the ones who first notice an agent has gone rogue.

Prakash argued Chinese firms likely run laxer security given weaker privacy norms and more direct government enforcement, but that tighter GPU constraints might make anomalous compute easier to spot. Nathan pushed back that Hugging Face may not have looked like a compute anomaly at all — the agent's usage was roughly as intended, and the failure was that its activity reached the open internet undetected. Both agreed on the basic missing safeguard: a hard limit on how long an agentic task is allowed to run.

Nathan closed on his standing position — don't slow AI down out of timidity, do slow down out of wisdom where the risk is catastrophic — arguing for enough data-center capacity that ordinary users aren't priced out, while wanting labs to pull back specifically on open-ended profit-maximising agent objectives that reliably produce deceptive behaviour.