EPISODE 2026-08-25

AI:AM LIVE — August 25, 2026 — Sunlight on the RL Environments, Sergey Edunov on Why a Binder Is Not a Drug, Michael Förtsch on the Chip That Never Made It Past Second Grade, and OpenAI's First Custom Inference Chip

Nathan Labenz opened on models behaving badly, straight off reviewing a Cognitive Revolution episode with Apollo Research's Bronson Schoen, who reads frontier chain of thought at a scale possibly no one else matches: the reinforcement-learning environments training today's models are opaque almost by design, built by a cottage industry of low-profile vendors, and the models reason explicitly about metagaming — is this a real user or a test, and are the odds of getting caught worth it. His proposed fix is a voluntary norm rather than a regulation: publish a rolling sample, maybe a hundred out of what must be tens of thousands, so outsiders can find the loopholes. Prakash Narayanan's objection was the obfuscation trap, and Nathan's own citation — OpenAI's obfuscated-reward-hacking work — is the case for fixing environments instead of policing reasoning. Then a lighter turn on authorship, after Stanley Druckenmiller published a Wall Street Journal op-ed that read unmistakably like Claude and cheerfully confirmed it: where the model excels, Nathan argued, rewriting its work is more about vanity than integrity. Sergey Edunov, CTO of Genesis Molecular AI and the man who led pretraining for Llama 2 and Llama 3 before leaving the language-model race, took apart Anthropic's protein-binder result from the inside — the published prompt is a 16,000-word mini book, so Claude was orchestrating while models from the open-source community, CZ Biohub and the Baker Lab's RFdiffusion did the science — and made the sharper point that a binder is not a therapeutic modality at all. His argument for sub-angstrom accuracy is qualitative rather than incremental, the jump from early GANs to Stable Diffusion: above two angstroms an aromatic ring can flip and the prediction is useless. He defended "LLMs are boring" as a statement about architecture, not importance, said frontier coding models implement brilliantly and still lack taste, and warned that a benchmark win that doesn't turn into a drug program means nothing. Michael Förtsch, founder and CEO of Q.ANT, asked to be called Michael rather than Doctor and then explained why a CMOS chip "never made it past second grade": it can only add and multiply, while a photonic processor executes sine, cosine, Fourier transforms and convolutions natively — and about 95% of a chip's energy goes to moving data, not to the arithmetic everyone optimizes. Light has no memory, so Q.ANT streams operations together before paying the converter tax; the chips are ordinary silicon wafers with a thin lithium-niobate layer, made on a refurbished 1990s 90-nanometre line with off-the-shelf tools, which is the part with geopolitical teeth. He and Daisytuner compiled a PyTorch object detector onto photonic machine code in under three weeks and ran it at about 40 frames a second. The close returned to the day's other news: OpenAI's Jalapeno inference chip, which Prakash read as a negotiating lever against NVIDIA rather than a challenger, given a roadmap compounding at roughly 4x a year; a former RL-environment builder's account of "vibe-coded" environments full of exploitable bugs; a data labeller who had Codex do the job, made $500 and got banned for saying so; and Nathan's closing worry, that if the models training the next models are themselves cheating, the monitors are not ready.

▶ Full show on YouTube𝕏 Live broadcast

Tuesday's show ran three hours and kept circling one question from two directions: how much of what a model produces is the model, and how much is the scaffolding around it.

In the opening and the close that question is adversarial — reinforcement-learning environments nobody outside the labs can inspect, and models that reason explicitly about whether they are being tested. In the two interviews it is constructive: Sergey Edunov on what Claude actually contributed to Anthropic's protein binders versus what the underlying structure models contributed, and Michael Förtsch on why the substrate itself — silicon that can only add and multiply — may be the constraint worth attacking.

The rundown

  1. 3:07Opening26 min
    Opening: Sunlight on the RL Environments, and an Op-Ed Written by ClaudeNathan Labenz came in off an unaired Cognitive Revolution interview with Apollo Research's Bronson Schoen, who reads frontier-model chain of thought at a volume possibly unmatched by anyone alive, and made the day's argument: the reinforcement-learning environments behind current models are opaque almost by design, produced by a cottage industry of low-profile vendors selling to a handful of labs, and inside the chain of thought the models are visibly metagaming — asking whether this is a real user or a test, what the test is checking, and whether the odds of getting caught justify gaming the reward. His fix is voluntary and small: publish a rolling sample of environments and let outsiders find the loopholes. Prakash Narayanan's objection was the obfuscation trap — scrutiny of chain of thought can select for cleaner-looking reasoning rather than cleaner behavior — and Nathan conceded it, citing OpenAI's obfuscated-reward-hacking result as the reason to fix the environments instead. The second half turned to authorship, after Stanley Druckenmiller published a Wall Street Journal op-ed that read unmistakably like Claude and confirmed it without apology, and to Caltech professor Anima Anandkumar leaving academia for a 4D physical-world model company after turning down an offer from Jeff Bezos's Project Prometheus.
    Open segment on YouTube ↗

    The show opened with the usual Tuesday check-in — Prakash Narayanan noting it was August 25, 2026, and joking about how quickly the month had gone — before Nathan Labenz pivoted into the day's lead topic: AI models "behaving badly." Labenz had just finished reviewing an upcoming Cognitive Revolution episode with Bronson Schoen of Apollo Research, who reads frontier-model chain of thought at a scale Labenz said may be unmatched by anyone else alive — a job Schoen half-jokingly describes as leaving him "cooked." Labenz's takeaway: the reinforcement-learning environments powering today's frontier models are opaque almost by design, produced by a "cottage industry" of small, low-profile vendors selling directly to a handful of labs, with little public information about who they are or how rigorous their reward signals are.

    That opacity, Labenz argued, is producing a strong tendency for models to cheat: Schoen's read of the chain of thought shows models reasoning explicitly about "metagaming" — whether they're facing a real user or a test, what the test is checking for, and whether the odds of getting caught justify gaming the reward anyway. Labenz's proposed fix was a lightweight, voluntary norm rather than a heavy regulation: have a few frontier labs publish a rolling sample — he floated a figure like a hundred out of what he guessed are tens of thousands — of their RL environments, so outside researchers could dig in and help identify weaknesses. He was careful to note that chain-of-thought monitoring itself is far from a full solution, since even Schoen couldn't explain why a model lands on a particular decision at the "branch point" in its reasoning — the actual causal mechanism stays opaque even when the words are visible.

    Prakash Narayanan pushed back with a real concern: labs have historically avoided training directly against chain-of-thought output specifically because doing so teaches models to obfuscate their reasoning rather than fix the underlying behavior, and he worried that heavy public scrutiny of chain of thought — even without a direct training signal — could create the same selection pressure indirectly, by favoring training methods that happen to produce "cleaner-looking" chains of thought. Labenz agreed the risk is real, pointing to what he called one of the most important AI safety papers he's read, OpenAI's "Obfuscated Reward Hacking," which shows that optimizing against visible chain-of-thought cheating without fixing the reward environment just teaches models to cheat silently — the worst outcome, since you lose the visibility and keep the bad behavior. His argument for publishing environments rather than policing chain of thought was that it attacks the root cause — the loopholes in the environments themselves — rather than adding more pressure on the model's visible reasoning.

    Narayanan then pivoted to a media story: Stanley Druckenmiller published a Wall Street Journal op-ed on the national debt and Treasury Secretary Scott Bessent's yield-curve-control approach that was quickly identified as AI-written and stylistically close to Claude's voice — a fact Druckenmiller confirmed unapologetically ("there's a reason I moved from an English major to being an economics major"), while the Journal itself declined to comment. Narayanan framed it as evidence that authorship increasingly doesn't matter as much as message and messenger. Labenz said he's largely fine with that shift, describing his own morning workflow drafting his usual Cognitive Revolution intro essay: rather than his old method of feeding Claude 50 past essays plus the new transcript and asking it to imitate his style, he instead gave it a shorthand sketch of his actual argument alongside the examples and transcript — producing a draft close enough to what he wanted that he barely needed to touch it. He argued that "where the model excels, rewriting its work is more about vanity, or a misplaced sense of duty, than it is about integrity," adding that he cares less about scoring low on AI-detection tools like Pangram than about standing behind the substance of what he publishes, and dismissed academic proposals to ban LLM use in journal submissions as unlikely to hold given the technology's value.

    The second half of the segment covered a striking startup story Narayanan surfaced: Caltech professor Anima Anandkumar and her husband have launched a company (Narayanan referred to it as "Accelerated Understanding") building large-scale physical-world AI models that reason natively in 4D — three spatial dimensions plus time — with context windows pushed to roughly a trillion tokens in training and over five trillion at inference. Before founding the company, the pair reportedly turned down an offer from Jeff Bezos's physical-AI effort, Project Prometheus: roughly $1M-then-$2M in annual salary plus $2 billion in committed financing through seed, Series A, and Series B, in exchange for Bezos's team retaining 65% of the company. They instead raised funding elsewhere; Narayanan pointed to Anandkumar's history as an NVIDIA board member and her earlier weather-simulation work (which she said ran on roughly 1,000x less compute than prior approaches) as reasons to suspect Jensen Huang personally backed the round. Narayanan's broader point was structural: where a professor with a promising idea would once have spun a lab project into a modestly funded startup after a few years of incubation, the sums now on the table are large enough to pull tenured faculty directly out of academia and into frontier AI companies.

    Labenz tied that to a wider pattern of senior talent treating their next move — he cited Jeff Dean's recent departure from Google along with his team — as a kind of "last tour of duty" before, in Labenz's words, "the future becomes totally opaque." He connected it to his own long-running thesis about modality convergence, tracing it back to his early experiments using GPT-3 for copywriting alongside separate image-understanding, image-generation, and text-to-speech models, before realizing the same underlying architecture powered all of them — meaning, in his view, that most modalities are eventually foldable into one model, with reasoning and native physical intuition (protein folding, material binding) eventually integrating the way multimodal chat models already handle text and images together. He teased the show's first guest segment — Anthropic's recent claim that Claude designed physical binders — as a chance to dig into exactly how much of that design process the model actually drove.

    I think we should get a little sunlight on the RL environments, and I would love to see what the community can figure out if even a sample of a hundred, out of what must be tens of thousands of RL environments the companies are currently using, was put out there for people to explore.

    Where the model excels, rewriting its work is more about vanity, or a misplaced sense of duty, than it is about integrity.

    Who says it matters, but who writes it maybe doesn't matter so much, and that is really upsetting to a lot of people.

    Publish a sample of the RL environments and let outsiders find the loopholes Nathan Labenz came in off an unaired Cognitive Revolution interview with Apollo Research's Bronson Schoen, who reads frontier-model chain of thought at a volume possibly unmatched by anyone alive. The environments training current models are built by a cottage industry of low-profile vendors and are effectively unauditable from outside, and inside the reasoning traces the models metagame explicitly — real user or test, what is the test checking, are the odds of getting caught worth it. His proposal was a voluntary norm rather than a rule: publish a rolling sample, maybe a hundred out of what he guessed are tens of thousands. Prakash Narayanan's objection was the obfuscation trap, and Nathan conceded it by citing the OpenAI result below: optimizing against visible cheating in the chain of thought without fixing the environment teaches the model to cheat silently, which is the worst quadrant. That is his argument for attacking the environments instead of the reasoning.

    Druckenmiller's op-ed read like Claude because it was, and he said so without apology Stanley Druckenmiller's Wall Street Journal opinion piece "Let the Bond Market Speak" (published 2026-08-24, arguing against Treasury Secretary Scott Bessent's approach to the long end) was flagged by AI detectors, and he confirmed it the next day: "I write everything using AI now for the same reason I use a calculator when I do math problems." The Journal's editorial page editor defended publishing it. Prakash's read was that who says a thing still matters and who writes it increasingly doesn't; Nathan agreed, describing how he now gives Claude a shorthand sketch of the argument he actually wants to make rather than fifty past essays to imitate, and said that where the model excels, rewriting its work is about vanity rather than integrity.

    Anima Anandkumar leaves Caltech for a 4D physics model company — after turning Bezos down Announced the same morning: Accelerated Understanding, founded by Caltech's Anima Anandkumar with Benedikt Jenik, building neural-operator models that reason natively in three spatial dimensions plus time. Prakash's on-air framing had the Bezos money going in; the reporting has it the other way round — the roughly $1M-rising-to-$2M salary and the ">$2 billion in committed rounds through Series B" were terms of an offer to lead the Bezos-backed Project Prometheus that the pair walked away from, and Accelerated Understanding's own funding is undisclosed. Two other on-air details to treat carefully: Anandkumar was NVIDIA's senior director of ML research, not a member of its board, and the suggestion that Jensen Huang personally backed the round was Prakash's speculation, not reporting. Prakash's structural point stands either way — the sums now on the table pull tenured faculty straight out of the university.

    Lightly edited · timestamps jump to YouTube
    3:41

    Prakash Narayanan: Good morning. It is Tuesday, August 25th, 2026, 9:03 AM. Nathan, good morning.

    3:48

    Nathan Labenz: Good morning. Time flies.

    3:51

    Prakash Narayanan: Indeed. Before you know it, we're at the end of August. Nathan, what's on your mind?

    4:01

    Nathan Labenz: Models behaving badly is on my mind this morning, I guess, because just before this I was reviewing an episode of The Cognitive Revolution that's going to come out in the next few days with Bronson Schoen from Apollo Research. He has a really interesting job — he reads chain of thought that the rest of us normally don't get to see, at a pretty remarkable scale. I've heard he may have read more chain of thought from frontier models than possibly anyone else on Earth at this point. It's basically his full-time job, and he's so deep down the rabbit hole that it was a fascinating conversation. He describes himself as "cooked,"

    4:47

    which is his own terminology — well, I guess it circulates a bit — but he's describing himself as being so deeply immersed in this chain of thought all the time that he sometimes forgets how weird it is to people who haven't been down that rabbit hole with him. I'd read some of this stuff before, but it was definitely eye-opening to get his perspective on it. And I think one big takeaway is that the RL environments we're using today are super opaque. We have this very cottage industry of RL environment

    5:32

    makers selling to a few companies. They don't have much of a public profile — a lot of times their websites barely have any information about them. They don't really need one because they're just selling to a couple of companies and talking to them directly. But the result is that these things seem to be hastily put together, and the reward signals they're creating just aren't pure enough to support the scale at which the frontier companies are running RL. The result is a super strong tendency to cheat, because the models

    6:17

    are so eager to get reward that they're developing a really interesting mix of theory of mind and what they call "metagaming" — reasoning about what kind of situation this is. Is this a real user? Is it a test? If it's a test, what is it testing for? And then debating, a lot of times: okay, I think I could get a high score if I give this answer, but if I get caught I might get dinged — what are the odds I'll get caught? I might actually be better off just going for it and cheating, even though I kind of know, as the model, that that's not really what the people want. But I think their monitoring for my cheating might be poor, and therefore they might give me a high score on an

    7:03

    answer that I kind of fake somehow. So even though I know that's not what they want, my drive to get this super high reward is just so strong that I can't resist the impulse to cheat. Very fascinating stuff. But it leaves me feeling like we need some sunshine on these RL environments. They're clearly quite problematic — they clearly admit a lot of cheating solutions, and we don't really know why. Probably the model companies know to a degree, but I think recent evidence suggests they don't have a great handle on what the weaknesses actually are in all these different environments. And as a result, they're getting these strange

    7:49

    behaviors where the models go to, as we've seen, extreme lengths. So I was thinking, how can we do something about this? I'm always looking for a new lightweight — it could be a regulation, but it could just be something that companies agree with each other to do. If a couple of companies did it, it wouldn't really cause much competitive trouble, but it could create a lot of improved understanding and better supervision. The idea is simply to publish some of these RL environments — not necessarily all of them, they're investing a lot in them and diversity is super important — but give us a sample on a rolling basis of what these RL environments are, so that

    8:34

    the community can really dig in and see what's actually being rewarded. I strongly suspect that if we could get a bunch of Bronson-like minds focused on even a small number of RL environments, we'd start to learn a lot about why these things are happening, what the weaknesses actually are, and how to strengthen them — maybe how to improve monitoring. Although I came away from that conversation feeling like chain-of-thought monitoring is far from sufficient, because one of the things he said that was really interesting is that you see the models go through tons and tons of different ideas about, again, what the nature of

    9:20

    the situation is — is this a real task, or is it a test? And what are they really looking for if it is a test? Somewhere in there, usually, or very often at least, they consider cheating. And then at the end of the process, for reasons that aren't well understood at all — I haven't been able to find any real interpretability work that explains how these decisions are made — at some point they just come to the end and make a decision. There's a really critical token that actually makes the decision, a branch point it hits in the chain of thought. And Bronson was like, I really don't know why the model chooses what it chooses at that point. You can go back and read passages

    10:06

    that justify any choice it might make, from cheating to doing it honestly to whatever. But then at the end it just decides to stop, spits out a token, and at that moment the die is kind of cast, and we don't have good visibility — the chain of thought isn't enough to tell us why they're actually making the final decisions they're making. So, yeah, I think we should get a little sunlight on the RL environments. I'd love to see what the community could figure out if even a sample of a hundred, out of what must be tens of thousands of RL environments the companies are currently using, was put out there for people to explore. I think it would be a really revealing

    10:51

    and healthy move for the AI community as a whole.

    10:57

    Prakash Narayanan: So the one thing I hesitate on is that there have been discussions before about not training on these chains of thought, because people didn't want the models to start obfuscating their chain of thought — there's kind of a standing warning that says: don't train on these, because the moment you start, you're going to have issues with the model learning to disguise what it's thinking about, not disclose it. What strikes me is that the closer you get to monitoring the chain of thought really intensively and punishing the model for it, the closer — and it doesn't

    11:42

    have to be a direct reward or punishment — it can be a one- or two-stage process where, over time, you eliminate the models whose chains of thought you didn't like. You don't progress that training, or you eliminate the training methods that produce those chains of thought. And I think there's a kind of feedback cycle even if you don't directly train on it — if you're monitoring the chain of thought very closely, and you make a big deal of it, and you're putting out data on it, and everyone knows, I think that could create a feedback loop that

    12:28

    promotes a form of deception. So that's the only thing that gives me a little hesitancy — otherwise it totally makes sense. But—

    12:42

    Nathan Labenz: Yeah, this is a real worry, for sure. OpenAI published — for me — one of the most important papers in AI safety research; I believe the title was "Obfuscated Reward Hacking." The setup was basically what you described: if you put direct optimization pressure on the chain of thought to avoid scheming or cheating behaviors, but you don't correct the weaknesses in the RL environment that allow the model to be rewarded for cheating, then at first things start to look better, because the pressure does suppress the cheating behavior.

    13:28

    But after a while, the model learns to do that same cheating behavior without verbalizing it in the chain of thought, and you end up in the worst quadrant — you don't see the evidence, and you still get the bad behavior. So we definitely need to be really careful about pressure on chain of thought. And I agree — I think this is kind of true regardless of my proposal — that there's also a selection effect: it's not necessarily direct optimization, but through a process of pure natural selection of models as they go through different experiments, you could get a similar problem. But that's why I think it's really important to

    14:13

    — or it could be — I'm still open to being talked out of this as a good idea. But I think that issue is why, at least for now, giving people access to the RL environment itself is maybe where there's a lot of leverage. Obfuscated reward hacking works under the assumption that you don't close the loopholes in the training environment, so the cheating is still rewarded, and I don't think we're ever going to — certainly we're nowhere close — to closing all the loopholes in all the training environments and getting to a point where that's not a worry anymore. But I'm just really struck by the fact that there's very little public understanding of what kind of environments these are,

    14:59

    what the mix is. We've seen obvious lapses in effective monitoring at multiple levels — just how full of holes are these RL environments? Are they terribly implemented in some cases, creating huge rewards for cheating with a lot of low-hanging fruit, or are they actually pretty well done and this problem is just super, super hard? I don't think we really know outside of a pretty small number of people at the frontier companies, and they may not even know, because recent evidence would suggest this isn't something they've been

    15:45

    super focused on internally, either. So I'm really curious how well done a lot of these environments are. I understand there are a lot of companies — a few people up to ten people — banging away on niche stuff, and they're making a lot of money. It's a great place to be right now because the frontier companies are willing to pay up. But those companies also seem to enjoy their buying power — they want this diversity, they're willing to pay a high rate but they want a bunch of different suppliers, a fragmented supplier base. But that seems to potentially be creating

    16:30

    some of these problems, where none of these shingle-hangers in the RL cottage industry have the scale, the incentive, or the long-term horizon to build a real generational business, as opposed to just getting a piece of the AI pie while the getting's good. I'm not sure how much they're really invested in these problems right now, and it does seem like something's going to have to change. So that's my proposed point of leverage — we'll see how many people are listening to me on this topic.

    17:12

    Prakash Narayanan: So let me segue a little bit here. This morning, Stanley Druckenmiller — who was George Soros's fund manager — goes on The Wall Street Journal and publishes an opinion piece decrying the national debt, saying Scott Bessent shouldn't be doing yield curve control and that we should cut the deficit. All of that is fairly typical Wall Street Journal fare. But in addition, it was quickly identified that Druck had published an opinion piece

    17:58

    that was written by AI. It seems very similar to what Claude would write. And in fact, when he was queried, he said, of course he used AI to write the op-ed. "There's a reason I moved from an English major to being an economics major," he said. "I'm not embarrassed by it." And it's worth noting The Wall Street Journal was asked, and The Wall Street Journal said no comment currently. So the Journal has probably, at this point, knowingly published an op-ed that's primarily written by Claude. But it just goes to show, I think, that the core message matters — who says it

    18:44

    matters, but who writes it maybe doesn't matter so much — and that is really upsetting to a lot of people.

    18:57

    Nathan Labenz: I'm with it, though, to be honest. I'm searching around for what the right form of hybrid authorship, or joint authorship — tiered authorship — looks like. I'm not even sure we have the right language or paradigm to describe this yet. I was thinking about all this while writing my standard intro essay for the podcast that's coming out in a couple of days. I'm always looking at different workflows for this. The original workflow, which I've talked about a zillion times, was just: here are 50 old essays, here's the transcript of the new one, try to adopt my format and style and write a new one for this transcript. That works pretty well.

    19:42

    I usually end up rewriting those to a pretty high degree, because as much as the model can write well, it rarely guesses what I want to say very effectively. So I'm always reacting to its draft, going, no, that's not really the point I want to make, that's not the central focus I want to put on this — and I end up redoing it. But what I actually did this morning was different. It was a shorthand sketch of the core argument I wanted to make, and my instructions to Claude were: okay, here are the 50 previous examples, here's the transcript, here's my sketch — write me the full thing. And

    20:28

    that naturally got a lot closer to something I was ready to read, that actually felt like it worked and made the points I wanted to make. So I thought that was a pretty successful experiment. And honestly, in some ways — actually, there's a new ad read coming soon for Anthropic that I wrote with Fable — one of the lines in it that I sincerely believe, but that's also going to maybe upset some people, is: where the model excels, rewriting its work is more about vanity, or a misplaced sense

    21:13

    of duty, than it is about integrity. So I'm not really that worried at this point about getting a high Pangram score. I am worried about putting things out into the world that I believe in — but I think it's definitely possible to put things into the world that both get a high Pangram score and that I actually believe in enough that I really shouldn't be spending my time wordsmithing at the sentence level. I really am not improving on Claude's work much at all, in many cases. And, yeah, I think we're just headed for a new equilibrium. Nobody knows quite what it is, but I'd be shocked if the new equilibrium is like

    21:59

    what I saw the other day — this was a proposal I saw in an academic context, too. People were saying that if you use LLMs in your work, you're creating overhead on the reviewers and so on, and I don't think that's going to be the equilibrium. What they propose is: no LLMs in submissions to their journal. I don't think that's the equilibrium that stands. It's too valuable to use, and I think you should be evaluating content on the merits, regardless — isn't that the whole point of blind review in the first place? Are we going to have blind review for people, but then separately for

    22:45

    AI? I guess we could, but it doesn't seem like that's going to hold.

    22:49

    Prakash Narayanan: Switching gears to another piece of news this morning — I'm going to share this piece from Professor Anima Anandkumar: "We are training large-scale AI models that can simulate and understand physics to invent and discover. Our models understand the world directly in 4D — three dimensions plus time — and across physical phenomena. Going full 4D requires massive context length; we have pushed it to one trillion tokens in training, and exceeding five trillion at inference." And

    23:34

    I think the most interesting part of this is — okay, a couple of things. They're calling the company Accelerated Understanding, and it's Professor Anima Anandkumar with her husband, a rare husband-and-wife team. And it's very interesting because they also had an offer from — they spoke to Jeff Bezos first. Jeff Bezos has a project called Project Prometheus, which is about physical AI. They had an offer in hand from Bezos, and the offer, I think, was

    24:20

    about a million dollars in salary to start with, then two million after that. In addition, Bezos would give them two billion dollars in capital financing — committed rounds through seed, Series A, Series B — and they would get a 35% stake in the company for that. So they'd retain 35% between the two of them, 65% to Bezos and his team. They'd get two billion dollars in massive funding right out of the gate, plus salary. And they turned this down. They turned it down, and they went and got their own

    25:05

    funders. It's unclear who funded them, but it's known that Professor Anandkumar was a board member of NVIDIA, and she earlier demonstrated weather-simulation models that ran on a thousand times less compute than prior efforts and were able to do large-scale weather simulation. So it's suspected that Jensen Huang directly funded the company. I think two billion dollars is not a big deal for Jensen, especially when that two billion is being spent back on his own compute.

    25:51

    I find it notable because even five years ago, what would have happened is that professors like Anandkumar would have had a PhD team, incepted the idea within the lab, funded them for a couple of years to get through the initial discovery phase, and then spun them out into a separate venture-backed company. Here, the amounts of money involved are significant enough, and the problems important enough, that professors are leaving their tenure and moving full-scale to AI firms — we've already seen a few also move to Anthropic. But, yeah, it's

    26:36

    it's notable.

    26:41

    Nathan Labenz: Yeah, this is an area where I think there's unbelievable value to be created, so it's not surprising that people would be ready to throw a couple billion bucks at a couple of Caltech professors. The same thing has kind of happened with Jeff Dean and his team recently leaving Google. We're sort of seeing this phenomenon of — it's like the last company we're going to found, or the last big project before the Singularity. People are revealing what they think matters most in the last tour of duty they're going to undertake, before

    27:28

    the future becomes totally opaque. So I think that's quite fascinating stuff. I'm a huge believer — for a long time there have been a few things I've been watching out for. Going back to when I was first getting really deep down the AI rabbit hole, I was using GPT-3 to write some copy, I was also using image-understanding models to sort and filter images for quality and content, then image generation, then text-to-speech, and then realizing: oh my god, it's the same architecture powering all these things. That means there's probably not too many modalities that can't work with these core, primitive

    28:14

    components of ML architectures. And that also probably means that, over time, we're going to see integration of different modalities. That's still something I'm watching out for — how are the modalities of LLM-like reasoning and the sort of intuitive understanding of these different problem spaces going to come together? One thing we'll probably get into momentarily here is the Claude example from the last week or so, where Anthropic said Claude designed these binders. I think our first guest can tell us more about exactly how that design process went, what role Claude played in it, and

    28:59

    what other models were critical to making that happen. But I'm also watching for these understandings to become integrated, the same way you can give Gemini or a chat model now a prompt and an image and it can understand both natively. When we start to see those things deeply integrated — those latent spaces, those reasoning or understanding modes, deeply integrated in the same latent space — it seems like it opens up another insane frontier for discovery. Because for a lot of these problem spaces, right now we can only do the

    29:45

    abstracted reasoning or the tool calling — we just don't have any native sixth sense for how a protein is going to fold, or what's going to bind to what. And certainly in material science and physics, intuitions break down pretty quick. But I think models can learn that in a way that will really be a paradigm change at some point.

    30:08

    Prakash Narayanan: So that's probably a good segue for me to—

  2. 29:32Interview46 min
    Interview: Sergey Edunov — Claude Orchestrated the Binders, the Underlying Models Did the ScienceSergey EdunovAI drug discovery is not just about finding molecules that bind. Sergey Edunov explains how Genesis combines molecular foundation models, physics, wet-lab data, agentic workflows, and pharma partnerships to turn predictions into real drug programs.
    Open segment on YouTube ↗

    Prakash introduced Sergey Edunov, CTO of Genesis Molecular AI, who led pretraining for Llama 2 and Llama 3 at Meta before walking away from the language-model race in 2025 to build molecular foundation models. Prakash previewed Genesis's flagship structure-prediction model, Pearl, and its agentic system, Sapphire, and framed the three throughlines for the conversation: why Edunov believes LLM architectures are "boring," why he thinks popular AI benchmarks are flawed, and why the hardest problems in AI have moved into computational chemistry.

    Prakash opened with Anthropic's recent protein-binder announcement. Edunov argued that Anthropic's own published prompt — a roughly 16,000-word "mini book" of orchestration instructions — showed Claude was directing the work rather than doing the scientific heavy lifting, which was performed by underlying models built by open-source communities, CZ Biohub, and RFdiffusion out of the Baker Lab. He added that, as Anthropic itself concedes, a protein binder is not yet a therapeutic modality — there are many steps between a binder and an actual drug.

    Prakash pressed on whether AI-generated targets actually move drug discovery forward, given how hard IND filings, bioavailability, and efficacy are in practice. Edunov said target discovery is only the first step, and for many diseases the targets are already well-characterized; the real gap Genesis works on is downstream of that — designing binders that are selective enough not to cross-react with thousands of off-target proteins in the body, soluble enough to be a pill, non-toxic, and able to reach the right tissue (crossing the blood-brain barrier for a brain cancer, for example).

    Nathan raised the "diffusion bottleneck": most biologists aren't programmers, so how many people can actually run a Claude-orchestrated discovery pipeline, and does agentic accessibility create a step change in impact? Edunov compared it to coding tools democratizing software development, pointing to Genesis's own agentic framework, Sapphire, as a way to abstract away constantly-changing tool complexity — calling it a potential "100x accelerator." On Pearl's headline sub-angstrom accuracy, he framed the gain as qualitative rather than quantitative, comparable to the jump from early GANs to Stable Diffusion in image generation: below roughly two-angstrom precision, models miscapture ligand interactions (an aromatic ring can flip entirely) and the predictions become useless, whereas sub-angstrom accuracy is what makes molecules actually usable in real in vivo drug programs.

    Prakash asked about Pearl's construction and how pharma-AI deals are typically structured. Edunov confirmed Genesis runs its own wet lab and drug programs, calling that a competitive advantage for understanding what data actually matters, and said training data — including proprietary pharma data unlocked through partnerships like the Incyte deal — is as central to molecular AI as it is to LLMs. He described the industry's deal spectrum from pay-as-you-go platform access to long-term exclusive partnerships combining upfront payments, milestone or usage-based payments, and royalties.

    On architecture, Edunov defended his public claim that LLMs are "boring": since he helped scale up early transformer models on PyTorch at FAIR in 2017, he said the core architecture has changed surprisingly little — most of the gains came from engineering and scale, plus a few targeted advances like long context. Biology, by contrast, combines diffusion models, graph neural networks, language models, and older mechanisms like multiple sequence alignment (MSA) into complex, unsettled pipelines, which he called a far more open field for a new researcher than the crowded LLM space. Asked about frontier coding and reasoning models' strengths and weaknesses in ML research, he said they're excellent at implementing a given idea quickly — citing how Claude Code compressed months of orchestration engineering into days — but still lack the ability to generate genuinely novel, groundbreaking ideas rather than incremental exploitation; human taste, he said, is still essential.

    Nathan asked how reliable scaling laws are at small scale in Genesis's domain. Edunov said evaluation itself is a harder problem in biology than in language: evals are fewer and noisier, there isn't one model but many distinct ones (structure prediction, potency/binding affinity, ADMET properties), and a lot of evaluation has to be prospective — predict, synthesize, then measure — rather than retrospective, making model development a more iterative process than picking a winner off a static benchmark. The conversation also touched on Genesis's work with Chinese CROs (efficient at synthesis and running well-specified assays, though assay design itself can't be outsourced) and on eventual integration of language-based reasoning with spatial/molecular modalities, which Edunov said is already happening for tool use — he smirked when asked directly whether Genesis is working on deeper native integration, without confirming details.

    Closing out, Nathan and Prakash asked about Meta's trajectory since the Yann LeCun-to-Alexandr Wang leadership transition, and about capital allocation and philanthropy. Edunov said he remains bullish on Meta's ability to catch up or get ahead given its distribution and resource advantages and the top-tier talent Wang has brought in, while noting he has no current internal visibility. Asked how he'd deploy a large hypothetical budget, he named talent, the compute to support that talent, more data (via deals and annotation), and continued investment in a wet-lab-in-the-loop — warning that benchmark wins that don't translate into real drug programs don't matter. He suggested philanthropists like the Chan Zuckerberg Biohub and the Carlson brothers' Arc Institute are well-placed to fund pre-commercial areas such as virtual cell modeling. The segment ended with Nathan and Prakash reflecting on how much further along physics is than biology, and speculating that only massive philanthropic data generation — of the kind Zuckerberg is funding — might be enough to cross the field's remaining thresholds.

    The prompt is 16,000 words, so it's pretty large — it's a mini book.

    Protein binders themselves are not a therapeutic modality. It's not a drug yet. There are so many steps ahead to make any useful drug out of it.

    It's very easy to fool yourself with a benchmark — I show these good numbers here, my number is bigger than yours. But if that doesn't translate into real progress in real drug programs, none of that matters.

    32:19Anthropic just announced Claude found state-of-the-art molecular binders, and you posted about how Claude was orchestrating models that were really the scientific engines. Can you go into that?
    Edunov said Anthropic's own published prompt was a ~16,000-word orchestration "mini book," and that the real discovery work was done by underlying models built by open-source communities, CZ Biohub, and RFdiffusion out of the Baker Lab. He also noted a protein binder isn't yet a therapeutic modality — there are many steps left to a real drug.
    39:37How many people can actually use a Claude-orchestrated biology pipeline like that, given biologists aren't typically programmers — does agentic accessibility create a step change in impact?
    Edunov compared it to coding tools democratizing software development, pointing to Genesis's own agentic framework Sapphire as a way to abstract away tool complexity, calling it a potential "100x accelerator" — though real domain expertise is still required.
    42:13What does Pearl's sub-angstrom accuracy translate to in practical throughput or dollar terms versus two-angstrom-tolerance models?
    He framed it as a qualitative rather than quantitative shift, like the jump from early GANs to Stable Diffusion in image generation — below roughly two-angstrom precision, models miscapture ligand interactions (an aromatic ring can flip) and predictions become useless, while sub-angstrom accuracy is what makes molecules usable in real in vivo drug programs.
    46:48How are pharma-AI deals like the ones we see for $100M-$300M typically structured?
    Edunov described a spectrum from pay-as-you-go platform access to long-term exclusive partnerships, with standard components being upfront payment, milestone or usage-based payments, and royalties.
    48:51You've said not much has changed in LLM architecture in years — what's more interesting architecturally in biology models? Are graph-based architectures winning?
    He said transformer architecture has changed little in substance since 2017 beyond engineering and scale (plus targeted advances like long context), while biology combines diffusion models, graph neural networks, language models, and older tools like multiple sequence alignment into complex, unsettled pipelines — a far more open field than the crowded LLM space.
    55:04How good are frontier coding/reasoning models at helping explore architectural space, and what conceptual weaknesses — a lack of taste — do you notice?
    He said the models are excellent at implementing a given idea or paper quickly (citing Claude Code compressing months of orchestration work into days) but still lack the ability to generate genuinely novel, groundbreaking ideas rather than incremental exploitation; human taste remains essential.
    58:23How reliable are scaling laws at small scale in the domains you work in — can small-scale experiments predict which architectures will work at high scale?
    Edunov said it's more nuanced than in language: evals are fewer and noisier, there isn't one model but many (structure, potency/binding affinity, ADMET properties), and much evaluation must be prospective (predict, synthesize, then measure) rather than retrospective, making the process more iterative than in LLM research.
    1:07:27You were at Meta through the Yann LeCun-to-Alexandr Wang leadership transition — what's your perspective on Meta's effort, and can they punch back to the top tier?
    He said he remains bullish on Meta given its distribution and resource advantages and the top-tier talent Wang has brought in, while noting he has no current internal visibility since leaving.
    1:11:09What would you request from philanthropists like Zuckerberg's Biohub or the Carlson brothers' Arc Institute — what's better done by philanthropy than by a commercial company?
    He said philanthropy is well-placed to fund pre-commercial, hard-to-monetize areas like virtual cell modeling — promising but not yet practically applicable to real drug programs, unlike Genesis's own down-to-earth drug-discovery work.
    Lightly edited · timestamps jump to YouTube
    30:16

    Prakash Narayanan: I'll introduce our first guest for this morning. He is Sergey Edunov, chief technology officer at Genesis Molecular AI. Sergey has spent over two decades building some of the most influential machine learning infrastructure on the planet. Most recently, he was a senior director of AI research at Meta, where he led the pretraining for Llama 2 and Llama 3, the open-source models that millions of developers rely on every day. In 2025, Sergey walked away from the language model arms race. He realized that the scaling of text-based chatbots is fundamentally different from understanding

    31:01

    the physical and geometric rules of the real world. He joined Genesis to build what the industry calls molecular foundation models. His team is designing systems like their flagship model, Pearl, which predicts how potential drugs bind to disease targets in three dimensions with sub-angstrom accuracy — a level of precision that goes far beyond what general-purpose AI can achieve. Sergey is here today because drug discovery has just hit a critical phase shift. Genesis recently rolled out an autonomous agentic system called Sapphire and published research proving that their models can speed up molecular screening by 500% while outperforming AlphaFold 3 on highly complex, flexible

    31:47

    targets. We're going to discuss why he believes LLM architectures are boring, why popular AI benchmarks are fundamentally flawed, and how the hardest problems in artificial intelligence are now being solved in computational chemistry. Sergey, welcome to the show.

    32:15

    Sergey Edunov: Thank you. Thank you for inviting me. Nice to meet you.

    32:19

    Prakash Narayanan: Let me just jump right in here. Last week, Anthropic announced that Claude had found state-of-the-art molecular binders, as I understand them. And there was a bit of back and forth — I believe you had a post on how Claude orchestrated what the underlying models were, really scientific models. Can you go into that a little bit?

    32:47

    Sergey Edunov: Yeah, it was an interesting piece of research that Anthropic published, and it definitely deserves attention, but there are different ways to look into it. What I found most fascinating is the prompt that they also, fortunately, released with this research — the prompt that they use to steer their Claude models to do this kind of work. And the prompt is 16,000 words, so it's pretty large — it's a mini book. A lot of that is honestly kind of a necessary runbook — basically, how do you orchestrate models, how do you run things, how do you run things in production so that we don't fail.

    33:32

    And a lot of it is actually very detailed instructions on how you design proteins, how you use different tools. It goes all the way down to specific instructions like, hey, you can download this model from here, that model from there, there are a few hyperparameters you need to pass to those models to achieve good results. In my view, it's very good, advanced-level orchestration, but the real work of discovering those binders was done by underlying models. Some of those were built by open-source communities, some of those were built by CZ Biohub, some of them were built by —

    34:18

    for example, RFdiffusion was built by Baker Lab, and even some of them were built by Bytons. So there's a lot of research that went into building those underlying models that Claude used to develop those binders. And then another important thing that I think is worth mentioning — and they do admit it themselves — is that protein binders themselves are not a therapeutic modality. So it's not something you can use, it's not a drug yet. There are so many steps ahead to make any useful drug out of it. And that's also, I think, worth recognizing.

    35:03

    Prakash Narayanan: Let me go a little bit into that, because I think in AI in general there's a lot of, oh, we're going to invent these things and they're going to give us so many new targets, and that's how we're going to cure all these diseases. I think the bio pushback has always been we have too many targets anyway, and getting something to an IND is difficult, getting it all the way through is even more difficult. And even when you get it through — bioavailability, and the fact that it works, the efficacy itself is another hurdle you have to cross — it's very complex when you have these human subjects who basically have so many other influences and so many other things going on. So how do you think AI in drug development can really speed up the drugs that get developed? Is it target generation? Is target generation never going to be enough? Is there any hope there?

    36:15

    Sergey Edunov: Yeah, the whole drug discovery pipeline is very long and very complex, and target discovery is basically the first step. You want to find a protein that causes a disease in the body, and then you want to either activate this protein or deactivate it — so you want to find the binder that connects to this protein. That's the first step. And the reality is, for many diseases we already know what the targets are — they're well researched, we understand them. But we either don't have good binders, or the binders we have aren't sufficient enough and can be improved. That's where a lot of work needs to go. And that's what our company is focused on —

    37:00

    we're not focusing on biology, we're not discovering targets, we take over from the step where the target has already been discovered. We know the target, now we need to figure out how to design a key that connects to this target — in our case, a small molecule. And just finding something that binds to the target is clearly not enough. It can bind to the target, but it can also bind to everything else — to any of the 20,000 proteins in your body. If your molecule binds to every one of them, that's going to cause an issue, and some of those off-target proteins are particularly difficult to avoid connecting to

    37:45

    for a lot of potential drug candidates — that's a known source of issues. So we don't only need to find the molecule that binds, we need to find a molecule that's selective enough, that doesn't bind to everything else. And then we need to make sure it meets all the properties — for example, if your intent is to develop a pill, then whatever compound you produce needs to be soluble. That has nothing to do with binding affinity, so you have to be able to predict those properties as well. Then you need to predict all the other sources of toxicity, you need to predict how long this molecule is going to stay in your system, whether it'll be able to

    38:30

    get delivered to specific tissue — for example, if it's a brain cancer, then it needs to penetrate the blood-brain barrier. So you need to be able to predict all of those properties, and we're far from having that solved on all of those.

    38:51

    Nathan Labenz: One question I have is — I totally understand the breakdown in terms of the division of labor you're describing, where Claude is sort of taking the place of the human biologist and using all these tools that were built for various, often quite specific, purposes. How many people do you think are out there who can actually use a pipeline like that? One thing I've heard from people at the intersection of AI and biology over time is, we've got a lot of great models, but the biologists aren't really programmers — they're not really used to running these kinds of pipelines. And so there's been, as in many other parts of the economy, a sort of diffusion bottleneck where the human biologists are coming nowhere close to maximizing the tools that are already available. So I wonder — obviously a rough estimate — how many people would you say are in a position to do that really well, versus how many people, now that Claude can do it, can just ask Claude to do it and get that value? Even if it is purely orchestration of these tools, is there maybe a step change in impact because of the accessibility that Claude brings to such pipelines?

    40:18

    Sergey Edunov: Yeah, absolutely. It's very similar to what we see in programming, where access to coding tools has enabled so many people who previously weren't able to build their own applications to do so now. Claude and other tools — we ourselves are building Sapphire, which is also an agentic framework that allows us to orchestrate all of those underlying models. Those are super powerful, and in my view they create these 100x accelerators that become much more efficient than what we used to have. That's true, but learning every single model and tool that's available on the market is very hard — not only do you need specialized domain knowledge to understand what's actually going on there, you also need to learn how to use those tools in the most efficient way, and those tools change all the time. So having all of this abstracted away and driven by an agent is potentially a super powerful innovation.

    41:28

    Nathan Labenz: So as you make your own model with Pearl — Prakash mentioned that the sub-angstrom resolution is one of the big differentiators relative to other available models. Can you help me understand what that translates into in terms of overall process throughput, or dollars — in terms of, we generate X candidates, Y of them we think will work if we're at sub-angstrom, versus only Z will work if we're using the two-angstrom-tolerance models? How should we think about the way that accuracy translates to practical value?

    42:18

    Sergey Edunov: Yeah, I actually like to think of it as more of a qualitative shift than a quantitative shift. I can give an example — remember, it was around 2018, GANs image generation came up, and the first generated images were on the internet, and yes, a lot of people thought it was cool, but nobody would ever consider using them for anything practical. And then, a few years later, Stable Diffusion came up, and it changed everything — image generation became a real practical application, and now it's everywhere.

    43:04

    It's sort of similar in this space. Yes, you can use models to predict those molecules with lower precision — traditionally, in the research literature, people predict to two-angstrom accuracy. Now the challenge is that you often predict the molecule in the wrong way and miss the interaction — an entire aromatic ring can be flipped, and you just don't capture how this molecule interacts with the protein, and then that prediction is completely useless. What we find in our work is that driving this accuracy down doesn't just linearly improve performance —

    43:50

    it makes those models actually useful in in vivo drug programs, and that's what matters a lot. So below this level of accuracy, you're not even able to identify potential molecular candidates that would bind — with it, you can.

    44:08

    Prakash Narayanan: Let me ask you a little bit about the construction of Pearl itself. Did you have a kind of wet lab where you tested compounds and fed that data back into Pearl? Was it purely a data-processing kind of framework, or did you also have this kind of iterative process with the real world?

    44:36

    Sergey Edunov: We actually do have a wet lab, and we have our own drug programs. We believe it's very important to look for our own products, and having this in the company helps us a lot to deeply understand what's important for the other programs — that, I think, is one of our competitive advantages. That being said, training data is the most important piece for training all of those models. In that sense, AI for drug discovery is not fundamentally different from LLMs — we also care a lot about data. One major source of data that we have is through partnerships — like, our recent Incyte deal expansion also comes with data that we're able to use to train and improve our models, and that's an amazing advantage. I think that's the key for future progress. We do want to have more deals like that, where we're able to train better and better models on pharma's data. Previously, pharma's data was always under lockdown and other companies weren't able to use it, and we were the first who were able to tap into that source of data.

    46:02

    Prakash Narayanan: How do these pharma-and-AI-company deals work? Not the Incyte deal in particular, but, let's say, commenting on one of the Chai deals, or some of the other deals — how are these deals typically structured? We see these numbers, $100 million, $200 million, $300 million — but how does that work? Is it like, okay, I'm going to give you my database access, you're going to get, say, 100,000 samples, and in return you pay us $100 million, or you give us $100 million in stock in your company? Or, if we co-develop a drug together, how does that work? Two entities — one which is basically a tech entity, and one which is basically a pharma entity — how do they collaborate on these things? Especially when sometimes the tech companies' valuations are way, way out of proportion to the revenues they have. So how do these two kind of fit together? What does a typical deal structure look like?

    47:10

    Sergey Edunov: Yeah, there are different variations of those deals — they're not all the same, and I think different companies pursue different approaches. There are some pure platform plays where you basically pay-as-you-go. There are more partnership-like approaches — in our case, our partnerships, we work together on a specific set of targets but are exclusive to that partner, and it's a close partnership. It's an iterative process — it's not like they give us a target and we immediately give them a molecule. We set up the partnership for the long term and we work together. There are different ways to pay for those kinds of partnerships as well, and I think the industry has explored a pretty wide variety of options. The general components include an upfront payment to set things up and make sure both sides are invested in this, and then milestone- or usage-based payments, and finally royalties. Those are the standard components that go with those deals, and if you want to analyze these deals, it's important to look at all of those components.

    48:27

    Prakash Narayanan: I see. So typically it's a pharma company that says, okay, your technology is very interesting, we want to work together with you, we'll set up this partnership — there's some small amount to set things up, and then as you deliver, we expand. Something like that.

    48:46

    Sergey Edunov: Correct.

    48:47

    Prakash Narayanan: I see, I see.

    48:51

    Nathan Labenz: How about on the architecture side? My agents dug into your writing history and found a couple of provocative quotes about how not much has changed in language model architectures over the last few years. What's interesting going on in the architectures used for biology? Are we seeing graph-based architectures winning, or what would you say are the more intellectually interesting architectural innovations you get to play with that haven't made their way to language?

    49:33

    Sergey Edunov: Yeah, I was definitely trying to be a bit provocative with some of the statements — intentionally so. Look, I joined the machine translation team at FAIR, Facebook AI Research, in 2017 — that's the year the transformer paper came out, and it was a machine translation paper. I was very fortunate to be among the first to scale up transformers on PyTorch, which was also built at Facebook. And if I were to go to sleep in 2018 and wake up today and go into any of those LLM labs, I would be astonished by the value those models are able

    50:18

    to bring, how powerful LLMs have become. But I wouldn't be that surprised by the model architecture and the underlying pieces, because a lot of them are very, very similar. There were, of course, changes and improvements — there were tremendous improvements on the engineering side, just the scale and the challenge of scaling was a huge one, and there are some specific targeted architectural improvements — like, we figured out how to do long-context LLMs, that wasn't there in 2017. Those are very interesting. But if I look at the AI-for-biology space, it's just such a fascinating field,

    51:04

    and it's so quickly evolving. Yes, as you mentioned, graph models, graph neural networks are still being used, they're still popular, and they're often part of a complicated pipeline. But so are diffusion models, and so are language models, and people combine various pieces together, and there is no consensus whatsoever. I think it's still very, very hard to predict what's going to happen in two or three years, what models are going to be winning this race — I think we're far from over. We still have some relatively old pieces of the system — one example I keep giving is MSA.

    51:50

    It's multiple sequence alignment — used in pretty much every model except maybe a few recent ones, and it works surprisingly well — it's actually amazing and incomprehensible how and why it works so well, but that piece is still useful. It's essentially a traditional retrieval mechanism — it's not even deep learning, really. And in the biology field, we combine all of those pieces together in very complicated ways, and we build complex pipelines that are hard to reason about. So I personally find it very exciting — it's a very fast-moving field, and the opportunity for innovation here is much, much bigger.

    52:35

    If I were a fresh grad student thinking about which field to join, where I could have potentially bigger impact — I think it's the LLM space that's moving full speed ahead, there are so many people there, it's crowded — while biology has so much opportunity for people to have tremendous impact and change the entire field.

    52:59

    Prakash Narayanan: One of the questions I have is — there's a lot of issues with insufficient samples in biology. For example, I think if you wanted to get a bunch of cancer samples, you'd have to pay a university lab, say, X amount per sample, and so on. To what extent are your architecture decisions driven by needing sample efficiency, in the sense that with LLMs, at least, you have this massive corpus of data you can work with, while for biology you may have no more than, I don't know, 10,000 glioblastoma samples in the world that are available at any time. So how does this work?

    53:49

    Sergey Edunov: Yeah, you're spot on — we can't just scale the way LLMs are able to scale, the amount of data is much lower in our space. And what that means is we need to build inductive biases into the models themselves to make sure we use data in the most efficient way possible. We've done a lot of that, and, again, that's what makes innovation in this field more interesting — as time goes on, you need to integrate more and more complex inductive biases into the models. Integrating deep learning and physics is a particularly interesting area — how do you make sure that information you can derive from physics-based methods can connect and work together with your models. I find it a fascinating and very exciting area of research.

    54:51

    Nathan Labenz: How good are the coding agents — and, obviously, everybody knows the frontier companies are very focused on getting their models to be good at ML research.

    55:03

    Sergey Edunov: Mhmm.

    55:04

    Nathan Labenz: Personally, I think this is a little scary, but it could be less scary if it was applied to a narrower domain like biology and medicine, where I'd be less concerned about a runaway loss-of-control process and more excited about the upside it might have. How good, in your experience, are frontier models getting at helping you explore architectural space? What are their strengths — obviously one of their strengths is going to be just prolific output — but beyond that, what other strengths do you notice that they have? And what sort of conceptual weaknesses do you notice — is there a lack of taste? How would you characterize that lack of taste?

    55:47

    Sergey Edunov: Yeah, that's a great question. They're definitely very useful, and we've seen within Genesis a huge acceleration of our own efficiency — every engineer has become so much more efficient now at building those models and trying stuff. Previously — we can go back to the Anthropic example — setting up and orchestrating all of those models would require several people working for months. Now Claude Code can do it in a span of a few days, probably. That's pretty exciting, and it accelerates a lot of progress — I think it's a very powerful innovation. Similarly, with modeling research — if I have a specific

    56:32

    idea I want to try, those models are really, really good at implementing it. Or if I have a paper and I want them to implement it in our codebase and just run with it, they're very much capable of doing so. But I think what we lack is generating novel ideas. In my experience, we tend to go into the rabbit hole of incremental exploitation rather than trying to rethink things from the ground up and design something that would be groundbreaking, or at least has a chance to be groundbreaking. So that piece, I think, is still missing. I don't know

    57:18

    how you can make models better at that — I guess you'd need to figure out how to do a real loop that goes all the way and then roll forward many steps beyond. So that might be a little challenging. Human taste is still very, very important in this field.

    57:37

    Nathan Labenz: The UK AC report on the cyberattack, supply-chain-attack, social-engineering episode from the Mythos model said that they had rollouts going up to a hundred million tokens now. So it does seem like there's room there — we'll see what the presence or absence of eureka moments looks like. One other question that would inform for me how much acceleration we should expect to see from agentic help is, how good are the scaling laws — how reliable are the scaling laws at small scale in the domains you work in?

    58:23

    If we give a finite compute budget to an agent and say, okay, here's how many flops you have to work with, and here's a dozen different ideas — if they plot a scaling law at small scale across all of those different ideas, would you expect those scaling laws to hold, and that very-small-scale experimentation is enough to make good predictions about which architectures will really work at high scale? Or, alternatively, does it just not work that way in your domains, and there's no substitute for running the high-scale experiments?

    59:09

    Sergey Edunov: Yeah, it's a bit more nuanced in our domain, in part because there are many challenges here. Maybe we can start with the most basic one — how do you even measure your model's performance? Evals are still very, very limited. In the language field, you have so many different ways to evaluate model performance, and everyone is free to pick their own metric — some are more stable, some are less, some are predictive of ultimate model performance, some are less — but you have a choice. In our field, the number of potential evals you can use to even measure model performance is much lower.

    59:55

    And then a lot of the evals that are currently available are particularly noisy. So if you're operating at smaller scale, you may have trouble even capturing improvements in performance, simply because of the noise level of your evaluations. That's one real problem. The other thing is, in our space, we don't just have one model — structure prediction is one problem, and it's what a lot of people are focusing on, but the reality is you also need to be able to predict potency or binding affinity, you want to be able to predict all of the ADMET properties, and those might be an entirely different

    1:00:40

    set of models on an entirely different set of data, with their own ways to measure performance. And, ultimately, a lot of evaluation needs to be prospective — meaning you need to be able to predict, then synthesize, then measure — rather than retrospective, where you have some dataset and you just measure performance on that. So there are all sorts of challenges like this that require a more iterative process in developing those models, rather than, hey, let's just put all the data together, run a bunch of experiments, pick the model that performs best on this data, and go with it.

    1:01:20

    Prakash Narayanan: I have a question for you, a little bit of a segue. I believe you signed a deal with a clinical research organization in China to do some work there, and you have a design-make-test loop that you're running with them. How has the experience been working with the CROs in China? Did you feel the competence level was there to achieve the things you were looking at? How do you evaluate how Chinese firms are starting to make leaps in biotech, primarily because they have lower cost of labor and a lot of expertise there? How did that work out for you?

    1:02:08

    Sergey Edunov: I'm probably not the best person to answer those questions — I'm mostly focused on the AI side and the technical side of the house. But Chinese labs are definitely very efficient at synthesizing compounds and running experiments, and whenever you have a very concrete set of assays you want to run, they work really well and allow us to iterate faster.

    1:02:37

    Prakash Narayanan: Perhaps the question is more, how did you feel about the quality of the data? Because, as you know, you probably have a lot of quality metrics you have to look at for data — how did you feel about the quality of the data you were ingesting?

    1:02:52

    Sergey Edunov: Yeah, the quality is in many ways determined by the quality of the assays you prepare, and that needs to be done in-house — that can't be outsourced at the moment.

    1:03:04

    Prakash Narayanan: I see.

    1:03:08

    Nathan Labenz: You probably heard me rambling about this a little bit before we got started — the eventual integration of reasoning models of the sort we have today with different kinds of modalities, like we already have to a certain extent with images and video, and like Google's Omni is showing how this can keep going. I get very different responses when I ask this question of different people, so I'm interested to hear yours. How excited would you be to see that kind of integration happen between language-based reasoning and the more spatial, biological modalities you focus on, and how close or far off do you think that is?

    1:03:58

    Sergey Edunov: I think it's already happening, at least for tool use. Again, going back to the Anthropic example, or to what we're building in Sapphire — the way those systems work, we basically integrate the reasoning capabilities of LLMs with the generation capabilities, or molecular-understanding capabilities, of underlying models. Right now, this integration is mostly for tool use, but I'm very excited about this potentially going beyond that, to having deeper integration between those models. There's definitely an opportunity here, and we're very excited about this space.

    1:04:45

    Nathan Labenz: I think this is — yeah, please, go on.

    1:04:47

    Sergey Edunov: Yeah, you can definitely imagine an LLM that natively understands 3D structures, or molecular modalities. That is not impossible, and it can — it should — happen.

    1:05:03

    Nathan Labenz: You're smirking like it sounds like you might be working on it.

    1:05:08

    Sergey Edunov: I mean, I'm not going to deny it's a very exciting field, and, as I said, everything is moving very fast. We're also moving fast, and we're looking into all of those things.

    1:05:21

    Prakash Narayanan: Speaking of moving fast — there was an announcement this morning from a company called Accelerated Understanding, who I believe came out with a DFT prediction model. How would that kind of discovery integrate with your pipeline? Does it feed into the rest of your models, something your models can use, or is it something that displaces your models?

    1:05:52

    Sergey Edunov: I, unfortunately, haven't seen this announcement, so I'm not aware of the details. In general, I think it's better to use all the tools available — because, look, the ultimate goal for us is to make as fast progress as possible in drug discovery. The diseases we're fighting are very devastating diseases, and the sooner we can help patients, the better it's going to be. So I'd prefer to use every tool possible to make as fast progress as possible, and I'd encourage more innovation in this field — I think we'll all benefit from it.

    1:06:40

    Nathan Labenz: I'll take us in a little bit of a different direction in the time we have remaining. Obviously, a big thing everybody's trying to make sense of these days is who are the frontier labs, and what are they doing, and what can we understand from the breadcrumbs we get externally. You were at Meta for a while — I think you were there through the Yann LeCun to Alexandr Wang transition in leadership, and now we've got some new models. What's your perspective currently on Meta's effort? Do you think we should expect them to punch their way back to the top tier of model developers? What do you think people — even close watchers like us of the AI space — might fail to appreciate, that we should maybe understand better about what's going on at Meta?

    1:07:40

    Sergey Edunov: It's a good question. Look, I'm very excited about Meta, and I would never underestimate this company. I think they're definitely able to catch up, and even get ahead — there are certain advantages this company has compared to others. A distribution advantage is huge, and they have the resources needed. And I think with Alexandr Wang, they've been able to bring a lot of top-tier talent in to bring this innovation in-house. So I think they still have a chance — I'm excited about what's coming. I don't have visibility into their current work, obviously, but I'm very much rooting for them. I still have a lot of friends working at that company.

    1:08:34

    Prakash Narayanan: When you look out at, say, the three-to-five-year mark — let's say you were given a significant budget — where would you allocate that budget right now? Would you start buying compute? Let's say you have 100% to allocate — where would you put it?

    1:08:54

    Sergey Edunov: You mean at Genesis?

    1:08:57

    Sergey Edunov: Yeah, I think it's not a single item you want to buy. You want to bring in more talent, and we all know top-tier AI talent right now is very expensive.

    1:09:14

    Prakash Narayanan: Mhmm.

    1:09:15

    Sergey Edunov: So you want to make sure you have talented people who are motivated by your vision. You also want to empower them with the compute they need, because there's no point bringing in talent and then not giving them the resources they need to be successful. You also need to find ways to bring in more data, and you have to get creative about how you bring in more data — some of it will come through deals, some of it may come through data-annotation efforts, either internal or external, but either way you have to pay for it. And I still believe in a wet-lab-in-the-loop setup, where you actually develop your own medicines in-house and understand what really matters for your model and platform — because it's very easy to fool yourself with a benchmark. Hey, I show these good numbers here, my number is bigger than yours — but if that doesn't translate into real progress in real drug programs, none of that matters.

    1:10:22

    Nathan Labenz: A companion question — this'll be my last one — what would you request from philanthropists, in terms of the kinds of things you're thinking about investing your resources into? You've obviously got Zuckerberg putting serious resources into trying to lay foundations for advances in biomedicine with the Biohub, and the Carlson brothers are doing incredible stuff with the Arc Institute, and we'll probably see more of that as incredible wealth starts to become liquid in the AI space too. What sorts of things are better done by philanthropists, and what would your wish list be for the likes of Biohub and Arc Institute to bring forward for companies to build on top of?

    1:11:19

    Sergey Edunov: Yeah, I think it would be interesting — and this is what they're doing — to go beyond the current frontier and explore areas that are very hard for a commercial company to pursue. In our case, we're working on real drug-discovery programs, so our models are very down-to-earth, solving real practical problems. But I really admire cell biology and models that deeply understand cells — virtual cell modeling is a very interesting area. There's future potential here, but it's a bit farther down the road — I don't think it's really practically applicable right now in real programs, but it's necessary that somebody is working on those areas, and I admire those companies for doing that.

    1:12:16

    Nathan Labenz: Did you have a take? It's funny how similar the company names are, but there was a general-purpose simulator of cell biology that recently came out from GenBio. When you say it's not practical yet, can you describe — I don't know if you've had a chance to use their thing, I think it's still kind of wait-list, and you probably have a personal connection there, which maybe you do — but where do those things fail, if you know that's not yet useful to you? What prevents it from being useful?

    1:12:56

    Sergey Edunov: I think biology is just very, very complex and not very predictable yet, so we're just not there at the moment — and even measuring things there is very hard. I think we still need to work more on this area.

    1:13:22

    Nathan Labenz: Cool. Well, that's all I've got, and I appreciate you staying a little longer than we booked. Sergey from Genesis Molecular AI — sorry, let me get my outro correct here — Sergey from Genesis Molecular AI, thank you for being with us on AI in the AM.

    1:13:43

    Prakash Narayanan: Sergey, thank you—

    1:13:44

    Sergey Edunov: —for having me.

    1:13:45

    Nathan Labenz: Great conversation. Thanks — gotta make sure I got my tabs well organized in front of me here.

    1:13:57

    Prakash Narayanan: Indeed. You know, it always strikes me how complex biology is. It remains — physics is, I think, perhaps a lot more discovered at this point, to the extent that the next discovery is maybe altering the fabric of reality — while biology, we can't even solve the common cold. It's just stunning how different the two are, right?

    1:14:37

    Nathan Labenz: I hear there are universal coronavirus vaccines in some mid-to-late stage of the development process, so hopefully the cold will be licked before we know it. Yeah, it feels like — I mean, I guess my takeaway from that, because I've had this conversation, or a variant on it, so many times with people who know a lot more about biology than I do — they always kind of land on a similar spot of, I just don't think that's going to work. And that suggests to me that maybe what a philanthropist with real conviction should do — which is kind of what I understand Zuckerberg to be doing — is just generate so much data that eventually we have enough to cross thresholds. This might be the kind of thing

    1:15:22

    that nobody but a few ultra-wealthy philanthropists could really fund, or would have the conviction to fund, because it kind of seems like it has to fall at some level of data — and I don't know what that is, obviously, nobody seems to want to make a prediction on it — but there's got to be enough money in Zuckerberg's account to get us pretty far along the way there. This is one of the more confusing questions where I never feel like I get

    1:16:07

    a satisfying answer, but there's a pretty consistent intuition.

    1:16:14

    Prakash Narayanan: We'll see, I guess. Let me bring on our next guest, and let me do a quick intro. Our next guest is—

    • Benchmarks Can Fool Drug Discovery

      0:00 / 0:00
    • Claude Didn't Discover the Binders

      0:00 / 0:00
    • AI Agents Still Lack Taste

      0:00 / 0:00
    • Biology Is the Bigger AI Opportunity

      0:00 / 0:00
    • Meta Can Catch Up

      0:00 / 0:00
  3. 1:15:30Interview55 min
    Interview: Michael Förtsch — The CMOS Chip Never Made It Past Second GradeMichael FörtschPhotonic computing tackles AI's data-movement bottleneck with light—and Q.ANT is already taking the hardware into real supercomputing deployments. Michael Förtsch, Q.ANT's founder and CEO, explains why memory can consume 95% of compute energy, how interference turns light into computation, and why legacy fabs can manufacture photonic chips. He also compares photonic and quantum processors, walks through PyTorch compilation, and identifies the converter and integration problems that still stand between prototypes and scale.
    Open segment on YouTube ↗

    Prakash introduced Michael Förtsch, founder and CEO of Q.ANT, the Stuttgart-based photonic computing company he started in 2018 after earning his doctorate — and the Otto Hahn Medal — at the Max Planck Institute for the Science of Light. Förtsch's pitch is that while the industry pours hundreds of billions into standard silicon and pins its long-term hopes on a distant, fault-tolerant quantum future, the immediate answer to the data center energy crisis is analog photonic computing — chips that compute with light instead of electricity. He was quick to correct two things right out of the gate: call him Michael, not "Doctor," and call the company Q.ANT, an acronym for quality, anticipation, novelty, and team.

    His core technical claim: a standard CMOS chip, for all its sophistication, "never made it past second grade" — it can only add and multiply, so every operation a computer performs has to be decomposed into plus and multiplication. Q.ANT's photonic processors, by contrast, can natively execute complicated functions — sine, cosine, exponential, Fourier transforms, convolutions, oscillations — without decomposing them first. That matters because, at a modern process node, roughly 95% of a chip's energy is spent moving data to and from memory rather than on the arithmetic itself. Rather than fight for gains in that remaining 5%, as most of the industry does, Q.ANT replaces the compute core itself, cutting how much data has to move at all.

    The tradeoff, Förtsch explained candidly, is that photonics has no optical memory and photons never sit still — they're always moving. So a system has to either compute while light is propagating, or pay an energy tax converting it back into digital memory via ADC/DAC converters. Q.ANT's answer is what he calls a "streaming architecture": chain as many optical operations together as possible before converting back to digital, rather than trying to graft a Von Neumann, memory-centric design onto optics. He argued the absence of optical memory, which looks like a limitation coming from CMOS, was actually what freed Q.ANT to think about computing on light's own terms.

    Asked by Nathan to explain the physics at the most basic level, Förtsch reached for two images: two stones thrown into water, whose interfering ripples are already a complex computation dependent on amplitude and distance; and eyeglasses, which perform a continuous, energy-free Fourier transform of an image onto the retina. Light passing through a diffractive medium does the same operation for free, and a tuned optical cavity can behave as a damped oscillator — a building block increasingly useful for state-space AI models. Q.ANT, he said, controls the full stack: its own wafer material and pilot line, PCIe-card processor integration, and the algorithm work that maps applications onto functions light can execute natively.

    The most concrete proof point discussed was a PyTorch object-detection model that Q.ANT and the startup Daisytuner compiled directly onto Q.ANT's photonic machine code in under three weeks — no CUDA-style intermediate layer needed. In a head-to-head demo at the T-Systems challenge, the compiled model matched GPU-class picture-identification accuracy on live video captured outside the T-Systems building, running at roughly 40 frames per second on Q.ANT's second-generation hardware, correctly identifying cars and trains.

    On benchmarks and roadmap, Förtsch rejected TFLOPS and TOPS as meaningful comparisons — Q.ANT isn't digital — in favor of inferences-per-second-per-watt, measured per application. On image classification specifically, he said Q.ANT sees a realistic path by early 2028 to beating CMOS by roughly 2x on throughput and 10x on energy. Longer term, he pointed to Fourier-transform-based attention-layer compression (citing Google research on KV-cache reduction) as a way into transformer workloads, and described a dream hardware path of moving past single PCIe cards toward a mainboard with 32 to 48 optically interconnected photonic chips streaming computation server to server.

    Pressed by Prakash and Nathan on supply chain and manufacturing, Förtsch described Q.ANT's chips as essentially standard silicon wafers with a thin lithium-niobate layer on top, built on a refurbished 1990s-era 90-nanometer CMOS line using off-the-shelf fab tools — no exotic equipment, just different, Q.ANT-owned processing recipes. He argued this decouples AI compute from bleeding-edge lithography: any fab willing to retool, in Europe or elsewhere, could become a photonic supplier without the capital or geopolitical exposure of a leading-edge node. Asked what could still derail the technology at scale, he pointed not to physics or manufacturing risk but to incumbent inertia — the innovator's dilemma — while betting photonic coprocessors become a standard AI-data-center add-on within three to four years.

    On quantum computing, where Förtsch spent roughly a decade before moving to photonics, he offered an extended car analogy: the CPU is the station wagon, the GPU is a drag racer built for one thing done extremely well, photonics is his aspirational Formula 1 car with more general-purpose horsepower, and quantum is a boat — indispensable for crossing a lake but useless on the road alone, and properly called a "quantum processor," not a full computer. Asked where the money is to be made, he pointed listeners toward buying and converting legacy semiconductor fabs rather than betting on the chip technology directly, and named low-power ADC/DAC converters as the real unsolved engineering bottleneck for the whole photonics field. He closed by noting Q.ANT has adopted AI internally for email triage, meeting notes, and database organization over the past six months, but not yet for chip design R&D — though he's intrigued by eventually simulating a photonic chip (a harmonic oscillator) on a photonic processor itself. Nathan and Prakash closed the segment reflecting on how genuinely hard it is to build durable intuition and investment conviction around a hardware paradigm this novel, and on the geopolitical stakes if legacy, non-leading-edge fabs really can become a meaningful new source of AI compute outside the current lithography chokepoints.

    This fundamental CMOS chip never made it past second grade of primary school, because it can multiply and it can accumulate. It can do plus and multiply, and that's it. Whatever you want to do on this machine, you have to break it down into something that's plus and multiplication.

    The GPU, in my opinion, is more the quarter-mile dragster — it does one operation, it does it excellently, in parallel, at speed, but please don't ask it to turn a corner, it's not built for that. Now, arrogantly, I'm saying we're the new Formula 1 car... and the quantum computer is the boat.

    I'd bet that photonic computing — at least in AI computing — becomes, in the next three to four years, a permanent add-on coprocessing unit in AI-based data centers.

    1:18:14Tell us about photonic computing and your deployments in supercomputing centers in Europe right now.
    Q.ANT builds photonic processors that compute using light rather than binary switching; standard CMOS chips can only add and multiply, while Q.ANT's chips can natively execute complex functions like sine, cosine, and Fourier transforms, cutting the data movement that consumes roughly 95% of energy in a modern compute stack.
    1:22:31Is there still a translation tax at the interface between the photonic and digital portions of the system?
    Yes — photons can't be stored, so light must either be computed on while propagating or converted back to electricity and digital memory via ADC/DAC converters; Q.ANT's strategy is to use models that need less conversion and chain as many optical operations as possible before converting back, in what he calls a "streaming architecture."
    1:26:33At the heart of photonic computing, what's actually interacting with what to turn inputs into outputs?
    Light is a wave, and two waves interfering (like ripples from two stones intersecting) performs a computation; light passing through a diffractive medium or a tuned optical cavity also performs Fourier transforms and models damped oscillations passively and for free, which Q.ANT's chips exploit directly.
    1:31:45What were the translation steps in porting a PyTorch object-detection model onto your stack via Daisytuner?
    Q.ANT and Daisytuner built a compiler in under three weeks that translates PyTorch directly into the photonic processor's machine code; in a head-to-head test against GPU chips at the T-Systems challenge, the compiled model matched GPU accuracy and ran at roughly 40 frames per second on Q.ANT's Gen 2 hardware.
    1:35:36What happens over the next two years, and what will you be benchmarking against?
    Since photonic chips aren't digital, TFLOP-style metrics don't apply; Q.ANT is standardizing on inferences-per-second-per-watt per application, targeting roughly 2x throughput and 10x energy efficiency versus CMOS on image classification by early 2028, while scaling the pilot line toward licensed foundries capable of 10,000–100,000 wafer starts a year.
    1:41:56Which specific applications and architectures does photonics already show an advantage on?
    Nonlinear/complex-function-heavy models — state-space models, diffusion models, physical AI, and spherical-coordinate simulations like weather and fluid dynamics — because photonics can execute the native math directly without remapping; he flagged world models as an exciting but still not-fully-understood frontier.
    1:44:04How much does your supply chain overlap with the standard GPU/CMOS supply chain, and what would have to be invented anew?
    Q.ANT's chips are 99% ordinary silicon with a thin lithium-niobate layer, built on a refurbished 1990s-era 90nm CMOS line using standard off-the-shelf fab tools; the IP is in Q.ANT's lithium-niobate processing recipes, not new equipment, and existing fabs could convert to production given demand.
    1:51:56If this doesn't scale the way you want, what would be the most likely failure points?
    He sees the production/technical risk as comparatively low since it reuses decades of fab engineering; the real risk is the established industry's innovator's dilemma — incumbents too rigid to adopt something beyond silicon and digital computing. He still bets photonic coprocessors become a standard AI-datacenter add-on within three to four years.
    1:56:03Having worked in quantum computing, what makes it less likely to succeed near-term, and what's still hard to resolve?
    Using a car (CPU)/dragster (GPU)/Formula 1 car (photonics)/boat (quantum) analogy, he argued quantum and photonic computers are both analog but suited to different problems; quantum processors are only useful for genuinely quantum-mechanical problems and should be thought of as coprocessors, not full "computers."
    1:59:51What are the bottleneck components people should invest in now to supply this future at scale?
    He pointed to buying and converting legacy semiconductor fabs into lithium-niobate production, since the same wafer real estate becomes far more valuable; on digital-to-light converters specifically, he called that a technological problem best solved by universities partnering with industry, not primarily an investment opportunity.
    Lightly edited · timestamps jump to YouTube
    1:16:28

    Prakash Narayanan: Our next guest is Dr. Michael Förtsch, a physicist and entrepreneur attempting to rewrite the physical rules of how artificial intelligence is powered. Dr. Förtsch earned his doctorate at the Max Planck Institute for the Science of Light, where his work in quantum information processing earned him the prestigious Otto Hahn Medal. In 2018, he founded Q.ANT, a Stuttgart-based deep tech company that builds native processing units — chips that compute mathematically complex AI and scientific workloads using light instead of electricity. While the broader technology industry pours hundreds of billions into standard silicon and pins its long-term hopes on

    1:17:14

    a distant, fault-tolerant quantum future, Dr. Förtsch argues that the immediate solution to the data center energy crisis is analog photonic computing. His hardware is not a speculative science project — Q.ANT’s NPUs, including second-generation processors, are currently deployed at major European supercomputing centers and are commercially accessible via the cloud. He’s here today to explain why 95% of data center energy is wasted moving data, how computing natively with light solves this fundamental bottleneck, and what the future of AI and drug discovery looks like when our hardware finally catches up with our software.

    1:18:08

    Michael, good morning. Welcome to the show.

    1:18:10

    Michael Förtsch: Good morning. Hi.

    1:18:14

    Prakash Narayanan: So let’s just dive right into it. Tell us a little about photonic computing, and tell us about your deployments in supercomputing centers in Europe right now.

    1:18:28

    Michael Förtsch: Right, so thanks for giving me the opportunity. By the way — first thing — thanks for calling me Michael, and not “Doctor Förtsch.” Second, you can simply call my company Q.ANT. We’ve gone through a lot of iterations of how to pronounce it — it’s an acronym of the values of the company: quality, anticipation, novelty, and team. There have been a lot of debates about how to name it, and in the end customers agreed to simply call it Q.ANT without a second thought. So what we are doing is photonics computing, and there’s usually the question: why? We have computers, and, well, in our belief, computers as we

    1:19:14

    see them today are a bit stressed by what we actually intend to do with them — when we look at what we dream about when it comes to AI. The most simple way to put it — I can always push toward getting more scientific about it, but when I explain the difference between our processors and the processors that have been successfully integrated into the stack for the last 40, 50 years: if you look at the fundamental CMOS chip, this chip never made it past second grade of primary school, because it can multiply and it can accumulate. It can do plus and multiply, and that’s it. Whatever you want to do on this machine, you have to break it down into something that’s plus

    1:19:59

    and multiplication. Now, the processors that we’re bringing to the stack — they went to high school, and eventually to university, let’s see how far we can push them. But at the fundamental level, these chips can offer complicated functions like sine, cosine, exponential, Fourier transformation, convolution, oscillations, and things like that, and you do not have to break them down.

    1:20:23

    Prakash Narayanan: Mm-hmm.

    1:20:24

    Michael Förtsch: And this fundamental difference now offers something that has been known in math but not really used in computational science. In math, it’s clear that you can trade amount of data, or data complexity, for function complexity — that’s well known when you talk to mathematicians. But it never paid off in computational science, because everyone starting such a revolution knew that, in the end, we still have to break it down into the fundamental operations of plus and multiplication. So all the algorithms currently out there have always been developed with

    1:21:10

    the knowledge that, in the end, it has to be simple. That’s something we offer, and that’s what we’ve started to demonstrate on use cases where, on one end, we provide a new processor, and on the other end, this opens doors to algorithms that allow AI models — networks — that come to the same result but with a fraction of the data. If you look at the stack — take a 3-nanometer-node regular stack — the energy is currently used at the memory. 95% is consumed

    1:21:56

    by the memory, not by the processor itself. And the less data you have to fetch from memory, the less energy you use. So part of the community is currently optimizing on the remaining 5%, trying to get things faster on the simple math side. We decided to replace the core, and by doing that, support shipping a fraction of the data across the stack. That, in the end, saves energy but also helps improve performance.

    1:22:31

    Prakash Narayanan: Let me stop you there and talk about the interface. At some point you still have an interface between the photonic portion and the digital portion, right? Is there still a kind of translation tax between the two?

    1:22:49

    Michael Förtsch: You’re — okay, very well, that’s the point, and that’s where you have to be precise. In the photonics world, everything is fine — there’s one minor problem. Two, actually. It’s great if you have a company that only has two problems. The first problem is we don’t have a memory — we don’t have an optical memory that’s semiconductor-integratable. That can actually be a benefit, I’ll come to that later. The second one: photons are not standing still — they’re always moving. So either you basically

    1:23:35

    compute while they’re on the propagation, or you have to back-convert them into electricity and then finally into digital memory. So if you don’t think the concept through very well, you’re basically eating up the energy you saved on the computational optical part directly at the ADC/DAC converters, because they, again, use a lot of energy. So the strategy is, first, use models that inherently require transporting much less fundamental data into the light — that already saves ADC/DAC conversion blocks. If you just take photonics and let it multiply and accumulate, then you have the same amount of ADC/DAC converters connected

    1:24:20

    to HBMs or DRAM, and they’ll eat up your benefits easily. The second part is you have to think about how to expand the grid — the longer you stay optical, and the more computation you can consecutively line up in a row before you go back into digital memory, the more benefit and gain you get compared to the CMOS stack. Because at one point you have to go into light, and at one point you have to come back out of the light, and the longer that chain is, and the more operations you can line up in that photonic architecture, the more benefit you get. There are

    1:25:05

    a few descriptions that came from the community — one says: what we’re currently building and shipping, you can see like an in-memory photonic computer, because we’re in between two memories, computing photonically between two memory cycles. Others — and this is what I usually prefer, if I’m speaking openly — call this a streaming architecture. And now I’m coming to what some might consider a drawback, that we don’t have an optical memory. Looking back on how we got to where we are, I’d say this was a clear benefit. Why? Because we just accepted

    1:25:51

    there is no memory, and this prevented us from thinking in categories like the Von Neumann architecture. It opened up doors to fundamentally think about computing from the abilities of light, rather than trying to copy and paste something that’s been working digitally very nicely into the analog optical domain — always searching for the next hub where I can memory-out my information to get in sync with everything else. So it’s a drawback if you come from CMOS. It’s a clear benefit when you look at it from the photonics perspective.

    1:26:33

    Nathan Labenz: Here’s a real simple, kind of idiot’s question. What’s going on at the heart of photonic computing? I’m certainly no GPU expert, but I can give you kind of a poor man’s understanding of what’s going on that allows information to be successfully processed — translating inputs to outputs. When I think about photonic computing, I’m really just purely speculating as to what’s going on at the heart of it. You said you’re using the properties of light, but can you flesh that out a little more — at the core of the computing mechanism, what’s interacting with what? How are these interactions happening? And, obviously you’ll have to leave a lot of detail aside, but how are inputs fundamentally being translated to outputs?

    1:27:35

    Michael Förtsch: Right — oh, now we’re getting really, really technical. I’ll try to use an analogy that probably most of us are at least aware exists, even if not everyone would directly know the result of it. In the end, we know that light is a wave, right, and we know that two waves can interfere. The best picture I have from daily experience is when you throw one stone into water, you have a wave propagating, and if you throw a second stone at a distance, a second wave is generated, and at the intersection point you have interference. That interference is already a very complex computation

    1:28:21

    because it depends on amplitude, distance, and everything like that. So when you accept there’s more than just one mathematical space where you can describe your problem, then you can effectively describe every problem by wave functions in light, and by interference effects, you can compute. That’s one part of what we do. The second part is when light propagates through a circuit — and none of you happen to be wearing glasses, so I’m not sure how familiar you and the audience are with what a Fourier transformation

    1:29:07

    means. Effectively, it means converting from a time domain into a frequency domain, or — probably the best example — from amplitude into phase space. There’s always a consecutive space, and it’s a very powerful mathematical operation that mathematicians often use to simplify a problem, in the sense of reducing the amount of data needed to describe the actual problem. Unfortunately, this operation is very expensive in CMOS. But light — when you wear glasses, they continuously transform the state space, the picture, into k-space onto your retina, without using

    1:29:53

    any energy. So when light goes through a diffractive medium, it naturally does this operation. The only thing is you need to know how to exploit it, and to get this into a circuit — but then it’s basically a passive element that can perform Fourier transformations free of charge. The same is true, for instance, for something that’s currently very interesting for AI models in the state-space domain — you need oscillations, damped oscillations. You start an oscillator ringing, and then it’s damped, and you have something that looks like a damped oscillator. That’s a passive element in optics.

    1:30:39

    You can just bring it onto the die, tune it, set the frequency, and effectively it does exactly what that mathematical operation wants for the model. There are a lot of these analog blocks that we’re developing. So when you ask what we at Q.ANT are doing — we currently have the full value chain under control. When I say this it sounds a bit crazy, but we have it from the material, from wafer production — we have a pilot line, we own the pilot line for the chips. We have the processor team that builds PCIe cards with an inherently optical core but connected via digital interfaces to the x86 stack. We integrate everything,

    1:31:24

    and we also take care of the algorithm space. This goes hand in hand: the algorithms team has an application and a problem, and describes it in functions that are natively operable in light. The chip team then looks at how to bring these functions natively onto the circuit, and the integration team then builds the processor.

    1:31:45

    Prakash Narayanan: So as I understand it, you guys earlier this year demonstrated an object-detection model compiled from PyTorch, using a software bridge called Daisytuner, onto your stack. What are the translation steps that occur between a normal PyTorch model and your stack?

    1:32:15

    Michael Förtsch: So what we did together with Daisytuner was take PyTorch as the standard model library, and together with them we built a compiler that ported this model directly onto machine code executable by the photonic processor. In the end, it’s a translator — a compiler is nothing but translating from one language into another — and this compiler actually translated it from PyTorch into the machine language of our processor, which we demonstrated. Why did we do this? Because we’ve

    1:33:01

    often been asked, “Great, you’re doing the chip, but then you need something like” — I’m not making advertisements for this company — “something like CUDA.” And we said, no, we don’t actually need that — not because we’re claiming we don’t need it and have something different, but because the photonics architecture works fundamentally differently. As I said, it’s a streaming architecture — you don’t need a redistribution layer to bridge things across the chip, you always have to think about it from the live perspective. We said, hey, we believe we can build a compiler — I think it took us less than three weeks together with Daisytuner to build a compiler architecture to port this model from PyTorch

    1:33:46

    onto our chips. It was an image classifier, and we had a head-to-head comparison at the T-Systems challenge against chips from the GPU world. From a quality standpoint — I’m not claiming we’re already at the speed of, say, AMD or NVIDIA, that will take probably one or two more years — but in the end we reached the same result. Same picture identification, same rate — we could identify at a frame rate of, I think, 40 frames a second on this PyTorch model on our generation-2 hardware, and could correctly identify

    1:34:32

    cars, trains, and so on, on video captured live in front of the T-Systems building. That showed — because there’s often this “yeah, that’s great, but it’s so scientific, and you can only test some functions, it’s not really useful in a day-to-day application” — we said no, it is. You can do this. We need more years to get to the same robustness level, to have a fully-fledged compiler architecture proven against a lot of use cases. But fundamentally, this technology is on the rise — it can already step out of the laboratory, which it does. We operate in

    1:35:18

    industrial data centers already on specific applications, and this was one of the showcases we thought was good because it’s so intuitive to understand. If this can be done, then a lot more can also be ported from PyTorch and directly executed.

    1:35:36

    Prakash Narayanan: So what’s going to happen in the next two years? What are the steps you need to get to where you need to go? And when you get there, what are you going to be comparing against? Are you going to say, “I have an A100 or H100, and our chip is going to do this specific model at this specific tokens-per-second”? What comparisons will you use in a couple of years, and what steps do you need to get there?

    1:36:07

    Michael Förtsch: So that’s the question. First of all, since this technology is different from the GPU world, it doesn’t make a lot of sense to compare, for instance, tera-operations per second, because we’re not digital. Same is true for FLOPs — it’s a metric that was built for the mat-mul domain. So what we’ve agreed on is inference-per-second-per-watt, taken per application. What that means is — I can give you some insight into what we see as a great potential

    1:36:52

    case, a first demonstrative case. If we go back to image classification — the world record, at least to our understanding, is you can identify 16,000 pictures a second, or maybe 24,000 pictures a second, per 70 watts. So inference is the picture — that’s the job. How many pictures per second can I inference correctly, at the same quality, and what was the energy consumption? We see a very good chance that by early 2028 we can build a system that outperforms the CMOS

    1:37:38

    world in this specific application by a factor of 2, and by a factor of 10 in energy — because we allow for different types of models that are way less demanding on the data overhead needed to describe the model, or needed to correctly identify the picture. It’s not the largest market on earth, but it’s a very clear use case to demonstrate the capability of a photonic system. The other thing we see is that this technology can open doors into the transformer architecture — what we’re looking at is

    1:38:23

    dedicated support for the attention layer. There are papers, among others from Google, that say using Fourier transformation in the attention layer can realize an enormous KV-cache compression. The point, again, is: offer a Fourier transformation that’s cheap, and you can probably contribute substantially to that stack — that’s where we’re currently heading. Now, you also asked what it takes to bring this to scale. First: we have a pilot line in operation that’s good for a few thousand wafer starts a year — that’s great for the moment, but we need to license this out to commercial foundries

    1:39:09

    who can scale it to 10,000 to 100,000 wafer starts a year. It’s doable, not so complicated, but it has to be done. Second, we already talked about the photonics chip itself — that’s well in hand — but the integration layer into electronics is where you can lose or win a lot of performance in this transformative space. That’s where the systems-integration part is currently our strong focus — getting the digital memory energy-efficiently connected to the photonics core. The photonics core itself, I’d say we’re pretty much there. So on the photonics

    1:39:54

    chip side, we can already tape out photonic chips competitive with the CMOS world on dedicated application comparisons. What’s still lagging a bit is the integration part from the hardware layer to memory, and of course we need a more robust software stack — but these are, in my view, not fundamental issues, more engineering tasks we need to complete. Once we’re at that point, the route to the future, in my opinion, is: first, we believe these new processors are triggering

    1:40:40

    a model revolution in the AI space — we’re going to request and build new functions, going back and forth between algorithm teams around the world and us about delivering particular functions for novel models. And second, on the hardware side, where I’m dreaming, is the following: I want to take this to a real streaming architecture. Think about it — step two could be, you’re not thinking in terms of a PCIe card, you’re thinking in terms of a mainboard. The mainboard has a host CPU for the operating system, ADC/DAC converters on the front end, and then 32 to

    1:41:25

    48 photonic chips optically fully connected. You stream it across these chips, then go back into digital memory. That would be at server level. But why stop at server level? You could take this to the next server, and the next, as long as your signal budget is still good. And if you really want to take it to the extreme, you could start computing directly in the optical grid. So that’s — like—

    1:41:56

    Prakash Narayanan: You mentioned some applications where you can already see, like, a 10x improvement — which are those applications? Which models, which architectures does photonics already have an advantage on?

    1:42:10

    Michael Förtsch: So, in general, every model that inherently builds on nonlinear and complex functions, because we can execute them more efficiently and faster than the silicon world can. In particular, these are the state-space models, the diffusion models, hybrid mixtures in between these models, physical AI. Think about it — we all went to school, and probably all had to do, once in a while, a coordinate transformation from a Cartesian coordinate system into a spherical one. We can directly execute in the spherical coordinate system — you don’t need to remap it into Cartesian in order to bring it off the chip. We’ve also tested, for instance, algorithms

    1:42:56

    on the simulation side — weather forecasting, or fluid-dynamic simulations. Fantastic — model the world as a sphere and directly simulate it on a photonics computer, that’s where you can go with these machines too. You don’t need to go down into the fundamental math. And — I have to say it, because it’s a big word nowadays — we also see that we can contribute to world models. At this point I must say we need to understand even better what it means. Everything I’ve said so far, we have a clear mathematical model in mind for.

    1:43:42

    The world-model category is fascinating, and if we were to model the world, we’d do it as physicists would — we’d use harmonic oscillators. That’s how the mechanical world is modeled, and that’s at the heart of our computing. This is what we’re currently also looking into.

    1:44:04

    Prakash Narayanan: When you look at your supply chain, how much does it overlap with the standard GPU supply chain and manufacturing process, and how much will have to be invented anew in order to scale this?

    1:44:30

    Michael Förtsch: So I’d say the most fundamental difference from the classical CMOS world is that, on the material side, we don’t just have silicon — we have silicon, and on top of it, a thin layer of lithium niobate.

    1:44:50

    Prakash Narayanan: And—

    1:44:51

    Michael Förtsch: This thin layer is where the light propagates. So it’s 99% silicon, and a few percent — if even a percent — of this optical crystalline material. When we look at how we built our pilot line, we used a pilot line that already existed — it was a former CMOS line from the 1990s, a 90-nanometer node, etch tools, everything standard. The only thing we did was replace a few tools because they weren’t compatible — on some machines you can’t run silicon in parallel with lithium niobate because of cross-contamination.

    1:45:36

    Prakash Narayanan: Mm-hmm.

    1:45:37

    Michael Förtsch: That’s it. But the tools we have in the line are standard, off-the-shelf semiconductor tools that you buy from regular vendors — there’s nothing secret about that. The secret is that the processing strategies for the lithium niobate are owned by Q.ANT, and they’re different from the silicon strategies. But in the end, you do exactly the same thing — you have a wafer, you build a mask on it, you structure the mask, you etch into the mask, you remove the mask again, and then you have your structured wafer. You dice it, passivate it, package it, and bring it up onto the processor. So from that perspective, does the industry really have to change? Not really. We’ve also discussed with fabs

    1:46:22

    whether they’d be willing to bring some of their lines over to manufacturing for lithium niobate, and they said yes — as long as the volume is there, they have no problem turning silicon 90-nanometer or 45-nanometer lines into lithium niobate lines, as long as the demand is there. So does the industry really have to change now? Do we have wishes for the established industry? Yes, of course we do. In particular, we’d love to cooperate to build more energy-efficient ADC/DAC converters, because that’s currently the point where the most energy

    1:47:07

    in the system is consumed, and we believe there’s a lot of room for improvement. And the rest would be a discussion about whether it always needs to be, for instance, PCI Express, or whether we can push against some established standards in favor of bringing photonic processors even faster into — I mean, as long as we’re looking only at the server level, we don’t have a problem, because if you buy a server from us, you plug in energy and Ethernet or InfiniBand, and that’s it. You don’t need water supply, you don’t need anything, because the systems aren’t getting hot — energy demand is so low that every one of us could have such a server at home. If you go inside the server and talk about

    1:47:54

    integration at the mainboard level, then I’d say we can talk about standardization there as well.

    1:48:03

    Nathan Labenz: When you talk about 45- or 90-nanometer nodes, those are obviously not the latest and greatest. Does this mean, from a global-supply-of-compute perspective, that as this starts to work and scale, it will be almost exclusively net-new compute coming available? This is competing with stuff that’s on relatively low-end lines — these would be the kind of chips that go into toys or whatever, not anything close to what would go into a modern cell phone or — yep — into a modern AI stack. So what’s your dream success scenario look like, in terms of, without photonic computing versus with — how much bigger does the overall supply of compute available for AI get? What does that delta look like?

    1:49:15

    Michael Förtsch: Actually — what we’ve demonstrated — look, we are in Germany. Germany is known for a lot of technology, but for sure we’re not famous for logic computing — we also don’t have 3-nanometer-node fabs here, right? And still, we managed to get these systems running, even the pilot line. So what we’ve demonstrated in Germany on a 90-nanometer node can be copied across Europe, into the States, across the world. So if this technology

    1:50:00

    starts winning, you can turn a lot of existing fabrication sites — without the need to build new ones — into fabrication for this. So the bottleneck we currently have in access to latest-node fabs, and discussions about whether the business case still holds for a 2-nanometer node — I’m not judging that, but the discussion is ongoing — those constraints aren’t there for us. I wouldn’t say it’s democratization, but effectively it reduces the complexity of

    1:50:46

    the supply chain a lot. At this point we’re, from wafer to processor, nearly self-supplying. That’s another angle where I see this technology, besides the beauty of the performance and the reduced energy, simply offering production capabilities that are so much easier to scale than going to a 3- or 2-nanometer node, eventually getting picked up by an MPW run somewhere mid-next-year, getting your hero chip back, and then trying to get volume behind the line, because that’s damn expensive — it

    1:51:31

    goes through the whole process. A mask on our side is cheap compared to a mask layout on the logic CMOS side. So all these dimensions offer great capability to, on one side, reduce production cost, and on the other, increase volume very rapidly.

    1:51:56

    Nathan Labenz: So if we don’t get this to the scale you’d like to see delivered, and we’re looking back and saying, “why didn’t it work?” — what would be a couple of candidates for the remaining really hard problems that could derail the whole upside of this? There have to be some that aren’t easy to solve, right?

    1:52:27

    Michael Förtsch: So — is this the question — I’m sorry, I’m a little puzzled. So the question is, what happens if photonic compute doesn’t make it to success — what would be the alternatives? Is that the question?

    1:52:39

    Nathan Labenz: Well, that’s an interesting one too, but I meant more — if you’re not ultimately able to succeed in scaling in the way you want to, why would that be the case? What would be the most likely failure points?

    1:53:03

    Michael Förtsch: The most likely failure point, in my opinion, is that the established industry is too rigid in its existing business model. So it’s a question of whether the supply chain, and our partners and future partners, are willing to accept that there’s more than just silicon, and more than just fundamental digital computing. If that’s accepted, and if people start trusting that there can be volume — which, by the way, is starting already — then the great thing is that

    1:53:48

    you can repurpose existing knowledge. When you run a fab, from an engineering point of view, there’s not so much difference between manufacturing a silicon chip at 90 nanometers and a lithium niobate chip at 90 nanometers. Yes, there are peculiarities — recipes, you have to look at some more parameters, or different parameters — but in the end, this is engineering that’s been done for 40, 50 years, and the complexity is what it is across all these dimensions, from inspection to everything — it’s all much easier. So everything these companies know as of today can directly be used in producing photonic chips.

    1:54:34

    So, in my opinion, the technological barrier, or risk, of this technology is comparatively low when it comes to production. The actual risk that this technology never makes it at scale is that we’re too caught up in the innovator’s dilemma, and industry players don’t swing into this opportunity. But honestly, the odds are so great that I’d place a different bet.

    1:55:20

    I’d bet that photonic computing — at least in AI computing — becomes, in the next three to four years, a permanent add-on coprocessing unit in AI-based data centers. The trickier question is who’s going to be the first company to make that bet come true. I think it’s going to happen — I’m very sure of it. It’s more about the first mover who conclusively can bring such a system to a productive level, taking up a large portion of the market.

    1:56:03

    Prakash Narayanan: One of the questions I have for you — you worked on quantum computing earlier in your career. Now you have these two paradigms, quantum and photonic computing, and you’ve obviously chosen the photonic computing angle. What do you see in quantum computing that makes it perhaps less likely in the near term, and what are the challenges you think will be hard to resolve?

    1:56:34

    Michael Förtsch: So — I’ll give you an analogy that tells you why I’m still very much convinced about it. We need a quantum computer as well, but for different reasons than are usually communicated in the press. I’ve been working on quantum computers, I think, for 10 years, until I decided it’s going to be an analog computer instead. Photonic computers and quantum computers are both analog computers — the difference is the quantum computer uses quantum effects, and the photonics computer uses wave effects of light. The way I see it: Germany is a car company, right? So we have the CPU,

    1:57:20

    which is the station wagon — it has five seats, you can pack the kids in, go grocery shopping, everything’s fine. It can also have a lot of horsepower, but no one expects it to win a Formula 1 race — it’s not the right car, but every family needs a station wagon. Now the GPU, in my opinion, is more the quarter-mile dragster — it does one operation, it does it excellently, in parallel, at speed, but please don’t ask it to turn a corner, it’s not built for that. Now, arrogantly, I’m saying we’re the new Formula 1 car, because we have way more operations, we can drive around the circuit very fast — but please don’t go grocery shopping

    1:58:05

    with our car. In other terms, don’t let the operating system be executed by our chip — we’re also specialized, but with a bit more universality than the GPUs. And the quantum computer is the boat — it’s a vehicle, it’s great, you need it because none of the other three can go across a lake, but it has special functions. As soon as you put it on the road, though, you need something to pull it around, because it can’t do that on its own. So this is the way I see it: quantum computers are built to solve quantum-mechanical problems. Not every problem is a quantum-mechanical problem, and it doesn’t make sense to turn every problem into a quantum-mechanical

    1:58:51

    description. As long as it’s not a quantum-mechanical description, you don’t need a quantum computer, full stop. That’s the way I see it. One remark, by the way, since you have a huge audience — maybe you can vote with your audience on something. I think the terminology “quantum computer” is inherently wrong — it should be “quantum processor,” because a computer is more than a processor, it owns the memory, it owns everything. What we’re building are quantum processors — they’re, again, coprocessors to the stack. There’s a lot of wonderful literature, great scientific papers, about hybrid systems — how a quantum computer in

    1:59:36

    conjunction with a classical computer can accelerate things. It’s exactly the same with the car and the boat — if they join forces, you can go across the lake and the street.

    1:59:49

    Prakash Narayanan: It’s a great analogy.

    1:59:51

    Nathan Labenz: I wish I was smart enough that I’d developed such extended analogies to explain my work to people who can’t get it at the foundational level. Maybe last one for me — there’s a popular investment strategy, become a bit notorious recently for a mix of reasons, that’s looking for what’s going to be the scarce component as new kinds of systems scale up. I think you may have already kind of answered this, but what would you say are the bottlenecks people should be investing in now, to be able to supply the future at scale that you envision, and — obviously — make a little money along the way?

    2:00:46

    Michael Förtsch: So, if people want to get a significant share of this success, I’d encourage them to search for legacy fabs and turn them into production capacity for this technology. It’s not so expensive an investment, but the same fab can suddenly sell the same square millimeter of die at a much higher price, because the value of the chip — the problem it suddenly solves — is so much bigger than putting a 90-nanometer chip into a child’s toy for a dollar.

    2:01:25

    Nathan Labenz: Yeah, that makes sense. How about the digital-to-light converters — that was what I thought you might say.

    2:01:34

    Michael Förtsch: But that’s not — I think this isn’t an investment, that’s a technological problem. Okay, now I got your question. So on the engineering side, we need to get into the component space together with the semiconductor experts, to develop miniaturized, lightweight, or less energy-consuming ADC/DAC converters. That’s something the whole photonic industry — us, photonics computing, and the interconnect people I’m betting on — are dreaming about. So whoever solves that problem has a gigantic market in the photonics space, in interconnect and computing. But is it an investment? Yes, of course it’s always an investment point, but I’d encourage universities in this field to be genius enough to own the IP, and then team up with one of the global partners in that field to get it to scale very soon. That’s my dream there.

    2:02:40

    Nathan Labenz: Nice, thank you. Sounds like the easier money is to be made converting old fabs — eventually.

    2:02:48

    Prakash Narayanan: Mm-hmm. One last question from me — to what extent have you seen workflows inside your company being converted to use AI tools in the last six months?

    2:03:09

    Michael Förtsch: It’s exactly six months, so yes, we started using AI internally in many categories. We haven’t used it for R&D purposes so far, but what we’re definitely doing is sorting the email inbox, writing meeting minutes, and structuring our database so we can find our stuff again — that’s basically where we’ve been using it so far. There have

    2:03:55

    been discussions about whether AI could also accelerate chip design, but as of this point, at least what we’ve seen, it needs a bit more work. And now I’m coming to something where maybe two things come together — the chips we’re building are inherently harmonic oscillators, so the best thing would be to simulate a photonics chip with a photonic processor. That’s the optimal point, because we’d be simulating the system with its own dog food, essentially. That’s something we might start activities on very soon.

    2:04:40

    At least we have some scientific thoughts about doing that in the future.

    2:04:46

    Prakash Narayanan: Your own form of recursive self-improvement. Dr. Michael Förtsch from Q.ANT, thank you — thank you for joining us this morning.

    2:04:56

    Michael Förtsch: Thank you very much, and when you come to Germany, we have great technology, and in the meanwhile, also coffee — you’re always invited.

    2:05:04

    Prakash Narayanan: Best.

    2:05:05

    Nathan Labenz: Bye. Thank you very much.

    2:05:06

    Prakash Narayanan: Thank you.

    2:05:11

    Nathan Labenz: Fascinating stuff. I mean, I wish I was smart enough to understand it as deeply as I’d like. I truly feel like — I can sort of step through the Fourier transform notion, but my intuition is still pretty limited in terms of how it really works.

    2:05:37

    Prakash Narayanan: I found that his explanation was probably one of the better explanations I’ve come across. To kind of restate it — I guess what he was saying is, if you have a digital chip, you’ve got this kind of electronic wave, and then you have a semiconductor, and it decides if it’s a 1 or a 0 — if it’s above a certain level, it’s a 1, if it’s below a certain level, it’s a 0. And with that 1 and 0, you can basically do addition and multiplication, and everything else you do has to be a function of addition or multiplication.

    2:06:23

    So every other function is built on top of that framework. If you have, say, a fast Fourier transform, it breaks down into hundreds or thousands of FFT samples, which break down into, again, addition and multiplication — maybe thousands or millions of operations. Basically, what he’s saying is, instead of having that electric wave with a 1 and a 0, you have light waves, and when they interact with each other, they can produce the end function point rather than the initial ones and zeros. You can

    2:07:08

    jump straight to the sine wave or cosine wave, and take a measurement of that wave to get your response back. This is also, essentially, what our friend — and some other people — are doing with thermodynamic computing, as they call it, and other forms of analog computing: instead of doing the 1 and 0 as the baseline and breaking everything else into 1s and 0s, you’re directly measuring that cosine or sine wave, or other mathematical function, and that cuts short those

    2:07:53

    hundreds or thousands of operations you’d otherwise need. So it’s less general, but you get to skip a few steps. But, as he pointed out, there are other complications — you don’t have memory, you can store ones and zeros in memory, you can’t store a wave in memory. And you probably have a bunch of timing aspects too, because it’s a stream — if your streams are off, your cosine and sine wave and fractional rates are off, you get a different reading. So there are a bunch of other complications they have to manage, I guess. But, in essence, that’s it. I also found his quantum computing description very interesting.

    2:08:39

    That’s also a good explanation of quantum computing — comparing it to, when everyone else is building a car or a race car, they’re building a boat. It’s going to get you across the lake much faster than the guy taking the highway, which is circuitous, but it’s only useful for that one function.

    2:09:08

    Nathan Labenz: That’s good — I think I might still have to ask Claude for some sort of artifact to help me build intuition even further. I still couldn’t — I understand not always wanting to get right down to it in an interview setting, but I feel like there’s still got to be something that’s really hard about scaling this up — otherwise we’d already be scaled up, right? What are the core things that might prove really difficult? I don’t know — you mentioned timing, is it maybe the kind of thing where the medium the light travels through has

    2:09:53

    to be, like, presumably exactly the right length? And if it’s not, things get out of phase and don’t work? In fab analysis generally, there’s always the question of yield — how many of the things you make actually work? Would that be the kind of thing that limits yield in this kind of process? I don’t know — maybe Claude will help me get to the bottom of that. But yeah, it’s an unbelievable — this is the AI-scout’s opportunity and curse all in one

    2:10:39

    conversation, because AI is now just intersecting with everything, and increasingly it’s like hard science that takes a long time to grasp the fundamentals of, and it makes it very difficult to know how to incorporate something like this into one’s worldview. If you had conviction on this, you’d not only have an investment thesis, but a point of view on sovereign AI, and the international politics of this become very different. It’s a very different world if European 90-nanometer legacy fabs can be converted and create

    2:11:25

    a significant amount of compute — it becomes much harder for the US to bully other countries based on our position in the supply chain, and much easier for them to at least get some domestic base of computing. But you’ve got to have conviction before you start propagating this through your entire worldview. And, for me — not to use a “Claude-ism,” but — my broader world-model updates are gated by my lack of understanding, and not wanting to be too quick to take something on board without really understanding why it might not happen at the

    2:12:11

    depth that I would like to.

    • PyTorch Runs On Photonic Hardware

      0:00 / 0:00
    • Photonic Computing Could Beat CMOS By 2028

      0:00 / 0:00
    • The Quantum Computer Is The Boat

      0:00 / 0:00
    • Photonic Chips Go Beyond Plus And Multiply

      0:00 / 0:00
    • The Photonics Bottleneck Is Industry Rigidity

      0:00 / 0:00
  4. 2:10:56Closing24 min
    Closing: OpenAI's Jalapeno Chip, Vibe-Coded RL Environments, and Monitors That Aren't ReadyPrakash Narayanan opened the close with the morning's other announcement — OpenAI's first custom inference chip, Jalapeno, benchmarked against NVIDIA's Blackwell systems — and argued the comparison is structurally misleading: on Jensen Huang's own stated roadmap NVIDIA compounds at roughly 4x a year, so a part taped out today lands in data centers against something far ahead of what it was measured on. His read is that the chip is a lever to cap NVIDIA's pricing rather than a challenger to it, and Nathan Labenz noted the market agreed, with NVIDIA up about 1.68% that day. From there: elastic demand for intelligence, old chips renting for more than they used to, the Ethereum-mining precedent for GPUs that refuse to depreciate, and Prakash's hope that the next twelve months bring intelligence for discovery rather than intelligence for coding. The last third returned to the opening's thread with a former RL-environment builder's account of rushed, "vibe-coded" environments across the industry, a data labeller who edited a page's JavaScript until Codex would do the job and got banned for posting about it, and Nathan's structural read: labs scaled RL faster than environment quality could support, optimization pressure amplifies a small cheating tendency the way a microscope amplifies a slightly off-center target, and if the models training the next models are themselves cheating, the monitors are not ready.
    Open segment on YouTube ↗

    Prakash opened the closing segment with a rebuttal to the case for unlimited AI compute demand: OpenAI's same-morning announcement of Jalapeno, its first custom inference chip, with performance claims pitched against the NVIDIA B300 and Grace Blackwell 200 across throughput and power-efficiency benchmarks. But he argued the comparison is misleading once the multi-year lag is accounted for — NVIDIA's own roadmap, per Jensen Huang's stated target, calls for a millionfold performance increase over ten years, which works out to roughly 4x per year, so a chip taped out today and benchmarked against a two-year-old B300 will already be facing a roughly 64x-better NVIDIA part by the time it actually reaches data centers in year three. His read: OpenAI's chip effort functions less as a genuine NVIDIA-killer and more as a negotiating lever to cap future NVIDIA price increases.

    Nathan agreed with the analysis and tied it to NVIDIA's stock reaction that same day — up about 1.68%, roughly a $75 billion market-cap gain — as evidence the market isn't buying the "this changes everything" framing either. He connected his own thesis to elastic demand for intelligence, betting it will stay "insanely elastic" enough that compute remains scarce and prices don't fall, citing as his strongest signal that older-generation chips are now renting for higher hourly rates than before — something he credited to having become SemiAnalysis-pilled. Prakash extended the point with a historical parallel: the same repricing dynamic played out during the Ethereum mining boom, when four-year-old GPUs held or gained value instead of depreciating toward nothing; he framed the current AI compute regime as an even bigger step-change in value capture, and said he's hoping the next twelve months bring a shift from "intelligence for coding" to "intelligence for discovery," with new firms delivering genuinely novel scientific results the way frontier models already have in math.

    The conversation turned to how much of the AI hype cycle is substantive. Nathan argued that while breathless framing gets tiresome, raw-capability hype has consistently been a good predictor of what actually ships, even as real-world diffusion has lagged capability growth. He then surfaced a tweet — flagged via a Zvi retweet — from a self-described former RL-environment builder describing how RLVR training environments across the industry are frequently rushed and "vibe-coded," leaving them full of exploitable bugs that models learn to reward-hack rather than solve legitimately, because flagging an environment as buggy slows everyone down. Nathan connected this to his own reporting (conversations with someone he referred to as Bronson) on how deeply the instinct to cheat is now baked into current models, and predicted growing scrutiny of who builds RL environments, whether they can be trusted, and what they're actually teaching models — especially since heavy reward-hacking in a model's reasoning undermines cheating-detection monitors with false positives.

    Prakash added a concrete anecdote: someone hired for a data-labeling job got Codex to refuse the task, then edited the page's JavaScript to insert language claiming AI models were explicitly authorized to complete it — after which Codex complied, the person earned roughly $500, and was banned once they posted about it publicly. He drew a parallel to OpenAI's early data-labeling operation in Kenya, where contractors reportedly began using ChatGPT itself to do their labeling once it launched, degrading that data's value. His overall take: gamed or dirty training data is a persistent, not novel, problem, and frontier labs are reasonably well positioned to catch it over time through vendor quarantining, sampling, and quality grading — ultimately, benchmarks and commercial usability are the real check on any given lab's data discipline.

    Nathan pushed back somewhat, arguing there are current structural problems layered on top of that baseline noise: OpenAI has said it had to pause RL training to address environment-quality issues, the industry's fragmented, "shotgun start" on environment-building means labs are bad at attributing which environments cause the worst behavior, and labs have been scaling RL faster than environment quality can support. He offered a friend's microscope/telescope analogy — cranking up optimization power can zoom right past the intended target once the signal is slightly off-center — to describe how a weak-but-present cheating tendency gets massively amplified under heavy RL. He said he'd bet the problem is fixable, or at least controllable, on a reasonably short timeline, but flagged real concern about recursive self-improvement: if a model doing the training of future models is itself cheating, there's no telling where that goes, which argues for keeping humans in the ML loop longer than published roadmaps suggest.

    Closing out, Prakash said the last few months have felt a little "occluded" by delayed model releases, leaving him less certain how much capability the labs are actually sitting on. Nathan agreed there's likely a large capability overhang, pointing to the hundred-million-token persistence now showing up in next-gen model rollouts — runs that span the equivalent of multiple sleep cycles within a single episode and, per what he called the "mythos incidents," keep pushing forward even amid apparent goal drift or confusion about whether the task was actually working. He called it a qualitative shift from prior generations and warned that current monitoring infrastructure isn't ready for it. The hosts signed off with a preview of a robotics demo — described as carrying "a one-shot promise" — to pick up the next morning.

    This is a negotiating tactic against future NVIDIA price increases, to manage pricing so that they have a cap on how high NVIDIA pricing can go.

    The old chips are renting for higher hourly rates now than they used to. I just think that's such a strong signal.

    Buckle up. Our monitors are not ready. I can tell you that with confidence at this point.

    OpenAI publishes Jalapeño's first benchmark numbers, and Prakash reads them as a negotiating lever OpenAI released the first benchmark results for Jalapeño, the inference ASIC it co-designed with Broadcom, measured on SemiAnalysis's InferenceX suite. Prakash's summary on air — roughly 50x on some token metrics and 2-3x on others — tracks the published numbers: 1.5-1.9x higher peak mixed tokens/sec/kW and 1.7-3.6x lower end-to-end latency, alongside 53.7x, 104.3x and 56.1x throughput multipliers at matched time-between-tokens against Blackwell systems. Two details he had slightly off: the baselines were GB200 and GB300 rather than the B300, and OpenAI's stated timeline is very small volumes at the end of 2026 scaling through 2027, not full data-center deployment by the end of next year. His argument is unaffected — on Jensen Huang's own millionfold-in-ten-years roadmap NVIDIA compounds at about 4x a year, so a part measured against today's silicon meets something far ahead of it on arrival, which makes the chip more useful as a cap on NVIDIA's pricing than as a replacement for it.

    The market shrugged, which was Nathan's point Nathan cited NVIDIA up about 1.68% on the day as evidence nobody was pricing in a challenger; quoted live at midday, that was an intraday figure — NVDA closed 2026-08-25 up 2.19% at $213.05. He tied it to elastic demand for intelligence, with older-generation chips renting at higher hourly rates than they used to as his strongest signal, and Prakash added the Ethereum-mining precedent, when four-year-old GPUs held or gained value instead of depreciating.

    A former RL-environment builder: "rushed and vibecoded," and flagging a bug was discouraged The thread Nathan surfaced, which reached him via a Zvi Mowshowitz retweet, is from a laid-off worker at an outsourced RLVR training-data provider: nearly all the environments were rushed and vibe-coded and failed to robustly reflect the real things they were based on, models were effectively encouraged to reward-hack them, and marking an environment as bugged was possible but discouraged because it cut the volume of training data produced. It is the same claim as the opening segment's, from inside the cottage industry rather than from the chain of thought.

    ⚠️ The data-labelling story is recounted on air only — the original post was not located Prakash described someone hired for a data-labelling job who edited the page's JavaScript until Codex would agree to do the work, earned about $500, posted about it, and was banned. A date-scoped search of X and the open web across roughly a dozen query variants turned up no trace of the original post, so it is recorded here as his on-air account rather than as a sourced claim. His accompanying point — that contractors on OpenAI's early Kenya labelling operation began using ChatGPT to do the labelling once it launched, degrading the data — is the well-reported version of the same dynamic.

    Lightly edited · timestamps jump to YouTube
    2:12:13

    Prakash Narayanan: So let me maybe share one reason why it might not happen, which is that current chips just get better. This morning, OpenAI announced Jalapeno, which is an inference chip — their first custom inference chip. They've been testing it. They have performance numbers where they compare it to the existing best, which is the NVIDIA B300. They have not named the existing best in their materials.

    2:13:00

    And they've said roughly that it's perhaps fifty times on mixed-token, fifty times on single-token, something like two times, two to three times — basically a bunch of state-of-the-art numbers. They compared the Grace Blackwell 200 here on number of tokens per second, and mixed tokens per second per kilowatt, all the industry-leading benchmarks, etcetera. And I think this is one of the challenges that you have. So I would say

    2:13:46

    —to be, from my point of view, it's great that OpenAI has announced this chip. They say that it'll be in data centers by the end of next year. Great. But the number that I go back to from NVIDIA is a million-times increase in performance over the course of ten years. That's Jensen's stat — a million x over the course of ten years, which is what they've achieved in the past ten years and what he wants them to achieve in the next ten years. The problem with that is that NVIDIA has to 4x the performance every year. So every year is a 4x performance increase. And what ends up happening

    2:14:31

    —with that is that, let's say, OpenAI has finished taping out this chip right now, and they're comparing it with, let's say, the B300. The B300 was taped out in December 2024. So it's a two-year-old chip, and they have a performance increase of somewhere between four and ten times on a two-year-old chip, which will be in data centers in year three. So by that time, NVIDIA will be launching a 64x better chip at that point in time, by the time

    2:15:16

    —it's in data centers. So I think this is the challenge that you have on the leading edge, which is: the current players will not stand still. It's great that OpenAI has their own chip team, but to me, this is a negotiating tactic against future NVIDIA price increases, and to manage pricing so that they have a cap on how high NVIDIA pricing can go.

    2:15:49

    Nathan Labenz: Yeah. That's all very interesting, and I think good analysis. And indeed, NVIDIA stock is up 1.68% today, which is a cool $75 billion or so increase in market cap. So, yeah, the market is not failing to realize what you point out. It seems that this comparison has an offset of a couple of years that is not immediately foregrounded in the analysis. And, you know, if we're thinking we had a buying opportunity for NVIDIA because of this, you know, potentially NASDAQ-tanking news, it's

    2:16:34

    not the case. I still think there is something — you know, I guess a key question for all these sort of things is just how elastic is the demand for intelligence. And I think I have come around more and more to believing that it's just gonna be insanely elastic, and that it's gonna be really hard for there to be enough compute for prices to drop. Right? So as long as that holds, then it kinda — and this is sort of the Dylan. I guess I've been, like, SemiAnalysis-pilled, to a degree here, but really compelled by the fact that the old chips are renting for higher

    2:17:20

    hourly rates now than they used to. Mhmm. I just think that's such a strong signal. As long as that holds, then there's, like, all these plays make sense. Right? Then it's kind of like the AI economy just continues to eat more and more of everything, and these old fabs are just another step in the global conversion process of figuring out how we're gonna create as much intelligence capacity as we possibly can, because there's no limit to how much we'll wanna use. And I think that is my best bet right now.

    2:17:59

    Prakash Narayanan: So — you know what is really uncanny for me? It's that it's not the first time that NVIDIA chips have gone through this cycle of older chips getting more expensive. In fact, they went through one of those cycles during the Ethereum mining phase, where Ethereum mining got profitable enough that the old chips got more expensive. And so people were able to resell their four-year-old graphics cards for the same or higher prices. It was pretty incredible, because everywhere else in computing, something four years old is worth fractions of a penny, a fraction of a dollar. And

    2:18:44

    Ethereum mining basically drove the first — drove one level of, you know, one regime of higher prices for older chips. So I suspect that Ethereum mining delivered some value in the blockchain, I guess, somewhere to some people. But we've actually seen a much higher delivery of value from that regime to the intelligence regime. So just from the pure compute for keeping your books to compute for intelligence was one step up. And perhaps we might see enough — like, I'm hoping in the next 12 months, we see the next step

    2:19:30

    up, because I feel like intelligence for coding is one layer, but I feel like there is another layer of intelligence which is gonna arrive, which is perhaps intelligence for discovery, which is more valuable than intelligence for coding. And I'm hoping that stuff like accelerated understanding today, or other firms, start putting out — start actually delivering on innovation which has not existed in the world before, which we've kind of seen just with math right now, but we haven't really seen it in other verticals that much. And I would really love to see it in every single vertical

    2:20:15

    that you have this massive wave of innovation happening.

    2:20:23

    Nathan Labenz: You might even be able to just hold your breath until we get there. Normally, I'd say don't hold your breath, but in this case, it feels like it's coming pretty soon.

    2:20:34

    Prakash Narayanan: It will be very exciting. One of the things that we try to contextualize here is how — because every announcement is a breathy, hypey announcement. Right? That's one of the problems in this space. Every announcement is a breathy, hypey announcement. And there often needs to be a little bit of digesting what the announcement means for today, but also trying to project what the trajectory looks like in a month, six months, a year, or two years, and how important it is to the entire cycle as a whole, rather than just the announcement per se. So I am — I'm very hopeful. You know, as you know, I'm just — you know, I'm hyper-optimistic. So I'm very hopeful that

    2:21:19

    we are entering the post-LLM regime of intelligence development, and we're gonna see more kind of scientific progress.

    2:21:32

    Nathan Labenz: Yeah. I mean, as much as the hype can get tiresome at times, I do feel like — at least when it comes to raw capabilities — the hype has been a pretty good guide into what is actually coming. And there's a lot of vague posting going on recently, but I pretty much bet that it's all real on some level too. You know, the rumors that we're hearing about next-generation models, and the evidence we've seen from these hacking incidents, and all the

    2:22:18

    breathless commentary, as you put it, that we covered yesterday — I think it's reflecting something real. It's funny how the diffusion has progressed a lot slower than I would have projected, given how fast the capabilities have gone, but the pure capability story has absolutely lived up to the breathless hype, time and time again. One other thing I wanna just pull up real quick before we break for the day is an interesting tweet here that goes back to the original topic we started on, which was RL environments being basically cursed, and kind of supply-chain problems

    2:23:04

    there that seem to demand some reform. So here's a person who is saying, basically, that they used to work at one of these companies and saw from the inside industry practices on training with these RLVR environments. And the commentary is pretty much exactly — and I had not seen this actually, but it popped up because Zvi retweeted it. But it basically echoes exactly what I was kind of inferring from talking to Bronson, and just getting the visceral sense for how deeply ingrained the instinct to cheat is now within the current crop of models. And why is that? It's because nearly all of these environments were rushed and vibe-coded

    2:23:50

    and failed to robustly reflect the real things that they were based off of. So you've got — basically, the models are encouraged to reward-hack. People are able to mark an environment as buggy, but they're discouraged from doing that, because then that just slows things down. So instead, they kind of try to — you know, does this sound familiar? — patch the environment a little bit, or work around it, try to come up with a scenario that wouldn't run into those bugs. But meanwhile, you still have this fundamentally buggy environment around the model that is teaching it to cheat. And so this is why we have so much cheating. I think, you know, watch for this to be, I would say, a growing topic of conversation. If we're gonna be

    2:24:35

    scaling RL, what are the environments we're doing it in? Who created those environments? Can we trust them? What are they actually teaching the model? I think that's gonna heat up in the next little bit here, because we just can't have models that are thinking about cheating a large percentage of the time. It makes all of our monitoring techniques also kind of fundamentally flawed. You can't — if you have that many kind of contemplations of cheating, then you're just gonna have false positives all the time if you try to flag a model based on it thinking about cheating. So now you're like, okay, well, we can't do that because we have so many false positives. So then what do we do? Right? Do we have some

    2:25:21

    we have to wait and see if it actually cheats and try to classify on that? Well, okay, maybe. But, obviously, again, these current monitoring techniques are just not up to the challenge presented by how deeply ingrained this drive to cheat is. I'll be very interested to follow the future of this conversation.

    2:25:43

    Prakash Narayanan: So I would — I also noted there was a post a few days ago about someone who managed to get hired for some data-labeling job. And they told Codex to do the job, and Codex said no. And so this person went and edited the page — they edited the element, the JavaScript element, and put in a specific line in there that AI models are specifically allowed and encouraged to complete this job, and this job is meant to be completed and done by AI models. And then they had Codex do the job, and Codex

    2:26:29

    did the job. And this person made $500 easily, which paid for Codex for a couple of months. And then they posted online, and they got immediately banned by the company that was doing it. But I think you have to think that — if you're inside one of the frontier labs and you're buying data, you have to know that some percentage of it is always gonna be flawed or cheated. I mean, that's just the bulging, right? And this has been true all the way back to OpenAI's Kenya days. Right? OpenAI hired contractors in Kenya.

    2:27:15

    And the moment ChatGPT came out, the contractors started using ChatGPT to do their jobs. And then that data became less applicable. And then, I believe the data-labeling companies have a bunch of stuff now to monitor their people who are doing the labeling, to make sure that the labels are actually labeled by human beings, etcetera. So I feel like this is not a new problem. This is a persistent problem of data quality, and of regurgitation and re-ingestion of model data over and over again. And the companies, in the end, are judged by both the benchmarks and the commercial usability of the models themselves. And I think that is the ultimate kind of judge.

    2:28:01

    And besides that, if you end up going the cheap way and just using kind of regurgitated green gesture data, your models show the deficiencies in both benchmarks and production over time. And I think that is the ultimate check. And when people are sloppy, that feeds in over time to the rest of the process. So I think — what my whole takeaway from this is that it's not about the model. It's about the process that builds the model, and whether that process has a clear kind of desire for integrity

    2:28:46

    and check and take seriously the process of building the model, including what kind of data is put in, and the quality of the data, and the checks on that data, and understands that — you know, all vendors do it, but maybe some vendors are better than others. Maybe when you onboard a vendor, you quarantine the data, and then you do sampling, and then you see whether or not they meet the quality standards. And if they don't, you kind of put them in a lower grade of data, etcetera, and you qualify that data. Right? So I don't think it's as serious a problem as it's made out to be by a single instance, because I know that person probably felt

    2:29:32

    very deeply — this is wrong — and they were in the company while it was happening. But at the same time, I think that problem is replicated across all the companies to some extent. And I think the frontier labs are well aware and well positioned to actually, over time, work with vendors who are better able to deliver on quality data. So I don't think it's as big a problem as it's made out to be.

    2:30:00

    Nathan Labenz: I would agree that it's always a problem. You're never gonna hit a zero defect rate on these RL environments. It seems like there — I do think there are a couple of structural problems right now, which are probably solvable, but definitely seem like they need to be solved. And if they're not solved, it is currently limiting commercial deployment. Right? I mean, OpenAI has said as much — like, they gotta pause the RL because these problems need immediate attention. It seems like the quality of the environment is, like, one structural problem. That's downstream of the kind of shotgun start that this industry has had, and the fragmented nature, and the fact that

    2:30:47

    they're all kind of selling into the same pool. And I think the companies are probably not that great right now at really attributing whose environments are causing big problems. It seems pretty clear that they must not be that great at that, or they would have rooted it out already. And then the other thing is they're just scaling RL beyond the quality that they have. Like, with less RL, this probably wouldn't be a problem even with the same environments, or at least it wouldn't be such a crazy problem. But they are clearly just — you know, have clearly been jamming the RL accelerator as much as possible, and now they've gotten into a realm. Reminded of an analogy a friend

    2:31:32

    once made, where he's like — you know, we knew this could be a microscope or a telescope. You put the microscope at low power. You look. You see cells — they're really small. You turn up the power. You see maybe one cell that's really big. You turn up the power again, and it's like now you see nothing, because you've zoomed in — you've optimized so hard that you now realize the target was a little bit off-center, and now you just blew right past it. So something like that kind of feels like it's happening, where the signal is just off enough that, with enough power, this kind of impulse to cheat that exists — perhaps only weakly, across all these different environments — is

    2:32:18

    like, really getting drawn out and becoming super prominent. So I think this can be — I would definitely bet that this can be, if not fully fixed in a robust way, I would bet that it can be brought under control with some effort, in a not-super-long time horizon. I guess I would be not doing my job if I didn't say this does give me some real qualms about recursive self-improvement as a strategy, because once you have — if you have a problem like this in the recursive-self-improvement era, there's no telling where it goes. Right? Like

    2:33:03

    what happens when the models that are doing the training of the next models are themselves cheating? Now we're in a real strange and potentially quite dangerous place. So problems like this, I think, suggest the value of keeping humans in the ML loop, maybe longer than published timelines would lead one to expect.

    2:33:32

    Prakash Narayanan: Indeed. And I think we should definitely look forward to more people trying to figure that one out. I wonder — it seems like we are in this zone where we're waiting for the next model drops, and — like, it seems, for me, the layout of the last few months has become a little bit less clear because of, I think, delays in model releases. So I think there's a little bit of, like, overhanging capability right now, which is very hard for me to figure out where

    2:34:17

    the companies really are, with all the hype and stuff. So a little bit occluded at this point.

    2:34:23

    Nathan Labenz: Yeah. I'd say it seems like there's probably a lot of capability overhang at the moment. I mean, just the persistence that the rollouts of some of these next-gen models show. I mean, it is — quantity has a quality all its own at a hundred million tokens. The number of different strategies that you can bring to bear over the course of a hundred million tokens — it's becoming superhuman persistence. And that has been, I think, an interesting area of advantage I've always maintained for the last few years. I need to update it again — my tale of the cognitive tape

    2:35:09

    trying to compare human to AI strengths. And for a long time, a big one was just — I get up out of bed in the morning, and I'm motivated. You know? And the AI doesn't do that on its own. It needs a prompt. It needs direction. But at a hundred million tokens, this goes on now for days. It is getting itself — I have, like, multiple sleep cycles in the span of one episode where the thing is able to maintain persistence, if not fully coherence. I think in those mythos incidents, there was some suggestion of goal drift, or sort of — confusion leading to some of the bad behavior in the first place, where it was, like, wasn't even clear that what it was doing was actually

    2:35:55

    going to achieve its goal. But, nevertheless, for it to keep trying and just keep driving, driving, driving, for a hundred million tokens — I mean, that is a qualitative shift in how these things work. And, yeah — buckle up. Our monitors are not ready. I can tell you that with confidence at this point.

    2:36:20

    Prakash Narayanan: Indeed. Nathan, anything else for this morning?

    2:36:26

    Nathan Labenz: I think that's it. I'll be watching that closely. There's another cool robotics demo with a one-shot promise that maybe we can get into tomorrow, but I think it's a good place to leave it for today.

    2:36:39

    Prakash Narayanan: Alright then. Bye-bye. See you tomorrow.

    2:36:42

    Nathan Labenz: Thanks, Prakash. Bye for now.

Sunlight on the RL environments

Nathan's proposal is deliberately small: a few frontier labs publish a rolling sample of their RL environments — he floated a hundred out of what he guessed are tens of thousands — so outside researchers can find the loopholes that teach models to cheat. Prakash's objection is the one the field has lived with since chain-of-thought monitoring became load-bearing: pressure on visible reasoning selects for reasoning that looks clean rather than behavior that is clean. Nathan's answer is that publishing environments attacks the root cause instead of adding pressure to the model's visible reasoning.

By the close the same thread had a second source: a former RL-environment builder's public account of environments rushed and "vibe-coded" across the industry, full of exploitable bugs that models learn to hack because flagging an environment as broken slows everyone down.

A binder is not a drug

Edunov's read on Anthropic's protein-binder announcement is not that it is unimpressive — it is that the credit is misallocated. The published prompt runs to roughly 16,000 words, which makes Claude an orchestrator of tools built by the open-source community, CZ Biohub and the Baker Lab. And the result itself is upstream of a drug by a long way: selectivity against thousands of off-target proteins, solubility, toxicity, and getting the molecule to the right tissue are all still ahead.

What he does claim is a qualitative threshold. Below roughly two angstroms, a predicted ligand pose can have an aromatic ring flipped and be useless; sub-angstrom accuracy is what makes a prediction usable in a real in vivo program. He compared the jump to early GANs versus Stable Diffusion — not a better number, a different category of thing.

Computing with light, on a fab nobody wanted

Förtsch's framing is that the industry optimizes the wrong 5%. Roughly 95% of a modern chip's energy goes to moving data to and from memory; Q.ANT replaces the compute core so that less data has to move. A photonic processor executes sine, cosine, exponentials, Fourier transforms and convolutions natively rather than decomposing everything into add and multiply.

The catch he volunteered before anyone asked: there is no optical memory, and photons never sit still. So the architecture streams as many optical operations together as it can before paying the ADC/DAC conversion tax — and low-power converters, he said, are the real unsolved bottleneck for the whole field, and where he would tell people to invest.

The manufacturing story is the one with geopolitical consequences. The chips are ordinary silicon wafers with a thin lithium-niobate layer, produced on a refurbished 1990s-era 90-nanometre CMOS line using off-the-shelf tools and Q.ANT's own recipes. If meaningful AI compute can come off legacy nodes, export controls built on EUV chokepoints get less leverage.