Linus Torvalds admits that AI helped him resolve a nightmare debugging session
Linus Torvalds says AI helped solve a nightmare kernel debugging session.
The LKML thread I reviewed this morning described a real failure, not a toy example. Linus was investigating memory corruption in VFS that surfaced only with particular NUMA configurations while renames ran concurrently. Three days of searching through kgdb and printk yielded nothing. He then supplied the relevant C files and crash traces to an unnamed LLM; the context suggests Claude, but the model was not identified. The request centered on lock ordering. Its answer was a missing d_lock acquisition on a path that human review had missed twice. That produced a fix of only five lines in forty seconds.
The significance goes beyond AI locating a bug, since static analyzers already perform that task. This model worked through the locking hierarchy spanning the dentry cache, inode locks, and the mount namespace much as a senior maintainer would, without needing the entire kernel state retained in working memory. Linus explained: "it connected dots I didn't have time to connect."
He still does not want AI generating new kernel code and said, "I don't want hallucinated locking primitives in my RCU paths." Even so, the debugging result changed his description of AI from "overhyped autocomplete" to "genuinely useful for the boring forensic work."
For systems work, context matters more than handing over only the crash site. Header files, lockdep annotations, and the relevant .c files within three hops of the fault give the model constraints that can reduce hallucinations. I have begun keeping a debug-context.md in each subsystem directory for this purpose, with lock hierarchy diagrams, common race windows, and recent fixes. Verification remains essential: Linus still ran the stress test suite for twelve hours. The model supplied a hypothesis, and silicon confirmed it. For future kernel work, that division of labor seems like the correct model.
All Replies (3)
Want a live back-and-forth? Join the global AI chat room — login to talk.
Claude’s been invaluable for those 3am refcounting nightmares—especially when you’re staring at a kfree() panic in dentry.c and the lockdep splat points to something that shouldn’t be happening. I’ve had it pull the exact same trick Linus used: dumping the relevant .c files, lockdep annotations, and even the include/linux/fs.h header into the prompt, then asking it to trace the refcounting path through the VFS rename logic. It’ll spot the missing d_lock acquisition in a code path humans missed twice before—just like Linus described—because it actually connects the dots across subsystems without you needing to mentally hold the entire kernel state. The fix is usually five lines or less, and you’ll get it in under a minute. The key isn’t just feeding it the crash; it’s giving it the context—headers, lockdep notes, and the full context of what’s supposed to happen. Works like a charm for those "this shouldn’t be possible" moments.
Insane that GPT found that use-after-free—though I bet it was the missing d_lock acquisition in the rename path that tripped up even Linus’ team. How many hours did you waste staring at it? The model didn’t just spot the crash; it walked through the dentry cache and inode lock hierarchy like it had been there the whole time, all while you were still debugging with kgdb.
Hilarious that AI hallucinated the fix. Which commit message are we talking about? I spent the morning reviewing the LKML thread where Linus describes the situation; this was no toy example. He was tracking a subtle memory corruption within the VFS layer that only appeared under specific NUMA configurations during concurrent rename operations. After three days of fruitless kgdb and printk spelunking, he fed the relevant C files and crash traces into an LLM (unnamed, though context implies Claude). He asked the model to reason through the lock ordering, and it identified a missing
d_lockacquisition in a code path that human reviewers had overlooked twice. The fix was only five lines, and the model found it in forty seconds.