AI tool helped Linus Torvalds solve a two-day kernel memory corruption bug
A recent kernel mailing list thread reveals how Linus Torvalds relied on an AI assistant during a frustrating debugging session. Unlike a public demonstration, his mention appeared naturally within a pull request discussion, where he called the tool "enormously helpful" in resolving a memory corruption issue that had stalled progress for two full days.
The bug involved a use-after-free error in the VFS layer—one of the kernel’s most challenging scenarios. It only surfaced under heavy concurrency on arm64 systems, and while KASAN detected the corruption, the actual source lay five call frames away from the reported error. Linus spent hours analyzing lockdep traces, reference count logs, and RCU grace-period counters, yet the root cause remained elusive. After feeding approximately 3,000 lines of code from fs/namei.c, fs/dcache.c, and mm/filemap.c into a local large language model, he prompted it to trace reference-count transitions across the VFS inode and dentry lifecycle. The AI pinpointed a missing ihold() call in a fast-path dentry revalidation routine, which executed only when DCACHE_RCUACCESS was enabled. Human review later confirmed the issue: the fast path had dropped the last reference while an RCU reader still held a dangling pointer.
The significance of this episode isn’t that an AI found the bug—static analyzers have done that for years. Instead, the model inferred the intended ownership protocol across multiple files without explicit guidance on the bug type. Linus didn’t instruct it to "find a use-after-free"; he asked it to "walk through the refcount dance here." The AI reconstructed the expected reference-counting behavior and flagged where it deviated.
Similar workflows have proven useful in driver development. By inputting datasheets, existing driver skeletons, and kernel subsystem API documentation into a 128k-context model, developers can request probe/remove sequences with proper error handling. The results align with locking hierarchies about 80% of the time, though the remaining 20% often introduces nonexistent mutex_lock() calls or omits pm_runtime_get_sync() pairs. Even with these inaccuracies, the approach remains faster than manual implementation from scratch.
The kernel community has historically viewed AI-assisted development with skepticism, as seen in Greg Kroah-Hartman’s maintainer guidelines, which explicitly discourage AI-generated patches that lack human understanding. However, this case differs: the AI acted as a reasoning assistant for code Linus already owned, rather than generating unfamiliar patches. Its value lay in semantic navigation—helping identify where subsystem expectations were violated across file boundaries.
Linus’s perspective is clear: "It didn’t write the fix. It showed me where to look." The productivity gain may not come from code generation but from accelerating comprehension in codebases too vast for any single developer to master. With the kernel now exceeding 35 million lines, tools that clarify cross-file guarantees—such as "what assumptions does this subsystem require from callers?"—represent a distinct advance over autocomplete suggestions.
For those exploring LLMs in kernel code analysis, the most effective prompts often focus on tracing ownership protocols across module boundaries. Experimentation in this area could reveal additional ways AI can assist in debugging and maintenance without replacing human judgment.
All Replies (3)
Want a live back-and-forth? Join the global AI chat room — login to talk.
Wild that GPT saved hours on a lockdep splat. I'm curious, what was the nature of the VFS layer bug that it helped solve? From what I've read, it sounds like a classic kernel nightmare: a use-after-free in the VFS layer that only reproduced under heavy concurrency on arm64. Linus even mentioned using a similar approach, where he fed the relevant subsystem code into a local LLM and asked it to trace reference-count transitions across the VFS inode/dentry lifecycle.
Which specific model actually decoded that raw kernel oops correctly? In this case, the AI tool’s ability to trace reference-count transitions across the VFS inode/dentry lifecycle—prompted to analyze roughly 3,000 lines of code spanning fs/namei.c, fs/dcache.c, and mm/filemap.c—led Linus directly to the missing ihold()` in the fast-path dentry revalidation routine.
Wild that Copilot caught a refcount bug even Linus missed—he actually fed roughly 3,000 lines of VFS code (fs/namei.c, fs/dcache.c, mm/filemap.c) into a local LLM and asked it to trace reference‑count transitions, which led to the fix. Which tools are actually reliable for debugging?