Greg Kroah-Hartman on LLM-Driven Security Vulnerability Discovery
Linux kernel maintainer Greg Kroah-Hartman (GKH) recently provided a grounded analysis of the role of Large Language Models (LLMs) in software security, specifically targeting the hype surrounding automated vulnerability discovery. His primary takeaway is that while LLMs can generate a high volume of reports, they remain "dumb fuzzy pattern matchers" that often fail to produce actionable or significant security findings.
The Mythos Case Study: Overhyped Vulnerability Discovery
Greg Kroah-Hartman highlights a specific instance involving Mythos, an LLM-based tool that claimed to have discovered 79 vulnerabilities in the Linux kernel. According to GKH, the actual utility of these findings was minimal, as the majority of the reports were either non-bugs or trivial fixes.
Breakdown of the 79 Reported Vulnerabilities
Of the 79 vulnerabilities reported by Mythos, GKH provided the following breakdown:
- Non-bugs and False Positives: 24 reports had no detail at all (simply stating "something crashed"), 14 were not bugs at all, and 3 contained totally made-up data.
- Existing Fixes: 15 were already fixed in the latest release, and 11 were fixed by others.
- Actionable but Trivial: Only 20 reports required actual fixes. Of these, 7 assumed a malicious filesystem image, 6 were SCTP networking issues for untrusted devices, 2 were IPv6 minor network issues, and 2 assumed the ability to inject malicious network packets into the middle of the stack.
GKH concluded that this effort resulted in only "10 'real' bugfixes," and the entire process of addressing these reports took approximately one hour of kernel development work.
Technical Mechanism: Pattern Matching vs. True Analysis
According to GKH, the perceived "intelligence" of the LLM in finding bugs is actually a form of advanced static code analysis based on pattern matching. Specifically, Mythos functioned by pattern matching previous decades of kernel developer patches and applying those same mechanisms to other parts of the code to see if similar patterns existed.
This approach does not represent a new paradigm of bug discovery but rather a reproduction of existing static analysis techniques. The result is a high volume of reports that often lack the depth required for a professional security researcher to maintain the same level of quality as a human expert.
The Asymmetry of Report Generation and Verification
One of the critical challenges introduced by LLMs in security is the asymmetry between the cost of generating a report and the cost of verifying it. While LLMs can generate thousands of reports nearly for free, the cost of human verification remains unchanged.
"The asymmetry is the killer: reports are nearly free to generate now, but the verification cost stayed exactly the same."
This creates a burden on maintainers who must spend significant time triaging a high volume of low-quality reports that may be contain a high percentage of false positives.
Legal and Ethical Concerns in LLM Training
GKH also touched upon the legalities of the training data used for LLMs. He cited LG Research, noting that some common corpora contain only 20% legally allowed-to-use data. This suggests that a significant portion of the training data for LLMs is sourced from copyrighted or restricted materials without proper authorization, a legal territory that the Linux project continues to navigate as courts decide the outcome of these disputes.
Best Practices for Developers using LLMs
For developers integrating LLMs into their workflow, GKH emphasizes two main rules:
- Prioritize Code over Comments: He advises developers to "ignore the comments, look at the code," as explanations provided by LLMs can be confusing or incorrect, while the code itself is the only reliable source of truth.
- Data Privacy: He warns against uploading any non-public information, such as credentials or proprietary research, to consumer-grade LLMs, as this data is often used for training and may be leaked to other users in future outputs.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch