Why Coding Is Not Solved: Limits of LLMs in Production Software
TL;DR – Coding Is Not Solved
LLM‑generated code can speed up creation, but the majority of software cost lies in non‑functional requirements (NFR) such as reliability, security, and scalability. Production‑grade systems still need human engineers to understand, verify, and be accountable for the code; AI cannot be held responsible.
1. The Core Argument: Code Generation ≠ Solved Engineering
- Creation is cheap, maintenance is expensive. The author notes that while LLMs lower the cost of writing code, the bulk of operational expense comes from maintaining and operating software at scale.
- Non‑functional requirements dominate. NFRs—security, reliability, performance, compliance—are rarely addressed by raw LLM output and require deep domain expertise.
- Logic and volume limitations. LLMs struggle with large contexts, logical consistency, and deterministic behavior. Their stochastic nature means they can produce syntactically correct code that still contains subtle bugs.
- Accountability gap. AI cannot be punished, fined, or held liable. Legal and organizational responsibility remains with the humans who ship the software.
"You cannot be responsible for what you can’t control either. That understanding is key to reasoning about system behavior and fixing it when the AI inevitably fails." – Alex Ewerlöf
2. Real‑World Risks Highlighted by the Author
| Risk Area | Why LLMs Fall Short |
|---|---|
| Healthcare, Finance, Aviation, Defense | Errors can cost lives or trigger legal penalties; AI lacks the rigorous verification pipelines required. |
| Security | LLMs may introduce hidden vulnerabilities; they cannot be audited like human‑written code. |
| Scalability | Performance regressions often emerge only under load; LLMs cannot anticipate all edge‑case interactions. |
| Legal Liability | No mechanism exists to hold an LLM accountable; the organization bears the blame. |
3. How LLMs Currently Work in Coding
- Prompt → Model – The developer provides a natural‑language spec.
- Generation → Harness – The model produces code, which is fed into a harness that runs compilers, linters, and test suites.
- Feedback Loop – Errors are sent back to the model (often via chain‑of‑thought or tool‑calling) until the code passes basic checks.
- Human Review – Ideally a developer inspects the diff, validates NFRs, and signs off.
The author stresses that step 4 is non‑negotiable for high‑risk domains.
4. Community Reactions on Hacker News
4.1 Points of Agreement
- @efficax argues that LLMs can augment reliability by automating exhaustive testing and fuzzing, but acknowledges that the author’s view seems based on limited hands‑on experience.
- @lordnacho distinguishes “coding in the small” (solved) from “coding in the large” (unsolved), echoing the need for human judgment on architecture and trade‑offs.
- @mstaoru sees the emerging baseline: engineers now spend most of their time guiding AI, not typing code, which aligns with the author’s claim that the role is shifting rather than disappearing.
4.2 Points of Dissent
- @brainless predicts a radical reinvention of programming, suggesting that LLMs will eventually replace current languages and frameworks.
- @manny_rat reports that in his day‑to‑day work, LLM‑generated code already requires minimal review and yields a 10× productivity boost, questioning the prevalence of “high‑risk” software.
- @jpadkins claims that for many internal tools, the agentic output meets all correctness standards, making code review optional.
- @bluegatty counters the “logic” critique, stating that LLMs are effectively trained on compiler feedback, making them good at producing compiler‑perfect code.
4.3 Nuanced Observations
- AI Overdose – Several commenters (e.g., @askonomm, @mywittyname) warn that over‑reliance on AI can erode developer skill and lead to massive, hard‑to‑review PRs.
- Accountability in Law – @hibikir points out that legal systems already hold organizations accountable for AI‑driven harms, contradicting the author’s claim that AI cannot be held responsible.
- Economic Perspective – @MatrixMan notes that the cost savings from AI‑generated custom software may outweigh quality concerns for many small teams.
5. Practical Takeaways for Engineers and Leaders
- Treat LLMs as assistants, not replacements. Use them to scaffold code, generate boilerplate, or explore alternatives, but always verify NFRs.
- Invest in harnesses and automated testing. A robust feedback loop (compiler → test suite → model) is the only way to catch deterministic failures.
- Maintain clear ownership. The three‑pillar model—knowledge, mandate, accountability—must stay with a human engineer, especially for regulated domains.
- Watch for AI‑overdose. Limit token‑budgeted generation, avoid massive PRs, and keep personal coding practice alive to prevent skill decay.
- Align incentives. Companies should price AI‑generated services realistically; overcharging for “human‑level” quality while using cheap AI will erode trust.
6. Future Outlook
- Model improvements (larger context windows, better reasoning) will reduce but not eliminate logical gaps.
- Domain‑specific agents (e.g., security‑focused LLMs) may bridge some NFR gaps, yet they will still need human oversight.
- Regulatory pressure is likely to increase, mandating audit trails and accountability for AI‑generated code in high‑risk sectors.
- Skill evolution – Engineers will increasingly become prompt engineers, AI‑orchestrators, and quality guardians rather than pure coders.
7. Conclusion
While LLMs have dramatically lowered the barrier to code creation, the essential challenges of software engineering—maintaining reliability, ensuring security, and bearing accountability—remain unsolved. The HN discussion reflects a split view: some see a near‑term shift in developer roles, others argue that the author underestimates current model capabilities. Regardless of stance, the consensus is clear: human expertise is still indispensable for production‑grade software.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch