2026-09-29 · updateAn LLM writing code is as good a software engineer as the test suite it runs against. SBSE found this in 2009.
- GenProg (Weimer, Nguyen, Le Goues, Forrest, ICSE 2009) repaired real bugs by mutating a program and keeping the variants that passed the tests.
- Qi, Long, Achour and Rinard (ISSTA 2015) showed most accepted patches passed by deleting the functionality the failing test exercised.
- Smith, Barr, Le Goues and Brun (FSE 2015) named it overfitting to the test suite and measured it getting worse as the suite got thinner.
- A coding agent is the same loop with a better mutation operator. It proposes, the tests judge, it keeps what passes. The judge did not change.
- The amount an agent can safely do to a codebase is set by the suite. A team that wants more from the agent gets it by writing tests.
- https://doi.org/10.1109/ICSE.2009.5070536
- https://doi.org/10.1145/2771783.2771791
- https://doi.org/10.1145/2786805.2786825
- https://doi.org/10.1016/S0950-5849(01)00189-6
- https://doi.org/10.1109/TEVC.2017.2693219
- #ai-coding
- #genetic-improvement
- #sbse
- #testing