Changing Where the Work Comes From Exposed 145 Defects in an AI Coding Pipeline
The first benchmark unit reported completion in 114 seconds, every gate green. Its hidden test refuted it. Twelve days of external targets filed 145 defects.
The first benchmark unit reported completion in 114 seconds, every gate green. Its hidden test refuted it. Twelve days of external targets filed 145 defects.
Hand-mining GitHub issues was the one step in an autonomous engineering pipeline that did not scale. LoopsBench replaced it with 112 tasks and a referee.