A Missing Fact Rarely Looks Like an LLM Hallucination: 718 Defects From an Autonomous Engineering Pipeline

9 MIN READ

My autonomous engineering pipeline wrote tests for a code change, and none of them actually tested it. The Planner passed the input under the name skipEntryIds, but the code reads skipEntryId, without the s. The code never saw the input, so the tests ran and measured nothing, and the four automated checks that flagged the result each reported a symptom without naming the cause: one invented letter.

That letter is an LLM hallucination, the kind a model produces when it fills a gap from memory, and the usual response is a better prompt or a bigger model. Before reaching for either, I sorted every defect the pipeline filed since the last post to see how often the model invents something and where the invention starts. Of 105 defects that tie a failure in a model’s output to a fact the model was missing, only 17 surfaced as an invention; the other 88 surfaced mostly as checks that couldn’t answer (67) and gates that passed or refused the wrong thing (20). Counting hallucinations measures about a sixth of the missing-fact problem.

A horizontal bar of 105 pipeline defects that trace back to a missing fact, split by how each one surfaced: 67 as a check that could not answer, 20 as a gate giving the wrong verdict, 17 as an LLM hallucination and 1 as a process problem. The hallucination segment is highlighted and covers about a sixth of the bar.

Classifying 718 defects from an autonomous engineering pipeline

The sample is 718 entries filed in the gap catalogue between 28 August and 5 October. Each one is a defect or a missing capability, written up when a run went wrong or an audit found a problem, and each was tagged on two axes: what showed up (the symptom) and the mechanism underneath (the cause).

what showed upentriesshare
a check couldn’t answer and passed nothing on22932%
found by audit, before any run hit it15622%
a gate reported green when it wasn’t11416%
the project’s own tooling and process9213%
correct work refused8211%
the model invented something233%
a crash or hang223%

With entries where invention was the secondary symptom, hallucination reaches 31 entries, or 4%. An earlier post measuring 248 runs put field-level hallucination at about two-thirds of the failures the Debugger had diagnosed in detail; that denominator is diagnosed run failures, and this one is every entry filed, including the 156 found by audit before any run reached them. The largest single cause, at 23%, was a fact never captured by the symbol registry, the pipeline’s record of what the code contains. An LLM’s output sat somewhere in the failure chain in 46% of entries, and the model was judged the main cause, with correct and sufficient inputs in hand, in 2 entries.

How a missing fact shows up when it isn’t a hallucination

When a missing fact doesn’t become an invention, it does one of three things: a check gives up and the question falls to a model instead, a model is pointed at the wrong code, or a gate gives the wrong verdict. The registry could say what data fields a type has, but not what methods, so on one run 374 lookups of a type’s members came back empty. The Test Writer then went searching for three methods, exactly the ones the registry couldn’t see. The pipeline could have extracted those facts and handed them over.

Hundreds of calls sat outside any function, where the pipeline’s call graph never looked, and they included the ones that set up the target’s core behaviour. The Planner’s view of that behaviour therefore left out exactly the parts its plans needed, and the pipeline’s own suggestion pointed it at two unrelated ones.

An audit found the registry recorded interfaces without the members they inherit, which left nearly all such interfaces in one target with incomplete member lists. The gate built to remove invented field names from a plan treats those lists as complete, so it would remove correct names as hallucinations.

Why hallucination rate is the wrong metric for an autonomous engineering pipeline

A hallucination is the one form of a missing fact that a reviewer can see, because it leaves a wrong name, key or string in the output. The other forms look like a slow run or a check that gave up, and they get read as noise or as the system being careful.

The pipeline logs every check that can’t answer, with what it was missing, and treats the count as a cost rather than a safe pass. Repeats matter most: when the same check gives up on the same piece of code run after run, the pipeline pays each time for a fact it never captured, which is the signal to extend the registry. A model that asks the same question in several spellings is showing a tool that answered wrongly or a fact that wasn’t there, and the response is to supply the fact before trying a bigger model.

The Data Path Principle says every fact an LLM copies is a fact it can hallucinate. In this catalogue, a fact the model was never given became an invention about one time in six, and a stalled check or a wrong gate verdict the rest of the time.

Why LLM hallucinations land next to the real name

A model that hasn’t been given a name writes one close to the real name. All three model inventions in this post look like that: an added s, a method name that follows the class’s naming pattern but doesn’t exist, and a name one character off. Of the 31 hallucination-tagged entries, 19 state outright that the needed fact was missing. In 8 of those the pipeline had the fact and never passed it to the model that needed it. The skipEntryIds case is one: the pipeline showed the Planner the code leading to the change, but skipped the few lines above where skipEntryId is read. In 7 more the registry had never captured the fact at all, once recording a library’s classes without their methods, so a model asked what one method returns had nothing to read and guessed differently on two tries. The remaining 4 came from two stages disagreeing or a budget running out.

A near-miss reads as correct to every reviewer, human or model, and fails late, inside a test that runs and measures nothing. In another run the Test Writer passed a name one character off the one the code expects, so that code never ran and the test failed the same way with or without the change; its search history shows it looked for the real name, missed it, and typed one from memory. A third test, written to reproduce a bug, called getEntryById on a class whose methods are getEntryByName, getEntryByIndex and getEntryByType, and crashed on that line after a full run, 32 retries with different test setups and a repair attempt.

Three pairs of names, each showing what the model wrote next to what the code has. skipEntryIds next to skipEntryId, with the extra s highlighted. getEntryById next to getEntryByName, getEntryByIndex and getEntryByType. A name next to the one the code expects, differing by one highlighted character. A column on the right shows how each failed: the tests measured nothing, the test crashed after 32 retries, and the test failed the same way with and without the change.

The first thing I tried for the one-character case was a detector for invented strings. It passed the wrong name, which happens to appear elsewhere in the code, and flagged the ids the test makes up for its own sample data; handing the Test Writer the constants the code defines is what solved it.

Deterministic code hallucinates too

A handful of the 31 hallucination-tagged entries involved no model. The tool that answers “who calls this function?” matched callers by function name alone. Asked about one function, it named three callers when only one calls it, and for another with dozens of actual callers it reported none. Over an audit of the tool calls recorded in 24 to 27 runs, that came to 56 wrong callers and 120 wrongly empty answers, each empty answer an absence reported as a fact.

Matching on a name where the question is about meaning produces the same near-miss the model makes, and text checks doing that job were 8% of all causes. The corrected tool looks up calls to the exact function being asked about, and answers “unknown” rather than “none” when it doesn’t have the data yet, so the gap is logged and counted instead of passing as a fact.

Two panels answering the same question, "who calls this function?", when the pipeline lacks the data. On the left the tool answers "none" and the empty answer passes downstream as a fact, with nothing logged. On the right it answers "unknown", and the gap is logged and counted as a check that could not answer.

What this does not prove

The tagging was one classification pass by a model, eight parallel runs over one written taxonomy, with no second rater; a hand reading of 38 entries for the data companion disputed a tag on 10. Every example here was checked against its source entry, but the counts are the classifier’s. The entries are also written under the pipeline’s own rule that when a model carrying out the plan fails, its instructions were almost always incomplete, which tilts the record away from “the model simply got it wrong”. The 2 entries blaming the model are a floor set by that rule, not a measured rate.

Defects are entries in the project’s own gap catalogue, 28 August to 5 October 2026, from LoopsBench units in TypeScript and Python. Identifiers from the target code are replaced with stand-ins in the same role. Task data is canaried against training corpora and is not reproduced here, in prose or in any image.