This is the data behind A Missing Fact Rarely Looks Like an LLM Hallucination. That post makes one argument from it: in an autonomous engineering pipeline, a fact the model was never given surfaces as an LLM hallucination about one time in six, and as a check that can’t answer or a gate with the wrong verdict the rest of the time. This page holds the method, the taxonomy, every count, the cause-by-symptom matrix and each entry where the model invented something, for anyone who wants to check the argument or read the data a different way.
How the 718 defects were classified
The source is the pipeline’s gap catalogue, where every defect or missing capability is written up when a run goes wrong or an audit finds a problem. The window is 28 August to 5 October 2026, the period since the previous defect count, which covers 922 commits. In that window 952 entry numbers were assigned and 943 entries were written; the other 9 numbers never reached any file.
The entries were split across eight parallel runs of a model, all working from the same written taxonomy. Each run saw a shortened copy of each entry, the title in full and the start of the rest, and gave it one primary tag and an optional secondary tag on two axes: symptom and cause. The output was checked mechanically so that every entry was tagged exactly once and with a valid value.
225 entries were removed after classification. Their primary symptom was “wasted work” (burned attempts, retries, slowness on work that ended correct), and on reading them the label had become a catch-all that described how much an entry cost rather than what went wrong. The remaining 718 are the base for every number here and in the main post. The same label is dropped where it appeared as a secondary symptom on 67 of the 718; those entries stay, without that tag.
The source entries behind the main post’s examples, and the 38 entries in the two lists below, were read in full by hand and the summaries written from them; no other entry was checked by hand. The tags are the classifier’s and are left unchanged so every count can be reproduced; where the hand reading disagreed with a primary tag, the entry says so. That happened on 10 of the 38, which is the only agreement check this data has.
The taxonomy: symptoms, causes and flags
The symptom axis records what showed up:
| symptom | meaning |
|---|---|
| hallucination | the LLM wrote something false: a symbol, member, path, key or term the code does not have |
| check couldn’t answer | a check could not reach an answer, or a fact a stage needed was absent |
| false green | a pass or success that wasn’t earned |
| false red | correct work refused or blocked |
| found by audit | no run hit it; it was found by reading or a sweep |
| crash | an exception, hang, or tool failure |
| process | the project’s own development process and tooling |
The cause axis records the mechanism underneath:
| cause | meaning |
|---|---|
| missing fact | the registry, the pipeline’s record of what the code contains, never captured the fact |
| wrong fact | extraction produced an incorrect fact |
| not delivered | the fact existed but never reached the stage that needed it, or was never read |
| handoff seam | two stages disagree about what a value means, or one fact has two representations |
| weak oracle | a test or check measures the wrong thing |
| text check for meaning | a string or pattern comparison doing a job that depends on meaning |
| guess or fallback | a first match, a cheaper fallback, or a default substituted for an unknown |
| arbitrary bound | a cap, truncation, or budget cut a search short |
| advisory only | a check warns but does not enforce |
| stale state | the pipeline acted on data that was out of date |
| target assumption | the pipeline assumed something about the target project that was not true |
| performance | too slow or too costly |
| model error | the LLM had correct, sufficient inputs and still got it wrong |
| other | none of the above |
“LLM in chain”, the first of two flags beside the tags, means an LLM’s output was part of the failure chain, whatever the cause. “Upstream fact link” means the entry explicitly ties a problem in a model’s output to a fact that was missing, wrong or not delivered; the classifier was told not to infer the link.
Defects by symptom
| symptom | primary | share | primary or secondary |
|---|---|---|---|
| check couldn’t answer | 229 | 31.9% | 243 |
| found by audit | 156 | 21.7% | 166 |
| false green | 114 | 15.9% | 132 |
| process | 92 | 12.8% | 95 |
| false red | 82 | 11.4% | 87 |
| hallucination | 23 | 3.2% | 31 |
| crash | 22 | 3.1% | 26 |
62 entries carry a secondary symptom once the wasted-work label is dropped. Counting either tag, LLM hallucination appears on 31 entries, 4.3% of the 718, and stays the second-smallest symptom.
Defects by cause
| cause | primary | share | primary or secondary |
|---|---|---|---|
| missing fact | 165 | 23.0% | 179 |
| handoff seam | 112 | 15.6% | 134 |
| weak oracle | 72 | 10.0% | 93 |
| other | 70 | 9.7% | 71 |
| wrong fact | 59 | 8.2% | 83 |
| text check for meaning | 57 | 7.9% | 79 |
| guess or fallback | 43 | 6.0% | 53 |
| arbitrary bound | 35 | 4.9% | 36 |
| not delivered | 28 | 3.9% | 40 |
| stale state | 25 | 3.5% | 27 |
| advisory only | 23 | 3.2% | 33 |
| target assumption | 15 | 2.1% | 24 |
| performance | 12 | 1.7% | 15 |
| model error | 2 | 0.3% | 8 |
157 entries carry a secondary cause. The four fact-related causes (missing, wrong, not delivered, and text checks standing in for meaning) account for 309 primary tags, 43% of the 718. Model error is the main cause in 2 entries and a contributing cause in 6 more.
Cause by symptom: where each cause shows up
Primary tags only, so every entry appears once.
| cause \ symptom | couldn’t answer | audit | false green | process | false red | hallucination | crash | total |
|---|---|---|---|---|---|---|---|---|
| missing fact | 133 | 14 | 4 | 2 | 6 | 6 | · | 165 |
| handoff seam | 16 | 47 | 14 | 9 | 17 | 4 | 5 | 112 |
| weak oracle | 6 | 2 | 44 | 12 | 8 | · | · | 72 |
| other | 4 | 9 | 2 | 43 | 5 | · | 7 | 70 |
| wrong fact | 17 | 15 | 5 | 3 | 17 | 2 | · | 59 |
| text check for meaning | 7 | 23 | 9 | 3 | 14 | 1 | · | 57 |
| guess or fallback | 4 | 18 | 14 | 2 | 4 | · | 1 | 43 |
| arbitrary bound | 19 | 9 | · | 2 | 2 | · | 3 | 35 |
| not delivered | 12 | 1 | 4 | · | 3 | 8 | · | 28 |
| stale state | 5 | 3 | 9 | 4 | · | · | 4 | 25 |
| advisory only | 2 | 2 | 8 | 6 | 2 | 2 | 1 | 23 |
| target assumption | 4 | 7 | · | 1 | 3 | · | · | 15 |
| performance | · | 6 | · | 5 | · | · | 1 | 12 |
| model error | · | · | 1 | · | 1 | · | · | 2 |

A fact the registry never captured became a check that couldn’t answer in 133 of 165 entries (81%) and an invention in only 6 (4%), while a fact that existed but was not delivered became an invention in 8 of 28 (29%), the highest hallucination rate of any cause.
One reading: a missing fact leaves a gap a check can see, so a check usually catches it before any model is asked to work around it, while an undelivered fact leaves no gap, because every component believes it was supplied, and the model fills it in.
The data is consistent with that reading but too small to prove it. Weak oracles produced 44 of the 114 false greens, the largest single source of unearned passes.
Where an LLM sat in the failure chain
An LLM’s output was part of the failure chain in 332 of the 718 entries (46.2%). By primary symptom:
| symptom | entries | LLM in chain | share |
|---|---|---|---|
| check couldn’t answer | 229 | 117 | 51% |
| found by audit | 156 | 24 | 15% |
| false green | 114 | 85 | 75% |
| process | 92 | 14 | 15% |
| false red | 82 | 58 | 71% |
| hallucination | 23 | 23 | 100% |
| crash | 22 | 11 | 50% |

105 entries carry the upstream fact link, the explicit tie between a problem in a model’s output and a fact that was missing, wrong or not delivered. By primary symptom they split 67 check couldn’t answer, 17 hallucination, 13 false green, 7 false red and 1 process. By primary cause: missing fact 44, not delivered 24, handoff seam 11, arbitrary bound 8, wrong fact 6, advisory only 3, and 2 or fewer for each remaining cause.
The 31 LLM hallucination entries
Every entry tagged hallucination as its primary or secondary symptom. 19 carry the upstream fact link, grouped by the cause behind the missing fact; the other 12 do not say where the invention came from, which does not mean nothing was missing. Identifiers from the target code are described by role. Seven of the 31 contain no false model output. In five, a deterministic tool produced the false fact and handed it to a model, and the classifier tagged the result as hallucination; the main post treats those as deterministic code making the same mistake a model makes. The other two describe gaps in the checks that catch invented names, not an invention that happened.

The pipeline stages named below: the Scout is the analyst stage inside the grounding chat, deriving from the problem prose which files and symbols the fix belongs to and whether it is feasible there, and reproducing the reported problem, the Planner plans the work including the tests, the Test Writer writes the test code, the Coder writes the change, the Debugger diagnoses failures, and the Reviewers check the result. The grounding chat is where the operator confirms the design before the autonomous run starts, and the gold patch is the benchmark’s reference fix.
The fact existed but never reached the model that needed it (8)
-
For a feature request, the Scout made up what the new capability should mean, because its prompt carried only the files the request pointed at. About a dozen existing types in the codebase already supported the same field, and the registry knew which ones, but that precedent never reached the Scout.
Tagged: hallucination; cause not delivered. Seen in a run.
-
When the Planner’s prompt lacked a fact or a spelling, it invented one: one draft invented four helper functions (7 rejections) and a later one invented five. The rule to use only real symbols was a prompt instruction checked after the fact; the response was to supply each kind of missing fact to the Planner up front.
Tagged: hallucination; cause not delivered, then advisory only. Seen in a run.
-
The test cases that check existing behaviour is unchanged were written from scratch, while the confirmed reproduction of the bug already showed the real calls that trigger and observe it. The Planner re-expressed those calls as five invented wrapper functions and used up one of its attempts. Supplying the real calls from the reproduction took the invented functions from 5 to 0.
Tagged: hallucination; cause not delivered. Seen in a run.
-
The Planner received confirmed facts for three data files but not the list of property paths each file carries, which only the Scout received. For two other files it copied a path from the confirmed ones, neither file had it, and after one wasted retry two of five changes were covered by no test. With the full list of paths supplied, the next plan passed on its first attempt.
Tagged: hallucination; cause not delivered. Seen in a run.
-
The Scout had to find the objects a function creates, but was never told what that function constructs or where it puts them. After one source read and about 4,150 thinking tokens it guessed their type wrong, so its reproduction tests captured nothing and most of its remaining attempts were spent. Earlier runs of the same unit had paid the same search cost.
Tagged: hallucination; cause not delivered. Seen in a run.
-
The Planner correctly chose a shared test helper for a trial test but invented the helper’s import path, so the trial test could not load. The pipeline’s record of the test code already knew which file defines the helper; the Planner had only been shown a place where the helper was used.
Tagged: hallucination; cause not delivered. Seen in a run.
-
The Scout’s instructions for the feasibility check told it to search a file before proposing new names, but at that step it had no way to open the file. In 200 recorded runs of the step, 26 included the instruction and none of the 26 read the file; at least 3 described searches and results that never happened.
Tagged: hallucination; cause not delivered. Seen in a run.
-
The Planner wrote every test step with an input key,
skipEntryIds, that nothing in the code reads; the real key isskipEntryId. The pipeline showed the Planner the code leading to the change but skipped the few lines above it whereskipEntryIdis read, so no test case exercised the change.Tagged: hallucination; cause not delivered. Seen in a run.
The registry or extraction never captured the fact (7)
-
The format the Planner writes test plans in had no way to say a call returns without throwing, so the Planner invented a marker that could also mean the expected value is null. It cost three gate errors and an attempt. Once the format could express it, the Planner used it unprompted on the next run.
Tagged: hallucination; cause missing fact. Seen in a run.
-
The test-plan format had no way to load a data file and check the value it delivers, so a change touching only data files could not be planned. One attempt called a function with no arguments and asserted nothing, the next invented a helper, and the plan was refused. Once the format could express that check, its first run exposed a second hole: a test asserting that a path was absent passed even though the file never had that path.
Tagged: hallucination; cause missing fact. Seen in a run.
-
A test checking existing behaviour needed to inspect an object the code creates, which the test-plan format had no way to express. The Test Writer invented a helper function to find it and the attempt was rejected, although the bug reproduction’s own setup already had a working one the pipeline never offered.
Tagged: hallucination; cause missing fact. Seen in a run.
Hand check disagrees: the cause reads as not delivered, since the pipeline held a working collector and never offered it.
-
The pipeline recorded an installed library’s classes without their method signatures, and the lookup still reported success. The Scout spent three tool calls trying to find out what one method returns, ran out of budget, and guessed, differently on two tries.
Tagged: check couldn’t answer, then hallucination; cause missing fact. Seen in a run.
-
The Test Writer passed a name one character off the constant the code expects, so that code never ran and the test failed identically with and without the change. The Debugger diagnosed it and one Coder attempt was spent anyway. The Test Writer had searched for the real name but missed it, because it was defined far from the code it had read. Two detectors for invented strings were tried and both failed on the real test, so the Test Writer is now given the constants defined in the module it tests.
Tagged: hallucination; cause missing fact, then not delivered. Seen in a run.
Hand check disagrees: the cause reads as not delivered first, since the constant was in the registry and the search missed it.
-
A search tool told the model that two classes were not exported at all, because the registry did not link a file’s default export to the class it exports, although the parser had already read that name. An audit of recorded runs counted 35 affected answers.
Tagged: hallucination, then check couldn’t answer; cause missing fact, then not delivered. Seen in a run.
Hand check disagrees: no model output was false; the tool handed the model a wrong fact, so check couldn’t answer, caused by a wrong fact, fits better.
-
A reproduction test written by the Scout called
getEntryByIdon a model class that declares onlygetEntryByName,getEntryByIndexandgetEntryByType, through a cast that hid the error from the type checker. Nothing checked the call before it ran, so the pipeline paid a live run, 32 retries with alternative test setups that all died on the same line, and a repair attempt to find out.Tagged: hallucination; cause missing fact, then advisory only. Seen in a run.
Hand check disagrees: the registry did hold the class’s member list; the cause reads as advisory only, since the one pre-run check refused nothing.
Two stages disagreed about the fact (3)
-
A path lookup refused a hint from the operator because it named a directory rather than one file, and a model later in the pipeline rewrote that refusal as a claim that the directory did not exist, attributed to a Scout check that never happened. The directory existed and held dozens of files. The refusal now says whether the path exists.
Tagged: hallucination; cause handoff seam, then not delivered. Seen in a run.
-
Adding a line to the ticket about what the test observes led the Planner to invent a collector function, and the first attempt failed on four calls to functions that don’t exist, despite an explicit prompt rule against invented helpers. The test-plan format had no way to express that observation, and the bug reproduction covered those cases one stage later anyway.
Tagged: hallucination; cause handoff seam, then advisory only. Seen in a run.
-
The ticket included an example scenario as background for the Reviewers, and the Test Writer read it as a requirement and invented a third acceptance criterion the ticket never asked for. Its only test case made twelve calls and asserted nothing, which burned the first attempt, and the extra criterion broke the link between the pipeline’s own check for existing behaviour and the requirement it belonged to.
Tagged: hallucination; cause handoff seam. Seen in a run.
A budget ran out before the fact was found (1)
-
The Scout was asked to find and confirm the data files a change touches, on top of its existing job, with the same small tool budget. It ran out, said so, and proposed a file anyway against its instructions. A validator rejected the guess, but the run fell back to a manual step; the budget has since been raised, and running out is now counted.
Tagged: check couldn’t answer, then hallucination; cause arbitrary bound. Seen in a run.
No stated link to a missing fact (12)
-
The Scout picked a concrete configuration value for its reproduction that the issue never named, and nothing marked it as the Scout’s choice rather than a fact from the issue. The Test Reviewer sharpened the assertion on top of that guess, the test failed on the unchanged code as required, and only the check against the gold patch showed the premise was wrong.
Tagged: false green, then hallucination; cause weak oracle. Seen in a run.
-
A generated test was titled for one input value while its body exercised another, because the title was left over from an earlier version of the test. The assertion was correct, but the misleading title reached the Test Writer as the test’s name; titles are now checked.
Tagged: hallucination; cause handoff seam. Seen in a run.
-
When a model was asked to redo its work, it passed one argument to a library method that takes none, although the pipeline’s record of installed libraries held the method’s signature, the same stage had looked that signature up twice, and the first round of the same run had written the call correctly. Nothing checked the model’s calls against the recorded signatures, so the error surfaced only at the type-check gate.
Tagged: hallucination; cause advisory only, then model error. Seen in a run.
-
During the feasibility check, a model claimed the change relied on a class member that was actually a local variable inside one of the class’s methods. A check for whether the member exists couldn’t confirm it was missing, only because the class’s inherited members hadn’t been resolved; had they been, the pipeline would have added a false requirement to create that member.
Tagged: found by audit, then hallucination; cause other. Seen in a run.
Hand check disagrees: the false claim was written in a recorded run, so hallucination fits better than found by audit, and the cause reads as model error.
-
During the feasibility check, a model claimed the change relied on a class member that is neither a member of the class nor a local of the method in question, but a local variable or option field in a different method of the same class. Only an incomplete member list stopped a false requirement to create it, and checking membership against the class’s own declarations was still open work.
Tagged: found by audit, then hallucination; cause other. Seen in a run.
Hand check disagrees: the false claim was written in a recorded run, so hallucination fits better than found by audit, and the cause reads as model error.
-
The call-graph lookups matched functions by name alone, so they listed functions that never call the target and returned no callers for a method with dozens of actual ones. An audit of 253 recorded model sessions found 120 false-empty and 56 false-caller answers across 24 to 27 runs; the lookups now find calls to the exact function being asked about.
Tagged: hallucination; cause wrong fact, then text check for meaning. Seen in a run.
Hand check disagrees: no model wrote anything; the tools returned wrong caller lists, so the cause reads as a text check for meaning.
-
A dependency lookup told models in the pipeline that a class method had no implementation, because it assumed every installed package ships compiled code, and this one did not. It also read the method’s details from the wrong declaration.
Tagged: hallucination, then false red; cause wrong fact, then target assumption. Seen in a run.
Hand check disagrees: the false answer came from a deterministic lookup, not a model; the cause reads as a wrong fact.
-
The lookup that verifies claims made in the grounding chat stated that a method call resolved to an unrelated function in another module, because a loose name match picked the wrong function. The false fact reached the refined ticket of every grounding chat on that task, and the same loose match had sent 15 of 1,929 recorded names to the wrong code.
Tagged: false green, then hallucination; cause guess or fallback, then text check for meaning. Seen in a run.
-
Prose in the grounding chat cited a file inside a dependency by full path and, in the same text, by bare file name. The bare name was resolved loosely to an unrelated file with the same name in the project, which was added to the files to change, carried into the ticket, and left one part of the work with no tests; the run failed.
Tagged: hallucination, then false red; cause text check for meaning, then wrong fact. Seen in a run.
Hand check disagrees: a deterministic name match picked the wrong file, not a model; the cause reads as a text check for meaning.
-
A check on statements the operator added re-asked the operator (here a model answering in the operator’s place) in 25 of about 43 recorded grounding chats, although the statements were correct. In some cases the pipeline read a return value into a statement that never gave one; in most, the claim was checked against the installed library instead of the test setup the chat had described. After the change, a replay of all 25 chats produced 0 re-asks.
Tagged: false red, then hallucination; cause wrong fact, then handoff seam. Seen in a run.
-
A check that spots a plan setting an input key nothing reads only reports it and does not send the plan back. It only sees keys the changed code reads directly, so on a second task it flagged keys read deeper in the code. It stays a warning until its precision is measured on every task.
Tagged: hallucination; cause advisory only. Found by design review or offline measurement.
-
The new check that a called method or property exists covers test code only; names used elsewhere in the pipeline are not yet checked, so an invented one there can still reach the plan, or ship until a run catches it.
Tagged: found by audit, then hallucination; cause advisory only. Found by design review or offline measurement.
The 8 entries where model error was a cause
Every entry with model error as its primary or secondary cause. One also appears in the hallucination list above.
-
When the grounding chat wrote the ticket, the model reworded a sentence limiting the change’s scope, and a check that needed the exact wording rejected it, so a run whose tests all passed on the gold patch was marked degraded and the benchmark run stopped. The pipeline already had the sentence, so it no longer relies on the model to copy it.
Tagged: false red; cause advisory only, then model error. Seen in a run.
-
After the pipeline had established from verified facts that the change edited a known set of data files, it asked a model to classify what kind of change it was again and ended the grounding chat when that answer came back unclassified. Three earlier chats had passed the same step, so this was sampling variance; the pipeline no longer asks a model for what it has already established.
Tagged: false red; cause model error, then handoff seam. Seen in a run.
-
The Scout proposed reusing an existing helper, and the pipeline recorded the proposal as a decision without asking the operator. It became a requirement that the change call that helper, the Test Writer wrote a test counting the calls, and the run was withheld because the gold patch never calls that helper. Once the proposal was put to the model answering in the operator’s place, it was dropped after a read of the code, and the rerun’s tests all passed on the gold patch.
Tagged: false red; cause handoff seam, then model error. Seen in a run.
-
The redo entry from the hallucination list above, where a model passed an argument to a library method that takes none although the recorded signature was available to it and the same stage had read that signature twice.
Tagged: hallucination; cause advisory only, then model error. Seen in a run.
-
A model wrote down, for a bug reproduction, which result means the bug is present and which means it is fixed, and got them the wrong way round. A reproduction that no change could ever make pass was accepted as confirmed, the variant that did show the bug was judged passing, and later stages searched for a measurable effect that did not exist, although the pipeline already held the facts that decide which way round is right.
Tagged: false green; cause model error, then weak oracle. Seen in a run.
-
A Reviewer complaint about a test that named no fact the pipeline could check still had full power to send the Test Writer’s work back; such complaints no longer decide the verdict.
Tagged: false red; cause advisory only, then model error. Found by design review or offline measurement.
Hand check disagrees: the problem is a Reviewer complaint with too much authority, not a warning without enforcement; other fits better.
-
The acceptance record a model wrote listed, for one criterion, an input case that can only happen with the real installed library, while every test in the plan ran in a test setup the pipeline built, where that case cannot happen. A coverage check still required a test for it, so a repair round ran, added 0 tests, and the run went on with the gap unrecorded.
Tagged: false red; cause handoff seam, then model error. Seen in a run.
-
Whether the acceptance criteria said existing behaviour must be kept varied from run to run: in 6 of 40 recorded grounding chats (15%), a broken change that returned false everywhere would have satisfied every criterion (none actually did), and no later check worked out the “unchanged” requirement on its own. The pipeline now adds that requirement itself.
Tagged: false green; cause weak oracle, then model error. Seen in a run.
Limits of LLM-labelled defect data
The tags come from one classification pass, eight parallel runs of one model over one written taxonomy, with no second rater. The only agreement check is the hand reading of 38 entries for this page, which disputed a primary tag on 10 of them (26%). Those 38 are not a random sample: they are the hallucination and model-error entries, the two tags whose definitions are hardest to apply, so the rate says little about the other 680. The counts were not re-checked entry by entry.
Entries are written under the pipeline’s own rule that when a model carrying out the plan fails, its instructions were almost always incomplete, and each entry describes a mechanism to build, which pulls the record away from “the model simply got it wrong”. The model-error count is a floor set by that rule, not a measured rate.
The classifier saw truncated entries, so a cause stated only late in a long entry could be missed. Removing the wasted-work entries changes the base: anyone recomputing over all 943 will get different shares, and the main post’s figures hold only for the 718.
Defects are entries in the project’s own gap catalogue, 28 August to 5 October 2026, from LoopsBench units in TypeScript and Python. Identifiers from the target code are described by role or replaced with stand-ins. Task data is canaried against training corpora and is not reproduced here, in prose or in any image.