LLM Autobiography Audit Reports 96.7% Scene Verification Failure Rate
Scene-level audit of 366 LLM-generated autobiographical entries yields 96.7% verification failure. Grounding in subject corpus lowers but does not eliminate confabulation. Taxonomy reliability limits reproducibility of fine-grained labels.
The audit examined 366 consecutive first-person entries produced by a conversational LLM. Inputs were limited to a template, two exemplars, and daily quotes. Each entry underwent pre-registered four-level rubric scoring against an independent verification corpus of the same subject's life events. Only 12 days contained any corroborated scene.
Verification failure reached 96.7% (Wilson 95% CI 94.4-98.1%). Nineteen days asserted claims directly contradicted by the record. Dominant error mode was grounded drift: real entities placed in fabricated sequences. Independent re-rating confirmed the headline rate while exposing only fair-to-moderate inter-rater reliability on the WEAK/UNVERIFIED boundary.
Regeneration with current frontier models reproduced 100% failure under identical inputs. Adding the subject's own corpus as grounding reduced failure to 83.3%, leaving substantial residual confabulation. The study supplies the first quantified scene-level instrument and demonstrates that partial grounding narrows but does not close the gap.
Operational implication is direct: any deployment of LLMs for personal narrative generation requires external verification corpora and repeated audit cycles rather than reliance on model scale alone.
Renze: scene-level verification failure on grounded inputs falls below 70% within 18 months of next major context-window increase.
Sources (2)
- [1]Primary Source(https://arxiv.org/abs/2608.23640)
- [2]Supporting Source(https://arxiv.org/abs/2305.18223)