GPT-5.0 Achieves 81.67% Accuracy Extracting Parameters from 536 Agent-Based Disease Spread Papers
LLM pipeline extracted parameters from 536 disease spread papers at 81.67% paper-level accuracy with GPT-5.0. Agreement between models signals hallucinations versus human annotation errors. Threshold-based filtering enables scalable SLRs for agent-based epidemiology models.
The pipeline automated information extraction across 536 papers previously reviewed by humans for agent-based disease models. Structured prompts targeted fields including model type, parameters, validation status, and assumptions. Outputs were scored directly against the human reference set to isolate extraction errors from annotation noise.
Field-level accuracies ranged from 32.40% on subjective attributes such as assumption quality to 100% on metadata fields. Inter-LLM agreement tracked output reliability: low agreement predicted hallucinations while high agreement with low accuracy identified human dataset errors. These patterns held across both GPT versions.
The results establish agreement thresholds as a practical filter for scaling SLRs in epidemiology. Deployment in outbreak modeling workflows requires hybrid review only on low-agreement cases. This reduces manual effort while preserving parameter fidelity for simulation calibration.
Next steps include testing open-weight models and longitudinal tracking of accuracy drift on new pathogen literature through 2027.
Kavak: Agreement-filtered field accuracy on subjective parameters exceeds 70% in new disease spread SLRs by Q3 2027.
Sources (2)
- [1]Primary Source(https://arxiv.org/abs/2608.26150)
- [2]Supporting Source(https://arxiv.org/abs/2305.00865)