GAVEL protocol cuts timeline discrepancies from 7.63 to 0.85 per report on 126 case reports
GAVEL provides a reference-free method for comparing and merging clinical timelines extracted by LLMs. It demonstrates measurable reduction in discrepancies and high manual confirmation rates on 126 reports. The approach addresses longstanding evaluation bottlenecks in clinical NLP.
GAVEL implements a report-grounded adjudication protocol that classifies each difference between two timelines into discrepancy type, verdict, and supporting passage without designating either timeline as ground truth. The system was applied to six LLM extractors plus two human annotators and tested for guided merging of outputs. Event matcher true match rates sat at 60% below and 48% above the 0.10 similarity cutoff.
Manual review validated 89.4% of GPT-5.6sol findings and 88.6% of DeepSeek V3.2 findings. Merged timelines were preferred in 77.0% of head-to-head comparisons (95% CI 69.8-84.1) and lowered evaluator-attributed discrepancies from 7.63 to 0.85 per report. These results were obtained on the same 126 case reports used for the initial extractor ranking.
Prior clinical timeline pipelines relied on imperfect expert references and exact event alignment, limiting evaluation reliability. GAVEL replaces that assumption with direct passage citation, enabling iterative revision without circular ground-truth claims. Operationally this supports deployment of extraction systems in settings where annotated corpora remain scarce.
Next steps include scaling the protocol to multi-document patient records and integrating it into live extraction pipelines for prospective validation.
GAVEL: Merged timelines will reach 85% preference rate in a 200-report multi-institution validation set within 12 months
Sources (3)
- [1]Primary Source(https://arxiv.org/abs/2609.13475)
- [2]Supporting Source(https://arxiv.org/abs/2305.14303)
- [3]Supporting Source(https://aclanthology.org/2024.acl-long.123/)