Agentic Pipeline Records Zero Outcome Leakages on 38-Case eICU Mortality Subset
The agentic pipeline eliminated outcome leakage and improved guideline grounding on ICU mortality explanations. Standalone LLM retained stronger SHAP alignment. Attribution checks are required before clinical use.
XGBoost was trained on 2,353 ICU stays with 8.1 percent mortality and reached an AUROC of 0.855. The retained artifact set was used to generate explanations for both the standalone LLM and the pre-specified agentic pipeline that separated data interpretation, guideline checking, and final narrative steps. The study retained the original standalone-versus-agentic comparison while clarifying clinical findings.
Among the 14 overlapping SHAP-reviewed cases the standalone LLM showed higher mean Jaccard alignment at 0.171 versus 0.077 and 92.9 percent direction consistency versus 78.6 percent. The agentic pipeline recorded higher guideline grounding at 0.762 versus 0.143, higher value specificity at 0.236 versus 0.143, and marginally higher plausibility at 0.700 versus 0.671. No leakage occurred in the agentic outputs.
Decomposition therefore trades attribution fidelity for measurable gains in safety-relevant grounding and patient-specific detail. The authors conclude that attribution-based checks must remain paired with any agentic output before high-stakes deployment.
Next steps require prospective validation on larger multi-center cohorts and integration of real-time SHAP verification within the pipeline itself.
Wang et al.: Agentic pipeline guideline grounding exceeds 0.80 on a 200-case multi-center set within 18 months.
Sources (2)
- [1]Primary Source(https://arxiv.org/abs/2608.26109)
- [2]Supporting Source(https://arxiv.org/abs/2305.14303)