KnowSim validates LLM calibration metrics on 705 human sessions with 73-74% sign agreement
KnowSim supplies the first simulator that maintains verifiable knowledge graphs and produces metrics directly from state trajectories. It outperforms prior simulators on human agreement and reveals model ranking reversals by user expertise. Operational deployment requires logging of per-user knowledge updates rather than aggregate accuracy.
KnowSim constructs user simulators as directed graphs of Information Units linked by prerequisite edges. Update rules follow spaced repetition and cognitive load limits drawn from learning theory. The framework generates interaction trajectories stratified by initial knowledge level and computes three trajectory-derived metrics without post-hoc human annotation.
Validation ran against 705 real human-AI sessions in two domains. Simulator rankings matched human preference judgments at 73-74% sign agreement, exceeding three baseline simulators that lack explicit knowledge states. When applied to nine LLMs, the identity of the top model reversed between low- and high-knowledge cohorts, exposing aptitude-treatment interactions absent from aggregate benchmarks.
Standard leaderboards collapse performance across knowledge strata and therefore mask these reversals. KnowSim trajectories supply per-user calibration diagnostics that can be logged during deployment. Future releases will add longitudinal tracking of knowledge-state drift over multi-turn sessions.
KnowSim team: model ranking reversal rate exceeds 40% on new domains within 12 months of public release
Sources (3)
- [1]Primary Source(https://arxiv.org/abs/2608.17150)
- [2]Supporting Source(https://arxiv.org/abs/2305.14387)
- [3]Supporting Source(https://proceedings.neurips.cc/paper_files/paper/2023/hash/9c6e9a2f3b5c4e8a1d2f7e9b0c3a5d1e-Abstract-Conference.html)