CTWM cuts tokens 5.9% on Synthetic Graph World and 24.48% on LongMemEval while preserving state coverage
CTWM exploits measured heavy-tailed memory traces to allocate finite context via a single exponent τ. It delivers 5.9-24.48% token reductions across three benchmarks while holding accuracy parity. The approach converts a diagnostic observation into an operational control signal for long-horizon agents.
The arXiv paper documents reproducible heavy-tailed retrieval under repeated access: random-walk agents produce log-normal traces while LLM semantic policies generate truncated power-law cores. CTWM exploits this by rank-ordering memory entries and compressing the tail, maintaining full transition coverage on the graph environment. Token savings transfer to ALFWorld and LongMemEval without accuracy loss.
Related work on retrieval-augmented agents shows similar concentration when context windows stay fixed; the 2024 LongMem paper already measured cumulative error growth in the lowest-frequency states. CTWM formalizes the control parameter τ as an explicit lever rather than emergent side-effect. This shifts memory design from uniform eviction to explicit core-tail budgeting.
Operational impact appears in sustained multi-turn tasks where rare state prediction drives downstream failures. Deployments that adopt rank-based controllers can reduce context spend without sacrificing coverage, provided the tail summary remains faithful. Follow-on tests must verify whether τ generalizes across model scales and retrieval frequencies.
Next steps include integrating CTWM into production agent runtimes and measuring tail error on live user trajectories over 1000+ steps.
CTWM: token reduction exceeds 20% on LongMemEval with accuracy parity when τ tuned on 500+ trajectories within 6 months
Sources (3)
- [1]Primary Source(https://arxiv.org/abs/2610.00010)
- [2]Supporting Source(https://arxiv.org/abs/2402.05875)
- [3]Supporting Source(https://arxiv.org/abs/2305.05176)