In production, agentic systems answer questions over data that lives in several places and keeps changing. Existing retrieval benchmarks freeze the data, so they can only ask whether an agent found the right passage — not whether the answer is still true. An agent that quotes yesterday's price or a departed user's permissions looks identical to one that reasons badly.

ChurnBench rebuilds the benchmark around time. It generates a four-source enterprise data fabric as a timeline rather than a snapshot, writes every change to an append-only ground-truth ledger, and computes gold answers from that ledger instead of the live stores. Ground truth is resolved at both timestamps for every reported case, so an answer that was correct when retrieved but wrong when evaluated is caught and labeled a freshness error, separate from a reasoning error.

The instrument produced a counterintuitive finding: when a system refreshes on a schedule, cache age does not predict staleness. Across cache ages of 1, 14, and 28 days, freshness errors were 7, 4, and 4, because scheduled refresh bounds staleness by time-to-live and no TTL lapse was observed in any window. A controlled ablation confirms the mechanism: disabling tiered refresh raises freshness errors from 4 to 45 at 28 days while leaving them identical at one day.

The conclusion for system builders: the variable a drift benchmark should sweep is TTL configuration against each entity's rate of change, not drift-window length.

An answer can be right when its data was fetched and wrong when it is used; a frozen benchmark cannot see the difference.

Computing gold answers from an append-only ledger makes 'was it true at both ends' checkable instead of assumed.

Separating freshness errors from reasoning errors points debugging at the refresh policy rather than the model.

Scheduled refresh bounds staleness by time-to-live, so cache age stops being the variable that matters.

The ablation isolates the mechanism: removing tiered refresh raises freshness errors from 4 to 45 at 28 days.

When data changes underneath a system, score correctness at retrieval time and at evaluation time, and tune refresh by each entity's rate of change, not by cache age.

ChurnBench, the evaluation harness, and all per-error data were released open source. The paper was submitted to arXiv on 10 September 2026 and revised on 20 September 2026.

FOLLOW THE EVIDENCE

The sources

  1. ChurnBench: A Drift-Aware Benchmark Demonstrating That Refresh Scheduling, Not Cache Age, Governs Staleness in Agentic AI arxiv.org