Megadose Built for builders and researchers.

Does an Agent's History Tell You When Compaction Will Hurt? A Modest, Bounded Effect on the TRACE Paired-Replay Corpus

· ArXiv · AI/CL/LG ·
Agent behavior before compaction only weakly predicted whether the summary would cause trouble afterward.

The paper tests 590 TRACE paired-replay compaction boundaries from AppWorld, comparing runs with the original pre-compaction context against runs with the summary. Harm is measured through the burden of next actions, including errors and repeated calls. The strongest reported trigger reached held-out AUROC 0.66, below a same-boundary replicate at 0.72. A frozen interpretable trigger avoided 21% of harmful boundaries while keeping 84% of compaction opportunities, but the release cannot show whether it beats a token-budget rule at matched retention. ArXiv · AI/CL/LG's note

score 4

Categories: Research