Megadose AI progress, ranked and analyzed.

TokenCast: Forecasting Token Consumption During LLM Agent Execution

· ArXiv · AI/CL/LG ·
TokenCast forecasts an agent run’s token bill as the run unfolds, without adding extra LLM calls.

The paper says token use for the same agent task can vary by more than 10x because tool feedback changes the path and accumulated context keeps raising later input costs. TokenCast represents each execution segment by its own cost and the context growth it causes, then composes segments into a running estimate. On SWE-bench Verified, its mean cumulative prediction time was 32.8 ms per run. Across four task suites and six agent models, it cut mean absolute error by 14.5% versus the strongest comparator, and used 21.3% fewer tokens than a fixed-budget policy in offline replay at matched completion.

ArXiv · AI/CL/LG's note

score 4

Categories: Research