Dating the Model: Hidden Dates in System Prompts Affect LLM Evaluation
A hidden current-date field in system prompts can move benchmark scores and even reorder model rankings.
The paper tests 9 recent LLMs across 6 datasets covering QA, math, code, and translation. It reports swings of up to 6% on MCQA, 14% on math reasoning, 7% on code generation, and 2.84 BLEU on machine translation from the date alone. Chain-of-thought and few-shot prompting did not remove the effect; chain-of-thought amplified it. HF Daily Papers' note
The paper tests 9 recent LLMs across 6 datasets covering QA, math, code, and translation. It reports swings of up to 6% on MCQA, 14% on math reasoning, 7% on code generation, and 2.84 BLEU on machine translation from the date alone. Chain-of-thought and few-shot prompting did not remove the effect; chain-of-thought amplified it. HF Daily Papers' note
score 6