Prompt Design at Scale: How Format, Instruction Count, and Context Length Shape Instruction Adherence and Hallucination in Large Language Models
Instruction adherence broke down completely once prompts carried 80 rules.
The paper reports two controlled experiments across five models using a synthetic corpus called the Book of Veyra. Formatting did not produce a reliable markdown advantage, and prompt placement sometimes mattered as much as format. Long-context recall held up through roughly 64k-128k tokens, then fell sharply depending on model and format. The study found no fabrication in its probes, but refusals rose near context limits. ArXiv · AI/CL/LG's note
The paper reports two controlled experiments across five models using a synthetic corpus called the Book of Veyra. Formatting did not produce a reliable markdown advantage, and prompt placement sometimes mattered as much as format. Long-context recall held up through roughly 64k-128k tokens, then fell sharply depending on model and format. The study found no fabrication in its probes, but refusals rose near context limits. ArXiv · AI/CL/LG's note
score 6