CreativeInstruct: Scalably Teaching LLMs to Balance Quality, Creativity, and Diversity
CreativeInstruct trains a post-tuned model to recover creative variation without giving up output quality.
The method adds special `[StartCreativity]` spans that steer generation toward more base-model-like creative behavior. The paper also proposes a graph-edit-distance diversity metric meant to catch narrative structure changes missed by lexical or semantic measures. On narrative generation, it matches or beats multi-model diversity baselines while using only one model at inference. Human annotators rated its outputs more creative than post-trained LLM outputs in 70.3% of cases, and GRPO on a CreativeInstruct checkpoint improved AMC and MATH results versus the same training on the post-trained checkpoint. ArXiv · AI/CL/LG's note
The method adds special `[StartCreativity]` spans that steer generation toward more base-model-like creative behavior. The paper also proposes a graph-edit-distance diversity metric meant to catch narrative structure changes missed by lexical or semantic measures. On narrative generation, it matches or beats multi-model diversity baselines while using only one model at inference. Human annotators rated its outputs more creative than post-trained LLM outputs in 70.3% of cases, and GRPO on a CreativeInstruct checkpoint improved AMC and MATH results versus the same training on the post-trained checkpoint. ArXiv · AI/CL/LG's note
score 4