FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth
Grok 4.6 posted the top reported score on FlavourBench’s 534 culinary choice tasks.
The paper tests 27 frontier model endpoints on ingredient substitution, pairing, and constraint prompts, scoring every three-ingredient portfolio from eight candidates. Its best point estimate was Grok 4.6 at 65.1, with the same ranking reproduced across several independently compiled panels and three public Epicure checkpoints. A small post-training study also found LoRA SFT on Qwen3-0.6B improved Epicure task performance by 13.3 points over a matched control. HF Daily Papers' note
The paper tests 27 frontier model endpoints on ingredient substitution, pairing, and constraint prompts, scoring every three-ingredient portfolio from eight candidates. Its best point estimate was Grok 4.6 at 65.1, with the same ranking reproduced across several independently compiled panels and three public Epicure checkpoints. A small post-training study also found LoRA SFT on Qwen3-0.6B improved Epicure task performance by 13.3 points over a matched control. HF Daily Papers' note
score 5