Megadose Built for builders and researchers.

FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth

· HF Daily Papers ·
Grok 4.6 posted the top reported score on FlavourBench’s 534 culinary choice tasks.

The paper tests 27 frontier model endpoints on ingredient substitution, pairing, and constraint prompts, scoring every three-ingredient portfolio from eight candidates. Its best point estimate was Grok 4.6 at 65.1, with the same ranking reproduced across several independently compiled panels and three public Epicure checkpoints. A small post-training study also found LoRA SFT on Qwen3-0.6B improved Epicure task performance by 13.3 points over a matched control. HF Daily Papers' note

score 5

Categories: Research