Megadose AI progress, ranked and analyzed.

CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data

· ArXiv · AI/CL/LG ·
CRAFT turns rubric criteria into a capability map, then uses the weakest nodes to steer fine-tuning data.

The paper treats each grading criterion as a probe of a specific model capability, clusters those probes into a hierarchical tree, and scores the target model across that tree. Low-performing nodes are selected at the level where the failure is clearest, rather than only by prompt or broad category. In tests across four open-source models, finance and legal domains, and 13 held-out benchmarks, CRAFT generally beat prompt-level clustering and random data generation after fine-tuning. ArXiv · AI/CL/LG's note

score 5

Categories: Research