Megadose AI progress, ranked and analyzed.

ESPO: Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize

· ArXiv · AI/CL/LG ·
ESPO reports higher accuracy than GEPA while cutting optimized prompt length nearly in half.

The paper says evolutionary prompt optimizers can accumulate rules without improving results. ESPO instead diagnoses training errors, diversifies candidate prompts, and uses bootstrap stability selection. Across seven public NLP benchmarks, it reports 74.67% average accuracy versus 70.91% for GEPA, with prompts 47% shorter. The authors also say ablations support the selection step: diversity without bootstrap selection reduced performance.

ArXiv · AI/CL/LG's note

score 5

Categories: Research