Megadose AI progress, ranked and analyzed.

How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure

· HF Daily Papers ·
The ranking mostly holds at the bottom, not at the top.

The paper audits an LLM prompt-structure evaluation across eight open model variants and finds the underlying measurements were unstable. Identical calls often failed to recover the same structure, and 72% of prompt-model cells were never node-set-perfect. Bootstrap checks made only the two worst ranks look firm, while middle and top positions shifted enough to weaken the table’s claims. The author also says reproducibility did not equal accuracy, and four tested endpoints were withdrawn within ten weeks, making the original study impossible to rerun as specified. HF Daily Papers' note

score 4

Categories: Research