EVISKILL: Grounding Skill Evolution in Replayable Evidence
EVISKILL keeps the evidence behind an agent’s skill edits replayable, so revisions can be checked before becoming persistent guidance.
The paper proposes Replayable Evidence Cards to preserve execution observations and the task contexts behind proposed skill changes. Its replay step re-executes targeted edits, using the result to verify or correct them. Supported edits can be kept provisionally across epochs, while global validation decides what enters the final skill. Experiments span three interactive benchmarks and six LLM backbones. HF Daily Papers' note
The paper proposes Replayable Evidence Cards to preserve execution observations and the task contexts behind proposed skill changes. Its replay step re-executes targeted edits, using the result to verify or correct them. Supported edits can be kept provisionally across epochs, while global validation decides what enters the final skill. Experiments span three interactive benchmarks and six LLM backbones. HF Daily Papers' note
score 4