Replicating TRACE: A Practitioner's Guide to Its Threshold and Particle Budget
The replication matches TRACE’s headline F1, but says that score mostly reflects lag-1 edge recovery under the benchmark’s skew.
The authors report mean per-sequence F1 of 0.90-0.91 at vocabulary size 1000 when the threshold is validation-selected, matching the cited TRACE result. Their analysis says the useful threshold tracks the truth margin rather than a fixed constant. At a single global threshold, TRACE recalls lag-1 true edges at 0.97-0.99, while longer-lag edges fall much lower. They also find F1 saturates from two particles at the selected threshold, even though the estimator itself follows N^(-1/2) convergence. ArXiv · AI/CL/LG's note
The authors report mean per-sequence F1 of 0.90-0.91 at vocabulary size 1000 when the threshold is validation-selected, matching the cited TRACE result. Their analysis says the useful threshold tracks the truth margin rather than a fixed constant. At a single global threshold, TRACE recalls lag-1 true edges at 0.97-0.99, while longer-lag edges fall much lower. They also find F1 saturates from two particles at the selected threshold, even though the estimator itself follows N^(-1/2) convergence. ArXiv · AI/CL/LG's note
score 3