Megadose AI progress, ranked and analyzed.

Auditable Long-Term Memory: A Deterministic Retrieval Chain Measured at 479/475 of 500 on LongMemEval-S

· ArXiv · AI/CL/LG ·
The paper reports high LongMemEval-S scores, but frames them as inspectable rather than settled benchmark wins.

The deterministic retrieval chain put all gold sessions in the candidate pool for 468 of 470 answerable questions and produced gold-complete packets for 462. Two 500-question passes with a Claude Opus reader scored 479 and 475 under GPT-4o judging, while a second judge scored both passes 472. The author notes scoring-prompt changes, reader setup, possible data-version differences, and lack of held-out evaluation or independent human adjudication. Materials for packets, scaffolds, reader outputs, judge verdicts, and controls are released for inspection and re-scoring. ArXiv · AI/CL/LG's note

score 4

Categories: Research