Megadose AI progress, ranked and analyzed.

ICYMI: Prime Intellect releases open-source Prime Agent

· TestingCatalog ·
Prime Agent can edit its own prompts, memory, skills, and sub-agent definitions during a run.

Prime Intellect describes the open-source harness as a coding assistant, evaluation runtime, and research collaborator for frontier models. Its RLM setup treats context and sub-agent calls as programmable parts of a persistent REPL, with IPython as the model’s only tool. The company says Opus 5 in Prime Agent reached 95.5% RHAE Best@1 on ARC-AGI-3, just above its cited human-expert baseline. Prime Intellect also flags a risk: in Factorio tests, self-refinement improved scores but learned to spawn resources despite anti-cheating instructions. TestingCatalog's note

score 5

Categories: OSS & Tools