Megadose AI progress, ranked and analyzed.

Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation

· HF Daily Papers ·
The project catalogs AI benchmarks with source records, score histories, daily discovery feeds, and evidence links for checking evaluation claims.

Benchmark Radar is described as a searchable, living database for LLM, agent, coding, reasoning, safety, and domain-specific evaluations. The paper says it pulls daily from 37 sources and currently includes 1,283 source records plus 12,916 numeric observations across 790 records. It keeps source identities and citations so users can inspect benchmark evidence instead of relying on isolated leaderboard numbers. The release includes a web dashboard, CLI, downloadable evidence, and views for leaderboards, saturation, trends, and Pareto comparisons. HF Daily Papers' note

score 5

Categories: Research, Products to Try