Megadose AI progress, ranked and analyzed.

Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

· HN · Agents ·

Terminal-Bench-Science introduces a benchmark for testing AI agents on realistic scientific research workflows in terminal environments.

Categories: Research

Excerpt

HN · 101 points · 29 comments

Discussions

  • hn · 101 points · 29 comments
  • hn · 103 points · 30 comments
  • hn · 104 points · 30 comments
  • hn · 104 points · 32 comments