Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

58 points | by matt_d 5 hours ago

13 comments