Ссылка
click to show
click to show
Terminal-Bench-Science launches 70 reproducible science tasks for AI agents
Terminal-Bench-Science released version 0.1 on August 28, featuring 70 tasks drawn from life sciences, physics, Earth sciences, mathematics, and engineering.
Each task provides a working environment with data and software; the agent must produce code, data, analysis, a simulation, or a proof that can be checked automatically by a reproducible test.
Out of 920 community submissions, 464 were approved for implementation, while 386 remained open for inclusion in the shared repository.
The team selects only workflows that admit a clear, reproducible verification method, echoing their view that ‘the bar for AI scientific capability is set by scientists, not model developers or data providers.’
In three independent runs per model across
🔗 Read original →
11 ·