27 August 2026
New benchmark reveals AI models struggle with complete scientific tasks
First reported
Deep Learning Weekly ran this on .
- FrontierChallenge tested AI models on 97 scientific workflows spanning six different fields, measuring whether models could finish entire tasks end-to-end.
- The best-performing models succeeded on only 20.6% of these complete workflows, despite often showing high scores on individual parts of tasks.
- Results show that current AI confidence scores and partial progress do not reliably predict whether models will actually deliver finished work.
How it was covered
Deep Learning WeeklyEditorial team
New benchmark of 97 end-to-end scientific workflow tasks across six domains found that best-performing models achieved only 20.6% pass rate, revealing that high partial scores and confident completion claims do not reliably indicate full task delivery.