Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

(March 2024) Antropic claimed Claude 3 Opus had "graduate-level expert reasoning" with GPQA results of around 60% showing a roughly phd level performance.

(Sept 2024) OpenAI claimed o1 was phd-level in their launch post.

You're kinda both wrong. :)



They claimed the model was PhD-level, but they never mentioned the university the model graduated from... :)




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: