Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

We need public benchmarks.

This is incredibly fast progress on large contexts and I would like to see if they are actually attending equally as well to all of the information or there is some sparse approximation leading to intelligence/reasoning degradation.



https://lmsys.org/blog/2023-05-10-leaderboard/

https://chat.lmsys.org/?arena

Claude by Anthropic has more favourable responses then ChatGPT


So I tried this prompt in their chatbot arena multiple times. Each time getting the wrong answer:

"Given that Beth is Sue's sister and Arnold is Sue's father and Beth Junior is Beth's Daughter and Jacob is Arnold's Great Grandfather, who is Jacob to Beth Junior?"


Is the right answer pointing out that Arnold might not be Beth's father, and so Beth Junior might be unrelated to Jacob?


I just tried it and gpt-3.5-turbo got it right.


ChatGPT3.5*

It's still below GPT4, but it is closer to 4 than 3.5




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: