Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

What does it say to the second question? I've found Claude is one of the worst models with regards to pop culture knowledge like this, even compared to the Chinese open ones. Just curious, not really relevant to the initial post but I don't pay for it so I only have access to Sonnet.

https://claude.ai/share/5e7e09b2-a75a-4024-b261-9a1a4e063a8b this is mostly hilariously wrong. wrong tie colors, they did not replace their bassist with a drummer, two completely made up albums, the rob cantor song it is thinking of is "shia labeouf", and a few fan behaviours i think it just made up



"I've found Claude is one of the worst models with regards to pop culture knowledge like this, "

Is that bad?

I want to use models for coding and reasoning capabilities, not pop trivia knowledge they can get with web search.


For most cases probably not. It's just something I like testing new models on sometimes, the pelican riding a bicycle benchmark probably isn't that useful either.


But that is just testing for encoded knowledge, the pelican riding requires some reasoning capabilities, but lost its surely usefullness a while ago.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: