Hacker Newsnew | past | comments | ask | show | jobs | submit | epolanski's commentslogin

Okay but...closed source harnesses change multiple times per week, sometimes per day.

And they also tune the inference of the models behind it.

You're using different tooling every day. Hard to benchmark.


From my experience DS 4.1 flash is a much more capable model than luna.

What's a command code plan?

That's because Opus 4.6 was the last good assistant model.

Everything after it might be more "intelligent" but is super tuned around end-to-end task (and related benchmarks), not to act as an assistant.

Now it's *you* being the assistant, reviewer, etc.


In one sense you're right.

In another one, Opus 4.6 level already solved 90% of my work-day tasks, so while better models have been instrumental into handling a higher % that does not mean that defaulting on cheaper models can't be good.

I run DS 4.1 flash daily, and then cross check with gpt-6-astra and I've nuked 90% of my AI monthly bill while having higher limits and better performance/intelligence than I did just at the beginning of this summer.


Why would they need to make anthropic or openai panic?

Every non tech company I know is using Gemini or copilot, because the same companies already were on Google or Microsoft suite and got those as extensions.

NotebookLM is way more popular in the real world than anthropic work or crap like that.

In business world contracts, data retention and procurements are more important than made up benchmarks only nerds care for.


Just watching the video grinds my nerves for how obnoxious it is.

Then write a note, stop recording me.

If you can't be bothered, it wasn't that meaningful or important.


The whole modern world is dystopian garbage.

Feels like I struggle to wake up to a good news since a decade.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: