In another one, Opus 4.6 level already solved 90% of my work-day tasks, so while better models have been instrumental into handling a higher % that does not mean that defaulting on cheaper models can't be good.
I run DS 4.1 flash daily, and then cross check with gpt-6-astra and I've nuked 90% of my AI monthly bill while having higher limits and better performance/intelligence than I did just at the beginning of this summer.
Why would they need to make anthropic or openai panic?
Every non tech company I know is using Gemini or copilot, because the same companies already were on Google or Microsoft suite and got those as extensions.
NotebookLM is way more popular in the real world than anthropic work or crap like that.
In business world contracts, data retention and procurements are more important than made up benchmarks only nerds care for.
And they also tune the inference of the models behind it.
You're using different tooling every day. Hard to benchmark.
reply