Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

i really hope it's just what Deepseek V4 does. Deepseek V4 is very cheap and highly performant

OpenAI tried to pull off the same trade secret thing with RL when they announced o1 and o3, aka "Compute time scaling". Then Deepseek revealed it with Deepseek R1.

Could also be something like Deepseek DSpark. Or using diffusion like DiffusionGemma as a draft model. The timing between the release of those, and this article, makes me think its maybe one or both of those things



deep down, i suspect they're all just drafting on implementations to llamacpp.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: