Hacker Newsnew | past | comments | ask | show | jobs | submit | aix1's commentslogin

What makes you conclude that?

Lack of releases relative to Gemini. Not that I expect them to take down the old versions, just unsure about new ones.

I just finished listening to A Man for All Markets (narrated by Thorp himself!)

Bitcoin derivatives arbitrage doesn't sound like something I'd personally be comfortable getting into, but I'd be curious to hear more about the types of things you're doing (to the extent you're happy to share).


Could someone explain the appeal of flow typing?

I can see how it can be useful to start with a broad type, e.g. a union, and narrow it down in a block. However, I don't quite get the opposite direction shown in their example (first an int, then a string, then a union).


Same rationale as flow valuing. Some people like values being reassigned, and some people like types being reassigned.

You might be reading too much into the union example. The checker just doesn't know if the middle block ran, so maybe it remained an int, or maybe it became a string.


Did you do a pixel diff against the original? The output could look plausible at a glance but contain hard-to-spot errors or inaccuracies.

Unchecked OCR can lead to amusing results: https://news.ycombinator.com/item?id=18374359

Are you saying the models are already autonomously constructing next-gen evals for themselves? (Which is what the GP is asking.)

Sure?

Doesn't everyone get their agents to construct evals it can't pass? There's nothing magical about this.


Would love to learn more about some techniques that "everybody" uses to do this well. So far, everything I've seen that meaningfully advances the frontier has been high-touch (involving human experts in one way or another).

It's fairly easy to describe a task that is slightly harder than an existing one.

For example if frontier models are able to one-shot a database query across 20 columns and 10 tables add one additional relationship then test. Keep doing this until the pass-rate drops below acceptable and now you have your new frontier eval.


I see, we're talking about different things.

My thought experiment was along the lines of "Let's say I'm Anthropic and I want to significantly improve my frontier model's performance on, say, theoretical physics research. How do I build a fully autonomous process capable of constructing an eval that's somewhat outside the current capability in some useful direction (decided by the autonomous process itself)?"

Would love to hear folks' ideas. :)


Also firmware updates (on the non-locked variants), no?

What other potential uses are there? Uploading books without WiFi is the only one that comes to mind.


As far as I can tell, firmware updates are only via WiFi.


One thing that jumped out at me is a slower-than-I'd-expect rate of adoption of Opus 5. If you're running Opus 4.8 and choosing not to migrate to 5, what's behind that decision?


These are GPU rental price futures, so you won't be investing in the hardware.

The pitch is that, if you know you'll need a certain amount of compute at a certain point in the future, you can lock in the rental price today.

https://www.cmegroup.com/media-room/press-releases/2026/8/11...


What have you been observing? Genuinely curious; I use Fable as my main model and haven't noticed any regression.


Yes they do. They have to file an S-1 with the SEC, which will be made public about a month before the IPO.

The S-1 has to include, among other things, three years of audited financial statements, plus interim statements (unaudited). It will cover both revenue and expenses, the latter breaking out things like cost of revenue, R&D, sales and marketing etc.

Based on the (unofficial but reported) IPO target date of late Sep to early Oct, the S-1 will have to be made public in a few weeks from now.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: