Hacker Newsnew | past | comments | ask | show | jobs | submit | ppeetteerr's commentslogin

CEOs are motivated by money and power. They will say anything to get either. If they truly wanted to slow down research, they would spend lobby money on the candidates who could enact laws.

Instead, they are partnering with some of the most disliked companies (x, meta) to further their goals.


Anthropic supported Alex bores who was one of the most AI safety pilled politicians https://www.theguardian.com/us-news/2026/jun/24/big-tech-new...

Good. Keep spending. Keep advocating.

There is a real possibility that either something you saw triggered the conversation and/or that you remember the ad because you spoke about it. It’s enough to linger on an ad for the platform to recognize your interest and serve you more ads like it.

Gruber was a good voice in the industry but this article misses the mark in a lot of ways.

A company the size of Anthropic would not voluntarily jeopardize their massive valuation if they didn’t feel the resulting output would maintain a similar level of quality as before. Is there a similar worry that their system prompt, which is injected at the start of every conversation also influences token generation in an artificial way?

If regulation will ruin Claude as a product, market forces will fill the void. There are also a ton of open weight models to choose from. It’s going to be okay.


Gruber is not much of a details man - he helped invent Markdown (to be lauded) but ghosted its standardisation. I would be fascinated to hear Prod John MacFarlane of UC Berkeley's opinion on it all given he was heavily involved in the push to get Markdown standardised.


I hate to be that guy, but I'm not reading a readme written by AI, nor using their tool. You don't care to put in a few hours to describe your project, I'm not going to bother learning about it.


Read through it an I'm curious whether setting the date and cmd on every system prompt call will cause the cache to invalidate.

I guess the cache would only be invalid if the day changed or the root directory, which would technically happen infrequently enough.


I get 95% or more cache hit rate with pi and DeepSeek or MiMo so it doesn't invalidate.

But I'll investigate how that works in a session. You got me curious.


Isn't Apple about to license some variation of this from google for on-device AI? Maybe it’s their sales pitch to Apple and then they will lock it down.


Engineers are not going anywhere, they are going to fill the spaces outside of coding that are still critical to shipping new product.

Here is a post that summarizes what I mean: https://substack.com/home/post/p-200064883


Traditionally those roles of providing "alignment" to overcome organizational inertia were held by program managers. So that seems like IC roles are indeed going away, to be replaced by more "technical program manager" type roles.


Maybe it's my experience, but TPMs were often responsible for coordinating large org-wide or cross-org initiatives. It's prohibitively expensive to have TPMs on anything smaller.

Since engineering with AI is still very technical, I would wager that software engineers would stretch into less technical areas of software development rather than TPMs stretching into technical areas. I only say this as someone with experience with AI and I see how easy it is to write bad code with AI if you're not aware of what it's doing.


Yes, size and performance are not only problems for local LLMs, they are problems for frontier LLM companies like OpenAI and Anthropic. The latter still lose a ton of money on inference and advances in efficient, performant models helps their bottom line.


The argument in the article is backwards. Evals test the stability and boundaries of a concept. They are not created before the concept has been prototyped (which the author acknowledges).

An eval is not somehow breaking silently due to some new capabilities in an LLM. It wouldn't be a good eval if it did. What it does is steer the LLM towards specific goals. If anything, an argument can be made that they restrict creativity and experimentation by narrowing goals.

If the argument is that evals need to written before some new behavior can be devised, that's incorrect. There are an infinite number of evals that test for things which cannot be done. Only when something has been demonstrated to work in a specific context, can an eval be written.


Most of these are addressed?


They are addressed but the core of the thesis is still wrong:

> This is the core problem: our entire evaluation infrastructure is structurally reactive. We measure the system after it has changed. We never predict the change.

That's kind of the point of evals.


How does this compare to using Claude Web with connectors to build the same feature?

On a separate note, READMEs written by AI are unpleasant to read. It would be great if they were written by a human for humans.


The main difference is that you have full control over this!


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: