Hacker Newsnew | past | comments | ask | show | jobs | submit | iot_devs's commentslogin

> Originally this was caused by an Istio sidecar pod reaching its concurrency limits and failing to auto scale correctly because of a misconfigured policy that watched host service but not sidecar limits.

I operated services at similar scale, and generally we use to put a bit of slack so that you would get an alarm when capacity goes up to 80%+ (or whatever number makes sense)

This allows to check, in the morning, after coffee, why the load balancer fleet didn't scale up automatically.

I am sure there is a good answer to why this is impractical, but it would be nice to know


I personally read the thinking traces to know if the model is on the right direction


I'm not saying they're not useful, of course they are. I am disputing that they are part of the agreed bargain between you and the proprietary LLM providers.

They explicitly do not promise reasoning traces. You (general you) agree to those terms and pay for that bargain anyways.


We "agree" to many things that are deeply unfair.


Yet we have the option to decide not to participate. That is an option.


I’m just driving by here but they bill by tokens — it’s a stretch to turn around and deny your right to see them. And it’s especially egregious when the tokens admittedly, routinely do the opposite of what you instructed.

But personally it’s not about right and won’t it’s just blatant bullshit.


And lawyers bill by 6-minute increments, yet that doesn't mean you get access to all of a law firm's internal discussions and notes about you and your case.

Just because you paid for the lawyer time/LLM tokens doesn't mean you get access to everything that happened within that time/tokens.


A writing app.

The overall idea is to work/write at an higher level of abstraction where you work with "ideas" (just small paragraphs).

You can move them around and basically restructure many/whole paragraphs by simply changing few words.

The final text is generated by an LLM.

LLM that has the whole context so you can ask it questions, but more importantly it can provide you feedback


I put it online and accessible for people to try! Any feedback is welcome!

https://structo.redbeardlab.com/


How is ironcalc different from other spreadsheets?


From Excel/Google sheets because ot is MIT/Apache 2.

From libre office because it is web first and also an engine first

It's written in Rust and has bindings to different languages. It's extremely light and fast


Absolutely!

Also if this scale anywhere close to LLMs in 2 years they will fold at super human speed and much better than me.


The real question is whether a hardware platform today would be good enough.

The model is software on a server or in the cloud, so the real question is are there any show stoppers from the hardware for going faster?


Physics. Right now the clothes folding is computationally limited, but assuming we get faster computers to get them folding faster, gravity and air resistance are still going to take their time.


The caption for the WAS aware clearly mention 52 US states.

Is that a typo or some quirk of radio amateur?


I believe that's a mistake. The page of the award mentions only 50 states: https://www.arrl.org/was


It seems to be a typo, unless the criteria for the award have changed: https://www.arrl.org/was


Or perhaps incredibly prescient!


I am afraid that this post is missing the biggest point.

Given a prompt (or task) how do I evaluate if it is a "simple" task that should be executed by a small model or if it is a complex one that may need a SOTA model?

Are you guys using heuristics? Which one?


Feel free to take a look at the docs for how routing decisions are made: https://role-model.dev/

When you use Pi with pi-role-model, Pi will include task and role metadata with its request; the role-model router runtime additionally holds benchmark and observability data, and a configured routing strategy. A composite of this is used to make the decision.

What you are pointing out is correct: making the actual decision and ensuring it is accurate can be difficult, which is why you need rich data as above, and also a model pool where each model is distinct, as I wrote in this post and in one of the comments below.


Do you have any example?


I would love if you could make some examples


What did you learn so far?

I am interested in a similar tool and it would be nice to skip some of the learning


Where to begin? Too much to reply to here. Fundamentally, coding agents are just a loop. The devil is in the details.

I will eventually get something in my blog about this project.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: