Is this equivalent though? The Laya article ends with “ Treat Laya as a fast foundation model to specialize, not as an omniscient zero-shot oracle.”
I have a dozen different things at work that are currently using LLMs as classifiers for different questions. I don’t have the time, data, or resources to fine tune a model for each of them.
I haven’t had a chance to plug in Jev yet (waiting on approvals), but if it has the general intelligence claimed in the press release, then Laya is in no way comparable for my use case, and whatever TypeSafe has done is a substantial innovation over the Laya paper.
Jev seems pretty cool! I just got access and have only gotten to do minimal experiments, but I love this general area of research and it fills a very real need.
I agree with you. I think the OPs pushback is emblematic of a larger reaction I've seen that is, at the very least, misinformed.
There are a lot of approaches that use a self-attention backbone for classifier-style outputs. You have structured generation libraries like SGLang and Outlines, but those basically give you guided generation on an autoregressive model. You also have a bunch of models that are non-autoregressive that try something similar. Older NLP stuff applies here, and there's newer stuff using diffusion transformers for this purpose.
But I don't think the Jev author has ever said that he's the sole human, alone in a vast sea of misguided researchers, who is interested in schema-guided classification? I think he said he found a novel way to train a model for this task that has much higher general intelligence at much lower cost than other approaches. Which is an exciting result with lots of applications if it bears out.
I think some people are just reflexively skeptical of anything that gets a lot of hype. Maybe that's fair. Things that are wildly successful and high impact also tend to get a lot of hype though, so it seems like a poor filter.
In a world of agents, doing a BERT run takes about 2 hours from having an empty folder. Just a thought you could consider. Once you've done the first you can do the rest of them before the end of the work day.
The domain is code analysis, all languages and frameworks. It’s b2b SaaS, so total volume is not incredibly high. And many customers have contract clauses that we don’t train on their data.
I’m not convinced we could train easily here, or that it’s worth the investment compared to (previously) spending fractional cents on Luna, or now paying even less on Jev. Especially given that these numbers are not meaningful to our margins.
In TS/JS you’re usually inlining the reducer fn, and there’s something hard to read/especially ugly about the comma after the bracket or arrow fn into the initalValue.
That said, when I’m reducing a list, I still use reduce.
You’re not wrong, but I think the alternative future is sama loses control, goes and raises a quick $100M, and hires all the profit-motivated staff to make something that looks like today’s OpenAI.
Dario already showed there was an appetite for this.
Right. Much like the alternative future is that someone considering stealing a public museum could instead build a building, hire away the staff, and fill it with art he acquired, instead of just stealing a museum that was created via charitable donations.
Late reply, but I think once you took away the profit motivated people and the money man, the “museum” would have been exactly that - a collection of antiquated artifacts. But I guess we’ll never know.
I’m not saying it was right or moral, just probably inevitable.
If you can afford to eat the loss, don’t insure it.
Obviously insurance companies make money on average for every policy. If they didn’t, they would go bankrupt. So on average, every time you buy insurance, you lose.
So insure only what’s required by law, or things that would throw your life off track if they vanished.
Yep. Which is also why I find vision insurance ridiculous. No it doesn't cover me going blind, it covers a pair of glasses at best, more like 25% of one.
Dental yeah, vision is just a scam. There are also "life insurance" plans that are basically tax-advantage investment/inheritance funds if you have like $50M+.
That’s nice in theory, but as these things get better and cheaper this kind of capability is going to drop from nation states to script kiddies. That future is coming, I don’t see any way around it.
We can round up all the bored teenagers we want, but it’s not putting the genie back. Better start adjusting our systems to account for it.
Ever read the Anarchist Cookbook? Anybody tech inclined with a hint of mischief in them, from a certain era, has. It's a list of all sorts of awful things you can do, mostly with household ingredients, and a few minutes. I think its overall impact on society was pretty much zero. Actually it may have been overall positive because I expect plenty of peoples first experience with things like thermite came from that book, and now there are all sorts of videos and neat experiments with such on sites like YouTube.
I think this is in part because most people, including awful, tend to be relatively morally inclined. But I also think because even with an LLM, doing things takes effort. And if you're willing to dedicate effort towards a task, there tend to be way more rewarding/gratifying things to do than try to hurt people. Countries tend to be excessively sociopathic because you have large scale 'intelligence' organizations who see their entire point of existence as being to engage in misdeeds.
I suspect people didn't start blowing up stuff because they understood that would be bad, harmful, and also very illegal. Everything computer related somehow seems to feel less real or consequential to some people. And AI doesn't have this compunction at all unless we make really sure it does.
When I was a mischievous kid, me and all my mischievous friends had our stories of learning how serious fire and explosions were considered by authority figures. We learned fast not to do that or the consequences would be grand. These were usually small fires or firework involved pranks. So, yeah I agree with this.
Computer stuff has generally always been a slap on the wrist in comparison. Maybe it’s more punitive now. But also, it’s one of those things that maybe you get in trouble officially but at home and behind the scenes you’re friends and maybe your dad are laughing and giving you high fives. So young mischievous kids will totally go there because they’re not afraid of punishment if it is minor and it gives them a notch on their belt. If they can take down Amazon.com website for a day, we all know that’s a massive financial implication, but it’s also a faceless mega corp and quite tempting if you can get the bragging rights with only risk of a small punishment. (Note; I don’t know what the current crime/punishment for this would be, and whether it’s small is very subjective).
It’s similar to how some people gravitate or succumb to the opportunity of white collar crimes. Embezzling $10m from a company almost makes sense in a situation where that only gets you 5 years max prison. If you hide it well, you simply serve your time, and then retire in comfort. I can see how that makes more sense or is tempting to people than slogging through a lifetime of low income job as a bookkeeper just trying to find a way to save for retirement.
Most of these people would never consider robbing a bank. First of all, it’s not a $10m dollar opportunity. Usually not enough for anyone to retire on, or live more than a year or two really. Second, it’s usually considered a much more severe crime and sentencing can be very long, I’ve seen 30+ years. Third, it’s much more risky to your person. Getting shot and dying is absolutely possible.
> Countries tend to be excessively sociopathic because you have large scale 'intelligence' organizations who see their entire point of existence as being to engage in misdeeds.
This theory interests me.
I'd love to understand how different individuals within intelligence orgs have reasoned about the morality of their actions.
The issue is that they just started forcing sign-ins on Old Reddit, allegedly because it's easier for bots to scrape. The ability to casually/anonymously peek at a post that answers your question is getting more rare.
I was at Dropbox from 2016-2020. We were certainly trying to build a sustainable business, but there was a major identity crisis. Were we consumer web? Buy Mailbox and build Carousel, then shut them both down.
Maybe we’re Notion/Evernote? Buy Hackpad, plow a ton of money into Paper (which was legitimately good), then quietly deprioritize it.
Maybe we’re actually some kind of enterprise document productivity suite? Buy HelloSign. Plow a bunch of money into a desktop app. Pull more plugs.
A lot of smart people were trying. We made a lot of bets (too many?). None of them proved to be a second act, and the competitors eventually caught up.
I intend to write a piece going deeper into the failures; I would love to chat if you're open to it. Also, not to say Dropbox sucked or anything, it is just that the broader strategy and industry structure make it hard; if anything, the success of the first product made it difficult to evolve the business.
Fair (I haven't been using the encrypted-reasoning systems, though this is common in open ones - I'm kinda surprised it's an option in encrypted ones too), though what they're doing here is cross-user replays in addition to cross-model.
AFAIK no provider guarantees compatibility of reasoning traces, even in the same model generation, and we've in practice seen most of the big LLM APIs throw errors indicating incompatibility (at least transiently) when switching models. The only stable solution right now is to just throw away reasoning traces whenever a model is switched.
There seems to be an obvious choice to make here, should you give the users to decrypt and use the COT that they did not generate themselves?
This is only required if you want users to be able to share things with everyone and you are going for the simplest implementation.
If not you could try to keep a record of keys associated with a user, then when a new request comes in look through to see if the user has a valid key to decrypt the COT.
For explicit shares, just add the key used in that one conversation to the users valid keys. For global shares use the global keys. But that's adding more complexity to the system.
It’s about being able to change models mid-task. For example, I want to be able to plan using Fable but implement the plan using Sonnet, and that won’t work if this is implemented.
Star Trek is a post-scarcity society, those problems don't exist there unless you're in the middle of a crisis and on emergency power.
LLMs briefly seemed like this too, after subscriptions made the SOTA models too cheap to meter, but before they walked back on that and introduced quotas...
Actually, that brings up a good reason they can't fix it. Fable falls back to Opus when the topic is too "unsafe". That behavior requires traces than can move between models!
For plan it's relatively easy, just make the plan the artifact. The point is to ingest knowledge with one model and use it in another, and that is not necessarily easily expressible in natural language.
I suspect that there are companies with internal proxies that load-balance across keys, and they didn't want to break that when adding encrypted reasoning.
I really don't understand why server-side storage of the trace isn't a viable approach here, with only a unique key flowing to the client and back. Does it have something to do with how backend load-balancing works?
Makes no difference. There is a policy as to whether to allow use of a reasoning trace in a given context. Whether that trace originates from authenticated ciphertext or a backend database is basically irrelevant.
Yes, this storage would be growing exponentially making the disk space and latency problems harder (add the disaster recovery/backups). I think the choice of using client side is not too bad if you ensure that its secured properly. Also the company can excuse itself from the liability of storing sensitive data on its servers, thats a big deal in itself to be compliant for enterprise audits
1. The down side is that it cannot be used across the clients even for the same user
2. Using the same encryption key was a bad choice here, a per user key would have solved this issue for sure.
Having thought about this a little more, it's clear that server-side storage is not compatible with Zero Data Retention (ZDR). However, in non-ZDR settings, it seems likely that the providers are capturing all that data anyway?
> a per user key would have solved this issue for sure
It would have helped with PII leakage, but not with plain-text trace extraction attacks, right?
Per user encryption key ties it with the user session (assuming you do authentication properly), no one else can access it. User being able to see the information is not really an attack vector in this case.
The compliance rules at times are outdated and people skirt around them by following the worded rule instead of the intent.
Encrypted state cookies solve real problems (server-side storage, latency, scaling) and are not the problem. The problem is insufficient binding of some of a session's encrypted state cookies and others -- insufficient binding of some session state to other session state. Here we have HTTP encrypted state cookies for identifying authenticate user IDs and maybe for identifying sessions / chats, while the reasoning traces are also encrypted state cookies but not HTTP cookies, and the latter are somehow not sufficiently bound to the former.
The fix is to either have per-user or per-session keys for encrypting reasoning traces, or write the user ID / account ID and maybe also session ID into the plaintext of the reasoning trace _then check that that matches the ones in the HTTP cookies when decrypting the traces_.
Honestly, GPT 5.6 Luna is worth a look. It’s a reasonably good implementer at a small fraction of the cost. $100 buys a heck of a lot of it at API pricing.
Not sure about the OAi Pro plan, doesn’t look like the 80% Luna price slash made its way into the quota system.
You could also try tuning down the effort level on Opus. It makes a huge difference in token consumption and you might be able to get away with lower than you’ve set
At the beginning of the pandemic all the mills cut production anticipating economic collapse that didn’t come. Prices spiked on high demand and low production.
Mills did ramp back up, but it’s unclear to me if they used it as a chance to do so slowly/preserve margins. Lumber never got close to pre-pandemic levels.
Tariffs probably also play a role here. About a quarter of US lumber comes from Canada, and barbed wire is just steel with a little bit of processing.
For products with little value add there’s not anywhere for the tax to be absorbed, and no real way for domestic producers to quickly scale up, even if they wanted to.
Part of the problem during Covid in the US was US lumber mills are set up to do trees of a certain size. If the trees are harvested late they're too big, so we had to send a lot of lumber to Canada where it could be processed.
I have a dozen different things at work that are currently using LLMs as classifiers for different questions. I don’t have the time, data, or resources to fine tune a model for each of them.
I haven’t had a chance to plug in Jev yet (waiting on approvals), but if it has the general intelligence claimed in the press release, then Laya is in no way comparable for my use case, and whatever TypeSafe has done is a substantial innovation over the Laya paper.
reply