i've been using adversarial critique and reviews for many planning, solution design and implementation steps inside workflows.
It's so effective and helps catching so many design flaws, implementations misses etc ... that i'm wondering how people manage to build complex/large projects with agents without this kind of process. Well, i actually built this thing because i couldn't get good results so i had to find a way.
I'm gonna open source the whole thing but it needs some cleanup, there's a basic landing page here https://kodfactory.com if anyone wants to be notified when it's released on github. Yeah i know, the world really needs another software factory :-)
although i initially thought it didn't make sense financially to run this kind of model locally, i did run the numbers and for heavy users this could justify buying $10k worth of hardware with a ROI over a few months, less than a year.
I was looking at my token usage, mostly from subsidized codex/grok subscriptions and i'm a somewhat heavy user. The thing is i would actually use even more tokens if it wasn't for the weekly quotas.
In the end, with a $10k investment and running this kind of model, estimating a 2x increase in token usage because i wouldn't have weekly quotas and comparing to glm api prices, this thing could pay for itself in less than a year.
Obviously i'm paying subscription price right now, so the math doesn't work. Although using local ai removes all weekly quotas. Keep a subscription to have access to frontier models for planning work, and local hardware + glm-5.3 flash for implementation, e2e testing, qa work 24/7.
It's not that crazy of an idea and the numbers aren't that bad.
You aren't going to get nearly as much token usage locally from DGX Sparks or even M5 Ultra (though it might be close, unsure would need to get my mittens on it to clarify).
You will get around 2-4 concurrent streams of aggregate tokens at best for such a model and around 0.5B output tokens per month assuming you use loops and run it when you are sleeping. That's 500 (per mill) * 0.5$ = 250$ only at most.
Then there is maintanence and efficiency costs due to electricity usage and such, any down time, etc.
You will be lucky if you can squeeze more than 200$ of value out of it in a month.
I don't think people should buy local hardware for money reasons, by the time you will pay off a 10K USD machine, 2-3K USD machine will catch up and beat it by a significant margin.
Unless your expectation is that we will be in hardware winter for the next 10+ years. At 200$ per month it will take around 200 * 50 = 10k, that is, 50 months, so around 4-5 years.
Again assuming you are making the most of your hardware somehow, very hard to do in practice.
I don't recommend people to use compute as investment or payoff thing, but if you have the money to burn and can afford it why not, maybe with some software optimizations it will be cheaper but then again Z.ai is currently offering 50% discount and providers will offer cheaper rates for sure.
But either way you will never be able to burn more than 200$ worth of token on a cheap hardware device, because inference becomes more profitable the more you scale it up, you have separate prefill and decode engines/systems, and a lot of nuance, but assume for every 10x increase in infra you increase margins by 5-10%.
So from 10K to 100K to 1M to 10M to 100M.. I don't think this curve continues beyond 100M but I have no idea about that scale unless some AI lab is interested in hiring me lol.
So a 100M infra will have ~30% better margins than you at 10K, then there is software optimizations but that's cheap enough, though some of it is only viable at scale.
Either way assume 10K is the price of privacy if you really want to buy it. Don't worry about making the most out of the usage, you will always be in a net loss but I would assume for you 10K doesn't matter.
I generally agree — go local for the hobby/tinkering, privacy, and control (ie not getting refused by an AI to defend and secure your own network and codebase; as HuggingFace has seen).
But whether you make a loss or not depends on how hardware prices and resell values go though.
I have spent ~$50K on local AI hardware. The market value of that hardware is about ~$80K right now.
So the maths is working out for me so far. I see it as a call option on compute.
> I have spent ~$50K on local AI hardware. The market value of that hardware is about ~$80K right now.
That's easy because no matter what IT hardware you bought, it's worth more now than it was two years ago. That's something that's unprecedented, never happened before, and as soon as we get flood gates open on ram manufacturing OR when the AI bubble pops, all IT HW deprecation norms will return and making a profit by buying something IT will vanish.
I had a GPU server four years ago. Had I not sold it like three years ago with 2x price I bought it, it would be likely something like 5x the price nowadays.
I really really miss filling my home rack with old enterprise stuff. All I want is this hardware winter to end.
yeah i mostly agree, especially compared to subsidized subscription cost.
But for a heavy user who has enough work to be done so that the box runs almost 24/7 at say 50tok/sec, the math gets interesting against API prices.
And it can be interesting compared to subscription in the sense that you don't have the quota anymore. That means there's probably a lot of things you're not doing because of the quotas that you could do now.
It depends heavily on the tok/sec obviously and the very best solution financially remains subscriptions. But the idea remains entertaining and not that disconnected from reality
At 50tps for single stream you are going to get 50 * 60 * 60 * 24 * 30 = 130M out tokens of GLM 5.3 Flash...
That's less than what 40$ at current API rates... So if you are willing to pay 200$ per month you will get much better limits paying API rates.
You can't run large Kimi K3 models on 10K worth of hardware either way, you need to spend like 50K USD minimum.
Just pay for the API rates or get a low cost provider that uses higher batching, you can get shittier tps but much better prices, probably go as low as 20$ for as much usage as you can ever get from a 10K USD machine from GLM 5.3 Flash...
The issue is nothing expensive runs on these devices and cheap stuff isn't worth running locally, eletricity costs ~12cents/kwh in us iirc, so at 330W M5 Ultra will burn around 8 * 0.12 = ~1$ per day extra in electricity so the electricity is going to cost you the same as the API rates(30$ per month).
I truly don't think you are accounting for the costs here properly. But again if money truly doesn't matter it's much better for privacy and better than paying one of the shady AI labs who are doing god knows what with your data.
Your point isn't lost on me, but a few other considerations:
1) Rates are theoretically discounted for GLM 5.3 Flash right now, by 50%.
2) Hardware costs have continued ascending with no sign of letting off, so it's unlikely that a DGX Spark depreciates to zero in one year.
3) Compare performance in terms of difficult tasks/$ over the last 6 months, 3 months, etc. Open weights are a ratchet. In terms of intelligence per $, a Spark is never going to be a worse deal tomorrow than it is today, at least until the entire platform is replaced or obsoleted.
71 days ago the best model you could run on two Sparks was an aggressive Q3 quant of Qwen 3.5 397B (AA 34). 70 days ago it was a mixed-quant of GLM 5.2 (AA 53). 30 days ago it was full fat DeepSeek 4 Flash (AA 53). Today it's GLM 5.3 Flash (AA57) and/or Qwen 3.8 Next (Unknown). Sometime this week it will likely become mixed-quant GLM 5.3 (AA 60).
So in < 80 days we have almost doubled the benchmark score. And that curve is still accelerating. If you view it as "cost per token of model vs API" then yes it's a bad deal. If you view it as "cost of task per $" then it has almost doubled in value in less than 3 months. All of this, imo, API and hardware, is still massively underpriced.
> 2) Hardware costs have continued ascending with no sign of letting off, so it's unlikely that a DGX Spark depreciates to zero in one year.
If someone told me that costs for X will keep increasing because they have been increasing rapidly in the last 1.5 years, but they have a history of continuously decreasing for decades before that.
I am not sure if I will take anything they say serious, I am not sure if it's HN or AI but people are delusional if they think compute costs will keep increasing from now on...
Either AI will be really good, hence compute and everything will materially depreciate or it won't be much better than it is today and token volumes will plateau compared to compute.
For instance the amount of token compute that's to come online in 6-12 months is several times what we have today...
Second 3) Compare performance in terms of difficult tasks/$ over the last 6 months, 3 months, etc. Open weights are a ratchet. In terms of intelligence per $, a Spark is never going to be a worse deal tomorrow than it is today, at least until the entire platform is replaced or obsoleted.
This is a bad take because again this assumes DGX Spark will not depreciate in price, we will have something better for far cheaper surely in the next couple years. M5 Max & Ultra are already arguably it, but will have to see.
> 71 days ago the best model you could run on two Sparks was an aggressive Q3 quant of Qwen 3.5 397B (AA 34). 70 days ago it was a mixed-quant of GLM 5.2 (AA 53). 30 days ago it was full fat DeepSeek 4 Flash (AA 53). Today it's GLM 5.3 Flash (AA57) and/or Qwen 3.8 Next (Unknown). Sometime this week it will likely become mixed-quant GLM 5.3 (AA 60).
This has nothing to do with DGX Spark's value, if models get cheaper the API costs also go down, this is not a defensible argument to cost to value.
Are people on HN really not thinking straight?
Tldr; no matter how you do the math compute is only getting more valuable because of a temporary crunch, don't expect this to continue permanently, sure you maybe able to time it and make money but so could you in stocks this is not for investments. Further second hand hardware sells for cheaper than sticker price, outside of a bubble..
And models getting cheaper == APIs getting cheaper == your hardware becoming worse value as your electricity & maintanence costs still remain.
I am not saying local models don't have their place but if someone is trying to use this logic to justify their purchase then I wish them all the best, as someone who is actively working on AI compute/inference/hardware stuff I personally don't have this level of courage.
But this is not a sound investment strategy that if something is going up and seems like it might keep going up, especially when investing in heavily depreciating assets like compute.
i'm not sure why people expect agents to one shot everything to perfection with just a prompt.
There's a reason why we talk about software development lifecycle, design, architecture, testing ... It's because it's been the most reliable way to build and ship software. We shouldn't expect discard this and expect agents to perform well outside of this.
I'm treating LLM agents as junior devs who happen to have vast knowledge of software engineering. As their team leader i make them go through planning, implementation, bug sweeping cycles using strict workflows. And it works quite well, i've been working on several large projects (1M+ LOC java,typescript,c/c++) and by any measure the projects are healthy. Sure the code isn't that beautiful, sure i'd have written things differently but it's pretty good nonetheless.
Shameless plug here: i've been also working on https://kodfactory.com, the code factory i've built to work on these large projects with workflows, reviews, etc ... I'm cleaning things up to open source it later.
I think it's perfectly fair to evaluate the tools based on how well they live up to the hype that is being pumped out by the sellers of said tools. If they want us to compare their products to a more measured, reasonable take then they can advertise them as that.
While some people are busy bickering about this, the rest of us are using these awesome new tools to get more work done in less time with higher quality than ever.
I don't care what the company claims, I just use the tool the way I want to. I work very closely with the AI. I'll tell it to plan a change, review the plan, then execute. Then I'll test the changes and have it fix whatever I'm not happy with one thing at a time. I'll specify in detail both what to do and loosely describe how to do it or if I'm not sure I'll ask it to plan the change then review the plan and ask for changes if I want them etc. I also review my own PRs before I submit them to colleagues.
This way I maintain full control of everything, it just saves me hours of googling, planning and typing code - which I do miss a bit but I can't really justify writing code myself when I can achieve the same thing just by loosely describing my idea instead. It also saves a lot of time debugging, I think I'm generally a pretty good programmer but the AI makes fewer mistakes than me. It'll often catch some logic error I made during planning and suggest a good alternative.
A lot of developers seem to give up control entirely and then complain that they're no longer in control. Trying for that 10-100x speedup doing weeks of work in a day. I'm happy doing one week of work in a day. There's a limit to how much I can oversee without compromising quality.
anthropic specifically brags about how good claude code is every annoucement of a new model. I will surrender that none of them claim its "to perfection", but IMO its implied because no one would claim that their model one-shots any issue to dog shit quality.
no where does this document suggest that codex can "one shot everything to perfection with just a prompt". It describes using a prompt plus agent skills (which are essentially many other prompts) to develop a playable game.. nothing about it being perfect or anything more than being in a playable state.
Oh please, you sound so disingenuous, the claim wasn’t that the documents contained a specific phrase. You’re moving the goal posts. Their name is literally a play on anthropomorphising the models, such as… the ceo going on tv shows and repeatedly saying the models may be conscious and they may start nuclear wars, etc.
Seems like motte and bailey fallacy. They say their models are good (the motte), therefore their models must one-shot everything to perfection (the bailey).
Besides, other people's claims about something doesn't give you license to abandon all critical thinking. Though it's evident they don't claim what you say they are.
That's irrelevant. Questions was why people assume somthing, and answer is because that's how it advertised.
To be clear that's not what I'm thinking, even Fable 5 produces some hilariously bad results under some conditions and sonnet 5 produced great results under others.
But you didn't provide the evidence for that. You shared some links and then admitted they didn't claim it.
It kinda seems like "because I think they're a little too positive about their product, I can set my expectations to anything I want and la-la-la it's their fault."
And I don't see the problem with agents building test scaffolding as they go. It might be too defensive at times, like testing a shell script you don't run often, but big deal. It's kinda cool imo, and it's trivial to make it stop.
I don't really think they advertise "Create your app idea in one weekend night" and assume the general public will mentally add "... but hire an experienced developer to supervise the process".
I've seen a lot of cursor app, about some PO that has an idea for an app in the morning, then asking her agent to make it when her commuting, when arrive at office the app is done
Claude has a /goal function that explicitly says it'll carry on working until it's done what you prompted it to do. That's exactly what a 'AI will zero-shot anything' believer is looking for. Behind the scenes it's really multi-shotting with generated prompts, but the user won't care.
Using AI is kayfabe. What I mean is, you create interaction patterns that resemble how humans work. This is because it is what the models are trained on but also because we've all been trained to interact in this way. So it manipulates you into providing more useful prompts.
But I don't really want to play a part in a simulation, trying to cajole my scene partners into saying the lines I need them to say. I want to use a tool the same way I would use any other tool. If this is AI it should just do the thing. Anything else is an imperfection of the technology.
But at the same time, language is a vague communication medium. We have a precise language for describing forms of computation, but that's code so we're back at square one. We still haven't nailed the right amount of follow up and correction and interrupt-ability of these coding agents.
And we may never figure it out. It may simply be impossible. But it doesn't mean this weird anthropomorphization of AI is something I want to do. If I wanted to be a manager, I would be a manager.
> language is a vague communication medium. We have a precise language for describing forms of computation, but that's code so we're back at square one.
That's why you tell it how to write the code, and then review the code to ensure it is what you wanted.
> We still haven't nailed the right amount of follow up and correction
I feel like I've got it under control. It's not really a problem at all to me. I just work closely with the AI. I do small tasks, I don't just have it generate thousands of lines at once. I tell it what to do and how, or give it some vague guidance and ask it to make a plan. Then review the plan, ask for some changes if necessary and execute. Then I go over all the changes, test them to ensure they work properly, have it fix any issues I find and so on.
I think the main problem with AI coding is people try to do too much. You can't keep a tight leash on it while also having it do a week's worth of work in an hour. I do one task at a time and I am heavily involved in it, deciding exactly how it's done. I micromanage the crap out of that thing. I write commits myself and I always review my own PR before submitting it to colleagues.
Works great. I get things done much faster than I used to, with better quality than before.
It's an impossible goal. "Perfection" is subjective and situational. It's basically impossible to specify a non-trivial task perfectly, and without a perfect specification you can't have a consistently perfect result.
Personally I find it much more efficient to give vague instructions and refine on the way, rather than trying to specify everything up front. With this workflow there is no such thing as a one-shot, I don't even know all the details of my intended result until I reach it.
Sure it can one-shot many small things, but for a larger feature it has to be a incremental process. Even when I have a Figma design to work from it never contains all the details like semantics of how various interactive elements work, edge cases etc.
> i'm not sure why people expect agents to one shot everything to perfection with just a prompt.
They do often enough that it's not a surprising event, depending on prompt quality, context available, ability for the result to be objectively judged and iterate on by the agent, etc. For frontiers on very high settings at least.
If you explicitly ask the agent to make the perfect architecture for the problem and write it down in to a spec and have the developer agents follow it they will. Its just that coding agents have a hard time coding at think about architecture at the same time.
I can one shot a prompt if I write down a nice spec file, Claude can do a lot in one shot. I test it every few months. With enough detail Claude will know what to do.
Opus 5 one shot an access virus B synth clone for me as a single page index.html that is more impressive than anything I have seen as a VST synth.
That is also because I have been obsessed with this synth for almost 30 years. I built clones of it 20 years ago in reaktor. I know how to spec out every aspect of this synth and I gave Claude a 150 page pdf on digital filter design too.
The results are far different than someone who has never used a virus prompting "make me an access virus B synth as a single html page".
We are calling both of these processes "one shot" but this is not even close to the same process.
I suspect this is the LLM discourse in a nutshell. People are using the same vocabulary for wildly different processes.
reply