The technology is an agent you name. You give it a face. Maybe tune the personality. And, it only gets useful as you grant it more and more access to your accounts.
The technology is produced by a public company. And, in this case, the company is useful context for this story since it has just recently agreed to pay about $18 billion to resolve a multistate lawsuit claiming it intentionally designed addictive platforms that harmed young people's mental health.
You cannot detach the technology from the company. It's definitely not a neutral artifact.
I agree with the sentiment of the post, and Github deserves every bit of it. Tried codeberg (which uses forgejo) and I am very happy with it personally.
Sal could not get to his cousin's house every afternoon to do math with her, so he recorded the sessions. She said she preferred the recordings, because she could rewind them and because he wasn't sitting there watching her not get it. He put them on the internet. Then he spent about fifteen years building an exercise engine, a "mastery" model, and a teacher dashboard around them, and told everyone who would listen that the point was to move the lecture out of the classroom so the classroom could be used for something better.
And the essay's premise is: this man wants your child to watch a video.
How many of those automated fixes were reverted? How many introduced a new bug? What's the false positive rate on the finding agents? The post has counts for everything that went right and nothing for what could go wrong.
Exactly, also it lacks a lot of context as to why they did that.
Possible (probable?) scenario:
- Marketing: "we found and fixed lots of bugs thanks to AI"
- Reality: the KPI is now to fix as many bugs as possible with the help of AI, so they used AI to search old and easy bugs in the backlog, and then fixed it manually
Huh? Maybe a third scenario is that AI helped fix the bugs? Like described in the actual post you are replying to? This level of conspiracy theory is getting a bit ridiculous.
I think you’re right, my base assumption is that the models can code and can fix bugs, and can code more in parallel and faster than humans at a lower cost.
If Google are tackling lower value bugs with AI the number is in a way inflated compared to some utility measure (fixing a smaller number of worse bugs could be preferable) but it’s still things fixed.
There's no reasoning with the folks who are anti-ai and claim LLM's cannot produce anything worthwhile. They are basically flat-earthers or anti-vaxxers at this point. There is literally no amount of evidence that will convince them. They will always say that LLMs cannot produce anything but garbage just like the anti-vaxxers who state vaccines just cause autism. They are both fucking morons who no one with any grip on reality should indulge.
At my company there's a lot of discussion about AI and complaints that people run out of tokens within a day, but zero results are shown. No measurable (or measured) gains. Or nothing that people are willing to talk about, in any case.
At my company the performance improvements channels has been exploding, with people claiming giant improvements in latency, throughput, and decreased cost of the services. The cost decreases itself is order of magnitude more (at annualized run rate) more than we pay for tokens. YMMV.
I definitely make use of AI but in my experience I almost always could have done it better myself, the places where I threw AI at the problem I didn't care about the results being good, only good enough.
When we see memory and compute requirements for version x+1 of software decrease instead of increase I will happily say AI is the oracle people proclaim it to be.
In $COMPANY, for the mid-yearly review, the employees were asked whether their AI usage was 1/ efficient, 2/ adoptive (integrated in the way they work) or 3/ transformative.
Saying “never used it” or “I tried and it was useless” was literally not possible.
You'd have to be a pretty stubborn software engineer to not find AI useful for anything at this point, unless they're making you use Github Copilot or something far behind the frontier. You don't have repetitive tests to write? You never need to write a script to run something?
If you were the CEO of Amazon, would you be setting up channels for people to talk about their AI failures? The general way technology is deployed is that we try to find ways to make it work, because those are the most interesting. We're not as interested in all the ways it doesn't work.
From my perspective, some people are trying to use AI in the same way somebody might use a laptop to paddle a canoe. Sure, you can do it, but it's not a good idea. The fact that it doesn't work well is not particularly interesting.
If I steelman your position, I guess the ideal repository would be a set of cases where it's known to work well and a set of cases where it's known to not work well. A little bit like ProtonDB or SteamDB, perhaps.
Sure it is. It's just like the .com bust: people didn't know what the internet was good for, so they did a lot of stupid stuff with it before settling on the most interesting uses.
No, AI is more like: people do know what it is good for, but that's immaterial because hundreds of billions of dollars of debt were poured into it, such that the only glimmer of hope of getting an ROI out of it is force-feeding AI everywhere, good or bad, and normal people are sick of that.
That's fair. I still think there big similarities (people with money seeing it as the future and investing heavily, even if in reckless ways), but I agree there's a forcefulness to it this time that wasn't present before, for the reasons you outline.
> If you were the CEO of Amazon, would you be setting up channels for people to talk about their AI failures?
I would set up channels to talk about AI, encouraging both successes and failures, with proofs required for both, and punishing people who intentionally misreport on either.
That's fine, lies are more harmful than anecdotes are beneficial. There's no forum to discuss the best IDE and people still manage, in reality there's no need for a forum around AI.
Well, you're at Amazon. Is the Dev Improvement discussion group still around? I feel like that was a pretty good place to learn about neat stuff (in my case, Emacs-related). I enjoyed the discussion.
I'd say that one is not really an issue. In 2012 the Pinkie Pie exploit chain already required chaining 6 bugs to lead to an exploit [1]. Since then we've seen chains requiring more than 10 bugs (!).
If you fix any one of those bugs, the exploit is non-functional anymore. Sorry out of luck.
So if, say, for every ten bugs you fix, you introduce two new ones then it's still a very net win. Unless of course it introduces a bug so bad it becomes a simple exploit not requiring a long chain of exploits.
But in the case of browsers we've only ever been moving to longer and longer chains of exploits required to pwn a browser.
A great many window of opportunities are closing for dark-side hackers / north korean intelligence etc.: there were probably exploit chains still open for exploitation in April that just got closed by Google.
If anything, besides the supply chains attacks in amateur-land, the world didn't stop working: projects (not just browsers but OSes too) are being hardened left and right.
Using AI to find potential bugs is an amazing use case and there really aren't many downsides.
> The post has counts for everything that went right and nothing for what could go wrong.
I'm not saying there aren't a few downsides but the benefits are just too good to ignore.
This is the thing the anti-ai zealots will never admit. Humans fucking suck at writing code. They talk about software development as if it's only ever performed by the top 1% of the top 1%. They never acknowledge that humans make mistakes. No. Humans create perfect code every fucking time while LLM's only produce slop. It's such a fucking mind-numbingly stupid position that I have to think they have never actually worked in an organization which produced code as a value. These fucking morons want to pretend a human never introduces a memory leak when it's plain as fucking day that vulnerabilities that these "expert programmers" introduce to software are rampant. But no. The anti-ai zealot likes to pretend that only LLMs ever produce bad code. The only bugs in existence are due to LLM slop and not the literal decades of fucking slop produced by humans without any LLM assistance. It's so fucking tired at this point. They will never be able to admit that current day LLMs are far beyond the median human programmer. Their opinions are fucking useless at this point. These idiots think The Daily WTF started with LLM code. They are basically anti-vaxxers wrt to their grasp on reality. No one should take them seriously to any degree.
Or people don't ask questions for which the answer is known to be "None, really."
I get that many don't like what LLMs are doing to the industry, but this is just incorrect reaction to a very specific benefit that's proven beyond doubt (Security hardening).
Accept it imo - LLMs are solving very large problems that have plagued software security.
Sorry it is your reaction that is really weird. These are legitimate questions, and I have definitely seen Claude finding the wrong cause and then implementing completely incorrect/irrelevant fixes, only to find that it didn't work and need to start over. Not saying humans don't do the same thing, but LLMs are far from perfect, and it would be delusional to only talk about successes.
I've done security work before, and I've done it now with frontier. I've seen the difference first hand, and as the OP and so many other articles show, so have the leading experts in the world.
So either they are lying, or you may not yet be seeing and experiencing what they are. If that's wierd, ok.
You did not address my points or those in OP but kept repeating yours (that are not even relevant) in a handwavy way.
And that's fine. You can continue to live in your bubble. Others simply have different experience, and your asserting "AI is perfect" on a forum is not going to matter in terms of everyone's own, first hand experience.
The "soul's crushing work" they are crying about is actually not enough of work or not enough interesting work. Not the sole volume. They sound like angry kids.
You're asking people to trust you and hand their codebase/IP to your tool while showing them exactly how you treat other people's code/licenses by "deciding" to not carry forward the GPL license.
This is coming from a cofounder at github, someone who probably knows precisely what the GPL is for. Whatever the legal merits, building on a GPL3 project's complete test suite and relicensing under MIT is not acting in good faith toward the original authors. I really find it disgusting and it makes me want to avoid gitbutler entirely.
I think you're saying that you don't believe in the freedoms to use the GPL licensed test suite for certain purposes which are explicitly allowed by the GPL.
You don't get to choose a license and then add extra terms to it when you don't feel like it's up to scratch. That's something explicitly not allowed by the GPL license.
> Where does the GPL say you have the freedom to relicense code or derivatives under MIT by fiat?
The first part of this sentence (where in the GPL) is unreached if the second part of it is unmet (relicense code or derivatives) which I contend it likely is. You're begging the question.
However:
> The output from running a covered work is covered by this License only if the output, given its content, constitutes a covered work
earlier:
> A “covered work” means either the unmodified Program or a work based on the Program.
It's that element that would be difficult to prove "work based on the Program"
Asking an LLM "here's a thing, rewrite it in Rust" is pretty clearly creating either a derivative work or a different form of the same work, just like asking a transpiler would.
There's no evidence that "here's a thing, rewrite it in Rust" is the technique Scott used here.
"here's a test suite, write code in rust that makes that suite pass" is reasonably supported by the article. That would likely not be a derivative work.
Ew. So it tells the LLM where the git source is for the thing they’re duplicating, but I don’t see instructions saying not to read or copy those files or algorithms.
I could have missed them. I didn’t read everything. I did some quick searches.
But the fact they’re not obvious is kind of troubling. Or that they didn’t just copy the tests and documentation for the LLM and not the source to prevent it from looking would hurt any case they had for clean-room privileges in my eyes, ignoring my other comment with concerns about using the tests at all.
If we assume an u licensed or MIT licensed test suite, an LLM could develop from that and documentation and you’d get something you could license MIT.
IMO, IANAL, etc.
And we’ll ignore the question of what the fact the LLM has certainly seen the git code during training means.
But the test suite would have to stay under the original license. And if you use a GPL test suite as they kernel to develop a program from can you license it non-GPL? I’d question that personally. Same acronyms above apply.
This is the exact thing I'm not sure about. See https://news.ycombinator.com/item?id=48470397 where I posit a simpler question: if a `test_sum()` function is copyrighted, does writing a `sum(a, b)` function infringe on the copyright of the software product that `test_sum()` is a part of. I'd say no. There's another part of the GPL that applies here:
> A compilation of a covered work with other separate and independent works, which are not by their nature extensions of the covered work, and which are not combined with it such as to form a larger program, in or on a volume of a storage or distribution medium, is called an “aggregate” if the compilation and its resulting copyright are not used to limit the access or legal rights of the compilation's users beyond what the individual works permit. Inclusion of a covered work in an aggregate does not cause this License to apply to the other parts of the aggregate.
So assuming that sum(a, b) is non-infringing and not combined to form a larger program (i.e. the tests aren't compiled into the grit code), then the GPL explicitly doesn't apply to this use
test_sum is assumedly relatively trivial. So as a lay person I’d expect some sort of obviousness test to apply. Like so much of the stuff in the Google/Oracle lawsuit.
But if you take all the individual tests used to test git as a whole, that seems far more unique. Seems like at that point you’re really having to duplicate the actual git internals, and that seems like it should be covered.
> test_sum is assumedly relatively trivial. So as a lay person I’d expect some sort of obviousness test to apply. Like so much of the stuff in the Google/Oracle lawsuit.
Feel free to extrapolate to the threshold where it's not and at that point apply.
> you’re really having to duplicate the actual git internals
Copyright covers the expression, not the method. So the Rust function:
The technology is produced by a public company. And, in this case, the company is useful context for this story since it has just recently agreed to pay about $18 billion to resolve a multistate lawsuit claiming it intentionally designed addictive platforms that harmed young people's mental health.
You cannot detach the technology from the company. It's definitely not a neutral artifact.