I also have found that AI has not lived up to many of its promises and have dialed back what AI gets control of. My projects were turning into unmaintainable messes. The people who say coding is solved aren't paying attention.
Yes. I started two projects with AI from scratch. Both abandoned, complete mess. Projects without AI are so easy to manage, maintain, add/remove features, etc. I use AI as a search engine on my projects instead of Google. I ask what's wrong with my code and change or improve it myself based on my experience.
I have the opposite experience. I don't have the patience sometimes to clean up my code and stick to coherent conventions and organization even though the will is there. With my AI projects I watch it and if the AI starts drifting I ask it to go through and look for convention/directory structure violations and it happily cleans everything up in about 10 or so minutes.
Are you in Python by chance? Python has a lot of crazy hidden/inexplicit/spooky action at a distance stuff (especially in the frameworks) that can make LLMs gunk up code by defensively programming or just burn context chasing data provenance
I've been a vibe-coding skeptic for years, but because of the math breakthroughs of the past few weeks I decided to experiment with the latest models on some test projects. They're a lot more capable than I thought they would be. I agree that it's easy to create an unrecoverable mess, especially when you're one-shotting a lot of features without detailed instructions. But I find that as long as I'm strict about the API boundaries and force the agent to work in small chunks, it's pretty effective. As one example, I got it to write an SVG renderer in a few hours (not the whole spec, but most of the path features and text rendering), which would have taken me at least a week just for the coding part, plus extra time to learn the algorithms.
I’ve had the best luck by spending quite a bit of time going over the big picture architecture up front and then diving into the modules to further refine the details, making sure to generate step-by-step chunks of work in Markdown format for implementation. I’ll spend literally a couple of days doing this before starting any coding.
Edit: My latest project is all GPT-6 Astra High. It takes a lot of steering to keep it from adding a bunch of, while useful, features that are not strictly enough to the point. That main issue is it’ll use a lot of extra tokens in the process!
Sure, happy to provide you with an example of how to hold it (turns out Steve was right) =D
https://github.com/NousResearch/hermes-agent is 99% (just a guess) LLM generated. 1140 closed pull requests this week. 1.5k closed issues. The github insights page for commits doesn't load for me presumably because it can't handle this scale of commits. But I estimate ~1K commits per day on average.
You're using quantity metrics to answer a quality question.
I had a look at the kind of issues that are reported at that project (there's 15k of them, so I can at best assess a couple). It looks like a complete mess: A lot of TOCTOU, concurrency, and resource mismanagement issues and edge cases that in a better-managed project would have been avoided by construction. They will now will likely be solved by more defensive programming, driving overall complexity ever upwards.
I'm reading y'all comments and it seems we're still finding our footing, and will for some time. I have the luck that I have access to pretty much all frontier and other models alike. I've been extremely negatively biased towards any LLM use in the start. Then the influx from juniors came, then the revolt of doing PRs on such slop came, then some structured methods how to do LLM work came, then vibe projects came, etc, etc. I literally have all the described experiences you've guys mentioned. From good, to bad, to ugly. It's like there's no one particular way about doing this, and no two projects share the same approach - just like ye olde times.
People, especially those who don't do software development, often conflate coding and software development.
The models aren't good at architecture and design. But they take direction on architecture and design and design well and can refactor code quite effectively. AI agents can absolutely be used to clean up vibe coded code bases once you figure out if the investment is worth it. The mess can be avoided if you give them sufficient guidance on architecture and design upfront.
That said, doing so purely in text form doesn't feel great right now. I've been thinking about UML lately. The problem with that was the roundtrip after the code was generated and then the implenetation happened. I don't necessarily think UML is the solution, but neither is walls of dense text.
>The models aren't good at architecture and design.
Can you elaborate to back up this claim? WHat exactly is your yardstick for "being good at SW design and architecture"?
Because I found the current SOTA AI models being great at that, much better than most average real-world devs. Is your yardstick just the John Carmack's of the world by any chance? Because most devs are not John Carmack. They are also not Linus Torvalds, they are not Stallmann.
Maybe your LLM experience is still stuck in the 2023 era of ChatGPT?
I've had the opposite experience. Spec driven development really makes things much easier to maintain. If you're just yolo'ing it and throwing prompts around things will fall apart fast.
As usual, when people post comments like this they never specify what model and when they used it. There is a huge difference between GPT 3 and 6 for example.
Late 2025 also had a step change when agents could largely code autonomously without handholding like previously, and to be honest it's not worth hearing opinions about AI from before that time, that's how significant the change was.
AI is a useful tool but it absolutely needs a lot of human guidance.
Unlike a compiler it won't give up at the first sign of trouble but that just means it left alone it will dig bigger and bigger holes.
Treat prompt engineering as a discipline and refine your technique. When it produces garbage throw out the work and start over until you figure it out.
Were you reading through the code it was generating to make sure the flow was intuitive and comprehensible for each PR? Projects only turn in to maintainable messes if you blindly merge in unmaintainable messy code.
Oh my god. I am sad to have to acknowledge this even if the models have “theoretically” gotten better. I am not a programmer but I am super opinionated with the design and structure of the code and need to make sure whatever I write/ask to generate and use for my own tasks is understood by me at least once. I wasn’t this way in 2023.
I am constantly reading/trying for an agent to help me write simple code without spending too much time. But so far nothing has worked. I wish someone figures it out otherwise, I personally would not be able to realize the AI agent productivity benefits that others are supposedly seeing.
If you specify a structure, pattern, or design, it will stick to that design after a code review phase. Just tell it what standards you have, and after two rounds, it will have written what you described.
Were you reviewing the changes? I've seen lots of projects with this problem. But if AI PRs have to pass the same review bar as any other PR then it shouldn't be an issue in theory (if you can actually maintain the discipline).
Isn't it tiring to keep up with considering the speed of the generated code? And if you want to be careful with the review you lose a significant part of the speed advantage.
It's the same as any other PR. Keep the changes small and contained to that feature. AI can do this if you instruct it to, there is no need to vibe code some 40k line monstrosity.
The people saying programming is solved couldn't program in the first place is what I've noticed. So to them it really does feel solved, things "work" and they don't have to learn what they don't know about programming.
I've been programming professionally for decades. LLMs are extremely useful. At this point if you haven't figured out how to get value out of them, you're either holding them very wrong or you're being willfully ignorant.
I think it depends which LLM tool you're using. If you're using an older, worse model (the kind that you can use for free), the experience is significantly more frustrating. On the other hand, I'd say that the current best models are very useful with a skilled operator.
There are a couple of notable counterexamples here (nobody sane thinks Carmack doesn't know how to program, for example), but by and large I agree with your observation. The people excited about programming with LLMs are, on average, people who weren't good at programming to begin with. Still, given that these counterexamples do exist I try to avoid painting with an overly broad brush for the sake of nuance.
You build your OS atop thousands of open source packages, many of which contain AI generated code. Are you going to audit them one by one and remove offending packages? What about the ones you won't remove because the OS would be irreparably broken?
It also doesn't answer the question of how they might even recognize LLM generated code in contributions to PopOS directly.
I have yet to understand how maintainers can't distinguish beyond (1) PRs that literally include Claude co-author notes or (2) low quality code contribution regardless of the creator.
This is a good idea, we should start a blacklist of open-source projects that are known to have used LLMs. There should be two universes of code, one for hand-typed code used by people who care about quality and one for slop used by those making trash.
Yeah but that's just bad code in general, no? You can make good code with LLMs, you just have to actually engineer it and give up some of the velocity; which is just a bigger version of the same problem we've always had (yes, I get that code review can't scale).
This is just really silly.
The more I think about this the more I think it's like self-driving cars. We have this expectation that self-driving cars MUST be 101% safe and never get into any accidents, ever, before the technology is worth adopting. LLMs are the same -- it's like we think if you can't one-shot a prompt and get perfect software out of it, it's failed. You can choose to spend time getting the LLM to refine the code it's written, review the architecture, come up with an actual engineering process around the LLM. Yes that means you'll be producing less code per time spent -- which is a good thing.
I had a PR in flight that got closed because of this. I had an issue with passwords in the network applet for the VPN and had used Claude to help me identify and then come up with a fix. I did spend a lot time handcrafting and making sure the quality was good, but I respect their decision and no hard feelings, but as someone who have struggled to find time and opportunity to contribute to open source it was a small set back.
I found your commit and your usage of AI seemed reasonable. It seems to me like your PR itself and the subsequent comments and correspondence was also human written.
I think a PR "in flight" shouldn't have been closed like that.
All this will do is push out developers like you that honestly disclose, and instead people will now just lie.
My guess: taking personal responsibility for the functionality, readability, and sanity of the change proposed, both atomically and in the context of the wider code base (adhering to existing conventions and patterns), to the best of the author’s ability.
"... many of the AI contributions were not planned and showed little understanding of the software architecture. So the team wants to "prioritize working on contributions from our own team and regular contributors.""
Sounds reasonable, even to avid LLM users, I suppose. You have to draw a line. This line is too simplistic, but it'll work, for now.
This would just push people to maintain their own fork. If I already have an agent to investigate and fix a bug and able to send a PR, the added cost of maintaining a local fork is minimum.
In fact I've start doing that myself. Sending PR and convincing the maintainer why the fix is necessary is just too much effort.
I wonder if the issue is mostly the code or the AI written PRs and people using AI to talk to the maintainers. I personally just ban anyone doing the latter, I don't want to talk to opus more than I already do lol
I don't see how this will survive the attacker/defender gap as ls get increasingly good at cyber security and finding 0 days... but maybe it's an obscure enough is it doesn't matter?
So you took every single line of open source code you could possibly get your hands on (using scrapers so violently dumb that they amount to a permanent low-grade DDoS) and spent billions of dollars to tune trillions of parameters, and the value you can offer is… “let us inundate you with bad code or else we’ll generate exploits for your software”.
Probably because pop_os and Cosmic are so niche and their market share so insignificant, that they're irrelevant to attackers and bad actors, when those now have much bigger fish to fry to make their time and effort worth it. See the Arch AUR attacks, for perspective.
I think even amongst the HN and Linux userbase, pop_os is still niche, let alone amongst normies who never heard about Linux. So they can afford take the high road and treat it like their personal sandbox, accepting only human written code.
But larger and more important projects like Fedora and Debian are more pragmatic with the fact that they'll have to accept AI written(but human reviewed) code, if they wish to keep up with the real world development and threats, as expressed by Linus Torvalds himself.
The thing is, the cat's out of the bag on this one now, especially in the field of pen-testing and reverse-engineering. AI can brute-force its way into projects in ways that beat even experienced researchers, so your only choice to keep up is to accept the use of AI generated fixes as a counter defense.
You only need one bad actor. For example, someone reading this thread could easily decide to start attacking it just because someone else said it wasn't worth it, as a personal challenge.
If that's your threat model then you shouldn't use any SW in exitance, FOSS or otherwise. In fact you shouldn't even go online, or even outside you own house, since one single bad actors exist everywhere. You can walk down the street and suddenly someone in a car runs you over(witnessed myself). And yet live goes on.
It doesn't have to be niche (or not) to have bad actors. That doesn't make it likely of course but it is possible, and increasingly more so in the age of AI when you can simply point it at multiple projects in parallel.
2) Secondly, keep in mind that distrowatch is not representative of linux userbase. MX-Linux kept showing up at the top spot for many years despite being niche.
3) And thirdly, TOP 5 DE isn't really an achievement when Linux DE market share is overwhelmingly dominated by KDE Plasma and Gnome as the majority shareholders, with XFCE and Cinnamon trailing. So Cosmic DE if it somehow made the no. 5 spot, would still be ignorable sub <1% market share, as I initially claimed, basically invisible to bad actors.
"The analytics", is the publicly available training dataset of the LLMs that share the same common opinion on that DE market share.
What do you expect exactly? Do you want me to now manually parse through terabytes of information at your whim for your own convenience? Sorry, but I'm not your personal unpaid servant.
If you wish to disprove me in the comment section, then you need to do the manual work and show us that the my quoted LLMs statistics are wrong. I'm not your personal errand boy to do your bidding, 'massa'.
Point at the public dataset then if it's so public. You really have no idea how the LLM is parsing the data and could easily be hallucinating. That you find being called out on this as some sort of insult is quite telling and frankly funny, what a strong reaction to someone asking for a basic source when the burden of proof is on you to prove (or at least show a source that) these stats are right, not on anyone else to disprove AI bullshit.
OK, but what's the sample size of that and who's measuring it? I never heard of that blog or took part in that poll. So how is that blog link the yardstick but mine is not?
<< Probably because pop_os and Cosmic are so niche and their market share so low, that they're irrelevant to attackers and bad actors, when those now have much bigger fish to fry.
That is such a weird statement that I am not entirely certain where to begin. PopOS is hardly niche. Its base are all fairly common components by linux standards. And, more importantly, attackers and bad actors may other considerations in mind than sheer population size -- just to point out the glaringly obvious.
<< I think even amongst the HN and Linux userbase, pop_os is still niche
I think rather than trying to disprove it, I think I should ask why you think that? If anything, PopOS annoys me because it is just a step before ubuntu ( and ubuntu is just windows at this point ). Maybe I am defensive, because my first real distribution ( that did not share disk with windows was popos )?
It also means security is not held as high and vulnerabilities not as much found. A simple 0-day may survive for years. Not much effort needed to have permanent access.
Its easier for an LLM to find vulnerabilities in a project with less usage, and those vulnerabilities will stay open longer, making them much cheaper to attack to keep the door open.
What's the point of attacking projects that almost nobody uses?
Do you think Netanyahu, Trump or Xi-Jinping are somehow secretly using Cosmic DE at home, to be worthy targets?
Bad actors have limited time, lives of their own and mouths to feed as well, so they concentrate their efforts where "the fish are" if they want to PWN someone for profit.
That's why Windows was the biggest target in the past for so long and why MacOS and Linux were ignored. Because most of the fish were on Windows.
Security employees, tokens and peoples' time are a finite resource too.
If you assume Mossad and NSA are Token-maxxing every FOSS project out there to cast as large as possible fishnet , then maybe using Mozilla and MacOS get you hacked too, maybe even visiting HN and commenting gets you hacked. Where this infinite paranoia argument end?
Using AI to find vulnerabilities doesn’t mean that you need to use AI to generate the code that fixes them. And you can still ask AI whether it thinks the fix is okay, as a second opinion.
Same trouble we have. Some clever person says to use AI agents for code review. 100kloc commit got flagged through on Friday. Taking this week off. Not my circus.
Faith is the problem. Extraordinary claims must stand up to scrutiny. They do not.
There are a whole lot of people (in tech) who truly hate AI and want nothing to do with it. Those people will flock to projects who take a stand against it.
it doesn't matter. it's delusional to think you can outcompete a thing for which solving a Millenium problem is just Tuesday. it's the anger phase of grief, nothing more.
This is straight up cult behavior. "You WILL assimilate or we WILL kill your project"
PopOS is not a product being sold by techbros trying to pump their stock like AI is, it's free open source software, it doesn't need to "outcompete" anything. If you don't like it, don't use it.
banning AI from PRs because you're swamped with too many low quality PRs, definitely. We pretty much are doing this with SQLAlchemy. If I'm going to have a small fix or improvement coded by an LLM (which I do all the time), I want to prompt the LLM directly, rather than having someone trying to pad their resume forward my communications onto their LLM via PRs. What's the point of that?
Reasonable, although I've taken a different approach. Either closing such PRs, or treating them as very detailed issues and having my own LLM build the actual fix.
My repos probably don't see as much traffic as SQLAlchemy though.
rejecting low quality ones should be the norm regardless of whether an AI or a human wrote them. the question is what would happen if you were swamped with high quality PRs? what will happen once you are? (that's probably a 2027 question!)
> If I'm going to have a small fix or improvement coded by an LLM (which I do all the time), I want to prompt the LLM directly, rather than having someone trying to pad their resume forward my communications onto their LLM via PRs. What's the point of that?
Which is why LLM PRs should just be issues (if there isn’t one already). Make the issue author a co-author on the PR. But let the maintainer actually oversee the LLM generated solution.
Are you in Python by chance? Python has a lot of crazy hidden/inexplicit/spooky action at a distance stuff (especially in the frameworks) that can make LLMs gunk up code by defensively programming or just burn context chasing data provenance
I’m saying it’s probably multiple factors and both you and GP are right.
Save your "you're holding it wrong" if you're not going to suggest how to hold it.
Cult speak escape hatches are intellectually lazy.
Edit: My latest project is all GPT-6 Astra High. It takes a lot of steering to keep it from adding a bunch of, while useful, features that are not strictly enough to the point. That main issue is it’ll use a lot of extra tokens in the process!
What was your process?
https://github.com/NousResearch/hermes-agent is 99% (just a guess) LLM generated. 1140 closed pull requests this week. 1.5k closed issues. The github insights page for commits doesn't load for me presumably because it can't handle this scale of commits. But I estimate ~1K commits per day on average.
There's a blog entry https://nousresearch.com/refactoring-hermes-with-1393-agents that details some work that was done by LLMs to refactor and improve the code.
I guess they know how to hold it?
I had a look at the kind of issues that are reported at that project (there's 15k of them, so I can at best assess a couple). It looks like a complete mess: A lot of TOCTOU, concurrency, and resource mismanagement issues and edge cases that in a better-managed project would have been avoided by construction. They will now will likely be solved by more defensive programming, driving overall complexity ever upwards.
The models aren't good at architecture and design. But they take direction on architecture and design and design well and can refactor code quite effectively. AI agents can absolutely be used to clean up vibe coded code bases once you figure out if the investment is worth it. The mess can be avoided if you give them sufficient guidance on architecture and design upfront.
That said, doing so purely in text form doesn't feel great right now. I've been thinking about UML lately. The problem with that was the roundtrip after the code was generated and then the implenetation happened. I don't necessarily think UML is the solution, but neither is walls of dense text.
Can you elaborate to back up this claim? WHat exactly is your yardstick for "being good at SW design and architecture"?
Because I found the current SOTA AI models being great at that, much better than most average real-world devs. Is your yardstick just the John Carmack's of the world by any chance? Because most devs are not John Carmack. They are also not Linus Torvalds, they are not Stallmann.
Maybe your LLM experience is still stuck in the 2023 era of ChatGPT?
Late 2025 also had a step change when agents could largely code autonomously without handholding like previously, and to be honest it's not worth hearing opinions about AI from before that time, that's how significant the change was.
Unlike a compiler it won't give up at the first sign of trouble but that just means it left alone it will dig bigger and bigger holes.
Treat prompt engineering as a discipline and refine your technique. When it produces garbage throw out the work and start over until you figure it out.
You build your OS atop thousands of open source packages, many of which contain AI generated code. Are you going to audit them one by one and remove offending packages? What about the ones you won't remove because the OS would be irreparably broken?
I have yet to understand how maintainers can't distinguish beyond (1) PRs that literally include Claude co-author notes or (2) low quality code contribution regardless of the creator.
They aren't saying that AI produces bad code or is terrible for the world in some way.
It's mainly just resulting in a lot of PRs that they don't have enough time to review or features they don't plan to add.
PopOS is Ubuntu with extra problems. Ubuntu itself is fine, but then PopOS adds weirdness.
Cosmic has been in beta for how long ?
Being hand written is no guarantee of high quality, just like using LLMs is no guarantee of low quality.
This is just really silly.
The more I think about this the more I think it's like self-driving cars. We have this expectation that self-driving cars MUST be 101% safe and never get into any accidents, ever, before the technology is worth adopting. LLMs are the same -- it's like we think if you can't one-shot a prompt and get perfect software out of it, it's failed. You can choose to spend time getting the LLM to refine the code it's written, review the architecture, come up with an actual engineering process around the LLM. Yes that means you'll be producing less code per time spent -- which is a good thing.
I think a PR "in flight" shouldn't have been closed like that.
All this will do is push out developers like you that honestly disclose, and instead people will now just lie.
Sounds reasonable, even to avid LLM users, I suppose. You have to draw a line. This line is too simplistic, but it'll work, for now.
In fact I've start doing that myself. Sending PR and convincing the maintainer why the fix is necessary is just too much effort.
I think even amongst the HN and Linux userbase, pop_os is still niche, let alone amongst normies who never heard about Linux. So they can afford take the high road and treat it like their personal sandbox, accepting only human written code.
But larger and more important projects like Fedora and Debian are more pragmatic with the fact that they'll have to accept AI written(but human reviewed) code, if they wish to keep up with the real world development and threats, as expressed by Linus Torvalds himself.
The thing is, the cat's out of the bag on this one now, especially in the field of pen-testing and reverse-engineering. AI can brute-force its way into projects in ways that beat even experienced researchers, so your only choice to keep up is to accept the use of AI generated fixes as a counter defense.
If that's your threat model then you shouldn't use any SW in exitance, FOSS or otherwise. In fact you shouldn't even go online, or even outside you own house, since one single bad actors exist everywhere. You can walk down the street and suddenly someone in a car runs you over(witnessed myself). And yet live goes on.
Because your comment didn't disprove that Cosmic DE isn't too niche for bad actors to get involved.
(I wasn't able to make it work on a scrap Dell I tried it on because the GPU was too old. Booted the USB key and COSMIC greeter failed to start)
1) Firstly, surce? My research according to Google Gemini 3.8 Pro shows top 5 DEs are as follows:
2) Secondly, keep in mind that distrowatch is not representative of linux userbase. MX-Linux kept showing up at the top spot for many years despite being niche.3) And thirdly, TOP 5 DE isn't really an achievement when Linux DE market share is overwhelmingly dominated by KDE Plasma and Gnome as the majority shareholders, with XFCE and Cinnamon trailing. So Cosmic DE if it somehow made the no. 5 spot, would still be ignorable sub <1% market share, as I initially claimed, basically invisible to bad actors.
"Way too defensive" how? By asking and bringing data for my PoV?
>GP replying to you was clearly trying to have a conversation
As am I, except I ask for, and also bring data to back up my PoV, instead of vague opinions.
>Chill.
Where am I not being chill?
Do you have a better source than GP?
>if needed post the actual source
Define "actual source"? In good faith, I mean.
Where else do you get this information that's, quote, "actual source"?
What do you expect exactly? Do you want me to now manually parse through terabytes of information at your whim for your own convenience? Sorry, but I'm not your personal unpaid servant.
If you wish to disprove me in the comment section, then you need to do the manual work and show us that the my quoted LLMs statistics are wrong. I'm not your personal errand boy to do your bidding, 'massa'.
That is such a weird statement that I am not entirely certain where to begin. PopOS is hardly niche. Its base are all fairly common components by linux standards. And, more importantly, attackers and bad actors may other considerations in mind than sheer population size -- just to point out the glaringly obvious.
<< I think even amongst the HN and Linux userbase, pop_os is still niche
I think rather than trying to disprove it, I think I should ask why you think that? If anything, PopOS annoys me because it is just a step before ubuntu ( and ubuntu is just windows at this point ). Maybe I am defensive, because my first real distribution ( that did not share disk with windows was popos )?
Do you think Netanyahu, Trump or Xi-Jinping are somehow secretly using Cosmic DE at home, to be worthy targets?
Bad actors have limited time, lives of their own and mouths to feed as well, so they concentrate their efforts where "the fish are" if they want to PWN someone for profit.
That's why Windows was the biggest target in the past for so long and why MacOS and Linux were ignored. Because most of the fish were on Windows.
Previously, time was the most precious resource. Now its tokens, and more cheaply at at.
If you assume Mossad and NSA are Token-maxxing every FOSS project out there to cast as large as possible fishnet , then maybe using Mozilla and MacOS get you hacked too, maybe even visiting HN and commenting gets you hacked. Where this infinite paranoia argument end?
Reminder that AI is quite stupid.
I'll trust the words of groups like curl (https://daniel.haxx.se/blog/2026/06/10/a-human-in-control/), Linux, and even the infamously anti-AI Gnome (https://blogs.gnome.org/mcatanzaro/2026/10/02/the-era-of-sof...) that AI is finding real vulnerabilities and you're your project a disservice by ignoring them.
Faith is the problem. Extraordinary claims must stand up to scrutiny. They do not.
PopOS is not a product being sold by techbros trying to pump their stock like AI is, it's free open source software, it doesn't need to "outcompete" anything. If you don't like it, don't use it.
My repos probably don't see as much traffic as SQLAlchemy though.
Which is why LLM PRs should just be issues (if there isn’t one already). Make the issue author a co-author on the PR. But let the maintainer actually oversee the LLM generated solution.
So basically you're saying you reject drive-by PRs.
I have a hard time knowing if anti AI is a mental illness or propaganda coming out of China.
Unironically.
Separately, why would anyone use a Debian based desktop OS? Your $11 Amazon mouse won't work. An Nvidia card won't work. Just use Fedora.