Unfortunately, I Think LLMs Might Be Useful for Development Now
A year ago, I thought the LLM bubble was going to completely burst under its own weight. Now, I'm not so sure. I still think the LLM sector is overvalued and destined for a big correction. But based on what I'm seeing in the software industry, I don't think LLMs are going to zero-- at least without some sort of intervention. So now my question is what do we do?
My Prior Skepticism
I've been doubtful about the usefulness of LLMs for software development [^1] since they first appeared on the scene. When GitHub initially released Copilot, I was unimpressed. At the time, it was just a slower and more error-prone form of code completion, inferior in every way to the incumbent options like IntelliSense.
Fast-forward a few years to today, and I grudgingly have to admit that things have changed significantly. I want it to all be marketing and hype, but as an engineer, I have to let my observations of reality win out over my personal feelings.
It Must Be Said
I would not be doing my ethical duty if I wrote about LLMs without taking the opportunity to also mention that:
- LLMs are an ecological disaster in terms of energy, water, and land use.
- Most LLMs are trained through wholesale theft of the work of authors, artists, and others.
- LLMs are being used by capitalists as a tool to break labor power.
- The web scrapers building training sets for LLMs are creating a huge burden for website operators.
- Most LLMs have been implemented without sufficient safeguards. As a result, they are being used as tools for fraud, abuse, harassment, privacy and security violations, and misinformation.
- LLMs are harming educational outcomes.
- LLMs are de-skilling people in a wide variety of industries.
And that's not even a complete list of the moral issues with LLMs.
What I'm Seeing in the Software Industry
What I'm seeing today is that LLMs have gotten quite good at performing a variety of development tasks. I've seen slop in the past, and I'm still seeing slop now. However, I'm seeing results that are not slop. I'm seeing skilled, experienced developers using LLMs are getting quite good results-- maybe even indistinguishable from what those developers would create without an LLM, but created in much less time.
I can give three recently anecdotes:
The first thing that gave me pause were code reviews from a tool called coderabbit. I was contributing to an open-source project that uses coderabbit as the first-line review on PRs before putting them in front of a human reviewer. I did my own self-review and thought I had very carefully considered all the cases. When coderabbit scanned the code, it found many edge cases and bugs that I had missed. At first, I was sure that the tool was wrong. I wrote unit tests to prove that my code was correct and coderabbit was wrong. But coderabbit was right and I was annoyed. As an engineer, I had to grudgingly admit that the resulting code was better for having been analyzed through that tool.
Secondly, I was at Elixirconf and catching up with developers I had not seen in years. Developers I've worked with in the past and know to be very careful and considerate engineers. Some of these people have given me the hardest, most thorough PR reviews that I've ever received. And then they showed me code they had written using LLMs. And the code was good. Not just functional, but also readable, maintainable, and efficient. Again the LLMs had even handled edge cases that I'm fairly confident most software engineers would have missed until a bug occurred in production.
Some of these folks let me shoulder-surf while they worked, so I know these weren't hand-picked best-case examples or random good luck. I had some eye-opening conversations about how these engineers were ensuring a high standard of quality while heavily using LLMs. In short, these folks were freeing up time that they used to spend on routine feature development, and reinvesting that time into quality and tooling. I was seeing devs putting LLMs at the center of a workflow and surrounding them with 100% branch-test-coverage, linting, code smell detection, fuzzing, property testing, benchmarking, and careful post-deployment monitoring. And yes, spending more time reading code. To be frank, I've never seen this level of quality assurance tooling in commercial development before (at least outside of life safety systems.)
Thirdly, I wanted to test this myself. For a few years, I've been testing the Hot New Model every few months. It's important to me to know what I'm talking about, so even while I thought the only thing was overblown, I wanted to be able to make strong arguments based on hands-on experience. Until recently, the results have always been bad and slow and I've been able to safely ignore the LLM trend for a few more months. This time, the LLM actually did what I asked it to do, the code was pretty good, and it completed the while task in about five minutes. And the more I invested in tooling improvements, the better it got.
Yes, I'm Still Seeing Slop
This is not to say all LLM-generated code is good. I'm still seeing a lot of slop out there. But the existence of slop does not preclude the existence of genuinely reasonable code output from LLMs. LLMs have lowered the barrier to creating minimally syntactically valid, but still truly awful code (slop). It is also seems that the ceiling been raised lately, and engineers who have developed the skills and tooling to use this stuff well are now able to produce acceptable code.
So What Now?
A year or too ago, I was seeing no genuine value from LLMs. With that perspective, I expected that eventually the bubble would burst more or less on its own. Even the most irrational market can't sustain a complete fiction forever. Eventually, some investors wake up and see the reality. They withdraw their capital, trying to get ahead of the inevitable. That dip causes more investors to wake up, and that becomes a cycle. Eventually, the market for a truly useless invention reaches approximately zero.
Now, I think the LLM sector is significantly overvalued, based on the investor assumption that LLMs are going to be in every aspect of life going forward. I don't think that's the case. Right now, startups are trying to cram these word-guessing-machines into your car and your toaster and a million other places where they are useless or worse. If there is some use for these things in the software industry, and maybe a few other places, then the true market value of the LLM sector isn't zero. Following market logic from there, I'm thinking we're headed for a major correction, but not a total implosion.
If that's the case, and if look at all the moral issues with LLMs, what do we do? If LLMs were going to zero more or less on their own, then abstention and waiting out the market would be a viable strategy. How do we live and work morally in a world where LLMs are an ongoing thing-- particularly for people like me in the software sector?
Boycott?
One option is obviously boycotting LLMs. I could simply refuse to use LLMs and accept the consequences. What if the industry reaches a state where not using LLMs in development makes you unhireable? Maybe you go independent and start your own software company. What if LLMs make it trivial to clone whatever software your company puts out at a much lower (dollar and time) cost? Do even the indies get pushed out?
For me personally, I guess that's a bummer, but survivable. Honestly, I used to work in the building trades and I could do it again. I could fuck off and become a carpenter or an electrician and be fine. And they could never stop me from simply programming for the love of the game.
But I think many people in the industry don't have the same options.
I'm also skeptical of the effectiveness of a million separate individual boycotts. Historically, "voting with your wallet" or "voting with your feet" have not been the most effective tactics, particularly as a way of disciplining employers.
Further, does it make sense to boycott the tool, or to fight against the abuses of the current form of the tool? After all, the luddites were not actually opposed to machines, but rather the consequences of machines such as economic precarity, child labor, pollution, and the poor quality of the produced product. Are we against LLMs per se, or are we against pollution, theft of creative work, the lack of meaningful safety standards, and the proliferation of slop? Is it possible, even theoretically, to disconnect these things?
What Might Work?
There are strategies that I think could work:
Organized labor action: I think unions and contract negotiations could make a difference here. We are even seeing some effective "work to rule" / malicious compliance campaigns working (e.g. employers pulling back on LLM budgets after employees intentionally blow token budgets).
Mass political pressure: in theory, it is possible that a mass popular revolt could pressure politicians into putting meaningful regulations on LLM use in place. There's already some grassroots pushback on data centers. I worry that without discipline and connection to a larger political struggle, the data center pushback is going to devolve into a million disconnected pockets of NIMBY-ism. If that happens, we're just going to end up with data centers in poor neighborhoods and no meaningful change.
Creating a better alternative and/or seizing-the-means: I don't know what this looks like, but I'm thinking about it a lot. I'm sure I'll write about it in the future. Could we create an alternative to petrol-guzzling, artist-impoverishing, unsafe-at-any-scale, monopoly LLMs? Is there something new and better we can hammer out of the thesis of LLM-maximalism and its antithesis in refusalism? It would have to look completely different from what we're being offered today. The master's tools will never dismantle the master's house and all that. However, is there some sort of Deleuzian way in which this force within capitalism can be radicalized against capitalism?
I'm not sure about any of this, to be clear. It's just something I'm thinking about a lot and grappling with. I'm trying to put together a coherent theory.
The one thing I know is waiting on the sidelines isn't it.
[^1]: and everything else