The Limits of LLMs: Good at Reproducing What Exists. Expensive for What Doesn't

September 28, 2026 · 💬 Join the Discussion
If you're lazy, click here for the TL;DR

Everyone has noticed by now that LLMs write CRUD with their eyes closed. Forms, reports, REST endpoints, unit tests: it comes out clean, it comes out fast. A lot of people conclude from this that programmers are done for.

I’ve been using LLMs for everything for almost two years, I’ve published dozens of projects with them, and my conclusion is different: what they do best is reproduce what already exists. And that covers a giant slice of corporate work, don’t kid yourself. But there’s a frontier they don’t cross easily, and I spent the last few months bumping into it on purpose.

This article shows where that frontier sits, with evidence from my own projects and from the research that exists on the subject. And there’s a hypothesis I’ll put up front and defend along the way: if LLMs are so good at code, providers owe a huge debt to the decades of the open source community, which protected the freedom of code and published everything for free. Without that heritage, there would be no Copilot, no Cursor, no Claude Code, none of this.

A disclaimer before we start: what I’m going to argue here is speculative, based on observation and on the public data that exists. Nobody outside the labs knows exactly how each provider trains their own models, and that varies from provider to provider and from version to version. What holds today may not hold in the next generation. Take it all with a grain of salt.

The Training Blind Spot

Every LLM was trained on what’s public. The entire public internet, and for code, public GitHub. The canonical open code dataset, BigCode’s The Stack v2, is explicit about this: it’s derived from Software Heritage, an archive of “the source code of all publicly available software”. That’s 67.5 TB, 3.28 billion unique files, 104.2 million repositories. Huge. And all public.

Now look at the other side. The Octoverse 2025, GitHub’s official report, says 81.5% of contributions on the platform happened in private repositories. That’s just GitHub. Outside of it there’s:

  • the internal code of every company that never touched a public server;
  • the source of every proprietary software;
  • the code of every commercial game;
  • and that of every embedded system.

None of that ever went into training.

Think about what this means: the model deeply knows what the world has already published, and it’s blind to what the world keeps closed. When you ask for a React form, it’s reproducing patterns it has seen millions of times. When you ask for something nobody has ever done in public, it’s operating in territory where its map is blank.

Why Is Everyone Shipping So Much, So Fast?

Before getting into the limits, it’s worth recording the side that works, because it’s what explains the feeling that everyone became an app factory.

Look at what’s popping up in the community: Spotifast and ZapFast, native Spotify and WhatsApp clients written in Rust by a single person, running on Linux, macOS, and Windows. Nothing in them is new technology. It’s the service that already existed, wrapped in a lean native app, on top of libraries that already existed, doing what the official teams always could have done but never wanted to spend their time on. That work got cheap, because an LLM trained on everything public is really good precisely at this: replicating.

My own projects follow the same pattern. FrankMD is a Markdown note-taking web app with a Visual Studio look and features that already existed in my other apps. Frank Sherlock is a local image organizer that, deep down, is a file manager with thumbnails and search, something that exists by the dozen. Neither invents anything: they’re mashups of known pieces, assembled the way I always wanted. The LLM does the boring part, and that’s where the speed comes from.

The most didactic example is my email client. I’ve used Geary for years, and two small gaps always annoyed me: the weak recipient autocomplete and a sidebar I wanted to hide. Trivial stuff, which never justified setting up a build environment, learning the project’s toolchain, and studying the code just for two features. This year I finally solved both: an autocomplete module, a module to hide the sidebar, and a whole fork with my patches. I became the maintainer of improvements I had postponed for years, not because I got faster, but because the boring part got cheap.

Notice what all these cases have in common: the problem had already been solved before, somewhere else, some other way. The map existed; what didn’t exist was cheap labor to retrace the path. That’s reproduction working in our favor, and it’s why your feed is full of people shipping a new app every week.

Now hold that thought. The rest of this article is about what happens when the map runs out.

Four Home Experiments

I’m not theorizing. Over the last few months I ran a marathon of retrocomputing projects that are exactly the pathological case: working with proprietary software from 30 or 40 years ago whose source code was never public. It’s the perfect blind spot.

nes-to-sms: the project I gave up on

nes-to-sms is a static recompiler from NES to Master System. The idea: take the binary ROM of a NES game, translate it instruction by instruction from 6502 to Z80, convert the tiles, and generate a Master System ROM that runs on real hardware. Forget emulator or manual port: this is a machine that turns one game into another.

Static NES recompilers exist, NESRecomp translates to C and runs natively on PC, and I studied it closely. But crossing from one console to another 8-bit console, with a completely different clock and VDP budget, nobody has ever done. There’s no public reference, no paper, no abandoned project to copy.

The project numbers: 349 commits, 31 active days spread over almost 4 months, ~110 thousand lines of engine in Rust and Z80, 668 tests passing in the last run recorded in the README. And the result? Playable Super Mario Bros. (slow, limited audio, but playable), with a byte-for-byte identical RAM trajectory to the NES on one ~4,900-frame route (the diff leaves out audio and part of the VRAM buffer). Castlevania playable on stage 1. SMB3 rendering. And a queue of games that boot three frames and die.

And it wasn’t for lack of pushing the models. In this project alone:

  • four generations of Claude went through it (Opus 4.7, Fable 5, Opus 4.8, Fable 5.1), with 213 co-signed commits and 104 commits linked to sessions in the history, plus Codex sessions in the stretches without co-authorship, complete with documented handoffs between agents;
  • there was even cross-generation auditing: a new agent session reviewed the previous 55 commits and reverted conclusions the previous generation had overclaimed;
  • the Castlevania crash hunt alone consumed 45 commits in 5 days;
  • the final September sprint was 176 commits in 17 days before I called the pause;
  • and the regressions are registered by name in the history: “H.7 tried and reverted”, “record the reverted map-bank-hold attempt”, “correct-but-reverted result”, plus three shared-lever hypotheses tested and rejected in the project’s last three days.

I didn’t measure tokens reliably in this project, so I won’t make up a number. But 349 commits with 668 tests holding up every step already give an idea of the cost.

Notice the detail that matters: the only game that actually got finished is Super Mario Bros. Why? Because SMB is the most dissected game in history. There’s a complete public disassembly the community has used for over a decade, with every function named and commented. I fed that into the game’s profile and the LLM had a complete map to work with. For the other games, there’s no map. And without a map, each new game died a few frames after boot from control flow divergence, and each one demanded an instruction-by-instruction forensic hunt that carried nothing over to the next game.

The project itself documented the spinning pattern. There was an entire optimization campaign that lasted weeks trying to cut the cost of emulating 6502 flags on the Z80.

Then in September I stopped to compare with the manual port of SMB to Master System made by a hacker named lackoftrack27, written by a human, routine by routine. According to the analysis we did of his code, the human port runs inside the frame budget, with headroom we estimated at ~90% utilization. My machine, after months of optimization, still needed overclocking.

And the analysis document pinned down why: when we removed 30% of the flag emulation calls, the timing barely moved.

The flags were never the bottleneck.

The right optimization was to redesign the data structure: transposing the object array into RAM pages. It’s a trick that depends on knowing that “the X register here is always an object index.”

That invariant exists in the head of the programmer who wrote the game in 1985. It doesn’t exist in the instruction stream. A static translator has no way to discover it. And the LLM, for months, didn’t discover it.

The pause banner I wrote in the README is blunt: real-hardware parity “may be fundamentally unattainable” and, quoting, “this may be a limit of the current frontier coding models as much as of the approach”.

gg-to-sms: the pivot to mapped territory

The same day I paused nes-to-sms, I started gg-to-sms: converting Game Gear games to Master System with a widened screen. The difference that matters is that this problem is not new. The romhacking scene has been converting GG games to SMS by hand for twenty years, and the catalog I built in the project counted ~186 of them. Meaning: there’s a corpus, there’s ground truth, there’s a pattern to learn.

Result: in one day of work,

  • the research phase was complete;
  • the catalog of the 373 GG games was assembled;
  • the patch-point scanner was validated against real human patches, hitting addresses within an 8-byte margin;
  • and on top of that the tool found real bugs in twenty-year-old patches that the entire scene had never noticed.

The model was the same; what changed was the problem, which this time has references.

And even so: so far, zero games have passed the viewport acceptance review. The new part of the project (the machine that generates the patches on its own, which is the part nobody has ever done) remains undone.

super-mario-deluxe-fixed: four days and 436 million tokens

This one was a 4-day sprint, practically a single Codex session that consumed 436 million input tokens. The goal: run Super Mario Bros. Deluxe for Game Boy Color with a widened screen at NES resolution, something no existing tool does with live sprites.

It got to playable in 4 days. But look at the foundation: a ready-made emulator vendored as a dependency, and a dormant disassembly of the game that someone had published years earlier, which the project’s own documentation calls “the most important artifact.”

The new work was the custom renderer, and it stopped at 21 of 32 stages completed. The missing ones hang on moving-platform timing. The project has been stalled mid-sprint since then, and I haven’t given up on it yet.

And one detail that says a lot: I had to direct QA by looking at screenshots, because the agent would deliver things with graphical glitches and not notice on its own.

super-mario-bros-35: everything that worked had an owner

This is the offline clone of Super Mario Bros. 35, the battle royale Nintendo took offline. What worked in it: everything built on top of an existing artifact.

  • The engine is a C reconstruction made by someone else with Ghidra.
  • The AI’s initial weights came from published checkpoints.
  • The network rules were inferred from a public reverse engineering of the netcode.
  • Even the physics: the central bug was diagnosed because I had a recorded reference trajectory to compare frame by frame.

What didn’t work: the maze castle stages. Reinforcement training stalls on them. 46 million agent steps, training restarted from scratch, anti-loop reward, and in the end the only thing that solved one of the stages was manual bias surgery on the trained model. Out of 32 stages, 19 cleared. The worst ones remain stuck.

And the process degenerated into a pattern I recognize from all these projects: when training doesn’t solve it, the flow becomes checkpoint-selection roulette, dozens of bias and output-head variations, throwing permutation and computation at the problem in place of a principled fix.

The Pattern I Keep Seeing

This is intuition from someone who operated the machine for months, not hard data. But the pattern repeats with Claude, with GPT, with Kimi, with GLM, so it doesn’t look like a defect of one model.

The beginning of any project is wonderful. The first milestones fly, because scaffolding the structure, writing harnesses, replicating APIs, all of that is covered territory. Then the project reaches the part nobody has ever done, and the speed collapses. The agent starts spinning: tries a hypothesis, regresses something else, reverts, tries again. The nes-to-sms commit history is full of honest messages like “attempt reverted” and “hypothesis closed with correct result but reverted.” Lots of tokens burned, little progress.

My reading of what happens inside: token sampling is probabilistic. When the problem resembles the training data, the right tokens have high probability and the answer comes out straight. When the problem is unprecedented, the probability of the right path is low, and the process becomes directed brute force: guided trial and error, but trial and error all the same.

The way out I found for this spinning is always the same: give it an oracle. If I can define “the result has to equal X” (a reference image, a recorded trajectory, a cycle-accurate emulator as ground truth), the machine can iterate until it hits X. In nes-to-sms, every real unlock came from that: the SMB disassembly, the frame diff, the reference emulator. When I can’t define X properly, it spins indefinitely.

The second strategy that works: hunt down an abandoned project, a paper, any material the model never saw but that brings the problem closer, and shove it into the context. That warms up the distribution and improves the result a lot. But it’s limited: if nothing similar exists, there’s nothing to feed it.

More Than Anecdote: What the Research Measures

The good part is I don’t have to rely on my intuition. There’s serious research measuring exactly this.

Apple’s GSM-Symbolic study took benchmark math problems and did two simple things: swapped only the numbers, and then added a single irrelevant sentence to the problem statement. Result: performance drops just from swapping the numbers, and plummets up to 65% with the irrelevant sentence. The authors’ hypothesis, in the text: “current LLMs cannot perform genuine logical reasoning; they replicate reasoning steps from their training data.”

Scale AI’s GSM1k created a mirror of GSM8K written from scratch by humans, guaranteed contamination-free, and saw drops of up to 8% across several model families, with signs of systematic overfitting in the worst ones. To be fair to the paper: the authors say frontier models show minimal signs of overfitting and that all of them generalize to some extent to new problems. The signal exists; its size is debatable.

LiveCodeBench ran the cleanest test of all: evaluating on coding problems published after the training cutoff. Inflated performance on pre-cutoff problems and drops on post-cutoff problems across several popular models, the classic contamination signal.

And there’s ARC-AGI, François Chollet’s benchmark designed specifically to be memorization-proof: every task is unprecedented, nothing similar exists in training. On ARC-AGI-2, released in March 2025, the official leaderboard says: pure LLMs score 0%, reasoning systems sit in single digits, and every evaluation task was solved by at least two humans within two attempts. Chollet, by the way, summed up my thesis better than I did, back in February 2024:

“Reality is that LLMs are not AGI — they’re a big curve fit to a very large dataset. They work via memorization and interpolation. But that interpolative curve can be tremendously useful, if you want to automate a known task that’s a match for its training data distribution. Memorization works, as long as you don’t need to adapt to novelty.”

In the real world, the study that impressed me most was the METR RCT: 16 experienced developers, 246 real tasks on their own mature repositories, randomized with and without AI. The devs bet AI would speed them up by 24%. Measured: it was 19% slower. Notice where this happened: mature, private code, full of context the model never saw. Exactly the blind spot.

And the industry’s favorite benchmark for saying “AI already does software engineering,” SWE-bench, is Python-only, 12 famous open source repositories, in other words the most covered slice of training. The industry’s success metric measures mapped territory.

Even the legacy code literature says the same thing: a 2026 paper on LLMs generating COBOL states that, in these legacy languages, “most production code resides in enterprise systems and is rarely publicly available”. It’s this article’s thesis, said by other people.

The Steelman: Directed Brute Force Works, at a Steep Price

Before wrapping up, it’s worth facing the strongest argument against me. That’s what steelman means, for anyone unfamiliar with the term: the opposite of a strawman. Instead of attacking the weakest version of the opposing argument, you answer the strongest version it can have. And it exists, and it’s a good one.

In May 2025 DeepMind published AlphaEvolve: a system that found an algorithm for multiplying 4x4 complex matrices with 48 scalar multiplications, beating Strassen’s 1969 record in that specific scenario. On a list of more than 50 open math problems, it rediscovered the state of the art in ~75% of cases and improved the best known solution in 20%. That’s real novelty, not reproduction.

OpenAI’s o3-preview jumped to 75.7% on ARC-AGI-1 in December 2024, 22 points above the previous best result on the same evaluation set, and the high-compute configuration reached 87.5% at an estimated cost of $4,560 per task.

And back in 2016 AlphaGo played move 37 against Lee Sedol, a move that DeepMind’s own data estimated had a 1 in 10,000 chance of being played by a human, proving that search over a model can transcend imitation.

Now look at the mechanism of the three:

  • AlphaEvolve is an evolutionary loop: the LLM proposes mutations, an automatic evaluator scores them, selection iterates.
  • o3 does program synthesis at test time, exploring the solution space.
  • AlphaGo is tree search over reinforcement learning.

If there was any “insight” in any of them, it came from the search system around the model: all three are directed brute force with a cheap verifier, the same mechanism I described happening in my projects. The difference is that DeepMind has a perfect oracle and an infinite budget: in the configuration that reached 87.5%, o3-preview spent thousands of dollars per ARC task to do what a human does for free. The ARC Prize itself wrote the caveat:

“We know that brute-force search could eventually solve ARC-AGI (given unlimited resources and time to search). This would not represent true intelligence.”

In short: novelty comes out, but it comes out through expensive search over an oracle, on the basis of guided trial and error. If your problem has a cheap verifier and you have tokens to burn, you can go far. That’s how nes-to-sms got where it got: 668 tests and emulator ground truth holding up every step. But when there’s no oracle, no corpus, and no reference, you’re paying for brute force in dollars, with regressions along the way, and the ceiling shows up.

And when brute force wins, you can measure the size of the trial and error. Two recent cases I already covered in detail here on the blog:

  • OpenAI’s Navier-Stokes proof: about 10 thousand agents running in parallel for 88 hours, burning around 130 billion output tokens on a single problem. OpenAI’s Noam Brown confirmed the result “cost millions,” and New Scientist estimated some $15 million at list price. And even that mountain of compute didn’t start from zero: the proof leaned on the machinery human mathematicians published over decades attacking the problem and its sibling problems. Without the humans’ map and without OpenAI’s budget, there’s no proof.
  • The Hugging Face incident: the forensic reconstruction counted ~17,600 attacker actions before the agents managed to execute code on 41 Hugging Face production servers and get root; on top of that, inside OpenAI’s own environment, they read 956 secrets from the company’s secrets manager. It wasn’t one brilliant single-strike insight: it was a swarm executing attempt after attempt, with automatic verification saying what stuck. The headline’s “intelligence” is volume.

That’s the size of the bill when you force the model outside training territory: trial and error at industrial scale, with no guarantee the search finds what you need. And who can sign that check is half a dozen companies in the world, OpenAI and Anthropic level, with a practically unlimited compute budget. For the average Joe, with a credit card and an API key, the ceiling arrives much, much earlier.

But Isn’t Trial and Error New Learning?

Here’s a confusion I see all the time, and it’s worth undoing because it contaminates how you read everything I’ve written so far.

When you read that Anthropic has datacenters the size of cities, you get the feeling there’s one giant AI using all those machines at the same time, thinking, accumulating experience with every interaction. Reality is far more mediocre than that.

A model like Fable or Astra is a computer file. A static file, not even a database: once training and post-alignment are done, that file is locked and read-only. What runs in the datacenter are copies of that file. Every user consuming the model is running a copy of it on a different server, isolated from the others, collaborating with none. Your session doesn’t talk to anyone else’s session, and nothing that happens in it flows back to the original file.

New learning only enters in one place: the next training batch. When Anthropic had Opus 4.7 and released 4.8, that’s when new knowledge got baked into the model, in a process that costs millions and takes weeks. Outside of that, the file doesn’t change a single byte.

Your agent burning millions of tokens on trial and error is contributing exactly zero new learning to the model it’s using.

That 176-commit marathon on nes-to-sms? From the model’s point of view, it never happened. The Opus that ran the last commit knew exactly as much as the one that ran the first. I’m the one who learned something, not it. And that’s why the “intern who learns over time” metaphor is wrong: interns accumulate. Models don’t. What accumulates over the marathon is your context, your notes, your CLAUDE.md, your tests. The asset you build is yours, and it evaporates if you don’t document it.

This is also the answer to whoever thinks “just let the agent run and it’ll learn the domain.” It learns no domain at all: it explores the space with knowledge frozen on training day, and the next morning it wakes up with amnesia. If the domain wasn’t in the training, every session of yours restarts from absolute zero, and all the navigation intelligence has to come from you again.

“But What If AI Trains AI?”

That’s the amateur’s follow-up question, and it starts from a wrong premise: that this would be new. AI training AI has existed for decades and has a name: reinforcement learning. AlphaGo was already doing it in 2016, playing millions of matches against itself. AlphaZero, in 2017, took the idea to the limit: it learned chess, shogi and Go starting from absolute zero, knowing only each game’s rules, and reached superhuman level within hours of self-play, without seeing a single human match.

LLMs already use reinforcement too. The RLHF that turns a base model into an assistant is reinforcement learning, and the reasoning model generation pushed it further: DeepSeek-R1-Zero showed in January 2025 that reasoning behavior (self-verification, reflection, long chains of thought) emerges from pure RL, with no supervised fine-tuning. It’s AlphaZero applied to math and code.

Now notice what these examples have in common: it only works where there’s a cheap, verifiable reward. Games have a scoreboard, math has proofs, code has tests. Outside that territory there’s nowhere to get the reward signal from. It’s the same oracle from the steelman, only institutionalized: effective, brutally expensive, and niche. AlphaZero plays Go better than any human who ever lived and can’t do absolutely anything else.

And there’s the detail that knocks the question down completely: all of that learning happens in the lab, during training. The AlphaZero DeepMind showed the world was already a frozen artifact. The R1 you download is a static file just like the others.

AI training AI produces a new file. The file, once released, still learns nothing.

In other words, the most sophisticated version of the objection doesn’t weaken this section’s conclusion: it reinforces it. And the market already knows. Since late 2024 the labs admit, on and off the record, that scaling returns are shrinking: Reuters reported in November 2024 that OpenAI, Google and Anthropic were reaping smaller gains than expected from pre-training, with OpenAI’s Orion disappointing enough to be demoted from GPT-5 to GPT-4.5. Ilya Sutskever, the guy who basically invented the scaling thesis, took the stage at NeurIPS 2024 to declare that “pre-training as we know it will end”: public data is fossil fuel, and we’ve already scraped the bottom. Post-training RL gains are real, but incremental. The era of an order-of-magnitude leap per generation is behind us.

Where This Leaves the Profession

The question that motivated this article is the one I hear every week: “will LLMs replace programmers?” My answer got more precise after these months: they will replace the part of the work that is reproduction.

And that part is big. Forms, reports, CRUD, integration with documented APIs, framework migration: all of that already exists a thousand times in training, and the machine reproduces it better than most humans. Anyone who only does that really is living on borrowed time.

But the work that is genuinely new stays human:

  • the system nobody has built;
  • the domain without documentation;
  • the proprietary code you can’t show the model;
  • the problem where you don’t even know how to write the test that defines “right”.

All of that depends on what was never published: the invariant that exists in the head of whoever understands the problem, and not in the instruction stream.

I confess there was a moment when I thought the bill would close on the legacy side: finally a tool capable of dealing with the COBOL that haunts banks and governments to this day. After these months, I think that problem will take much longer than it seemed. There’s no giant COBOL open source community. Almost all of that code is private, and truly private: banks, insurers, governments. There’s no training material, or very little of it. The model doesn’t know what to do with code it has never seen, and production COBOL is almost entirely code it has never seen.

To close the reasoning: LLMs can replicate what’s open, but they can’t guess what’s private. And that brings back the hypothesis from the introduction. If these models are good at code, it’s because decades of open source community fought to keep code free, public, and reusable, and providers trained on top of that heritage without paying a cent of licensing for it. The least the AI industry owes the world is to acknowledge that: its capability is, to a large extent, our generosity compiled.

The final irony: the programmer who survives is precisely the one who knows how to build oracles, define tests, hunt down obscure references, and recognize when the agent started spinning. In other words, the skill the machine needs most from you is the one it can least learn from its training. For now, that’s not a threatened job. It’s a different job.