AI Is Running Out of Internet, and It's Starting to Eat Itself

cover
26 Aug 2026

The web gave language models their first picture of humanity. The next generation increasingly risks being trained on a copy of a copy of that picture, and every copy leaves something behind.

In 1969, the American composer Alvin Lucier sat down inside a room with a microphone, a loudspeaker, and a deceptively simple idea about what would happen if a recording were forced to keep listening to itself.

He read a short prepared text into the microphone and recorded his voice on tape. Then he played that recording back into the same room, captured the playback as a new recording, and repeated the process again and again, with each generation becoming the source material for the one that followed.

At first, Lucier was still unmistakably there. You could hear the pace of his speech, the words he’d chosen, and the particular hesitations in the voice that produced them. But the room was there too, adding its own acoustic character every time the sound traveled from the loudspeaker, crossed the physical space, reached the microphone, and returned to tape.

Some frequencies fit the room more naturally than others. Those frequencies were reinforced during the first playback, reinforced again during the next one, and then handed forward as though their growing dominance had always belonged to the original performance. Meanwhile, the finer details that made Lucier's speech intelligible began to recede.

The consonants softened. The words blurred. The meaning disappeared. Eventually, almost nothing recognizably human remained, and the recording had become a field of resonant tones determined largely by the shape of the room surrounding the voice.

Nothing had broken, and no outside saboteur had introduced the distortion. Every machine in the process had done exactly what it was asked to do. The loss emerged because each new generation treated the previous generation as though it were still the original.

Lucier turned that loss into art.

The LLM Cartel is preparing to turn the same broad mechanism into training data.

The first recording was the human internet

The current generation of LLM AI was built by absorbing an extraordinary record of human expression. Books, articles, websites, forums, code, research papers, reference works, conversations, and all the glorious disorder in between gave the models enough examples to learn the statistical relationships that make their responses feel coherent.

That first pass through the human internet was enormous, but it was never infinite. People write new things every day, but nowhere near fast enough to feed models this hungry. And the models can't just swallow whatever they find. The words have to be reachable, legal to use, varied, and clean enough to teach the model something instead of just padding it out.

Researchers at Epoch AI have run the numbers, and if things keep going the way they are, the models will have chewed through just about all the useful human writing the public web has to offer sometime between 2026 and 2032. That doesn't mean the internet suddenly runs out of sentences. It means the scaling strategy reaches the point where another model can’t simply find another internet's worth of fresh human language waiting behind the one it already consumed.

The difference matters because the LLM Cartel built its progress story around abundance. More data helped produce broader capability, so the natural answer to the next model was more data again. But once the available human corpus has been scraped, filtered, deduplicated, licensed, and repeatedly reused, more no longer means discovering a second untouched civilization of writers. And so that means it will increasingly have to manufacture the material internally.

And so the recording gets played back into the room.

Synthetic data isn't poison

This is the point where the argument often becomes much sloppier than the research allows. Synthetic data’s not automatically bad data, and a model doesn’t collapse just because another model helped produce part of its training material.

Carefully generated examples are already useful for teaching systems to follow formats, solve particular classes of problems, practice rare situations, and improve smaller models with output selected from stronger ones. Simulations produce valuable data where real examples are scarce or dangerous to collect, while human review, external verification, and deliberate filtering all change the quality of what survives into the next training run.

The danger begins when generated material is mixed into the corpus indiscriminately, its origin becomes invisible, and a model starts learning the world from statistical choices made by an earlier version of the same kind of system. At that point, the copy is no longer supplementing the original. It’s gradually replacing it.

That difference is a huge story, because a generated sample doesn’t preserve every possibility contained in the original or even the originals that produced it. Common patterns appear often. But unusual combinations, minority expressions, rare events, and awkward exceptions appear less often. Sample the model once and some of those edges get missed. Train the next model on that sample, and the stuff that fell out starts to look like it was never there in the first place.

So, the average moves toward the center. Then the center becomes the next average. And so on, and so on, and so on…

The edges disappear first

In 2024, researchers from Oxford, Cambridge, Imperial College London, the University of Toronto, the Vector Institute, and the University of Edinburgh published the central warning in Nature. They called the process model collapse, a degenerative effect in which generative models trained recursively on generated data begin to forget the true distribution the original data represented.w

The early stage is the most important because it doesn’t necessarily look like collapse at all. Low-probability events disappear first, which means the model may continue producing smooth, competent, entirely readable output while becoming less able to represent the uncommon cases sitting at the edges of human experience.

Then the losses compound. The next generation learns from a world already narrowed by the previous one, adds its own sampling and approximation errors, and hands an even thinner version forward. Over enough generations, the model drifts farther from the original distribution while concentrating around a smaller set of dominant patterns.

This isn't the chatbot forgetting the beginning of a long conversation, and it isn't a model producing a single strange answer. It’s a training-level failure across generations as the system loses contact with the full shape of the source material because every new version is learning from an echo that’s already discarded part of it.

Lucier's room made certain frequencies louder because the architecture favored them. Recursive model training does something conceptually similar with probability. The common becomes more common, the unusual becomes less visible, and whatever the earlier model distorted returns as material the next model is asked to trust.

By the time the words dissolve, the real damage happened several generations earlier, when the system still sounded perfectly fine and nobody noticed which parts of the human signal had stopped returning.

The internet is becoming part of the room

There’s another problem the laboratory experiments didn’t need to create, because the open web is now creating it on its own. Model-generated writing is flowing back into websites, product descriptions, search results, social posts, customer support pages, code repositories, academic submissions, and content farms, often with no form of label separating it from human work.

That doesn't mean every synthetic sentence is weak, false, or useless. It means that the provenance of the next training corpus becomes harder to establish. After all, a model collecting text from the public web increasingly encounters language produced by models built from earlier collections of that same web, which turns the training environment itself into Lucier's room.

The obvious defenses all make sense. Preserve original human data. Track where generated material came from. Keep verified reality in the training mixture. Pay for high-quality private corpora. Use human experts to evaluate outputs. Generate synthetic data for a defined purpose rather than treating machine volume as a substitute for human variety.

But the reality at the very base of every defense also reveals an economic shift underneath the technical one. The scarce resource is no longer text by itself. It’s text with a trustworthy relationship to the world that produced it, and that resource costs more to identify, license, verify, and protect than a web page waiting to be scraped.

The industry's old advantage was that humanity had already paid to create the first recording. The next recording comes with invoices attached.

Learning shouldn't require another internet

This is where Vertus begins from a different dependency. It isn't an LLM waiting for another web-scale training run, and it doesn't treat intelligence as a frozen statistical picture that has to be replaced wholesale whenever the system’s expected to know more. Vertus describes what it’s built as a Cognitive Reasoning Superintelligence, with persistent, structured knowledge that develops through real interactions in the world where the intelligence is being used.

The plain-English version is straightforward. What Vertus learns about a person, project, domain, entity, or relationship doesn’t disappear when the conversation ends, and it doesn’t return merely as a transcript pasted into the next prompt. The knowledge remains organized and becomes a part of the context from which the system forms a new cognitive structure for the next problem.

Vertus calls that persistent structure Cognitive Memory. Its significance here is not that memory replaces external evidence or makes bad information harmless. Every intelligence still depends on the quality of what reaches it. The difference is that development no longer depends entirely on compressing a larger public corpus into another fixed model and hoping the next compression preserved everything that mattered.

A research system develops deeper knowledge of the research domain through the work it actually performs. A business system accumulates structured understanding of the business it’s actually serving. The next question begins with what’s been learned from those direct relationships, while its problem-specific neural topology is generated around the work now sitting in front of it.

That’s a fundamentally different learning loop. One architecture tries to improve by feeding another generation a larger representation of what language already looked like. The other develops by keeping what reality teaches it and reorganizing that knowledge as new problems arrive.

Neither approach eliminates the need for human knowledge, because that was never the objective. The question is whether human knowledge remains a living source the intelligence continues to learn from, or becomes a finite deposit that must be mined, copied, compressed, and eventually replaced by its own exhaust.

Eventually, the room is all you hear

Alvin Lucier knew exactly what would happen when he placed his voice inside that loop. The disappearance was the point of the piece, and every performance made the fullness of a particular room audible by allowing it to overwhelm the person who had started the sound.

The AI industry doesn't have that luxury. A training system that gradually loses rare knowledge, minority patterns, inconvenient exceptions, and contact with original human expression hasn't discovered an interesting new composition. It's mistaken its own narrowing echo for the world.

This is why the internet problem is larger than running out of websites to scrape. The first danger is scarcity. The second is recursion. And the third is that recursion remains almost invisible while the output still sounds polished enough to pass for some sort of original voice.

But none of this is a law of nature. It's the fate of one particular design, the fixed LLM model that has to keep eating a bigger copy of the past to learn anything new. A system built the other way, one that keeps learning from real people and real work as it goes, never has to climb back into the loop in the first place. That's the whole point of building a cognitive intelligence that stays in contact with the world instead of a recording of it.

A room has a sound of its own. So does every model. A cognitive reasoning intelligence is the one built to keep listening past its own walls, to keep learning from real people and real work instead of an echo of what it already said. That's how the human signal keeps coming through.

The alternative is a machine slowly turning up the sound of the room until that's all that's left. And when that happens, we won't hear humanity through the machine at all.

We'll hear the architecture talking to itself.

This story was distributed as a release by Jon Stojan under HackerNoon’s Business Blogging Program.