Twenty years ago, if you asked someone what reasoning meant, they would describe a process that happened inside a human mind. If you asked about judgment, they would talk about wisdom earned slowly through experience. Autonomy meant the freedom of a person to choose their own path. Consciousness meant being awake in the world.
Today those same words appear in product announcements.
A model can reason. An agent can act autonomously. A system can make judgments. A machine may even possess a form of consciousness.
The words did not change slowly over centuries. They moved in less than a decade. And almost nobody noticed it happening.
The Short Version (read this first)
Languages have always changed. Nobody consented when awful stopped meaning awe-inspiring, or nice stopped meaning foolish. Word change is natural. That’s exactly why what’s happening now is different.
In 1986, computers borrowed words from the world they found: bug, chip, boot, cursor. By 2006, the internet took web, virus, stream, search. Inconvenient, but recoverable. The changes were slow enough for culture to absorb and contest them.
What’s happening in 2026 has a feedback loop with a speed and scale that previous language evolution never approached. The printing press narrowed dialects over centuries. Broadcast media standardised accents over decades. The internet compressed that to years. AI compresses it further still and uniquely, the system generating the outputs is the same system being trained on them. The loop is internal in a way it never was before. AI models trained on human writing are now generating content that feeds the next generation of training data. The mirror is no longer passive. It now participates in shaping the reflection.
The words being annexed now are not borrowed from gardens and kitchens. Reasoning. Consciousness. Judgment. Autonomy. These are the words we use to describe what makes us human. They’re being quietly reassigned by an industry that needs them to describe its products.
Model collapse research suggests the possibility of significant linguistic narrowing, and while modern training pipelines actively work to mitigate it through data filtering and human data mixing, the risk scales with the volume of synthetic content entering the ecosystem. The mitigation is reassuring. The trajectory of synthetic content growth is not. Maybe linguistic diversity narrows. Minority languages fade. Unusual patterns wash out. What may survive is whatever was most popular, most reproduced, most indexed.
No villain planned this. No ministry ordered it. It’s linguistic natural selection, with a fitness function measuring popularity, not truth.
The dictionary is rewriting itself. And the words we need to object are already on the shortlist for reassignment.
The Full Argument
There is a comfortable story we tell about language and power. Someone in authority, a government, a corporation, an ideology, decides to change what words mean. They push the new meaning through institutions, media, and education until the old meaning fades. Orwell described it precisely in 1948: a Ministry of Truth, methodically rewriting the past, controlling the present through the careful management of vocabulary.
It is a story with a villain. And villains, at least, can be resisted.
What is happening to language right now is more unsettling, because it has no villain. No ministry. No agenda. Just an enormous statistical mirror, reflecting humanity back at itself, and in doing so, quietly changing what we are able to think.

The First Appropriation: Technology Borrows Your Words
Technology has always helped itself to the dictionary. In 1986, when the personal computer arrived in homes for the first time, it came with a peculiar vocabulary assembled from the world around it. A bug was something crawling in the garden. A chip was something you ate. A boot was something you pulled on before stepping outside. A cursor was a person who used bad language. These words were borrowed because metaphors are cheaper than inventing new ones, and then the metaphors hardened into definitions.
By 2006, the internet had completed what the PC started. Web was no longer something a spider made. Virus no longer meant catching the flu. Apple and Blackberry stopped being things you ate. Nobody asked permission. Nobody offered compensation. The words simply changed hands.
You might reasonably object: languages have always changed without consent, and they have always had feedback loops. The printing press standardised spelling through repetition over centuries. Broadcast media flattened regional accents through exposure over decades. The internet compressed that to years. These were all self-reinforcing systems.
That objection is correct. And it makes what is happening now easier to understand precisely because it fits a pattern — while breaking it in one critical way.
Language has experienced feedback loops before. The printing press standardised spelling through repetition, and broadcast media flattened accents through exposure. But those loops still depended on human authors. What is new is not the loop itself. It is that language is now being generated in industrial quantities by systems trained on language itself. The loop is internal. The author has left the room.
This is not simply a faster version of what came before. It is a qualitatively different arrangement — because for the first time, one of the participants shaping language is not human.
The shift is not that language changes. Language has always changed. The shift is that language now has a non-human participant — one that generates, reinforces, and trains on its own outputs, with no cultural memory, no lived experience, and no stake in what is lost.
That is the interface that has changed. And the interface, as it has always been, is language itself.
The Second Theft: Your Words Trained the Machine
Here is where two distinct but connected issues converge. It is worth being precise about them.
The first is a matter of copyright and labour. The AI systems now reassigning your language were built on your language. Large language models were trained on the internet, on decades of human writing scraped from blogs, forums, articles, books, comment sections, creative writing platforms, and social media posts. The collective output of human thought, poured into a digital ocean over thirty years, was used to teach a machine how to sound like a person. Most of it without consent. Much of it without compensation. That is an active legal and ethical dispute, with court cases in progress in multiple jurisdictions.
The second issue is different, and deeper. It is not about who owns the words. It is about what happens to meaning when a machine inherits language statistically rather than experientially.
The model learned your vocabulary, your sentence rhythms, your ways of explaining and arguing and comforting and creating. It learned what words mean by reading what you wrote when you thought you were simply talking to other humans. And now it uses that inherited fluency to generate content at scale, content that is already finding its way back onto the internet, into the training pipelines of the next generation of models.
The mirror is no longer just reflecting humanity. It is beginning to reflect itself.

Model Collapse: When the Mirror Loses the Signal
Researchers have begun documenting a phenomenon called model collapse. While its long-term severity remains actively debated, the early evidence is consistent enough to take seriously.
When a model trains heavily on synthetic data, on text generated by previous models rather than by humans, something subtle begins to happen. The rich, contradictory diversity of human expression starts to wash out. Statistical edges fade: minority languages, regional dialects, unusual phrasings, niche expertise, counterintuitive arguments. What remains is the centre of mass, the most common patterns, the most frequently reproduced constructions. A greatest hits album with all the interesting B-sides removed.
The consequences for linguistic diversity are already visible in current models. Languages with smaller digital footprints, Welsh, Swahili, Tamil, most of the world’s 7,000 languages, are dramatically underrepresented in training corpora dominated by English-language Reddit, Wikipedia, and news. A Welsh speaker using an AI assistant is not getting a mirror of Welsh linguistic culture. They are getting a translation layer over an English-language statistical centre of gravity. The model did not decide to marginalise Welsh. But the outcome is marginalisation nonetheless.
Nobody planned this. No single actor decided that future AI should speak in a narrower register. It emerges from the mathematics of training at scale, which is precisely why it is harder to resist than deliberate manipulation.
Survival of the Fittest — But Fittest for What?
Here the absence of a villain becomes the argument’s sharpest edge, not a weakness to be explained away, but the central provocation.
When political or corporate actors manipulate language, when disruption is rebranded from a warning into a virtue, when flexibility becomes the polite word for job insecurity, when contested political terms get loaded with new ideological freight, you can trace the campaigns, name the actors, argue back. The manipulation has an address.
Model collapse has no address. It is linguistic natural selection operating on a synthetic ecosystem, where the environment was shaped by whoever got online first, wrote the most, and got indexed. The fitness function is not truth, or nuance, or representativeness. It is popularity. Reproduction rate. Statistical dominance.
And crucially: the feedback loop means the process accelerates. Each generation of model outputs feeds the next generation of training data. The convergence is not linear, it compounds. The ratchet turns, and the diversity of human expression that was the original raw material gradually narrows into something more uniform, more central, more statistically average.
Orwell imagined a Ministry of Truth rewriting the dictionary. The stranger and more vertiginous reality is that the dictionary may rewrite itself, through sheer weight of repetition, with no ministry required, and no one to hold accountable when the signal is gone.

What This Means for How We Think About AI
The practical consequence tends to get lost in more technical discussions.
The language we use to think about AI has been substantially shaped by AI, and by the industry building it. We debate whether machines understand, whether they know things, whether they can reason, using words that have been pre-loaded with assumptions about the answers. The vocabulary of the conversation has been set by the people with the strongest interest in a particular outcome.
This is not a conspiracy. It is something more mundane and harder to counter: a fast-moving industry that needed vocabulary quickly, borrowed what was available, and moved on before anyone examined what was quietly assumed in the borrowing.
Alignment, a word borrowed from engineering, now pressed into service as the term for one of the most complex ethical challenges humanity has ever faced, may be doing its job too efficiently. It makes an impossibly difficult problem sound like a calibration issue. Hallucination makes a fundamental epistemic failure sound like a charming quirk. Intelligence in artificial intelligence has been doing philosophical heavy lifting since 1956 that nobody formally authorised it to do.
A century from now, historians may describe this moment less as the arrival of artificial intelligence than as the moment language acquired a second kind of speaker. For thousands of years, every sentence in human history came from a mind that had experienced the world through a body. That constraint quietly shaped the entire architecture of knowledge, argument, storytelling, and explanation. Now, for the first time, sentences are being produced at scale by systems that learned language without ever inhabiting the reality it describes. The technology may evolve, the models may change, but that shift, the entrance of non-experiential speakers into the linguistic ecosystem, may turn out to be the most consequential change of all.
The Only Honest Ending
Here is the demand this argument leads to, stated plainly:
We need independent, multidisciplinary oversight of how AI systems define, use, and reinforce language, before the feedback loop completes another cycle. Not because anyone is acting in bad faith. But because the absence of bad faith does not prevent bad outcomes, and the ratchet does not pause while we deliberate.
Previous generations could afford to let language evolve at its own pace. The feedback loop changes that calculus entirely. By the time linguistic monoculture becomes undeniable, the training data that could have prevented it will be gone, overwritten by the outputs of models that were already converging.
The dictionary is rewriting itself. The question is whether we are paying attention while it still contains the words we need to object.





