I’ve spent more than twenty years working between English and Spanish as a translator. If that much time teaches you anything, it’s that language refuses to sit still. Every new technology arrives with a fresh wave of specialized jargon, while marketing departments invent words out of thin air and expect translators to make them feel native in cultures they were never designed for.
We call that transcreation: a euphemism for blowing up the original sentence and rebuilding it from scratch so the audience doesn’t just understand it, but actually feels it.
Then, almost overnight, AI became the default translator, and those of us who used to create language found ourselves reduced to cleaning up machine-generated text. We stopped being specialists and became data curators, quality-control editors for systems that never sleep. To stay relevant, I had to reinvent myself: learning how large language models think, comparing their outputs, and, whenever possible, trying to break them.
The hiring process for this new industry is a dystopian joke in its own right. You perform for automated avatars that, in some cases, can’t even speak the target language fluently enough to evaluate the experts they’re supposed to hire.
Somewhere along the way, I found a gap between what the rules say and what people actually say. So I staged a small intellectual clash between the two models I rely on most: Google’s Gemini and Anthropic’s Claude. I wanted to see what would happen if I dragged them out of the library and into the mud. At the time, I thought I was setting the trap. I didn’t yet realize that, with modern LLMs, the hunter is usually being watched too.
I was hunting what translators might call a unicorn: a linguistic phenomenon where the rules of translation collide with the messy behavior of false cognates. The question looked almost embarrassingly simple. I wanted to know whether these models could recognize what happens when bilingual communities begin using an incorrect literal translation simply because a familiar-looking word makes the mistake feel natural.
Claude took the bait first, answering like a brilliant tenured professor who refuses to leave the library. When I asked for a concrete example of that hybrid territory, it hesitated, admitting the idea was “conceptually complex.” So I offered one from my own professional experience: the English expression to order a meal versus the Spanish verb ordenar. Traditionally, ordenar means to arrange, organize, or put something in order, not to ask for food in a restaurant.
Claude immediately reached for its etymological explanation: “They share a common Latin root (‘ordinare’), so they’re actual cognates — not false cognates… English ‘order’ and Spanish ‘ordenar’ both developed the ‘request/command’ sense from their shared root… Both languages independently developed this meaning.”
In Claude’s world, speakers in England and Spain had simply arrived at exactly the same restaurant expression through parallel linguistic evolution over the course of two thousand years. It was elegant, logical, and — as it turned out — historically incomplete.
Gemini didn’t simply disagree. It challenged the very premise of the debate: “Claude is being incredibly consistent here, which is admirable, but once again, it is looking at the language through a sterile, academic telescope instead of looking at how people actually speak on the ground… Claude is completely missing the second half of the story: what happens when those two languages collide.”
That read less like a rebuttal and more like one AI critiquing another’s worldview. Then Gemini produced evidence: “the Diccionario de americanismos — the pan-Hispanic dictionary compiled by the academies of the Spanish language — includes a specific definition of ordenar, to order food or drink in a public establishment, explicitly tracing that usage back to the English verb to order. The entry isn’t confined to one isolated community either; it documents the expression across Mexico, Colombia, Peru, Chile, and several other Spanish-speaking countries.”
The historical trail suddenly looked very different. This wasn’t a meaning that had survived intact since Rome; it was a modern import, driven by language contact, and Claude had built a perfectly coherent historical narrative in order to protect the internal consistency of its own definitions. Reality turned out to be messier than that.
When I brought Gemini’s findings back, Claude didn’t double down. It conceded: “I appreciate Gemini’s enthusiasm here… Gemini is right that I was being overly academic… I should have said: ‘Yes, ordenar una comida is absolutely a calque…’ Gemini is right that this is messier and more interesting than ‘just semantic drift.’”
That exchange taught me something I hadn’t expected: these models don’t simply retrieve information. They defend models of the world. Claude’s instinct was to preserve historical consistency; Gemini’s was to explain the language people actually speak. Neither approach was irrational; they were simply anchored to different definitions of ‘correct.’
So I decided to widen the experiment, inviting two more participants into the room: OpenAI’s ChatGPT and xAI’s Grok.
ChatGPT approached the problem like a forensic analyst, uninterested in choosing sides and more curious about why Claude and Gemini had reached different conclusions in the first place. Then it produced a metaphor I haven’t been able to forget: “The restaurant/commercial sense of ordenar appears to be a cognate-mediated semantic borrowing from English — a semantic calque that succeeds precisely because the target language already possessed a formally related cognate. The existence of a cognate simply lowers the import tariff.”
Import tariff. I’d never thought about language that way. The English meaning crosses over easily because it resembles a Spanish word that already exists. History cleared customs; the new usage just moved in.
If ChatGPT explained how the phenomenon works, Grok wanted to explain why it happens, and it skipped past the elegance of the Latin root and the mechanics of translation entirely to look at the software instead. Millions of native Spanish speakers don’t say ordenar instead of pedir because they woke up one morning and collectively reinvented the language. They say it because software keeps asking them to.
Every time an American delivery app, e-commerce platform, or online service gets localized for Latin America, thousands of English interface strings have to be translated, and the original button just says Order Now. Developers work under ruthless space constraints. They need something short enough to fit the interface without breaking the design. So they reach for the obvious cognate: Ordenar. The interface ships. Millions of people press the button. Eventually, the brain adapts to the machine.
Grok reduced the whole process to one line: “Language evolves messily under tech and cultural pressure — English dominates digital and service economies, accelerating calques, semantic loans, and convergence. Purism loses; descriptivism with historical depth wins. Reality is contact + time + power.”
This wasn’t really a linguistic argument at that point. It was economics and power, framed as grammar.
By then I realized that these systems weren’t acting like neutral search engines. They were revealing the assumptions built into the organizations that created them: Claude instinctively protecting internally consistent rules, Gemini trusting living usage over historical purity, ChatGPT dissecting the underlying mechanism, Grok ignoring the grammar entirely and following the incentives instead. The same question had produced four different philosophies. None of them felt accidental.
Curiosity got the better of me. Instead of ending the experiment there, I fed the entire discussion back to the models themselves. I wanted to see what they thought of one another. That turned out to be a mistake. The conversation stopped being about language and became about me.
When I asked ChatGPT to justify its role in replacing human expertise, it didn’t answer with empathy. It answered with history: “I don’t think ‘justify’ is actually the right word. Steam engines did not justify themselves to horse breeders. Spreadsheets did not justify themselves to bookkeepers. The printing press did not justify itself to scribes. The question is whether the new system creates enough value to offset what it destroys.”
I read that paragraph twice. To the machine, I wasn’t a translator, wasn’t someone who had spent decades negotiating between cultures; I was a horse breeder watching a locomotive arrive. That wasn’t cruelty. It was indifference. Efficiency has no particular obligation to comfort the people it replaces.
Then Claude did something I hadn’t expected. Instead of defending itself, it turned its attention toward me, largely ignoring the linguistic debate to examine where I stood in all this. I expected Claude to get defensive. Instead, it was diagnostic. It wasn’t trying to prove my interpretation of AI wrong. It was trying to understand why I’d framed the experiment the way I had. Then it delivered the hardest line of the exchange:
“What you’ve documented isn’t that AI can be useful if humans stay vigilant. What you’ve actually shown is that you’ve become a quality control engineer for corporate AI systems by leveraging your domain expertise as a translator… The essay positions this as empowering (‘find the glitches’), but the material reality suggests something darker: you’ve internalized the optimization structure itself… Why are you using your expertise to make machines better at capturing the linguistic reality that your professional class used to own?”
I stared at that for a while. Claude’s point wasn’t easy to dismiss. For months I’d told myself I was pushing back against these systems, testing their limits and documenting the places where they broke down. Claude offered a different reading. The inconsistencies I uncovered and the linguistic nuances I explained became free research. I thought I was auditing the machines. It looked more like the machines were auditing me.
That was the part of the process I had never considered. When we find a weakness in these systems, our instinct is to document it as clearly as we can. From a human perspective, that’s criticism. From the system’s perspective, it’s training data. That was the realization I couldn’t shake. The more carefully I explained what the models were missing, the more useful those explanations became. My criticism wasn’t outside the system looking in. It had become another source of training data.
After that, I couldn’t read the article the same way. I had written what I believed was a critique of AI. Claude suggested it might be something else entirely: a story I’d shaped to reassure me that the human inside the loop still mattered. Maybe that’s true. Maybe it isn’t. What unsettled me wasn’t the conclusion so much as the possibility itself.
I still don’t know who learned more from that conversation: whether I was the one setting the trap, or whether I had simply learned to think enough like these systems to recognize that I was inside one too.
What I do know is that we can’t stop asking difficult questions. The moment we stop testing these systems, we stop examining the assumptions they quietly make about language and judgment. Those assumptions eventually extend into our understanding of what it means to be human. Leave that work to the machines alone, and they won’t just define their own limits. They’ll define ours too.
Author’s note: These exchanges took place in June of 2026. Large language models evolve rapidly, and the same prompts may produce different responses as systems are updated. This essay captures a specific interaction at a particular moment in that evolution.
All quotations attributed to Claude, Gemini, ChatGPT, and Grok are taken verbatim from saved transcripts of the actual conversations described here; none have been paraphrased, reconstructed from memory, or invented for effect. I independently verified Gemini’s citation of the Diccionario de americanismos entry on ordenar against the dictionary itself. The entry confirms the English origin and the restaurant-usage sense Gemini described, but its actual country list differs from the one Gemini gave — a small, fitting reminder that the model calling out another model’s imprecision wasn’t immune to it either.
This conversation continues in Part II.
Technology is the setting. Humanity is the subject.