How AI Mistranslates Terminology: Why ChatGPT Picks a Different Word for the Same Concept Across Your Strings
Terminology consistency is one of those localization problems that doesn't announce itself — it just quietly erodes user trust until someone leaves a one-star review explaining that your app "feels unfinished." AI translation tools like ChatGPT are fast and often accurate at the sentence level, but they have a structural blind spot: they have no memory of what they called the same concept three strings ago.
This post is part of a series on AI translation failure modes in mobile apps. If you've already read about placeholder corruption and structural file breakage, this is the third pillar: semantic drift. The kind of failure that passes every structural check and still ships broken.
Why LLMs Have No Terminology Memory
Every call to the ChatGPT API — or any LLM API — is stateless by default. The model has no awareness of what it translated two minutes ago. Feed it account_settings_title at 2:00 PM and edit_account_button at 2:01 PM and it treats them as completely independent problems. There's no internal state saying "I already decided 'Account' maps to 'Cuenta' in this app."
This matters more than it might seem. Language models generate output by sampling from a probability distribution. The same source string, translated twice, can produce different outputs — not because the model is broken, but because that's how sampling works. With temperature above zero (which is the default in most integrations), you get variation. For creative writing, that's a feature. For UI strings, it's a reliability problem.
Even when you batch strings into a single prompt, the problem doesn't fully go away. LLMs handle long contexts imperfectly — attention degrades, earlier constraints get forgotten, and a glossary instruction you put at the top of a 200-string prompt may have lost its grip by string 150. This is well-documented behavior, not an edge case.
The practical result: unless you're actively enforcing terminology at the infrastructure level, you will get drift. Not sometimes. Every time, at scale.
The 5 Most Common Terminology Drift Patterns
1. UI nouns with multiple valid translations
Take the English word "Account." In Spanish, you have at least three defensible options: Cuenta, Perfil, and Cuenta de usuario. All three are correct. All three will appear in an AI-translated app if nothing constrains the choice. Users navigating from a login screen to a settings screen to a profile page will encounter a different word each time for what is functionally the same concept.
2. Action verbs that map to several target-language options
"Submit" in Spanish can be Enviar, Confirmar, or Aceptar — and all three appear in the wild. In German: Senden, Absenden, Bestätigen. Each translation is technically defensible in isolation. Together, they make the app feel inconsistent at best, confusing at worst.
3. Brand-adjacent terms re-translated instead of kept verbatim
Feature names, plan names, and product-specific terminology should almost never be translated — they should be kept as-is or substituted from an approved list. An LLM doesn't know that "Pro Plan" is a brand term and not a generic phrase. It will translate it. "Plan Pro" in one language file, "Piano Pro" in another, "Professioneller Plan" in a third. Your brand consistency is now gone, and no CI check will catch it.
iOS Human Interface Guidelines specify "Tap." Android Material Design uses "Tap" as well. ChatGPT, working without platform context, will sometimes output "Click" or "Press" in translated strings — not because it's wrong in an absolute sense, but because it doesn't know you're targeting a touchscreen and it has no platform convention awareness baked in. This issue compounds in translated strings where the AI is choosing both the verb and its target-language equivalent simultaneously.
Spanish distinguishes formal (usted) from informal (tú). German distinguishes Sie from du. These are not interchangeable. An app that uses tú on the onboarding screen and usted in the settings flow will feel strange to native speakers — not broken, just subtly off, like someone switching between formal and informal modes mid-conversation. LLMs default to whatever register fits the local context of the string, with no awareness of the register decisions made elsewhere in the app.
A Concrete Before/After: strings.xml Terminology Drift at Scale
Here's a realistic Android strings.xml snippet showing five strings that all reference the user's account — translated by five independent ChatGPT calls:
<!-- Before: AI-translated, no glossary enforcement -->
<resources>
<!-- String 1 -->
<string name="nav_account">Cuenta</string>
<!-- String 2 -->
<string name="profile_header">Tu Perfil</string>
<!-- String 3 -->
<string name="account_settings_title">Configuración de la Cuenta de Usuario</string>
<!-- String 4 -->
<string name="edit_account_button">Editar Perfil</string>
<!-- String 5 -->
<string name="account_deleted_message">Tu cuenta ha sido eliminada. Todos los datos de tu perfil se han borrado.</string>
</resources>
Five strings. Three different words used for the same underlying concept (Cuenta, Perfil, Cuenta de Usuario). String 5 manages to use two of them in a single sentence.
Now here's the same snippet after glossary-enforced translation, where the pipeline has been told: Account → Cuenta (always, no exceptions):
<!-- After: glossary-enforced, memory-aware translation -->
<resources>
<!-- String 1 -->
<string name="nav_account">Cuenta</string>
<!-- String 2 -->
<string name="profile_header">Tu Cuenta</string>
<!-- String 3 -->
<string name="account_settings_title">Configuración de la Cuenta</string>
<!-- String 4 -->
<string name="edit_account_button">Editar Cuenta</string>
<!-- String 5 -->
<string name="account_deleted_message">Tu cuenta ha sido eliminada. Todos los datos de tu cuenta se han borrado.</string>
</resources>
Consistent. Every instance of "account" resolves to Cuenta. The sentence in string 5 now reads uniformly.
The important thing to understand here is that the fix isn't "write a better prompt." You can tell ChatGPT to use Cuenta for "Account" in the system prompt, and it will comply — until it doesn't. In a file with hundreds of strings, context window effects, variation in how the term appears (possessive, plural, embedded in a longer phrase), and sampling randomness all work against you. The fix has to be structural: the glossary has to be enforced at the pipeline level, with post-translation validation that checks for violations before the file lands in your repo.
Why You Won't Catch This in Code Review
The developers reviewing your PR are almost certainly not fluent in Spanish, German, or whichever language you're shipping. They'll look at the diff, confirm the XML is well-formed, spot-check a few strings, and approve. Structural correctness is easy to verify. Semantic consistency across 300 strings in a language you don't speak is not.
Lint and unit tests don't help here either. Your XML parser doesn't know that Perfil and Cuenta mean different things. Your CI pipeline validates format, not meaning. As we've discussed in the build-time locale validation post, most automated checks operate at the structural layer — missing keys, malformed plurals, broken placeholders. Semantic consistency requires a different category of check entirely.
So it ships. And a few weeks later, you get a review from a native Spanish speaker who says something like "La aplicación usa palabras distintas para la misma cosa, parece que está mal hecha." Which it is, in a sense — just not in any way your development process was designed to catch.
This is the same failure mode described in why AI translation loses context between strings: the AI was confident, the output was plausible, and nothing in the review pipeline had the information to flag it.
The Fix: Glossary Enforcement and Translation Memory
Two mechanisms work together here, and they're worth understanding separately.
Translation memory stores previously approved translations for specific segments. When the same source string appears again — in a future release, a new feature, a different screen — the system reuses the approved translation rather than generating a new one. This handles exact and fuzzy matches and is the primary defense against drift across releases and across features.
A managed glossary is a curated list of key terms and their approved translations per language. Account → Cuenta (es), Submit → Enviar (es), Settings → Configuración (es). When any string containing "Account" is translated, the glossary term is injected or enforced — either by modifying the prompt automatically, by post-processing the output, or by flagging glossary violations for review.
The critical distinction is where this enforcement happens. Putting glossary instructions in a prompt is a hint. Enforcing glossary compliance at the pipeline level — with validation that rejects or flags non-compliant outputs — is a constraint. The former is better than nothing; the latter is what actually works at scale.
This is exactly the problem that purpose-built localization platforms are designed to solve. Tools like GetTranslated.AI maintain translation memory and glossary enforcement as first-class pipeline features, not as prompt-engineering workarounds. If you're evaluating platforms, safe AI translation covers what that looks like in practice — including how glossary violations are surfaced before files land in your repo.
Practical Checklist: Spotting Terminology Drift Before Release
You don't need a full platform integration to start catching terminology drift. Here's what you can do today:
1. Export and diff translations for key terms across all strings
Write a script that extracts every string containing a specific source term (e.g., "account") from your English base file, then finds the corresponding translated strings in each target locale file. Print them side by side. If you see more than one unique translation for a term that should be consistent, you have drift.
# Rough bash approach for Android strings.xml
grep -n "account" res/values/strings.xml | awk -F'"' '{print $2}' | while read key; do
echo "=== $key ==="
grep "name=\"$key\"" res/values-es/strings.xml
done
It's not elegant, but it will surface inconsistencies fast.
2. Grep for expected glossary terms in output files
If you've decided that "Account" maps to "Cuenta" in Spanish, write a check that confirms Cuenta appears wherever "account" appears in source strings, and flags occurrences of Perfil or Cuenta de usuario in those same positions. This can be a simple Python script or a CI step.
3. Route at least one native-speaker review pass per language per release
This is the check that catches everything the scripts miss — register inconsistency, awkward phrasing, brand terms that got translated when they shouldn't have. It doesn't need to be exhaustive. A 30-minute review from a fluent speaker focused specifically on key UI flows will catch the majority of high-visibility issues. The goal is to find the problems before your users do.
Conclusion
AI translation is fast. It covers ground that would take human translators days or weeks. That's real and worth having. But speed without consistency produces apps that feel amateur to the people who matter most: native speakers of your target languages.
The terminology drift problem is structural. It's not a prompt quality issue or a model quality issue — it's a fundamental consequence of stateless inference applied to a problem that requires stateful memory. The fix is equally structural: glossary enforcement and translation memory, applied at the pipeline level, validated before merge.
AI speed is real. AI consistency is not — unless you enforce it.
Next up in this series: how to build a pre-translation glossary and translation memory that survives team turnover, including what to do when the person who made the original terminology decisions has left the company and the glossary lives in a Notion doc that nobody updates.