GetTranslated.AI
AI Translation & Automation

6 Ways Claude and ChatGPT Silently Corrupt Your .stringsdict and Android Plurals (And Why You Won't Catch Them Until Production)

6 Ways Claude and ChatGPT Silently Corrupt Your .stringsdict and Android Plurals (And Why You Won't Catch Them Until Production)

You pasted your plural resource files into an LLM, the output looked reasonable, your CI passed, and then Arabic shipped with only two plural forms instead of six. Nobody noticed until a user in Cairo filed a bug report. Here's exactly why that happens — and five other ways it happens that you probably haven't caught yet.

Pluralization is the single highest-risk surface for one-shot LLM translation. The rules are language-specific and non-obvious, the file formats are deeply structured and unforgiving, and LLMs are confidently wrong in ways that are completely invisible at compile time. If you're using raw LLM output in your plural resource files without a validation layer, you're carrying bugs you don't know about. For a broader look at how AI translation goes wrong in mobile resource files, see our overview at /solutions/ai-translation-problems/.


Quick Primer: Why Plurals Are So Hard for LLMs

Before getting into failure modes, it's worth establishing why plurals break LLMs in particular — as opposed to regular strings, where the failure surface is smaller.

The Unicode CLDR defines up to six plural categories: zero, one, two, few, many, and other. English uses two: one (singular) and other (everything else). That two-form mental model is baked into essentially every LLM's training distribution, because English dominates. When you ask an LLM to translate plural strings into Arabic, which uses all six categories with complex rule boundaries, the model's default behavior is to emit one and other — because that's what it's seen in the overwhelming majority of training examples.

The three main plural file formats each have their own structure that makes this worse:

iOS .stringsdict is a nested plist XML format. Each pluralized key requires an NSStringPluralRuleType sub-entry, and every required CLDR category for the target locale must be present as a child key. The nesting depth is significant, and the format is strict about structure even though it's technically XML.

Android strings.xml with <plurals> uses a <plurals> block with quantity attribute values corresponding to CLDR categories (zero, one, two, few, many, other). The set of required quantities depends entirely on the target locale. You can ship an incomplete set and aapt will not complain.

ICU MessageFormat (used in React Native with i18next, react-intl, and similar) embeds plural logic inline: {count, plural, one{# item} other{# items}}. The # is a shorthand for the formatted count value, and the syntax is sensitive to bracket matching and keyword placement.

For deep dives on each format individually, see iOS pluralization with .stringsdict, Android plurals explained, and React Native pluralization with ICU.


The 6 Failure Modes

1. Missing Plural Categories

This is the most common failure and the one that ships most quietly.

Arabic requires all six CLDR categories: zero, one, two, few, many, other, with specific numeric range rules for each. An LLM translating English plural strings will almost always output only one and other for Arabic, because that's the English pattern — and often the only pattern in its training data for the target format.

On iOS, a .stringsdict missing required categories will silently fall back to other for every count. On Android, same behavior. Your Arabic users see the grammatically wrong plural form for every count except those matching other. You won't catch this in QA unless your testers manually verify plural rendering at counts like 0, 1, 2, 3, 11, and 100 — in Arabic.

<!-- What the LLM outputs for Arabic in strings.xml -->
<plurals name="message_count">
    <item quantity="one">رسالة واحدة</item>
    <item quantity="other">%d رسائل</item>
</plurals>

<!-- What Arabic actually requires -->
<plurals name="message_count">
    <item quantity="zero">لا رسائل</item>
    <item quantity="one">رسالة واحدة</item>
    <item quantity="two">رسالتان</item>
    <item quantity="few">%d رسائل</item>    <!-- 3–10 -->
    <item quantity="many">%d رسالة</item>   <!-- 11–99 -->
    <item quantity="other">%d رسالة</item>  <!-- 100+ -->
</plurals>

Languages with this risk: Arabic, Russian, Polish, Czech, Slovenian, Welsh, Irish — essentially any language outside the simple one/other model.


2. Broken .stringsdict XML Structure

.stringsdict files are plist XML, and the nesting structure is precise. Each pluralized string requires a specific hierarchy: an outer <dict>, a NSStringLocalizedFormatKey entry, and a nested format specifier dict with NSStringPluralRuleType and the category keys.

LLMs frequently re-indent this structure, collapse nested <dict> elements, or simply flatten the hierarchy when reformatting output. The result is a plist that parses as valid XML but fails NSPropertyListSerialization at runtime — or, worse, parses successfully but loses the plural semantics entirely because the structure no longer matches what the runtime expects.

<!-- Correct .stringsdict structure -->
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
    <key>%d items selected</key>
    <dict>
        <key>NSStringLocalizedFormatKey</key>
        <string>%#@item_count@</string>
        <key>item_count</key>
        <dict>
            <key>NSStringFormatSpecTypeKey</key>
            <string>NSStringPluralRuleType</string>
            <key>NSStringFormatValueTypeKey</key>
            <string>d</string>
            <key>one</key>
            <string>%d item selected</string>
            <key>other</key>
            <string>%d items selected</string>
        </dict>
    </dict>
</dict>
</plist>

<!-- What an LLM sometimes produces — collapsed, invalid -->
<dict>
    <key>%d items selected</key>
    <string>%d Elemente ausgewählt</string>  <!-- entirely wrong format -->
</dict>

This one actually can crash your app, depending on where the malformed plist is loaded and how aggressively your error handling is.


3. Placeholder Drift Inside Plural Branches

Pluralized strings often contain format specifiers (%d, %1$d, %s, %@, {count} in ICU). In a well-formed source file, every plural branch contains the same placeholder — possibly formatted differently, but present and correct.

LLMs rephrase. When rephrasing across plural branches, they frequently drop or change placeholders: %1$d becomes %d in one branch and disappears entirely in another because the model "completed" the sentence without needing a number. On Android, a missing %d in a quantity="few" branch means the getString(R.plurals.foo, count, count) call at runtime either crashes or silently omits the count from displayed text.

<!-- Source (English) -->
<plurals name="files_remaining">
    <item quantity="one">%1$d file remaining</item>
    <item quantity="other">%1$d files remaining</item>
</plurals>

<!-- LLM output for Polish — placeholder drift in 'few' -->
<plurals name="files_remaining">
    <item quantity="one">%1$d plik pozostał</item>
    <item quantity="few">kilka plików pozostało</item>  <!-- %1$d dropped -->
    <item quantity="many">%1$d plików pozostało</item>
    <item quantity="other">%1$d pliku pozostało</item>
</plurals>

For more on how placeholder corruption propagates across translation pipelines, why placeholders break during translation is worth reading before your next release.


4. Wrong Plural Rules Applied from a Different Language

Slavic languages have complex plural rules, and they differ from each other in ways that aren't obvious. Russian, Polish, and Czech all use more than two categories, but the numeric ranges that map to each category are different. Russian and Polish are often confused by LLMs — both use one, few, many, other, but Russian's few applies to numbers ending in 2–4 (excluding 12–14), while Polish's rules differ in their interaction with the many category.

An LLM that has seen more Polish localization examples than Russian ones may apply Polish-style rules when generating Russian translations, or vice versa. The category key names will be correct, but the strings themselves will be written for the wrong numeric ranges — so the grammatical form is wrong even though the structure looks valid.

This is almost impossible to catch programmatically. It requires a native speaker reviewing the plural branches with knowledge of the underlying CLDR rules. Languages at risk include Russian, Polish, Czech, Slovak, Ukrainian, Serbian, and Croatian.


5. ICU Syntax Mangling in React Native

ICU MessageFormat syntax is terse and positionally sensitive. The # shorthand expands to the formatted count value within a plural branch. Bracket depth must balance. The plural keyword must follow the variable reference directly.

LLMs mangle this in several distinct ways: they move the # outside the branch (breaking expansion), they nest translated text inside the outer {count, plural, …} wrapper rather than inside branches, they add spaces in the wrong places, and they occasionally translate the keyword plural or other into the target language.

// Correct ICU MessageFormat
"{count, plural, one{# unread message} other{# unread messages}}"

// LLM corruption: # moved outside branch, text misplaced
"{count, plural, one{unread message} other{unread messages}} #"

// LLM corruption: keyword translated (French — 'autre' is not valid)
"{count, plural, un{# message non lu} autre{# messages non lus}}"

// LLM corruption: bracket imbalance
"{count, plural, one{# message non lu} other{# messages non lus}"

i18next will throw a parse error for some of these, but the bracket-imbalanced form will sometimes parse silently and just render incorrectly. If your React Native tests don't exercise pluralized strings at multiple counts in every locale, this goes undetected.


6. Quantity Attribute Mistranslation in Android XML

This one is almost too simple, but it happens. The quantity attribute in Android <plurals> takes CLDR keyword values: zero, one, two, few, many, other. These are not translatable — they're format keywords.

An LLM translating an Android strings.xml file will sometimes translate these attribute values along with the string content. For French, quantity="other" might become quantity="autre". For German, quantity="few" might become quantity="wenige". The result is XML that aapt will either reject during build (best case) or accept and then silently ignore at runtime because the quantity value doesn't match any known category.

<!-- Correct -->
<plurals name="notifications">
    <item quantity="one">%d notification</item>
    <item quantity="other">%d notifications</item>
</plurals>

<!-- LLM output for French — quantity attributes translated -->
<plurals name="notifications">
    <item quantity="un">%d notification</item>      <!-- invalid -->
    <item quantity="autre">%d notifications</item>  <!-- invalid -->
</plurals>

Whether aapt catches this depends on your build configuration and API level. In some configurations it throws a resource compile error. In others it silently drops the invalid entries, which means you fall back to an empty string — also not great.


Why These Bugs Survive CI

Most CI pipelines validate XML and plist well-formedness. They do not validate semantic correctness of plural resource files. The gap between "valid XML" and "correct pluralization" is where all six of these failure modes live.

Specifically:

  • xmllint and plist validators check structure, not CLDR category completeness. A .stringsdict with only one and other for Arabic is well-formed XML. It's just wrong.
  • aapt / aapt2 will catch invalid quantity attribute values in some configurations, but will silently accept an incomplete set of categories for any locale.
  • swiftc does not validate .stringsdict content at all. It compiles successfully whether your Arabic plural file has two forms or six.
  • Runtime fallback behavior means that incomplete plural sets don't crash — they produce wrong output. Wrong output is much harder to catch in automated testing than crashes.
  • QA testers rarely manually verify plural rendering across the full count spectrum in every locale. Testing "the app shows the right string for 5 files" in English doesn't tell you anything about Arabic at count 11.

For a comprehensive look at what your CI pipeline should actually be checking, build-time locale validation: 8 checks your CI pipeline should run before every release covers the tooling side in detail.


What a Safe Plural-Translation Pipeline Looks Like

The core principle is: never let the LLM infer plural structure. Treat structure as a fixed constraint and use the LLM only for the text content within each branch.

Pre-translation: enumerate required categories per locale

Before any translation happens, generate a template for each target locale that lists every required CLDR category. CLDR data is publicly available and machine-readable. For Arabic, your template has six category keys. For Japanese, it has one (other). The LLM receives a template with empty slots — it fills text, not structure.

Prompt constraints: lock keys and semantics explicitly

Don't paste a .stringsdict file into a chat interface and ask for a translation. Instead, supply explicit instructions:

Translate the following plural strings into Russian.
Return ONLY the translated text for each category key listed.
Do NOT add, remove, or rename any keys.
Do NOT translate keys — they are format identifiers.
Preserve all format specifiers (%d, %1$d, %@) exactly as they appear.
Do NOT translate the words: one, few, many, other, zero, two, plural.

Keys to fill:
- one: (e.g., for count = 1)
- few: (e.g., for counts 2–4, 22–24, etc.)
- many: (e.g., for counts 5–20, 100–119, etc.)
- other: (fallback)

Source (English):
- one: "%d item selected"
- other: "%d items selected"

This dramatically reduces surface area for structural corruption, though it doesn't eliminate it. You still need validation.

Post-translation validation: automated checks

A post-translation validator should check:
1. Category completeness against the CLDR-required set for each locale
2. Placeholder parity — every placeholder present in the source appears in every translated branch
3. XML/plist well-formedness and structural integrity (not just valid XML, but correct nesting depth)
4. ICU bracket balance and keyword integrity for React Native strings
5. quantity attribute values are CLDR keywords, not translated text

Human review layer: spot-check high-risk locales

Automated checks catch structural problems. They can't catch semantically wrong plural forms in the right categories. For Arabic, Polish, Russian, and other complex-plural languages, native speaker review of plural branches — with explicit documentation of which numeric ranges each branch should cover — is the only reliable catch for failure mode #4.

This is the same argument we make in /solutions/safe-ai-translation/: LLM output in localization should be treated as a first draft that requires structured validation, not as finished output.


Conclusion

To summarize where each failure mode hits hardest:

Failure Mode iOS Android React Native
Missing plural categories High risk High risk High risk
Broken .stringsdict structure Critical N/A N/A
Placeholder drift High risk High risk High risk
Wrong Slavic plural rules High risk High risk High risk
ICU syntax mangling N/A N/A Critical
Quantity attribute mistranslation N/A Medium risk N/A

None of these failures require bad luck to hit. They're the default behavior of one-shot LLM translation applied to structured plural files without a validation layer. The LLM isn't malfunctioning — it's doing what it does, which is generating plausible text that matches the patterns it's seen, not validating locale-specific plural completeness against CLDR specifications.

The practical action item before your next release: run your current plural resource files through a validator that checks CLDR category completeness for each target locale. You will almost certainly find something, especially in Arabic, Russian, or Polish. /solutions/ai-translation-problems/ is a reasonable place to start if you want to understand the full scope of what structured validation covers versus what raw LLM output leaves on the table.

Plural bugs are quiet bugs. They ship, they stay, and they're grammatically wrong in ways that native speakers notice immediately and that nobody on your team ever sees in code review.

Ready to localize your app?

Get started free — no credit card required.

Start Translating →