Build-Time Locale Validation: 8 Checks Your CI Pipeline Should Run Before Every Release
Most localization CI setups have a confidence problem: they verify that translation files exist, then ship. That's not validation — it's attendance-taking. This post covers the eight checks that catch what presence checks miss: broken placeholders, incomplete plural forms, invisible encoding corruption, and the specific class of silent failures that AI-assisted translation introduces into your pipeline.
Why Existing CI Checks Are Not Enough
There's a meaningful difference between "the strings file is present" and "the strings file is correct." Most pipelines enforce the former, few enforce the latter, and production incidents live in the gap.
The failure modes are well-documented at this point. A malformed strings.xml doesn't throw a build error — Android silently falls back to the default locale, and the first signal your German users get is an English interface. A placeholder mismatch in a runtime-formatted string doesn't fail at compile time; it crashes on the specific screen, for the specific locale, the moment a user hits it. These are the bugs that make it through code review, through QA (which typically runs in the default locale), and into production.
AI translation tools introduce a newer and subtler variant of this problem. The output looks correct to a human reviewer who doesn't speak the target language, passes a "file is present" check, and still contains structural corruption: escaped sequences rendered as literal characters, duplicate keys, zero-width Unicode characters injected mid-string, or placeholders silently dropped because the model decided they were optional context. These failures are largely invisible until runtime. The patterns are consistent enough across platforms that they warrant their own class of CI check — and that's exactly what several of the checks below address.
If you've already read our posts on why localization fails in CI even when it works locally and string extraction automation pipelines, this post closes the loop: formatting correctness and runtime switching mean nothing if broken strings ship before those mechanisms ever run.
The 8 Checks
Check 1: Placeholder Parity Validation
Every placeholder in a source string — %s, %d, %1$s, {name}, %(count)d — must appear in every translation of that string. This is the single most common source of runtime crashes in localized apps, and it's also the easiest to catch automatically.
The failure mode is straightforward: a translator (human or AI) drops a positional argument or renames a named placeholder. At runtime, your string formatting call receives the wrong number of arguments and either crashes or renders garbage.
Shell one-liner for CI (Android strings.xml):
# Extract placeholder sets from source and target, diff them
python3 - <<'EOF'
import re, sys, xml.etree.ElementTree as ET
def get_placeholders(value):
return set(re.findall(r'%(?:\d+\$)?[sdf]|%\(\w+\)[sdf]|\{\w+\}', value))
source = ET.parse('app/src/main/res/values/strings.xml')
target = ET.parse('app/src/main/res/values-de/strings.xml')
source_map = {el.get('name'): el.text for el in source.findall('string')}
target_map = {el.get('name'): el.text for el in target.findall('string')}
errors = []
for key, src_val in source_map.items():
src_ph = get_placeholders(src_val or '')
tgt_ph = get_placeholders(target_map.get(key, '') or '')
if src_ph != tgt_ph:
errors.append(f"FAIL [{key}]: source={src_ph}, target={tgt_ph}")
if errors:
print('\n'.join(errors))
sys.exit(1)
EOF
For iOS .strings and React Native JSON, the same logic applies — adjust the parser. Tools like i18n-check and twine have built-in placeholder validation modes worth evaluating before rolling your own, though neither handles every placeholder syntax out of the box.
Pluralization is where CLDR complexity meets CI negligence. Most teams validate that a plural key exists; almost none validate that it has the right number of categories for each target locale.
Arabic requires six plural categories (zero, one, two, few, many, other). Russian requires three (one, few, other). Japanese requires one (other). Shipping an Arabic build with only one and other defined means four of six quantity cases fall through to undefined behavior.
Android strings.xml — valid Arabic plural:
<!-- values-ar/strings.xml -->
<plurals name="items_count">
<item quantity="zero">لا توجد عناصر</item>
<item quantity="one">عنصر واحد</item>
<item quantity="two">عنصران</item>
<item quantity="few">%d عناصر</item>
<item quantity="many">%d عنصرًا</item>
<item quantity="other">%d عنصر</item>
</plurals>
iOS .stringsdict — validating NSStringLocalizedFormatKey completeness:
<key>items_count</key>
<dict>
<key>NSStringLocalizedFormatKey</key>
<string>%#@items@</string>
<key>items</key>
<dict>
<key>NSStringFormatSpecTypeKey</key>
<string>NSStringPluralRuleType</string>
<key>NSStringFormatValueTypeKey</key>
<string>d</string>
<key>one</key>
<string>%d item</string>
<key>other</key>
<string>%d items</string>
</dict>
</dict>
In CI, pull the required categories from a CLDR data file keyed by locale code, then diff against what's defined in your resource files. This is a hard failure if the build targets production; a warning is appropriate only in development branches where a translation is actively in progress. For more on the edge cases here, see our deep-dives on Android plurals and iOS .stringsdict.
Check 3: XML/Plist Structural Integrity
This check costs almost nothing to run and catches a surprisingly high volume of failures. xmllint and plutil are available in every standard CI environment with zero additional dependencies.
# Android strings.xml
find app/src/main/res -name "strings.xml" | while read f; do
xmllint --noout "$f" || { echo "FAIL: $f"; exit 1; }
done
# iOS .strings plist (binary or XML)
find . -name "*.strings" | while read f; do
plutil -lint "$f" || { echo "FAIL: $f"; exit 1; }
done
The structural breaks most commonly introduced by AI translation tools fall into four categories: unclosed tags (the model truncates output mid-element), escaped quotes rendered as literal characters (\" becoming " in the final XML), duplicate keys (the model generates an entry it already produced), and malformed CDATA sections. A well-formed XML check won't catch semantic problems, but it will catch all four of these before they ever reach a device. See platform-specific corruption patterns for a more complete taxonomy.
Check 4: Key Coverage Diff
A strings file can be structurally valid, placeholder-correct, and still missing 40% of your app's strings because a new feature shipped without anyone sending those keys for translation. This is the "silent omission" problem, and it's endemic in teams that manage translation as a manual step rather than a pipeline artifact.
The check is conceptually simple: every key in your source locale file must exist in every target locale file.
Python diff script (works for Android, iOS, and React Native with minor parser changes):
import xml.etree.ElementTree as ET
import sys, os, glob
SOURCE = 'app/src/main/res/values/strings.xml'
TARGET_PATTERN = 'app/src/main/res/values-*/strings.xml'
source_keys = {el.get('name') for el in ET.parse(SOURCE).findall('string')}
source_keys |= {el.get('name') for el in ET.parse(SOURCE).findall('plurals')}
exit_code = 0
for target_file in glob.glob(TARGET_PATTERN):
locale = target_file.split('values-')[1].split('/')[0]
target_keys = {el.get('name') for el in ET.parse(target_file).findall('string')}
target_keys |= {el.get('name') for el in ET.parse(target_file).findall('plurals')}
missing = source_keys - target_keys
if missing:
print(f"FAIL [{locale}]: Missing keys: {sorted(missing)}")
exit_code = 1
sys.exit(exit_code)
For React Native JSON, the same logic applies with json.load() replacing the XML parser. Android Lint's MissingTranslation rule covers this natively if you're already running Lint in CI — but it only fires if you haven't suppressed it in lint.xml, which many teams have done to quiet false positives and then forgotten about.
Check 5: Untranslated String Detection
This one is subtle and worth explaining carefully. An untranslated string — where the translated value is identical to the source English — is not always a bug. Brand names, URLs, and technical terms are frequently left in English intentionally. The problem is when everything is left in English, which is a common failure mode for AI translation when it encounters proper nouns, code-adjacent strings, or technical jargon it's uncertain about.
The check needs an allowlist:
import json, sys
ALLOWLIST = {"app_name", "support_email", "privacy_url", "analytics_event_name"}
with open('translations/en.json') as f:
source = json.load(f)
with open('translations/de.json') as f:
target = json.load(f)
warnings = []
for key, value in source.items():
if key in ALLOWLIST:
continue
if target.get(key) == value:
warnings.append(f"WARN [{key}]: translated value identical to source")
if warnings:
print('\n'.join(warnings))
# Soft warning — don't exit 1 unless threshold exceeded
if len(warnings) > 10:
sys.exit(1)
Whether you fail the build or emit a PR comment depends on your tolerance for false positives. A threshold approach (fail if more than N strings are identical to source) works better than a binary check in practice. The allowlist should be committed to the repo and reviewed as part of new feature PRs, not managed out-of-band.
Check 6: Encoding and Whitespace Validation
This is the one that bites teams the hardest because it's completely invisible in code review. A zero-width joiner (U+200D), a non-breaking space (U+00A0), or a UTF-8 BOM injected by a translation tool produces a strings file that looks correct in every text editor, passes XML validation, and renders incorrectly on device — or not at all, depending on where the character lands.
Detect UTF-8 BOM:
find . -name "*.strings" -o -name "strings.xml" -o -name "*.json" | while read f; do
if file "$f" | grep -q "BOM"; then
echo "FAIL: UTF-8 BOM detected in $f"
exit 1
fi
done
Detect zero-width and non-breaking characters:
grep -rP "[\x{00A0}\x{200B}\x{200C}\x{200D}\x{FEFF}]" app/src/main/res/values-*/strings.xml \
&& echo "FAIL: Invisible Unicode characters detected" && exit 1
# Hexdump spot-check for a specific file
hexdump -C translations/de.json | grep -E "c2 a0|e2 80 8[bcd]|ef bb bf"
These characters are almost never introduced intentionally. When they appear, the source is almost always a translation tool or AI system that copied text from a rendering context where those characters were meaningful and didn't sanitize before writing output.
Check 7: String Length / Truncation Risk Flagging
Languages expand. German and Finnish are the canonical examples — a German translation of an English UI string is typically 20–35% longer. A hardcoded maxWidth or single-line UILabel that fits the English string comfortably will truncate the German equivalent in ways that your QA team, running the app in English, will never see.
This isn't a hard failure — it's a signal that needs to reach a designer or engineer before the release, not block it outright. The right implementation is a PR comment with a report.
import json
THRESHOLD_RATIO = 1.4 # flag if translation is >40% longer than source
with open('translations/en.json') as f:
source = json.load(f)
for locale in ['de', 'fi', 'pl', 'ar']:
with open(f'translations/{locale}.json') as f:
target = json.load(f)
for key, src_val in source.items():
tgt_val = target.get(key, '')
if src_val and len(tgt_val) / len(src_val) > THRESHOLD_RATIO:
print(f"WARN [{locale}][{key}]: {len(tgt_val)} chars vs {len(src_val)} source chars "
f"({len(tgt_val)/len(src_val):.1f}x)")
If your design system uses tokens that include maxLength constraints for specific string keys (some do, especially in component-driven design systems), pipe that data into the check so the threshold is per-key rather than a global ratio. That gets you genuinely actionable output instead of noise.
Check 8: RTL Locale Smoke Test via Pseudo-Localization
The cheapest way to catch RTL layout regressions in CI is to never let them accumulate. Pseudo-localization — generating a synthetic locale that forces RTL layout and expanded string lengths — lets you run a snapshot diff on every PR against the previous release baseline.
The workflow is: generate a pseudo-RTL locale from your source strings, build the app with that locale active, run UI snapshot tests, diff against baseline. Any new hardcoded LTR constraint that slipped into a feature branch shows up as a snapshot delta.
Pseudo-locale string generator (React Native JSON):
const fs = require('fs');
const source = JSON.parse(fs.readFileSync('translations/en.json'));
const pseudo = {};
for (const [key, value] of Object.entries(source)) {
// Wrap with RTL markers, expand length by ~30%
pseudo[key] = '\u200F' + value.split('').reverse().join('') + '\u200F';
}
fs.writeFileSync('translations/ar-pseudo.json', JSON.stringify(pseudo, null, 2));
This is intentionally lightweight — the goal isn't linguistically accurate Arabic, it's surfacing layout breakage. For a comprehensive pre-release RTL validation process, the RTL pre-release checklist covers the full scope of what CI alone can't catch. The testing failures specific to RTL languages post is also worth reading alongside this check.
Putting It Together: A Sample CI Configuration
Here's a GitHub Actions workflow combining all eight checks. The structure distinguishes between hard failures (exit 1, blocks merge) and soft warnings (report-only, posted as PR comments).
name: Localization Validation
on:
pull_request:
paths:
- '**/strings.xml'
- '**/*.strings'
- '**/*.stringsdict'
- '**/translations/*.json'
jobs:
locale-validation:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Install dependencies
run: |
sudo apt-get install -y libxml2-utils
pip install lxml
# HARD FAILURES — block merge
- name: Check 1 - Placeholder parity
run: python3 scripts/check_placeholders.py
- name: Check 2 - Plural form completeness
run: python3 scripts/check_plurals.py
- name: Check 3 - XML/Plist structural integrity
run: |
find app/src/main/res -name "strings.xml" | \
xargs -I{} xmllint --noout {}
find . -name "*.strings" | \
xargs -I{} plutil -lint {}
- name: Check 4 - Key coverage diff
run: python3 scripts/check_key_coverage.py
- name: Check 6 - Encoding validation
run: |
grep -rP "[\x{00A0}\x{200B}\x{FEFF}]" \
app/src/main/res/values-*/strings.xml && exit 1 || true
# SOFT WARNINGS — post as PR comment, don't block
- name: Check 5 - Untranslated string detection
id: untranslated
run: python3 scripts/check_untranslated.py > untranslated_report.txt 2>&1
continue-on-error: true
- name: Check 7 - String length report
id: lengths
run: python3 scripts/check_lengths.py > length_report.txt 2>&1
continue-on-error: true
- name: Check 8 - Pseudo-RTL snapshot diff
id: rtl_snapshot
run: |
node scripts/generate_pseudo_rtl.js
# Run snapshot tests with pseudo locale
npx jest --testPathPattern="snapshot" --locale=ar-pseudo > snapshot_report.txt 2>&1
continue-on-error: true
- name: Post warning summary as PR comment
if: always()
uses: actions/github-script@v7
with:
script: |
const fs = require('fs');
const reports = ['untranslated_report.txt', 'length_report.txt', 'snapshot_report.txt']
.filter(f => fs.existsSync(f))
.map(f => `**${f}**\n\`\`\`\n${fs.readFileSync(f, 'utf8').slice(0, 1000)}\n\`\`\``)
.join('\n\n');
if (reports) {
github.rest.issues.createComment({
issue_number: context.issue.number,
owner: context.repo.owner,
repo: context.repo.repo,
body: `## Localization Validation Warnings\n\n${reports}`
});
}
A few notes on this configuration. First, the paths filter on the trigger is important — you don't want this workflow running on every backend change. Scope it to files that actually change when translations are updated. Second, the hard/soft distinction is deliberate: failing the build on a string-length warning creates alert fatigue and gets the job disabled within a week. Post the soft warnings as PR comments so they're visible without being blocking. Third, if you're using a managed translation platform rather than raw AI output, the platform's validation API can replace several of these checks with a single authenticated call — particularly for placeholder parity and structural integrity, where platform-validated output removes the need to detect corruption you shouldn't have received in the first place.
Conclusion
If you're starting from zero, the highest-value checks to implement first are 1 (placeholder parity), 3 (structural integrity), and 4 (key coverage). These three catch the crashes and silent fallbacks that cause the most production impact, they're fast to run, and they're easy to make hard failures without generating noise.
Checks 2, 5, 6, and 7 are the next tier — important, but they require more tuning to avoid false positives that erode trust in the pipeline. Check 8 requires snapshot infrastructure that some teams don't have yet, but it's the only check that catches RTL layout regression automatically before manual testing.
This validation layer complements the runtime behaviors covered elsewhere in this series — locale switching, formatting correctness, and plural resolution — by ensuring that what gets bundled into the release is structurally sound before any of those runtime paths are exercised. None of the runtime-level defenses matter if a malformed strings file is already in the app.
For a broader view of what can go wrong before you even get to CI, Localization Testing 101 is worth reading alongside this. And if your team is weighing whether the infrastructure overhead of running all of this in-house makes sense at your current scale, the real cost of maintaining localization in-house has numbers worth reviewing before you commit to building the full pipeline yourself.
The goal of all eight checks is the same: make the CI pipeline catch what code review can't see, before production catches what QA missed.