Make translation quality measurable before release
Translation quality assurance is the operating system that turns a subjective request—“make it sound right”—into a repeatable release decision. A practical workflow defines the intended audience and risk, separates production from independent review, records errors by type and severity, samples intelligently, tests content in its real context, and gives one named owner authority to approve or stop release.
For buyers, the goal is not a universal quality score. It is evidence that the translation is fit for its specified purpose. A medical instruction, product interface, legal notice, campaign headline, and internal knowledge-base article do not carry the same consequences or require the same checks. This guide shows localization managers, procurement teams, and global content owners how to set acceptance criteria, compare vendor LQA reports, and avoid paying for a vague promise of “native quality.”
Translation QA versus proofreading
Proofreading is one activity inside a broader quality system. It usually catches target-language spelling, grammar, punctuation, consistency, and obvious presentation problems. It may not verify every meaning against the source, test variables in a live interface, confirm market-specific facts, or decide whether an error is severe enough to block release.
A complete workflow can include source preparation, terminology approval, translator self-check, bilingual revision, automated checks, linguistic quality assessment (LQA), functional or in-context testing, correction, regression review, and final approval. ISO 17100 addresses resources and core processes for translation services. ISO 11669:2024 gives buyers a framework for project specifications, needs analysis, risk assessment, and workflow communication. Neither reference replaces a project-specific acceptance brief.
Define these roles before production:
- the translator produces and self-checks the target content;
- the reviser compares source and target and confirms meaning, terminology, and purpose;
- the LQA evaluator scores a defined sample or delivery against an agreed model;
- the in-context tester checks the rendered product, page, document, or media;
- the release owner accepts residual risk and signs off the deliverable.
One person may cover several roles on a small project, but the handoffs and evidence should remain visible. Independent review is especially valuable when the content is regulated, safety-related, public-facing, or difficult to reverse after release.
Build an error taxonomy buyers can use
An error taxonomy is a controlled list of issue types. Start small enough that reviewers can apply it consistently. A useful buyer-facing model often includes:
- accuracy: mistranslation, omission, addition, untranslated content, or incorrect relationship to the source;
- terminology: wrong approved term, inconsistent term, or domain-inappropriate wording;
- language: grammar, spelling, punctuation, syntax, fluency, or unnatural phrasing;
- style and register: wrong tone, voice, formality, brand convention, or audience level;
- locale and compliance: wrong date, number, unit, address, legal reference, cultural convention, or required notice;
- functional and formatting: broken tag, placeholder, link, line break, truncation, layout, directionality, or file behavior.
The W3C Internationalization Tag Set 2.0 defines interoperable localization quality issue types and provides fields for issue type, comment, and severity. Its categories are a useful reference, but a project should use only the granularity reviewers can distinguish reliably. The European Commission Directorate-General for Translation’s 2024 evaluation information pack offers a public example of mapping requirements to error codes and evaluating a random sample.
Severity must describe impact, not reviewer irritation. A practical three-level model is:
- critical: could cause safety, legal, financial, privacy, or serious reputational harm; corrupts a required function; or makes the content unusable;
- major: changes meaning, misleads a user, violates an important instruction, or materially harms the task, even if the user can continue;
- minor: a limited language or presentation defect that does not change the main meaning or block the task.
Add examples for each content type. A wrong decimal in a dosage can be critical; an inconsistent button label may be major if it sends users to the wrong action; a missing nonbreaking space may be minor unless it breaks a legally required format. Avoid automatic severity based only on category.
Turn errors into an acceptance rule
ISO 5060:2024 discusses analytic evaluation using error types and penalty points to produce an error score and quality rating, and it explicitly covers sampling and evaluator competence. Buyers can use the same logic without pretending that one score works for every project.
For example, assign 1 point to a minor issue, 5 to a major issue, and make any confirmed critical issue an automatic hold. Normalize weighted points by reviewed words, segments, screens, or another stable unit. A threshold might read: “No critical errors; no more than 10 weighted points per 1,000 reviewed words; all functional failures corrected; all major corrections regression-tested.” These numbers are illustrative. Set them from risk, pilot evidence, and stakeholder tolerance rather than copying a vendor template.
Also define how repeated errors count. One wrong term repeated 80 times may reflect one root cause, but it can still affect 80 user interactions. Report both occurrence count and unique root cause. State whether source defects, preference changes, and out-of-scope rewrites are excluded from the vendor error score while still being tracked for resolution.
Use sampling without hiding risk
Full human review is not always economical, but “10% sampling” is not a strategy unless the selection method is documented. Choose the sampling unit, population, size, selection method, and escalation rule before review begins.
A robust plan combines:
- a random sample for an unbiased view of common content;
- a risk-based sample of safety text, pricing, claims, legal notices, calls to action, high-traffic screens, and new terminology;
- coverage across files, translators, languages, content types, and production batches;
- a rule that expands review when critical or clustered major errors appear.
For a 100,000-word release, a buyer might review a random sample from every file plus 100% of safety warnings, numbers, and newly approved terms. If a critical error appears, or major errors exceed the agreed threshold in one stratum, the team can hold the release and expand that stratum to full review. The plan is defensible because it connects sample design to a decision, not because the percentage sounds large.
Do not combine languages into one average. A passing global score can hide a failing small market. Report sample size, error counts, severity, normalized score, and release status by language and by high-risk content group.
Calibrate reviewers before scoring production
Reviewer disagreement creates noise, delays, and unnecessary rework. Calibration aligns interpretation before the score affects payment or release.
Give reviewers the same small set of representative segments, including deliberately ambiguous cases. Ask each person to mark the span, category, severity, correction, and rationale independently. Then compare results and resolve differences against the brief, glossary, style guide, and product context. Update examples in the error guide and repeat until the team can apply the key distinctions consistently.
Measure agreement at several levels: did reviewers flag the same issue, select the same category, and assign the same severity? A single percentage can conceal the source of disagreement. Record adjudication decisions in a living decision log so the same edge case is not debated in every batch.
Protect against reviewer bias. A reviewer should not rewrite correct content merely to match personal preference. Require a cited instruction, meaning defect, audience problem, or functional impact for scored errors. Track “preference” separately so useful stylistic suggestions do not distort acceptance metrics.
Test the translation in context
A linguistically correct string can fail after it is placed in the product. In-context testing checks the delivered experience rather than an exported bilingual file.
For software and websites, test text expansion, truncation, line wrapping, buttons, links, variables, plural forms, gender, sorting, search, right-to-left behavior, date and number formats, and screenshots of high-risk flows. For documents, inspect tables, page references, headers, fonts, diagrams, and print or PDF output. For subtitles, check timing, speaker context, line breaks, reading load, and on-screen text. For marketing, confirm the final headline, visual, CTA, landing page, and market claim work together.
Use stable identifiers so every issue connects to a string, segment, page, screen, timestamp, or file location. Capture evidence with a screenshot or rendered file, not only a free-text comment. After correction, regression-test the affected context and nearby content; a fix can introduce a new truncation, tag error, or inconsistent term.
Our guides to human review in translation workflows, translation versus MTPE, and website localization versus translation help teams choose the right production method and testing depth before they set QA thresholds.
Require a decision-ready LQA report
A useful report is not a spreadsheet of comments with no conclusion. Ask the vendor to provide:
- project, version, language, files, evaluator, and evaluation date;
- approved specifications, glossary, style guide, and taxonomy version;
- total population, sample size, selection method, and risk strata;
- every issue with stable location, source, target, category, severity, rationale, and proposed correction;
- counts and normalized scores by language, category, severity, and content group;
- critical and systemic findings, root causes, corrective owner, and due date;
- retest evidence, unresolved exceptions, and a clear pass, conditional pass, or fail recommendation.
Procurement should compare how providers define acceptance, select samples, calibrate reviewers, manage appeals, and prove corrections—not only their headline score. Ask to see a redacted sample report. A provider should be able to explain one issue from detection through adjudication, correction, regression check, and release decision.
When comparing two proposals, normalize the deliverable rather than the day rate. Suppose Provider A reviews 10% of the words but excludes screenshots, terminology verification, adjudication, and retesting. Provider B reviews 7% randomly, adds complete coverage of high-risk strings, includes reviewer calibration, and retests every major correction. The smaller headline sample may provide stronger release evidence because its risk coverage and escalation path are explicit. Ask both providers to price the same source volume, languages, risk strata, report fields, correction rounds, and approval responsibilities. Then compare the expected cost of an accepted release, not only the cost of the first review pass.
Treat the pilot as a test of the measurement system as well as the translation. Check whether reviewers can find the same issues, whether severity is justified consistently, whether reports can be filtered by language and risk, and whether corrections close cleanly. If the pilot produces unresolved disputes or cannot trace a fix back to a stable segment, scaling the same workflow will scale uncertainty rather than quality.
Release checklist for localization owners
Before approving publication or deployment, confirm:
- the current source, scope, audience, market, and content risk are documented;
- the correct target locale—not only the language—is identified;
- terminology, style, product references, and do-not-translate items are approved;
- translator, reviser, evaluator, tester, and release owner responsibilities are clear;
- the taxonomy, severity definitions, sample plan, thresholds, and escalation rules were agreed before scoring;
- automated checks for numbers, tags, placeholders, terminology, and completeness ran successfully;
- required bilingual review and risk-based LQA are complete;
- the final rendered content passed in-context and functional tests;
- every critical and major issue has a disposition, owner, and retest record;
- changes after LQA were regression-tested;
- the final report identifies residual risk and the authorized approver;
- source files, approved targets, reports, decisions, and version history are retained according to policy.
Quality assurance works when it produces an auditable choice: release, release with documented exceptions, or hold for correction. Ask prospective partners for a sample LQA report and documented acceptance criteria. Smart Language Service can help design a risk-based translation quality assurance workflow, localize and review the content, and deliver language-by-language evidence that supports the final release decision.

