The Fluency Trap: Translation Became Free, Checking It Did Not

Somewhere tonight, a person will leave a hospital holding a sheet of paper they can read perfectly and should not trust.
The paper will look right. Clean typography, the hospital's logo in the corner, sentences running in grammatical order in the reader's own language. It will have been produced in under a second at a cost that rounds to nothing, and handed over by a clinician who cannot read a word of it, to a patient with no way of comparing it to anything. Everyone will believe the document is complete, because it looks complete, and looking complete is what this technology does flawlessly.
We know what these documents can contain, because researchers have checked. In 2021, a team led by Breena Taira at the University of California, Los Angeles took twenty commonly used emergency department discharge statements, ran them through Google Translate into seven widely spoken languages, and had native speakers evaluate the output. Their paper in the Journal of General Internal Medicine contains two findings that ought to have ended the argument about unsupervised machine translation in clinical settings. The first is statistical. The second is a pair of words: in the Armenian output, ibuprofen was rendered as “anti-tank missile”, and in Chinese, the anticoagulant Coumadin came out as “soybean”.
These are not subtle failures. A bilingual reader would spot them in a second. The problem is that no bilingual reader was ever going to look, because the entire economic point of the exercise was to avoid paying one.
That is the shape of the thing. Over roughly three years, the cost of producing a translation has fallen by close to three orders of magnitude. The cost of finding out whether it is correct has not moved, because verification still requires a person who reads both languages, sitting down with both documents, and thinking. Generation collapsed. Verification did not. Everything that follows, from the restructuring of a seventy billion dollar industry to the question of who is answerable when the instructions are wrong, falls out of that single asymmetry.
The Arithmetic That Broke in Only One Direction
Google's Cloud Translation service prices standard neural machine translation at twenty United States dollars per million characters, the first half a million each month free. DeepL's professional interface has been offered at around five and a half dollars per million. A thousand words of English is roughly six thousand characters, so translating it costs about twelve cents at the top rate and about three at the bottom. Professional human translation of the same thousand words sits, in 2026, at one hundred and fifty to three hundred dollars, with rate surveys clustering human work between ten and thirty cents per source word.
Generation is effectively free. But consider what verification requires and the picture inverts. To know whether a thousand-word translation is correct, a qualified bilingual reader must read the source, read the target, and compare them for meaning, register, omission, addition and negation. That is not faster than translating. It is often slower, because the reviewer must reconstruct the translator's reasoning as well as the author's. And it requires precisely the scarce human capability the machine was bought to replace.
Verification now costs more than production by a factor of roughly a thousand. If a factory could stamp a component for a tenth of a cent but testing each one cost a hundred dollars, nobody would call the testing optional. They would redesign the process, or admit on the record that they were shipping untested parts. Translation took the third option, never written down: shipping untested parts while behaving as though they had been tested, because they look exactly like the tested ones.
A second scarcity sits beneath the first. The binding constraint is not money but qualified bilingual attention. Only so many people can competently review a discharge instruction in Armenian, a tenancy notice in Tigrinya, or a pesticide warning in Hmong. Machine translation multiplied the volume of text requiring review by orders of magnitude while leaving the pool of reviewers where it was. The bottleneck is a population, not a budget line.
Failure With No Symptoms
The asymmetry persists because its consequences are invisible at the point of delivery. Machine translation does not fail the way software fails. It does not crash, return an error code, garble characters or leave a field blank. It fails by producing a well-formed sentence that means something other than what the source said.
The canonical example comes from a 2019 study in JAMA Internal Medicine by Elaine Khoong, Eric Steinbrook, Cortlyn Brown and Alicia Fernandez, who passed one hundred sets of emergency department discharge instructions comprising 647 sentences through Google Translate into Spanish and Chinese. Eight per cent of the Spanish sentences and nineteen per cent of the Chinese sentences were inaccurate. More to the point, two per cent of the Spanish and eight per cent of the Chinese carried potential to cause clinically significant harm.
One sentence read, in English, “hold the kidney medicine until you have a chance to speak with your kidney doctor”. The Chinese output instructed the patient to keep taking the kidney medicine, and the Spanish said much the same. “Hold”, in clinical English, means stop. The machine took it in its ordinary sense of retain and produced a fluent, grammatically flawless instruction to do the opposite of what the prescriber intended. Elsewhere, “aortic aneurysm” became, in Spanish, an “evacuation of the main blood vessel”.
The pattern is not confined to hospitals. A 2010 study in Pediatrics by Iman Sharif and Julia Tse examined 76 Spanish-language prescription labels generated by the computer programmes New York pharmacies were actually using and found errors in half of them, including the notorious collision in which the English word “once” reads in Spanish as eleven. In immigration proceedings, translators with the volunteer organisation Respond Crisis Translation told Rest of World in 2023 that machine translation of Pashto and Dari was corrupting asylum claims, in one case rendering the first person singular as a plural throughout a personal narrative, so that an individual persecution story read as a collective complaint. At least one Afghan claim was rejected after such errors.
No surface signal attaches to any of these. A negation reversal reads exactly like a negation preserved: same length, same cadence, same register, same institutional wrapper. That is what distinguishes machine translation from almost every other failure mode in consumer technology. The output carries no evidence of its own unreliability, and the evidence that would reveal it exists only in a document the recipient does not have.
Consider the chain. The clinician cannot read the target language and assumes the system did its job. The organisation that licensed the engine sees an aggregate quality score, if anything. The patient sees fluent prose on hospital letterhead and assumes that if it had not been checked, it would not have been handed over. Nobody holds both halves of the comparison, and the appearance of completeness does all the work verification used to do.
The older failure mode is instructive by contrast. In January 1980, an eighteen-year-old named Willie Ramirez arrived unconscious at a Florida emergency department. His Cuban family said he was “intoxicado”, meaning ill from something he had eaten. It was taken to mean intoxicated. Ramirez was treated as a drug overdose, his intracerebellar haemorrhage went undiagnosed for two days, and he was left quadriplegic; the settlement was estimated at around seventy-one million dollars over his lifetime. That failure had a location: a room, a decision, a defendant. When the same error comes from an engine invoked automatically inside an electronic patient record, it has no location at all.
A Thirty-Nine Point Gap Inside One Product
The Taira study's statistical finding should trouble anyone who believes machine translation is solved. Across seven languages, the overall meaning survived in 82.5 per cent of cases, 330 out of 400. That headline is meaningless, because the variation underneath it is enormous.
Spanish came out at 94 per cent accuracy. Tagalog at 90. Korean at 82.5. Chinese at 81.7. Vietnamese at 77.5. Farsi at 67.5. Armenian at 55. A thirty-nine point spread between best and worst inside a single product, through a single interface, with a single set of expectations attached. Farsi had a further problem no accuracy metric captures: directional text handling rendered some of it illegible.
Nothing in the interface communicates that spread. The menu presents Armenian and Spanish as equivalent options, and output arrives with the same speed, formatting and absence of caveat. An administrator enabling automated discharge translation is not choosing between a reliable service and an unreliable one. They are enabling both at once, with no signal for which patients get which.
This is language-tier inequality, and it maps onto each language's digital footprint. Languages with vast parallel corpora perform well; those spoken largely by populations who were never a lucrative localisation market perform badly. So the patients most likely to receive an unreliable translation are those least likely to have an alternative: recent arrivals, smaller diaspora communities, speakers of languages for which the local health system has no on-call interpreter at three in the morning.
The gap is not closing, because progress concentrates where the data is. A study in JMIR Formative Research in January 2026 by a team at Beth Israel Deaconess Medical Center tested Claude Sonnet 3.5 on Spanish translations of emergency discharge instructions and found the output essentially clinically acceptable, with evaluators scoring completeness and severity at 5.0 on a five-point scale and meaning at 4.9. Genuinely impressive, and a result about Spanish. It is also no longer isolated, which is the part of this argument that has changed. Meanwhile a framework paper in npj Digital Medicine in September 2025 by Ivan Lopez, David Velasquez, Jonathan Chen and Jorge Rodriguez recommends deploying first into languages with abundant digital resources, naming Spanish and Portuguese, and warns that underrepresented languages such as Quechua and Yorùbá need further testing first. The authors are right. But notice what that implies: the languages validated first needed validating least, and those where risk is highest get served last, or with no validation at all.
The strongest evidence against the case being made here arrived in Academic Emergency Medicine in 2026, and it deserves stating at full strength rather than filing under limitations. Giovanni Rodriguez, Patricia Hernández and colleagues ran a blinded non-inferiority study on fifty-three real emergency department discharge instructions of between one hundred and five hundred words, taken exactly as the clinicians wrote them, original spelling and grammar errors left in. Google Translate and ChatGPT-4o rendered them into Spanish, Brazilian Portuguese and Simplified Chinese, against professional translations of the same text. Professional medical interpreters, blinded to which was which, scored 477 unique translations for fluency, adequacy, meaning and severity against a non-inferiority margin of half a point.
Both machines came out non-inferior to the professionals on most domains. In Spanish and Brazilian Portuguese they matched professional translation on adequacy, meaning and severity, falling short only on fluency. In Simplified Chinese they matched it on all four. The frequency of clinically significant errors did not differ significantly by translation method.
That has to be conceded without qualification, because it is true and it matters. In those three languages, against genuinely messy clinician-written source text rather than tidied research prose, current machine translation is not measurably more dangerous than a professional human being. The 2019 and 2021 numbers quoted earlier are real, and they were real about the systems they tested, but those systems are two generations back. Nobody should now cite eight per cent Spanish inaccuracy as a description of what a modern engine does to a Spanish discharge sheet. It is not what happens.
What the study does not do is rescue the position it appears to demolish, and both reasons are visible in its own design. Start with the languages. Spanish, Brazilian Portuguese and Simplified Chinese are among the largest localisation markets on earth and among the best-resourced pairs in existence. The study says nothing whatever about Armenian at 55 per cent, or Farsi at 67.5, or any language whose speakers were never worth building a corpus for. It does not narrow the thirty-nine point spread. It measures the top of it, carefully, one study later.
The second reason goes to what non-inferiority establishes. That clinically significant errors did not differ by translation method is not a finding that machine translation is safe. It is a finding that professional human translation also carries a clinically significant error rate, and that the machine has drawn level with it. Nobody who commissions professional translation of consequential documents believes a single pass is sufficient, which is why the profession has review stages, standards describing them, and indemnity behind them. Both arms of that study were then read by trained evaluators. Every translation in it was checked. The subject here is the document nobody checks at all, and non-inferiority to an unverified human baseline is not a warrant for skipping verification. It moves the argument off provenance and onto verification, which is exactly where it belongs. The question was never whether a machine or a person produced the text. It is whether anyone competent read it afterwards.
There is direct evidence on both counts, from a study that measured what the non-inferiority design left out. In npj Digital Medicine in October 2025, Ryan Brewster and a large multidisciplinary team took paediatric discharge instructions into six languages, Arabic, Armenian, Bengali, simplified Chinese, Somali and Spanish, and compared three routes: ChatGPT-4o alone, professional linguist translation, and human-in-the-loop, meaning a machine draft post-edited by a professional linguist. Forty-two evaluators scored the output, twelve professional linguists, sixteen clinicians and fourteen bilingual family caregivers.
The tier gap turned up intact in a frontier model. For Spanish and Bengali, ChatGPT-4o came close to professional quality. For Armenian it scored 2.4 on overall quality against 3.6 for professional translation, a deficit of more than a full point on a five-point scale, and Somali and simplified Chinese also fell significantly below the professional standard. Whatever has been fixed since 2021, this has not been. Note that the two studies disagree about Chinese, non-inferior on every domain in one and materially below professional standard in the other, on different document types, against different comparison standards, judged by different evaluators, a year apart. That disagreement is not a scandal. It is a measurement of how thin the evidence is.
The other finding is the one this whole argument turns on. Human-in-the-loop did not merely rescue Armenian, it beat the professionals, scoring 3.9 against their 3.6. Across all six languages it was also the faster route, averaging 7.1 minutes per document against 16.8 for professional translation. Post-editing a machine draft was better than either route alone and roughly twice as fast as translating from scratch. That is what working looks like, and it is worth being precise about what it costs. It still requires the qualified bilingual reader. It makes that person faster. It does not make them unnecessary.
Every such study is itself an act of verification, and inherits the same cost structure. Beth Israel validated one language pair, one document type, one institution, which tells you nothing about Armenian, consent forms, or next quarter's model update. The evidence base is built at human speed against a deployment surface expanding at machine speed.
The Instruments That Cannot See the Errors That Matter
The obvious response is to automate the checking. If a machine can produce the translation, surely a machine can grade it. That is the promise of quality estimation, and what most enterprise deployments rely on when they route some segments to review and let others through.
Two pieces of 2026 research suggest the hope is misplaced exactly where it matters. In August, Serge Gladkoff, Angelika Vaasa, Sue Ellen Wright, Ingemar Strandvik and Lifeng Han released a peer-reviewed paper titled “Looking under the Wrong Lamppost”, accepted for the ninth International Conference on Natural Language and Speech Processing in September, arguing that automated quality estimation suffers structural rather than incidental limitations. Their conclusion is unusually blunt: segment-level scores should not be used as a standalone basis for routing, release or review bypass in production. They document overfitting, distribution collapse, failure to generalise across domains, and blindness to cohesion and coherence, which are not properties of individual segments at all. Their proposed direction is not better scoring but making human review cheaper rather than pretending it is unnecessary.
The clinical version arrived two months earlier. At the American Medical Informatics Association's Amplify conference in June 2026, William Mundo, Elizabeth Goldberg and Yanjun Gao presented work on AI-generated translation of emergency discharge instructions and found direct disagreement between automated evaluation and clinician judgement. Automated metrics suggested high quality. Clinicians found errors in precisely the meaning-sensitive areas carrying clinical consequence: mistranslated medication instructions, omitted return precautions, negation errors of the “take with food” variety, confusion between daily and twice daily, mild discomfort turned into severe pain. Automated metrics alone, they concluded, should not drive deployment.
Read together, the escape route closes. The class of error automated estimation misses is not a random sample. It is disproportionately the class that reverses instructions, drops warnings and changes doses, because those errors are small, local, fluent and invisible to systems trained to reward fluency. The scoring machinery is best at catching failures a human would catch instantly, and worst at the failures only a human can catch at all. You cannot build a verification layer out of the competence that produced the problem.
What the Market Did With the Savings
To see where value goes when a commodity becomes free, read the 2026 Nimdzi 100, the annual ranking of the world's largest language industry providers. It sizes the market at 72.6 billion dollars and reports growth figures that read like a diagram of stratification.
The hundred largest providers collectively grew by 1.1 per cent. The top ten grew by 3.6 per cent. Providers ranked fifty-first to hundredth contracted by 4.3 per cent, the first negative growth recorded for any segment since 2021. Ten providers entered the ranking who were not there the year before, largely through acquisition; Propio Language Services vaulted into third place by acquiring the interpreting company CyraCom. Industry optimism, on Nimdzi's own confidence rating, fell from 6.5 to 6.2.
That distribution is the signature of a market whose commodity has moved to machines. If you sell volume translation, your product competes with something priced at twelve cents per thousand words, and no cost base survives that. If you sell verification, liability, domain expertise, certification and the ability to put a name and an indemnity policy behind a document, you are selling the one thing machines did not make cheaper. The middle contracts because the middle sold the thing that became free.
From inside the profession, the same restructuring reads differently. A study submitted to arXiv in June 2026 by Yujun Wang, Ehud Reiter, Shimei Pan, Steffen Eger and Wei Zhao analysed 79,286 social media posts from Reddit, Facebook, Bluesky and Mastodon between 2019 and 2025 across four stakeholder groups. The arc in translator discussion tells the story in three data points. In 2021 they were talking about computer-assisted translation infrastructure, translation memories and termbases. By 2023 it was integrating AI into those workflows. By 2025 it was automation and quality assurance, which is to say they had stopped discussing how to translate and started discussing how to check machines. Translators held the most critical stance of any group. The communities are not even arguing about the same variables: where developers frame reliability as a hallucination rate, everyone else frames it as verification and oversight.
The professional consequence is a job description inverted without being renegotiated. Translators are increasingly employed not to produce text but to verify it, the more cognitively demanding half, and are paid at post-editing rates set on the assumption the machine already did most of the job. The standards infrastructure sees the distinction even where the market does not. ISO 17100 covers human translation services; ISO 18587 covers post-editing of machine output and requires post-editors to meet the same competence levels as translators under ISO 17100. Same competence, different pay grade, different assumption about how much of the output anyone read.
Nobody Ever Decided This
The most striking feature of this migration is that almost nowhere can you find the meeting where it was approved.
Organisations did not convene a risk committee, evaluate error rates by language, define a threshold above which human review became mandatory, and sign off. They enabled a feature. A content management system offered automatic localisation; a support platform added a translate button. A free allowance covered the first half a million characters a month, enough to translate a small organisation's entire public documentation set without generating an invoice that needed a signature.
Research commissioned by DeepL and reported in 2026 found enterprises investing heavily in language AI while still running critical global operations through manual, unintegrated workflows, which tells you the adoption pattern is not a strategy but a scattering of local decisions. Marketing copy, support documentation, product listings, internal policy, contract summaries and patient-facing instructions crossed the same threshold at different moments, authorised by different people, none asked to consider the whole. A system nobody decided to build is one nobody feels responsible for.
The Chain in Which Everyone Is a Bystander
Try to identify the responsible party in a concrete case and the exercise becomes a study in evaporation.
A patient is discharged with machine-translated instructions containing a reversed medication direction. Is the model provider liable? Its terms disclaim high-stakes use and it never knew the text was clinical. The record vendor? It integrated a service at a customer's request. The hospital? Arguably yes, the most promising thread, but it will point to the impossibility of reviewing every document in every language, and to the absence of any regulatory instruction telling it where the line sits. The clinician cannot read Armenian. The patient received a document from a hospital in their own language and read it correctly. The document was wrong.
Occasionally an institution refuses to accept this diffusion. In June 2018, a United States district court in Kansas suppressed evidence in United States versus Cruz-Zamora, in which an officer used Google Translate on a patrol car laptop to obtain consent to search a vehicle from a driver who spoke very little English. The officer typed “can I search the car?” The Spanish output, translated back, asked something closer to “can I find the car?” The court found the literal but nonsensical rendering meant the driver had not given unequivocal consent. That is a court locating responsibility with the party who chose to deploy the technology, rather than the person who received its output.
Cruz-Zamora is instructive precisely because it is unusual. It required an adversarial process, a defence lawyer, a back-translation and a judge willing to inspect the mechanism. Almost no machine-translated document gets that treatment. The vast majority arrive where there is no adversary and nobody whose job it is to check: a discharge sheet, a benefits letter, a tenancy notice, a warning label, a consent form. The error surfaces, if ever, only through the harm it causes, by which point the chain back to a mistranslated clause is nearly impossible to reconstruct.
The industry sells professional indemnity insurance precisely because translation carries assignable liability when a professional does it. A named translator working under ISO 17100 is insurable because the responsible party is identifiable. A model invocation inside a content pipeline is not, and the risk does not disappear when it becomes uninsurable. It transfers to the person holding the paper.
The Law Points Two Ways at Once
For a case study in regulatory incoherence, consider the current United States position on language access.
Section 1557 of the Affordable Care Act, whose final implementing rule took effect on 5 July 2024, requires covered health entities to provide qualified interpretation and translation, and specifically requires that machine-translated material be reviewed by a qualified human translator where accuracy is essential or the source is complex or technical. That addresses the verification gap directly: you may use the machine, but not unsupervised where it matters.
Eight months later, on 1 March 2025, Executive Order 14224 designated English the official language of the United States and revoked Executive Order 13166, the 2000 order requiring federal agencies to plan for meaningful access for people with limited English proficiency. It directed the Attorney General to rescind guidance issued under it, and the resulting Department of Justice memorandum of 14 July 2025 encouraged agencies to rescind their own guidance, to consider English-only services, and, explicitly, to use artificial intelligence to reduce translation costs. In July 2026, formal rescission notices for Title VI language access guidance appeared in the Federal Register.
The policy environment now simultaneously requires human review of machine translation in health settings and recommends machine translation as a cost-reduction strategy across federal programmes. Underlying civil rights law has not changed: Title VI still prohibits national origin discrimination, and recipients of federal funding still owe meaningful access. What has changed is the guidance telling them how, so the question of what meaningful access means when the translation is machine-generated is unanswered by anyone with authority to answer it.
Europe's failure has a different texture. The General Product Safety Regulation, applicable since 13 December 2024, requires products to be accompanied by instructions and safety information in a language easily understood by consumers in the member state where they are sold. That is an outcome requirement, usually preferable to a process requirement. But it is silent on how the translation is produced and on who must verify the text is in fact understood rather than merely present. A machine-translated warning that reverses a prohibition satisfies the formal requirement while defeating its purpose. The EU Artificial Intelligence Act adds transparency obligations from August 2026, including machine-readable marking of synthetic content, though whether a translation engine running over patient instructions falls inside its high-risk regime is a matter for lawyers rather than settled fact. A small joke is buried here: some widely consulted online reference versions of the AI Act's own articles carry a notice explaining that the translations are machine-generated and not the official European Parliament versions.
In the United Kingdom, NHS England published an improvement framework for community language translation and interpreting on 27 May 2025 that is unusually honest about the gap. It records annual NHS spend at around 75.5 million pounds against an estimated 250 to 300 million needed to meet actual demand, and warns that translation apps, while convenient, carry risks to accuracy and patient safety. It commits NHS England to national guidance specifying when AI is suitable, which tools are approved, how accuracy is verified across languages, and the governance frameworks including indemnity and responsibility. That last clause is an admission: who is responsible when an AI translation harms an NHS patient had not been settled. It is being written now, after the tools are already on the wards.
What Recipients Are Actually Being Handed
Strip away the institutional framing and answer the first question plainly. What does it mean to receive information machine-translated at a fraction of human cost and never checked by anyone who reads both languages?
It means being handed a document carrying the full social signalling of institutional assurance and none of the substance. Every cue on that page, the letterhead, the formatting, the fluency, evolved in a world where producing text in a language required a person who knew it. Those cues were reliable because they were expensive. Fluency was a costly signal of competence. Machine translation made fluency free, so the signal conveys nothing, and nobody has told the recipients.
That is usually asserted about lay recipients, and it would be easy to dismiss as condescension: of course a patient cannot tell, they are not a translator. But the non-inferiority study tested the claim without setting out to. Its blinded evaluators were professional medical interpreters, reading in their own working languages, paid to assess translation quality, and aware that some of what they were reading was machine output. They frequently took the machine translations for professional work. If the signal has stopped carrying information under those conditions, for those readers, there is no version of this in which it still carries information for a patient reading a discharge sheet in a corridor. Fluency is not weak evidence of human competence. It is no evidence at all.
It means epistemic risk has been transferred to the party least able to bear it. Previously the burden of ensuring correctness sat with the institution producing the text, because the institution employed the translator. It has been silently relocated to the reader, who is expected, without being told, to treat official documents in their own language as provisional. Those readers are disproportionately the people with fewest resources to seek a second opinion.
It means informed consent is being quietly hollowed out. Consent requires comprehension of accurate information. A patient who fully understands a fluent instruction that has reversed the source has comprehended something, but not the thing they were meant to consent to. Their signature attests to a document nobody on the institutional side has read.
And it means we have built, without deciding to, a two-tier information citizenship. Speak a high-resource language and the machine serves you at close to professional quality. Speak Armenian and the same interface, the same button, the same institutional promise delivers something accurate slightly more than half the time in controlled testing. Both are told they have been given information in their own language. Only one of them has.
An Accountability Regime That Would Actually Bite
The second question is harder. Who bears responsibility when the system producing the translation is more capable than any system available to tell whether it worked?
The honest starting point is that no technical fix resolves this. The verification gap is not an artefact of immature models but a structural property of a task whose correctness can only be assessed by someone holding both languages. Better models will narrow the error rate. They will not create the reviewers. So the regime must be built out of allocation of responsibility rather than engineering, starting by reversing the current default.
That default should be simple. The organisation publishing a translated document warrants it, regardless of how it was produced. Not the model provider, not the platform, not the clinician, and emphatically not the recipient. If a hospital hands a patient a discharge sheet in Armenian, it stands behind that sheet exactly as if a staff translator had written it. This single rule does most of the work, because it puts the cost of unverified translation onto the party that captured the savings rather than the recipient who absorbs the risk.
The rest follows. Risk tiering should be organised by consequence class rather than document type. A marketing headline and a dosing instruction are both text; only one can put someone in intensive care. Content that can change a medical action, alter a legal obligation, forfeit a right, or carry a safety warning belongs in a tier requiring independent bilingual review before release, attributable to a named professional. Everything else can run unverified, provided it is labelled as such.
Per-language performance disclosure should be mandatory and visible in the interface. A product performing at 94 per cent in one language and 55 per cent in another should not present those as equivalent menu options. Deployers in regulated settings should publish language-specific error rates for their own content domains, sampled by qualified reviewers, and suspend automated release for pairs below a defined floor. The Lopez framework recommends the healthcare version: extend Joint Commission oversight to machine-assisted translation, build a shared clinical translation corpus, and update Section 1557 guidance with language-specific benchmarks. Language-specific is the operative phrase, because aggregate accuracy conceals the inequality that matters.
Provenance labelling should appear in the target language, in the document, in terms a recipient can act on. Not a disclaimer in eight-point type, but a plain statement: this text was produced automatically and has not been checked by a person who reads both languages; if anything here concerns your medication, contact us. The harm mechanism is a false impression of human authorship, and labelling is the cheapest counterweight.
Procurement is the most immediately available lever. Contracts in regulated settings should specify ISO 18587 full post-editing as a floor for consequence-bearing content, name the accountable reviewer, require logging sufficient to reconstruct which engine produced which document on which date, and prohibit reliance on automated quality estimation as the sole basis for bypassing review. The Gladkoff paper gives that last clause its evidentiary basis. The Brewster results give the first clause its own: full post-editing was not a compromise between speed and safety but better than either route alone, and faster than commissioning a translation from scratch. The expensive part of translation is no longer the translating. It is the reading.
Finally there is a public-goods problem no single organisation will solve. Armenian performs at 55 per cent because of a shortage of high-quality parallel data, and that shortage persists because nobody has a commercial incentive to fix it. Building open, domain-specific corpora for lower-resource languages in the areas where errors are most dangerous, clinical instructions, legal notices, safety warnings, is unglamorous infrastructure that health systems, regulators and standards bodies should fund jointly. It is cheap relative to the litigation it would prevent, and the only intervention that closes the tier gap rather than documenting it.
None of this restores the old equilibrium, in which producing a translation and trusting one cost roughly the same. What a workable regime does instead is make the residual risk visible, priced and owned, so that an organisation choosing to publish unverified translations makes that choice explicitly, in writing, with its name attached.
The Sentence That Reads Perfectly and Says the Opposite
Return to the instruction that started this. Hold the kidney medicine until you have a chance to speak with your kidney doctor. In the source it is a stop order. In the output, in two of the world's most widely spoken languages, it became a continue order, in prose so natural that a native speaker would have no cause to question it.
That sentence is the whole problem in miniature, and it explains why the failure mode has been so hard to take seriously. Every intuition we hold about text quality formed under an assumption that no longer applies: that fluent, contextually appropriate prose in a language implies someone competent in that language was involved. It was a reliable inference for the entire history of writing. It became false around three years ago, and our institutional design has not caught up.
The organisations that migrated to machine output did nothing obviously reckless. Each decision looked like a straightforward efficiency gain, and the aggregate was never assembled by anyone. The people receiving the documents are not careless either; they read correctly and trust appropriately, given every signal available. The models are not malfunctioning. They perform the task they were built for well enough that the residual failures are undetectable by anyone in the chain.
That is what makes it a trap rather than a scandal. Nobody behaves unreasonably from their own vantage point, and the system as a whole produces outcomes nobody would defend if they could see them. The technology has become powerful enough to make the question of whether it worked unanswerable at the point of use, and we have responded by not asking. The gap will not close on its own, because the cheap thing keeps getting cheaper and the expensive thing is expensive for reasons that have nothing to do with technology. The only variable under our control is who is answerable for the difference.
At present the answer is the person holding the paper, who does not know there is a question.
Sources and References
- Breena R. Taira, Vanessa Kreger, Aristides Orue and Lisa C. Diamond, “A Pragmatic Assessment of Google Translate for Emergency Department Instructions,” Journal of General Internal Medicine, 5 March 2021: https://link.springer.com/article/10.1007/s11606-021-06666-z
- Elaine C. Khoong, Eric Steinbrook, Cortlyn Brown and Alicia Fernandez, “Assessing the Use of Google Translate for Spanish and Chinese Translations of Emergency Department Discharge Instructions,” JAMA Internal Medicine, 25 February 2019: https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2725080
- Giovanni Rodriguez, Patricia Hernández, Christopher Kirwan, Lisette Dunham and Sayon Dutta, “Comparative Evaluation of Machine Translation Accuracy of Emergency Department Discharge Instructions: A Non-Inferiority Study,” Academic Emergency Medicine, volume 33, issue 4, 2026: https://onlinelibrary.wiley.com/doi/abs/10.1111/acem.70289
- Ryan C. L. Brewster et al., “Evaluating human-in-the-loop strategies for artificial intelligence-enabled translation of patient discharge instructions: a multidisciplinary analysis,” npj Digital Medicine, 24 October 2025: https://www.nature.com/articles/s41746-025-02055-6
- Nimdzi Insights, “The 2026 Nimdzi 100,” 2026: https://www.nimdzi.com/nimdzi-100-2026/
- Yujun Wang, Ehud Reiter, Shimei Pan, Steffen Eger and Wei Zhao, “Beyond Accuracy: Community Perspectives on Machine Translation,” arXiv:2606.09655, 8 June 2026: https://arxiv.org/abs/2606.09655
- Serge Gladkoff, Angelika Vaasa, Sue Ellen Wright, Ingemar Strandvik and Lifeng Han, “Looking under the Wrong Lamppost: On the Limitations of Automated Translation Quality Estimation,” arXiv:2608.03577, 4 August 2026, forthcoming in the Proceedings of the 9th International Conference on Natural Language and Speech Processing (ICNLSP 2026), Trento, Italy, September 2026: https://arxiv.org/abs/2608.03577
- University of Colorado Anschutz Medical Campus, “Researchers Examine Safety Risks in AI-Generated Translation of Emergency Department Discharge Instructions,” 23 June 2026: https://news.cuanschutz.edu/dbmi/ai-generated-discharge-instructions
- Carreras Tartak et al., “Evaluating Spanish Translations of Emergency Department Discharge Instructions by a Large Language Model: Tool Validation and Reliability Study,” JMIR Formative Research, 12 January 2026: https://formative.jmir.org/2026/1/e79676
- Ivan Lopez, David E. Velasquez, Jonathan H. Chen and Jorge A. Rodriguez, “Operationalizing machine-assisted translation in healthcare,” npj Digital Medicine, 30 September 2025: https://pmc.ncbi.nlm.nih.gov/articles/PMC12485017/
- Gail Price-Wise, “Language, Culture, And Medical Tragedy: The Case Of Willie Ramirez,” Health Affairs Forefront, 19 November 2008: https://www.healthaffairs.org/do/10.1377/forefront.20081119.000463/
- Iman Sharif and Julia Tse, “Accuracy of Computer-Generated, Spanish-Language Medicine Labels,” Pediatrics, May 2010: https://pubmed.ncbi.nlm.nih.gov/20368321/
- Rest of World, “AI translation jeopardizes Afghan asylum claims,” 19 September 2023: https://restofworld.org/2023/ai-translation-errors-afghan-refugees-asylum/
- TechCrunch, “Judge says 'literal but nonsensical' Google translation isn't consent for police search,” 15 June 2018: https://techcrunch.com/2018/06/15/judge-says-literal-but-nonsensical-google-translation-isnt-consent-for-police-search/
- The White House, “Designating English as the Official Language of the United States,” Executive Order 14224, 1 March 2025: https://www.whitehouse.gov/presidential-actions/2025/03/designating-english-as-the-official-language-of-the-united-states/
- Federal Register, “Notice of Rescission of Guidance to Federal Financial Assistance Recipients Regarding Title VI Prohibition Against National Origin Discrimination Affecting Limited English Proficient Persons,” 14 July 2026: https://www.federalregister.gov/documents/2026/07/14/2026-14128/notice-of-rescission-of-guidance-to-federal-financial-assistance-recipients-regarding-title-vi
- American Translators Association, “Section 1557 of the Affordable Care Act and Language Access: Who, What, How”: https://www.atanet.org/client-assistance/blog-section-1557-of-the-affordable-care-act-and-language-access-who-what-how/
- International Organization for Standardization, “ISO 18587:2017 Translation services, Post-editing of machine translation output, Requirements”: https://www.iso.org/standard/62970.html
- TÜV SÜD, “ISO 17100 and ISO 18587 Certifications, Translation Quality and Machine Translation Standards”: https://www.tuvsud.com/en-us/services/auditing-and-system-certification/iso-17100
- NHS England, “Improvement framework: community language translation and interpreting services,” 27 May 2025: https://www.england.nhs.uk/long-read/improvement-framework-community-language-translation-and-interpreting-services/
- EU-OSHA, “Regulation (EU) 2023/988 on general product safety,” applicable from 13 December 2024: https://osha.europa.eu/en/legislation/directive/regulation-2023988eu-general-product-safety
- EU Artificial Intelligence Act, “Article 50: Transparency Obligations for Providers and Deployers of Certain AI Systems”: https://artificialintelligenceact.eu/article/50/
- Google Cloud, “Cloud Translation pricing”: https://cloud.google.com/translate/pricing
- The AI Journal, “Manual translation processes still stifling enterprises despite surge in AI spending, finds DeepL research,” 10 March 2026: https://aijourn.com/manual-translation-processes-still-stifling-enterprises-despite-surge-in-ai-spending-finds-deepl-research/
- National Immigration Law Center, “Trump Administration's Attempts to Dismantle Language Access Do Not Erase Civil Rights Law”: https://www.nilc.org/articles/trump-administrations-attempts-to-dismantle-language-access-do-not-erase-civil-rights-law/

Tim Green UK-based Systems Theorist & Independent Technology Writer
Tim explores the intersections of artificial intelligence, decentralised cognition, and posthuman ethics. His work, published at smarterarticles.co.uk, challenges dominant narratives of technological progress while proposing interdisciplinary frameworks for collective intelligence and digital stewardship.
His writing has been featured on Ground News and shared by independent researchers across both academic and technological communities.
ORCID: 0009-0002-0156-9795 Email: tim@smarterarticles.co.uk
Listen to the free weekly SmarterArticles Podcast








