Why Your App Is a Legally Unqualified Translator
Machine translation gaffes provide endless online amusement, but when these digital tools enter the sterile confines of a hospital, the laughter stops. A mistranslated symptom or a misunderstood diagnosis is a high-stakes vector for patient harm, especially for speakers of low-resource languages.
"The 2024 Final Rule from the U.S. Department of Health and Human Services on Section 1557 of the Affordable Care Act is explicit"
We tend to frame this as a new technological problem, but the legal system has already rendered its verdict. The most catastrophic precedent for language-access failure, the $71 million malpractice settlement for Willie Ramirez in 1980, had nothing to do with a computer. It hinged on a single word. His Spanish-speaking family said he was "intoxicado". For he who knows a little Spanish, that means "poisoned" or "ill" from something Willie ingested. A bilingual but unqualified hospital staffer translated this as "intoxicated", leading physicians to treat a drug overdose while the patient's actual condition, a brain hemorrhage, went undiagnosed, leaving him quadriplegic (or Tetraplegic, if we use the Greek root).
The legal precedent was set: Using an unqualified interpreter is an act of negligence. A free, general-purpose MT tool is the digital equivalent of that same unqualified staffer. The assumption that simple bilingualism equals qualification is the core of the error, whether the language is Spanish or Tongan.
A 2025 study on healthcare viability for low-resource languages from UC Berkeley cataloged the grim failure modes when using public MT systems for Amharic and Tigrinya. The errors were not subtle. The word Antibiotics was consistently translated as Insecticide in Tigrinya. Antivenom became Anti-nutrients. In physician-patient communication, the simple instruction Strict Non-Weight Bearing was translated as Strict Weight Bearing. It was a direct, life-altering inversion of a medical order.
This is the machine's version of the Intoxicado error. The system, lacking any medical or real-world context, seizes on the wrong synonym. Masses in the sense of tumors becomes Collection, and Stool becomes Chair. One of the most alarming examples of this context-blindness was an English idiom: "Aim to wean down this medication" was literally translated into Amharic as "Aim to cut off breasts". Well, this isn't a problem limited to Africa. Translating such idioms into Hawaiian or Māori with a machine can be just as perilous.
In fact, this problem of linguistic disparity is baked into the technology itself. Public MT models are trained on the Internet's firehose of high-resource languages, like English and Spanish. For languages outside this dominant corpus, the tools are brittle and unreliable. A 2021 study by Breena R. Taira from Olive-View UCLA Medical Center assessing Google Translate for emergency department discharge instructions found a 94% accuracy rate for Spanish. This figure, while not perfect, seems serviceable.
The same study, however, found the accuracy for Armenian translations plummeted to 55%. So, the technological risk is not distributed equally. It falls heaviest on the speakers of low-resource languages who are, by definition, the most vulnerable. This is a recurring theme in our own linguistic work, where capturing the medical and cultural context for languages like Fijian or Samoan is a puzzle that automated systems are unequipped to solve.
The failure, however, now extends beyond mistranslation. The new generation of smarter models has introduced an entirely new, almost sinister, category of error: Fabrication. Or Hallucination, as we put it in NLP circles. A 2024 article by Thom Holwerda in OSNews.com detailed how his family, reviewing a Swedish medical record translated into Dutch by Google Translate, was horrified to read that the patient was "combative and non-cooperative". This sentence, which caused his family quite some distress, was not a bad translation. It simply did not exist in the original Swedish document. The machine, in its next-token prediction function had hallucinated a phantom sentence of clinical slander. Imagine the confusion if a Tahitian medical record suddenly contained such hallucination: "The patient was medivacʻd by drone to the local hospital".
On the opposite end of the spectrum, other models help by processing medical notes to remove private information. A 2025 study on this de-identification process from the University of Alberta found that AI models consistently over-redacted notes. This help removed clinically meaningful information, and physicians reviewing the automated edits rated 70-80% of these removals as having a high-impact on patient care.
For decades, healthcare providers have operated in a gray area, using these tools because they were fast, free, and, often, the only option available. That gray area has now vanished. The 2024 Final Rule from the U.S. Department of Health and Human Services on Section 1557 of the Affordable Care Act is explicit. This federal regulation now mandates that if machine translation is used for critical, complex, or technical language (a category that includes virtually all patient communication, right?) it must be reviewed by a qualified human translator. With this rule, the unsupervised use of a free translation app for patient care has moved from the category of "Risky" to "Legally Non-compliant". The legal exposure, once theoretical, is now concrete.
The technology will undoubtedly continue to be part of the clinical workflow. But the belief that it could, or should, replace human linguistic and cultural expertise in high-stakes environments has been shown to be pure myth. The machine can process words but it comprehends nothing, and yes, sometimes, gets too creative. The accountability, as Willie Ramirez's $71 million settlement proved four decades ago, remains stubbornly, and expensively, human. The question for organizations is no longer whether to adopt these tools, but how they plan to manage the new altitude of risk they have inherited.
The legal precedent was set: Using an unqualified interpreter is an act of negligence. A free, general-purpose MT tool is the digital equivalent of that same unqualified staffer. The assumption that simple bilingualism equals qualification is the core of the error, whether the language is Spanish or Tongan.
A 2025 study on healthcare viability for low-resource languages from UC Berkeley cataloged the grim failure modes when using public MT systems for Amharic and Tigrinya. The errors were not subtle. The word Antibiotics was consistently translated as Insecticide in Tigrinya. Antivenom became Anti-nutrients. In physician-patient communication, the simple instruction Strict Non-Weight Bearing was translated as Strict Weight Bearing. It was a direct, life-altering inversion of a medical order.
This is the machine's version of the Intoxicado error. The system, lacking any medical or real-world context, seizes on the wrong synonym. Masses in the sense of tumors becomes Collection, and Stool becomes Chair. One of the most alarming examples of this context-blindness was an English idiom: "Aim to wean down this medication" was literally translated into Amharic as "Aim to cut off breasts". Well, this isn't a problem limited to Africa. Translating such idioms into Hawaiian or Māori with a machine can be just as perilous.
In fact, this problem of linguistic disparity is baked into the technology itself. Public MT models are trained on the Internet's firehose of high-resource languages, like English and Spanish. For languages outside this dominant corpus, the tools are brittle and unreliable. A 2021 study by Breena R. Taira from Olive-View UCLA Medical Center assessing Google Translate for emergency department discharge instructions found a 94% accuracy rate for Spanish. This figure, while not perfect, seems serviceable.
The same study, however, found the accuracy for Armenian translations plummeted to 55%. So, the technological risk is not distributed equally. It falls heaviest on the speakers of low-resource languages who are, by definition, the most vulnerable. This is a recurring theme in our own linguistic work, where capturing the medical and cultural context for languages like Fijian or Samoan is a puzzle that automated systems are unequipped to solve.
The failure, however, now extends beyond mistranslation. The new generation of smarter models has introduced an entirely new, almost sinister, category of error: Fabrication. Or Hallucination, as we put it in NLP circles. A 2024 article by Thom Holwerda in OSNews.com detailed how his family, reviewing a Swedish medical record translated into Dutch by Google Translate, was horrified to read that the patient was "combative and non-cooperative". This sentence, which caused his family quite some distress, was not a bad translation. It simply did not exist in the original Swedish document. The machine, in its next-token prediction function had hallucinated a phantom sentence of clinical slander. Imagine the confusion if a Tahitian medical record suddenly contained such hallucination: "The patient was medivacʻd by drone to the local hospital".
On the opposite end of the spectrum, other models help by processing medical notes to remove private information. A 2025 study on this de-identification process from the University of Alberta found that AI models consistently over-redacted notes. This help removed clinically meaningful information, and physicians reviewing the automated edits rated 70-80% of these removals as having a high-impact on patient care.
For decades, healthcare providers have operated in a gray area, using these tools because they were fast, free, and, often, the only option available. That gray area has now vanished. The 2024 Final Rule from the U.S. Department of Health and Human Services on Section 1557 of the Affordable Care Act is explicit. This federal regulation now mandates that if machine translation is used for critical, complex, or technical language (a category that includes virtually all patient communication, right?) it must be reviewed by a qualified human translator. With this rule, the unsupervised use of a free translation app for patient care has moved from the category of "Risky" to "Legally Non-compliant". The legal exposure, once theoretical, is now concrete.
The technology will undoubtedly continue to be part of the clinical workflow. But the belief that it could, or should, replace human linguistic and cultural expertise in high-stakes environments has been shown to be pure myth. The machine can process words but it comprehends nothing, and yes, sometimes, gets too creative. The accountability, as Willie Ramirez's $71 million settlement proved four decades ago, remains stubbornly, and expensively, human. The question for organizations is no longer whether to adopt these tools, but how they plan to manage the new altitude of risk they have inherited.
Huri Translations
Tel. +689 89 205 483
[email protected]
PO BOX 365 Maharepa
98728 Mo'orea
French Polynesia
N°TAHITI 876649