The Permeable Privacy of Language Models
It seems we can hardly go a day without hearing of some new man-made oracle, a large language model ready to automate our workflows, debug our code, and perhaps even soothe our existential dread. This rush to integrate AI into every facet of our digital lives comes with a chorus of reassurances about privacy, a notion that feels increasingly quaint, like a historical curiosity from an analog era. Yet, the question of privacy isn’t cast in some dusty relic but actually set in a very live wire.
"the granular linguistic knowledge needed to teach an LLM"
Cast your mind back to 1987, when a newspaper published a list of 146 films a Supreme Court nominee had rented, a minor scandal not for its salaciousness (there was none) but for its simple invasion of a private space. That small intrusion was enough to spur the US Congress into action and create laws to protect the sanctity of one's video rental history. Today, the local video store has been replaced by a global network of streaming platforms and digital services, and the threat is not a nosy journalist but an automated system that never sleeps, one that shares our digital footprints with unseen partners for purposes we never agreed to. At least not explicitly.
The heart of this digital tension lies in a deceptively simple acronym: PII, or Personally Identifiable Information. We might think of PII as a fixed list (name, address, social security number) but that tidy definition is a phantom. Legal frameworks like GDPR and HIPAA try to pin it down and offer foundational definitions of data that can directly or indirectly identify a person. Yet, people’s own understanding of PII is far more fluid and, frankly, more attuned to reality. For most, PII is a process, something that evolves as data accumulates over time, as disparate pieces of information are linked together by technological affordances, and as our digital lives become more pervasively exposed.
Shif
Even a "spam" email address, created for anonymity, can become identifiable through repeated use and association with other details, like a credit card number used for a single purchase years ago. The courts themselves are wrestling with this moving target. In a recent case, judges had to determine whether a combination of a users Facebook ID and the title of a video they streamed constituted an impermissible disclosure of PII. Their conclusion, which leaned on what an "ordinary person" could deduce, tells you everything you need to know about the ambiguity at play. If we humans, with our laws and our ordinary common sense, cannot draw a bright line around what is and isn’t PII, what chance does a machine have?
Shift and
Here lies the paradox at the center of the AI revolution, these LLMs are being pitched as both the poison and the antidote. Frameworks are being built to de-identify clinical records and automate the redaction of government documents (with or without auto-pens), with some models reporting accuracy scores as high as 99.55% in detecting PII in unstructured texts. Such figures are waved about like talismans, meant to ward off the fears of regulators and the public. But the devil, as always, is in the details, that sliver of a percentage point where things go wrong.
Shift and
These models, trained on gargantuan but non-specific datasets, exhibit what might be called a regularity bias. They are excellent at spotting patterns they have seen before, but they are easily confused by ambiguity. A recent study built a benchmark of human names that resemble non-human entities and the results were a comedy of errors: LLMs regularly mistook the name Italys for the country Italy, even when the text used the pronoun "she" and misclassified a name like Adomite as a mineral. The recall rate for these ambiguous names dropped by a staggering 20-40% compared to more common names.
This is a gaping hole in the privacy net we are told is being woven around us. If a state-of-the-art model struggles with ambiguous English names, its performance on names from low-resource languages, such as those spoken across our Pacific Islands, will be nothing short of abysmal. A name in Tokelauan or Niuean is almost certain to be misinterpreted by a model that has never encountered it, likely being classified as a location, an object, a flower, an animal, or simply ignored altogether. This failure to see what is right in front of it is only the beginning of the problem. The very architecture of these LLM-powered systems creates new vulnerabilities.
This is a gaping hole in the privacy net we are told is being woven around us. If a state-of-the-art model struggles with ambiguous English names, its performance on names from low-resource languages, such as those spoken across our Pacific Islands, will be nothing short of abysmal. A name in Tokelauan or Niuean is almost certain to be misinterpreted by a model that has never encountered it, likely being classified as a location, an object, a flower, an animal, or simply ignored altogether. This failure to see what is right in front of it is only the beginning of the problem. The very architecture of these LLM-powered systems creates new vulnerabilities.
Shift a
Private conversations with a chatbot become a permanent, exploitable record of our preferences, habits, and secrets, a record that can be leaked through everything from the model's own reasoning traces to insecure tool usage and side-channel attacks that analyze server response times. Worse still, these powerful tools can be weaponized. Malicious actors can instruct an LLM to perform automated profiling, systematically scraping a person's public online activity to infer sensitive attributes and de-anonymize them. They can be used to launch large-scale, hyper-personalized social engineering campaigns, crafting phishing emails so convincing that they bypass our natural skepticism. One recent incident saw fraudsters use a combination of phishing and deepfake video to trick an employee into transferring US$25 million.
Sh
So where does this leave us? On the one hand, we have the immense utility of LLMs to process information at a scale and speed previously unimaginable, and on the other, a cascade of privacy risks that current models seem ill-equipped to handle. The process for disclosing government information in countries like Japan, for example, requires a meticulous, character-level masking of confidential data based on the contextual interpretation of legal texts. Lawyers using these tools risk breaching the bedrock principle of client confidentiality, as any PII included in a prompt to an external LLM could be considered an unauthorized disclosure.
Shift
So the need for a solution is palpable, isn't it. This is where the path forward narrows, moving away from the fantasy of a single, all-powerful AI (Yes, think AGI) and toward a more pragmatic, collaborative model. Generic LLMs, with their built-in biases and blind spots, are not enough. Preparing them for specialized, high-stakes environments must go beyond surface level in order to understand language and culture. It requires the creation of bespoke benchmarks to test for weaknesses, the careful fine-tuning on domain-specific data, and the rigorous auditing of their outputs.
This is the quiet work that happens at the intersection of technology and the humanities, the domain of specialized language experts. A company like ours, with its focus on the great linguistic diversity of the Pacific Islands, is positioned to translate by contextualizing culture. We can provide the granular linguistic knowledge needed to teach an LLM the difference between a person’s name and a place name in Fijian. We can build the datasets that reflect the realities of legal or medical practice in Papua New Guinea.
This is the quiet work that happens at the intersection of technology and the humanities, the domain of specialized language experts. A company like ours, with its focus on the great linguistic diversity of the Pacific Islands, is positioned to translate by contextualizing culture. We can provide the granular linguistic knowledge needed to teach an LLM the difference between a person’s name and a place name in Fijian. We can build the datasets that reflect the realities of legal or medical practice in Papua New Guinea.
Shift
We can serve as the human auditors who verify that a model tasked with redacting sensitive information from a court document in Palauan does its job without error. Without this layer of specialized human expertise, organizations operating in linguistically diverse regions are walking in a minefield. They are trusting a tool that has been shown to fail in predictable ways, exposing themselves and their clients to legal, financial, and reputational ruin.
Shift
The journey ahead with these new technologies is not about finding a perfect, autonomous solution that removes human judgment from the equation. It seems, rather, to be about forging a more symbiotic relationship between human expertise and machine capability. The question is not whether we will entrust our information to these models, for that ship has already sailed. The more pressing question is how we will guide them, how we will teach them the intricacies they cannot learn from a generic web scrape, and how we will hold them accountable when they inevitably fall short. The future of AI and privacy may lie in a council of specialized collaborators, where everyone lends their voice to a more responsible and context-aware technological ecosystem.
Huri Translations
Tel. +689 89 205 483
[email protected]
PO BOX 365 Maharepa
98728 Mo'orea
French Polynesia
N°TAHITI 876649