Making the Case for Precise Pacific Orthography
On the social media, a Tahitian teacher, puzzled by orthography, asked a simple question: Is it Māuruuru or Māuru'uru? The query, seemingly minor, pulls the pin on a century-long grenade of linguistic conflict. It exposes the rifts between academic standardization efforts and legacy writing systems, a battle being fought across the Pacific Islands that leaves learners confused and languages compromised in the digital age.
"We are forced to become linguistic arbitrators, decoding ambiguous texts before we can even begin to translate"
As our very own Tamatoa Audouin explains, the speaker's uncertainty is a direct byproduct of the so-called Turo or Raʻapoto spelling system (named after prominent theologian Duro a Raʻapoto). This graphy, as Edgar Tetahiotupa's 2004 paper details, operates on the flawed assumption that any two consecutive, identical vowels are inevitably read with a glottal stop. The system therefore just writes haapiiraa, forcing the reader to already know the word is pronounced as Haʻapiʻiraʻa [haʔapiʔiɾaʔa]. A learner, seeing the academic Māuruuru, applies this Turo logic and hallucinates a glottal stop that isn't there.
Tamatoa's acidic response gets to the heart of it: Such a system prevents learning the language from writing because it supposes that people know from the outset how the word is pronounced. It’s a closed loop, a key that only works for those already inside the room. This is precisely why the Marquesan Language Academy's 2000 decision to adopt this very French-looking writing system is seen by linguists as such a baffling, retrograde step.
This is not a new fight, though. Scholars have been trying to pin down Polynesian phonology for a century, often with clumsy or overly elaborate results. Take anthropologist J. Frank Stimson's 1928 proposal in the The Journal of the Polynesian Society. Working on a comparative dictionary, he found the existing spellings so inadequate for philology that he devised a technical system with nine distinct diacritical marks.
He had marks for consonants lost with a glottal stop (‘oe from koe), consonants lost without a glottal stop (ia from kia), and three different marks for vowel length depending on whether the vowel was inherently long (mā), fused from two non-long vowels (noäʻtu for noa atu and vaïʻho for vai iho), or fused from vowels where one was already long (manâʻha for manā aha and ʻuâʻta for ʻua āta). Stimson's system was a tool for the lab, not the street (or the church) but it underscores a fundamental point: The sounds, the history, and the phonemic facts of the language need precision, and orthographies that ignore them are, for scholarshipʻs sake, useless.
For, the real-world damage of these economical or contextual systems is now being assessed. A paper from NAACL 2025 in Suzhou, China, entitled "Døñ’t Tòůčḥ Mý Dïąçŗıtīcs", lays out the high cost of ambiguity. Researchers Gorman and Pinter found that stripping diacritics (or even just encoding them inconsistently in Unicode) has adverse effects on modern NLP models. When fed diacritic-rich languages, tokenizers required between two and four times as many tokens to process the text versus its stripped version. This is a direct computational tax worth every penny in the bank (for what is left, Uncle Sam).
A 2025 paper from EMNLP 2025 (苏州市, China) on Hawaiian orthography tells us that simple statistical n-gram models outperformed neural networks (think Llama and the likes) at the task of restoring missing glottal stops and macrons from Old Hawaiʻi nūpepa, which can certainly be attributed to poor data used to train the neural nets. Their most advanced tools are failing because they are being trained on compromised, ambiguous, user-generated data. This is the direct inheritance of letting flawed orthographies fester.
The story of Samoan diacritics is a case study in how this happens. A 2015 paper by Eseta Magaui Tualaulelei et al from University of Hawaiʻi traces the current chaos to three competing historical sources. First was the missionary Rev. George Pratt, who in the 1860s used diacritics sparingly, only adding them to prevent blatant ambiguity, like distinguishing laʻu (my) from lau (your). He trusted his fluent audience to fill in the blanks.
Second was the linguist George Milner, whose 1966 dictionary used marks comprehensively, arguing that words in a dictionary are divorced from their context and need full phonemic detail. Third was the post-independence Samoan government itself, which in 1969 essentially prohibited the use of diacritics in schools to end the confusion. Good Gracious. This policy created several generations of native speakers educated to see the marks as irrelevant. The result: Three incompatible systems all in use at once, a perfect recipe for the intolerable inconsistency J. Frank Stimson complained about back in 1930.
Oh yes, this inconsistency is a serious obstacle. The Samoan paper's authors call it the most serious issue for the diaspora, where learners in classrooms are heavily dependent on written materials and cannot guess from context like a native speaker. The lack of a standard has real-world consequences that create irritation and unintended offense when, for example, a chiefly title like Matāʻafa is mispronounced because its spelling is ambiguous. The name is pronounced differently in Lotofaga versus Palauli, and the diacritics are the only things that show which is which. And the same ambiguity is left intact between Avera village in Rūrutu (Averā) and Avera village in Raʻiātea (ʻĀvera).
This is where, in our own work as translators, the academic collides with the practical. We are forced to become linguistic arbitrators, decoding ambiguous texts before we can even begin to translate. When we encounter new medical or technical terms, the absence of a clear, phonemic orthography makes our task of rendering them accurately that much more fraught.
We are left with a situation where some languages are hobbled by a phonotactic problem, a mismatch between sound and symbol. A 2018 survey of Central Pacific languages by Kie Zuraw from UCLA found that trochaic shortening, the famous rule from Fijian, is poorly attested. The language doesn't actually shorten vowels in the way the textbooks claim. Instead, languages like Tongan solve the phonological puzzle using Breaking to split a long vowel into two syllables, as in ma(áma) (lamp, light).
This strategy, unlike shortening, preserves the underlying vowel length, making it a stable and well-attested system for learners to acquire. This may be why Tongan Accent, Albert Schütz's 2001 paper, argues that it is better understood not by counting syllables but by measures, which are phonological units that combine in a hierarchy. Fale (house) and lahi (big) are two separate measures, fale.lahi. But nófo (stay) and the particle ni shift tone to combine into one measure: nofó ni.
An orthography that obscures languages, like the so-called Turo system, may serve as a shorthand for those already fluent, but it is a dead end for everyone else. It is a gate slammed in the face of learners. It is a poisoned well for the digital tools that will determine the language's future accessibility. The academic spellings, like that of the Fare Vānaʻa (Tahitian Language Academy), are not academic because they are difficult. They are academic because they are precise. They are built for transmission, for preservation, and for learning. A writing system, after all, should be a map to the language, not a riddle.
Tamatoa's acidic response gets to the heart of it: Such a system prevents learning the language from writing because it supposes that people know from the outset how the word is pronounced. It’s a closed loop, a key that only works for those already inside the room. This is precisely why the Marquesan Language Academy's 2000 decision to adopt this very French-looking writing system is seen by linguists as such a baffling, retrograde step.
This is not a new fight, though. Scholars have been trying to pin down Polynesian phonology for a century, often with clumsy or overly elaborate results. Take anthropologist J. Frank Stimson's 1928 proposal in the The Journal of the Polynesian Society. Working on a comparative dictionary, he found the existing spellings so inadequate for philology that he devised a technical system with nine distinct diacritical marks.
He had marks for consonants lost with a glottal stop (‘oe from koe), consonants lost without a glottal stop (ia from kia), and three different marks for vowel length depending on whether the vowel was inherently long (mā), fused from two non-long vowels (noäʻtu for noa atu and vaïʻho for vai iho), or fused from vowels where one was already long (manâʻha for manā aha and ʻuâʻta for ʻua āta). Stimson's system was a tool for the lab, not the street (or the church) but it underscores a fundamental point: The sounds, the history, and the phonemic facts of the language need precision, and orthographies that ignore them are, for scholarshipʻs sake, useless.
For, the real-world damage of these economical or contextual systems is now being assessed. A paper from NAACL 2025 in Suzhou, China, entitled "Døñ’t Tòůčḥ Mý Dïąçŗıtīcs", lays out the high cost of ambiguity. Researchers Gorman and Pinter found that stripping diacritics (or even just encoding them inconsistently in Unicode) has adverse effects on modern NLP models. When fed diacritic-rich languages, tokenizers required between two and four times as many tokens to process the text versus its stripped version. This is a direct computational tax worth every penny in the bank (for what is left, Uncle Sam).
A 2025 paper from EMNLP 2025 (苏州市, China) on Hawaiian orthography tells us that simple statistical n-gram models outperformed neural networks (think Llama and the likes) at the task of restoring missing glottal stops and macrons from Old Hawaiʻi nūpepa, which can certainly be attributed to poor data used to train the neural nets. Their most advanced tools are failing because they are being trained on compromised, ambiguous, user-generated data. This is the direct inheritance of letting flawed orthographies fester.
The story of Samoan diacritics is a case study in how this happens. A 2015 paper by Eseta Magaui Tualaulelei et al from University of Hawaiʻi traces the current chaos to three competing historical sources. First was the missionary Rev. George Pratt, who in the 1860s used diacritics sparingly, only adding them to prevent blatant ambiguity, like distinguishing laʻu (my) from lau (your). He trusted his fluent audience to fill in the blanks.
Second was the linguist George Milner, whose 1966 dictionary used marks comprehensively, arguing that words in a dictionary are divorced from their context and need full phonemic detail. Third was the post-independence Samoan government itself, which in 1969 essentially prohibited the use of diacritics in schools to end the confusion. Good Gracious. This policy created several generations of native speakers educated to see the marks as irrelevant. The result: Three incompatible systems all in use at once, a perfect recipe for the intolerable inconsistency J. Frank Stimson complained about back in 1930.
Oh yes, this inconsistency is a serious obstacle. The Samoan paper's authors call it the most serious issue for the diaspora, where learners in classrooms are heavily dependent on written materials and cannot guess from context like a native speaker. The lack of a standard has real-world consequences that create irritation and unintended offense when, for example, a chiefly title like Matāʻafa is mispronounced because its spelling is ambiguous. The name is pronounced differently in Lotofaga versus Palauli, and the diacritics are the only things that show which is which. And the same ambiguity is left intact between Avera village in Rūrutu (Averā) and Avera village in Raʻiātea (ʻĀvera).
This is where, in our own work as translators, the academic collides with the practical. We are forced to become linguistic arbitrators, decoding ambiguous texts before we can even begin to translate. When we encounter new medical or technical terms, the absence of a clear, phonemic orthography makes our task of rendering them accurately that much more fraught.
We are left with a situation where some languages are hobbled by a phonotactic problem, a mismatch between sound and symbol. A 2018 survey of Central Pacific languages by Kie Zuraw from UCLA found that trochaic shortening, the famous rule from Fijian, is poorly attested. The language doesn't actually shorten vowels in the way the textbooks claim. Instead, languages like Tongan solve the phonological puzzle using Breaking to split a long vowel into two syllables, as in ma(áma) (lamp, light).
This strategy, unlike shortening, preserves the underlying vowel length, making it a stable and well-attested system for learners to acquire. This may be why Tongan Accent, Albert Schütz's 2001 paper, argues that it is better understood not by counting syllables but by measures, which are phonological units that combine in a hierarchy. Fale (house) and lahi (big) are two separate measures, fale.lahi. But nófo (stay) and the particle ni shift tone to combine into one measure: nofó ni.
An orthography that obscures languages, like the so-called Turo system, may serve as a shorthand for those already fluent, but it is a dead end for everyone else. It is a gate slammed in the face of learners. It is a poisoned well for the digital tools that will determine the language's future accessibility. The academic spellings, like that of the Fare Vānaʻa (Tahitian Language Academy), are not academic because they are difficult. They are academic because they are precise. They are built for transmission, for preservation, and for learning. A writing system, after all, should be a map to the language, not a riddle.
Huri Translations
Tel. +689 89 205 483
[email protected]
PO BOX 365 Maharepa
98728 Mo'orea
French Polynesia
N°TAHITI 876649