Cannibalizing Human Knowledge
Generative AI copyright battles are the industrial scaling of a linguistic heist begun in the 1990s. From the sentence-pair databases of Translation Memory to the probabilistic hallucinations of Large Language Models, the erosion of translator rights remains the primary vector of data extraction.
"are we just rearranging deck chairs on a sinking epistemological ship?"
This article traces that lineage to argue that current legal frameworks are ripe for a fundamental reboot to address the realities of Indigenous Data Sovereignty and the commodification of syntax and verb.
The current hysteria regarding copyright in the age of generative Artificial Intelligence is often framed as a sudden, unprecedented disruption precipitated by the arrival of Large Language Models. Rigorous historiography of the language services industry however suggests otherwise. We posit that the contemporary legal conflicts surrounding AI training data are not sui generis events. They are the industrial culmination of a process of linguistic dispossession that began decades earlier with the advent of Translation Memory (TM) technology. We serve as witnesses to the fact that the translation industry acted as a testing ground for the legal theories and economic models now deployed by global technology giants.
To have a grasp on the coming of the current dispute, one must first dissect the technological ecosystem of the Translation Memory. This tool dominated the industry from the mid-1990s through the 2010s . It fundamentally altered the ontology of translation. It shifted the craft from a linear act of creative writing to a non-linear process of database management. TMs operate on a principle of explicit storage and retrieval rather than the probabilistic generation of modern LLMs. When a translator segments a text, the software stores the output alongside the original. This creates a repository of so-called Translation Units.
The ubiquity of these tools masked a shift in the nature of the work. The translator ceased to be the sole author of a coherent text. They became the operator of a dynamic database. The pivotal moment in the commodification of translation data occurred when TMs began to be conceptualized as tradeable corporate assets more than productivity aids. Platforms like TM Marketplace emerged by 2007 to explicitly market these assets. Agencies and corporate clients asserted ownership over the TM files generated during a project. The economic logic was seductive. If the client paid for the translation of a manual, they believed they owned the memory of that translation.
This practice collided with the fundamental tenets of copyright law. A translation is by definition a derivative work. The creator of a derivative work holds a copyright in their contribution provided it is original. Yet translators were systematically stripped of these rights through contractual coercion. The primary legal mechanism used to alienate translators from their data was the "work made for hire" doctrine in U.S. copyright law. This doctrine creates a legal fiction where the author is the entity that commissioned the work rather than the person who created it.
A major point in the TM era complicates the ownership narrative. We must distinguish between the copyright in the content and the copyright in the database. The US Supreme Court ruling in Feist Publications, Inc. v. Rural Telephone Service Co. established a high bar for database protection. The Court ruled that "sweat of the brow" does not merit copyright protection. A compilation is only copyrightable if it possesses originality in its selection or arrangement. Most TMs are comprehensive collections of all translations done for a specific project. There is often no creative selection involved. Consequently, legal scholars argued that the TM structure might not be copyrightable under U.S. law.
The industry bypassed this uncertainty through contract law. Non-disclosure agreements defined TMs as proprietary information. This strategy of using Terms of Service to capture data rights was a direct precursor to the strategies used by social media and AI companies today. The transition from Translation Memories to Machine Translation and eventually LLMs was a fast, gradual evolution. The data amassed in TMs became the fuel for the next generation of automation. TMs were the ideal training data because they were high quality, aligned, and domain-specific.
The reuse of TMs for training MT engines represented a secondary commodification. Translators who had surrendered their TMs to agencies found their work used to train the very machines that threatened to devalue their labor. This reuse represented a fundamental betrayal of the spirit of copyright law. The focus shifted from the expressive value of a text to its informational and statistical value. This conceptual shift laid the intellectual groundwork for the "fair use" arguments seen today.
In fact, the rise of Large Language Models represents the industrial scaling of these dynamics. The LLM era relies on the copyright doctrine of Fair Use to justify the ingestion of the entire available web. Tech companies must argue that their use of copyrighted data is transformative. Several high-profile lawsuits have begun to test these arguments. The New York Times v. OpenAI strikes at the heart of the memorization issue. OpenAI has contended that The Times is not harmed by AI outputs generally but is conflating its claims with broader market harms caused by so-called pink slime organizations.
Now, while the commercial translation industry fights over economic rights, Indigenous communities in the Pacific are fighting a parallel battle over cultural rights and Traditional Knowledge. Their experience is the starkest evidence of the dangers of uncontrolled data extraction. LLMs frequently hallucinate when translating Indigenous languages because they lack a grounded understanding of cultural context. Why? Because there is simply not enough available data to train AI models. In Hawaiian culture, the concept of Kaona refers to hidden or layered meanings in poetry and speech. AI models trained on surface-level text are blind to Kaona. They translate literally and effectively erase the cultural archive encoded within the ʻŌlelo.
Samoan society is structured around the Matai system. The vocabulary changes entirely depending on the rank of the person being addressed. AI models process text as statistical tokens without social awareness and the models frequently fail to master this hierarchy. They may address a high chief with commoner language or mix registers in a way that is socially offensive. Research has documented specific error categories where AI inserts unjustified words or uses an inappropriate language register. For a Samoan user, an AI model that cannot distinguish between a Susuga and an Afioga is an insult.
The most egregious example of AI harm is the hallucination of Whakapapa in Māori contexts. Whakapapa is the foundational principle of Māori identity and genealogy. AI models have been documented fabricating genealogy by assigning the wrong ancestors to tribes. Because the models generate text based on probability rather than fact, they construct plausible-sounding but entirely false connections. And that is a violation of Tapu (Taboo in its original sense).
The conflict between Te Hiku Media and tech giants is a definitive case study in Indigenous resistance and resilience. When OpenAI developed its Whisper model, it claimed to have trained on Māori and Hawaiian data. Te Hiku Media suspected this data was scraped without consent. Traditional agencies attempted to purchase Te Hiku's data for low sums. Te Hiku recognized that selling their data would result in a model owned by a foreign corporation being sold back to Māori people.
In response, Te Hiku developed the Kaitiakitanga License. This license is grounded in the Māori concept of guardianship. It asserts that data is a Taonga (a cognate of the Hawaiian Kaona we mentioned above). The license demands benefit sharing and provenance tracking. This framework challenges the axioms of the TM/Copyright model and rejects the idea that data can be alienated from its creator via a contract.
Well, the struggle is seen across the Pacific Islands. For example, the Tongan language possesses three distinct social registers. AI models trained primarily on commoner text struggle to generate the regal terminology required for formal contexts. In Fiji, the concept of Vaka-Tūraga defines respectful interaction and we have yet seen countless misused renditions of the concept.
So, we see that the TM era established the normative precedent that linguistic labor could be alienated from the laborer, all legally. Once the industry accepted that a translator did not own their TM, it accepted the principle that the value of linguistic data accrues to the aggregator. The LLM era just expanded the scale and scope as generative text models unfolded. The aggregator is no longer a translation agency but a global tech platform that acquired their training data by hard-to-track methods. The laborer is no longer just a professional translator but anyone who has ever written text online.
To see how confusion has spread to courts, in Kadrey v. Meta, the judge called the theory that a model is a derivative work nonsensical without evidence of exact replicas. This distinction between the training process and the model weights suggests that current copyright law may struggle to classify a neural network as a copy. However, in Thomson Reuters v. ROSS Intelligence, a court rejected a fair use defense because the AI created a competing commercial product. This precedent suggests that courts may be less forgiving when AI serves as a direct market substitute.
It is safe to say that Translation Memories were indeed a precursor to the copyright issues facing LLMs today. TMs pioneered the atomization of text and prepared the ground for the statistical tokenization of Neural Machine Translation networks and next-token prediction now used by LLMs. Legally, the TM era normalized the circumvention of creator rights through the Work Made for Hire doctrine while economically, TMs entrenched a model of extraction.
But that's not all: If LLMs are trained on AI-generated content, the quality of the memory degrades. Indigenous examples of Whakapapa hallucination are early warning signs of this epistemic decay. The specialists speak of Cannibalism, for if the future of human knowledge is an LLM trained on the output of other LLMs, the derivative work becomes a copy of a copy, with output quality dropping dramatically until breakdown. The most robust solution may lie in adopting models like the Kaitiakitanga License. If creators retained a fractional stewardship interest in their data tokens, the extractive model would be forced to become a collaborative one. But is the current technology prepared, let alone willing to go this path? Not sure.
Courts in 2025 are deciding whether the logic of the database over the creator will become the constitutional law of the digital age. The evidence indicates that without a fundamental rethinking of data dignity, the dispossession that began with a sentence pair will end with the full automation of culture itself, and subsequent collapse. The question: Does the current legal infrastructure have the requisite elasticity to contain this phenomenon or are we just rearranging deck chairs on a sinking epistemological ship?
The current hysteria regarding copyright in the age of generative Artificial Intelligence is often framed as a sudden, unprecedented disruption precipitated by the arrival of Large Language Models. Rigorous historiography of the language services industry however suggests otherwise. We posit that the contemporary legal conflicts surrounding AI training data are not sui generis events. They are the industrial culmination of a process of linguistic dispossession that began decades earlier with the advent of Translation Memory (TM) technology. We serve as witnesses to the fact that the translation industry acted as a testing ground for the legal theories and economic models now deployed by global technology giants.
To have a grasp on the coming of the current dispute, one must first dissect the technological ecosystem of the Translation Memory. This tool dominated the industry from the mid-1990s through the 2010s . It fundamentally altered the ontology of translation. It shifted the craft from a linear act of creative writing to a non-linear process of database management. TMs operate on a principle of explicit storage and retrieval rather than the probabilistic generation of modern LLMs. When a translator segments a text, the software stores the output alongside the original. This creates a repository of so-called Translation Units.
The ubiquity of these tools masked a shift in the nature of the work. The translator ceased to be the sole author of a coherent text. They became the operator of a dynamic database. The pivotal moment in the commodification of translation data occurred when TMs began to be conceptualized as tradeable corporate assets more than productivity aids. Platforms like TM Marketplace emerged by 2007 to explicitly market these assets. Agencies and corporate clients asserted ownership over the TM files generated during a project. The economic logic was seductive. If the client paid for the translation of a manual, they believed they owned the memory of that translation.
This practice collided with the fundamental tenets of copyright law. A translation is by definition a derivative work. The creator of a derivative work holds a copyright in their contribution provided it is original. Yet translators were systematically stripped of these rights through contractual coercion. The primary legal mechanism used to alienate translators from their data was the "work made for hire" doctrine in U.S. copyright law. This doctrine creates a legal fiction where the author is the entity that commissioned the work rather than the person who created it.
A major point in the TM era complicates the ownership narrative. We must distinguish between the copyright in the content and the copyright in the database. The US Supreme Court ruling in Feist Publications, Inc. v. Rural Telephone Service Co. established a high bar for database protection. The Court ruled that "sweat of the brow" does not merit copyright protection. A compilation is only copyrightable if it possesses originality in its selection or arrangement. Most TMs are comprehensive collections of all translations done for a specific project. There is often no creative selection involved. Consequently, legal scholars argued that the TM structure might not be copyrightable under U.S. law.
The industry bypassed this uncertainty through contract law. Non-disclosure agreements defined TMs as proprietary information. This strategy of using Terms of Service to capture data rights was a direct precursor to the strategies used by social media and AI companies today. The transition from Translation Memories to Machine Translation and eventually LLMs was a fast, gradual evolution. The data amassed in TMs became the fuel for the next generation of automation. TMs were the ideal training data because they were high quality, aligned, and domain-specific.
The reuse of TMs for training MT engines represented a secondary commodification. Translators who had surrendered their TMs to agencies found their work used to train the very machines that threatened to devalue their labor. This reuse represented a fundamental betrayal of the spirit of copyright law. The focus shifted from the expressive value of a text to its informational and statistical value. This conceptual shift laid the intellectual groundwork for the "fair use" arguments seen today.
In fact, the rise of Large Language Models represents the industrial scaling of these dynamics. The LLM era relies on the copyright doctrine of Fair Use to justify the ingestion of the entire available web. Tech companies must argue that their use of copyrighted data is transformative. Several high-profile lawsuits have begun to test these arguments. The New York Times v. OpenAI strikes at the heart of the memorization issue. OpenAI has contended that The Times is not harmed by AI outputs generally but is conflating its claims with broader market harms caused by so-called pink slime organizations.
Now, while the commercial translation industry fights over economic rights, Indigenous communities in the Pacific are fighting a parallel battle over cultural rights and Traditional Knowledge. Their experience is the starkest evidence of the dangers of uncontrolled data extraction. LLMs frequently hallucinate when translating Indigenous languages because they lack a grounded understanding of cultural context. Why? Because there is simply not enough available data to train AI models. In Hawaiian culture, the concept of Kaona refers to hidden or layered meanings in poetry and speech. AI models trained on surface-level text are blind to Kaona. They translate literally and effectively erase the cultural archive encoded within the ʻŌlelo.
Samoan society is structured around the Matai system. The vocabulary changes entirely depending on the rank of the person being addressed. AI models process text as statistical tokens without social awareness and the models frequently fail to master this hierarchy. They may address a high chief with commoner language or mix registers in a way that is socially offensive. Research has documented specific error categories where AI inserts unjustified words or uses an inappropriate language register. For a Samoan user, an AI model that cannot distinguish between a Susuga and an Afioga is an insult.
The most egregious example of AI harm is the hallucination of Whakapapa in Māori contexts. Whakapapa is the foundational principle of Māori identity and genealogy. AI models have been documented fabricating genealogy by assigning the wrong ancestors to tribes. Because the models generate text based on probability rather than fact, they construct plausible-sounding but entirely false connections. And that is a violation of Tapu (Taboo in its original sense).
The conflict between Te Hiku Media and tech giants is a definitive case study in Indigenous resistance and resilience. When OpenAI developed its Whisper model, it claimed to have trained on Māori and Hawaiian data. Te Hiku Media suspected this data was scraped without consent. Traditional agencies attempted to purchase Te Hiku's data for low sums. Te Hiku recognized that selling their data would result in a model owned by a foreign corporation being sold back to Māori people.
In response, Te Hiku developed the Kaitiakitanga License. This license is grounded in the Māori concept of guardianship. It asserts that data is a Taonga (a cognate of the Hawaiian Kaona we mentioned above). The license demands benefit sharing and provenance tracking. This framework challenges the axioms of the TM/Copyright model and rejects the idea that data can be alienated from its creator via a contract.
Well, the struggle is seen across the Pacific Islands. For example, the Tongan language possesses three distinct social registers. AI models trained primarily on commoner text struggle to generate the regal terminology required for formal contexts. In Fiji, the concept of Vaka-Tūraga defines respectful interaction and we have yet seen countless misused renditions of the concept.
So, we see that the TM era established the normative precedent that linguistic labor could be alienated from the laborer, all legally. Once the industry accepted that a translator did not own their TM, it accepted the principle that the value of linguistic data accrues to the aggregator. The LLM era just expanded the scale and scope as generative text models unfolded. The aggregator is no longer a translation agency but a global tech platform that acquired their training data by hard-to-track methods. The laborer is no longer just a professional translator but anyone who has ever written text online.
To see how confusion has spread to courts, in Kadrey v. Meta, the judge called the theory that a model is a derivative work nonsensical without evidence of exact replicas. This distinction between the training process and the model weights suggests that current copyright law may struggle to classify a neural network as a copy. However, in Thomson Reuters v. ROSS Intelligence, a court rejected a fair use defense because the AI created a competing commercial product. This precedent suggests that courts may be less forgiving when AI serves as a direct market substitute.
It is safe to say that Translation Memories were indeed a precursor to the copyright issues facing LLMs today. TMs pioneered the atomization of text and prepared the ground for the statistical tokenization of Neural Machine Translation networks and next-token prediction now used by LLMs. Legally, the TM era normalized the circumvention of creator rights through the Work Made for Hire doctrine while economically, TMs entrenched a model of extraction.
But that's not all: If LLMs are trained on AI-generated content, the quality of the memory degrades. Indigenous examples of Whakapapa hallucination are early warning signs of this epistemic decay. The specialists speak of Cannibalism, for if the future of human knowledge is an LLM trained on the output of other LLMs, the derivative work becomes a copy of a copy, with output quality dropping dramatically until breakdown. The most robust solution may lie in adopting models like the Kaitiakitanga License. If creators retained a fractional stewardship interest in their data tokens, the extractive model would be forced to become a collaborative one. But is the current technology prepared, let alone willing to go this path? Not sure.
Courts in 2025 are deciding whether the logic of the database over the creator will become the constitutional law of the digital age. The evidence indicates that without a fundamental rethinking of data dignity, the dispossession that began with a sentence pair will end with the full automation of culture itself, and subsequent collapse. The question: Does the current legal infrastructure have the requisite elasticity to contain this phenomenon or are we just rearranging deck chairs on a sinking epistemological ship?
Huri Translations
Tel. +689 89 205 483
[email protected]
PO BOX 365 Maharepa
98728 Mo'orea
French Polynesia
N°TAHITI 876649