header
  • leqo
  • How the Courts Are Approaching LLM training with Copyrighted Works

How the Courts Are Approaching LLM training with Copyrighted Works

For those of us dedicated to the precious linguistic heritage of the Pacific Islands, our work in translation now brings us face-to-face with important legal discussions about how language models are taught, particularly concerning the rights of creators. Recent decisions from the courts exposes how US judges are grappling with concerns raised about the training of these powerful LLMs.

Description of image

"outline the legal framework for addressing copyright claims related to the training of LLMs"

The cases are laying down principles that could significantly influence how our concepts of intellectual property interact with the rapidly advancing field of Artificial Intelligence in the years ahead.
Several recent legal cases illustrate the thoughtful way the courts are trying to find a balance between fostering technological progress and safeguarding the rights of those who create. In January 2025, Judge Eumi K. Lee of the Northern District of California chose not to grant an immediate halt to Anthropic's use of song lyrics to train its AI, Claude, in the case of Concord Music Group, Inc. v. Anthropic PBC – hundreds of thousands – raised practical concerns for the court about how such an order could be effectively implemented and monitored.

Crucially, Judge Lee drew a distinction between using copyrighted material as input to train the AI, and the AI then generating output that essentially copies copyrighted works. While the parties had found common ground on the issue of output, the fundamental question of whether using copyrighted material for training qualifies as "fair use" – a legally recognized exception to copyright – remains an open and complex issue.

The case of Tremblay v. OpenAI sheds light on how the courts are managing the process of each side gathering information in copyright disputes involving AI. In January 2025, Magistrate Judge Robert M. Illman established boundaries for this "discovery" process, limiting the amount of time each side could spend on formal questioning. He also directed OpenAI to provide training data for models they are currently developing, including one known as "Orion".

Judge Illman's order appears to reflect an effort to be equitable: Making sure that the individuals bringing the lawsuit have access to the information they need, without placing an undue burden on the AI company. He declined the plaintiffs' requests to delve into OpenAI's broader business strategies, suggesting that the financial records and strategic documents already provided were sufficient for the purposes of the case.

Similarly, in Kadrey v. Meta Platforms, Magistrate Judge Thomas S. Hixson examined Meta's responses to direct requests for admission of certain facts. He found Meta's answers to be insufficiently clear, particularly where they acknowledged that "some" of their training data included text from "some" copyrighted books. Judge Hixson emphasized that Meta needed to offer a more precise indication of the extent to which they were admitting or denying these facts.

However, the court did not grant the plaintiffs' request to review documents under the argument that Meta might have engaged in unlawful activity to obtain training data. Judge Hixson expressed concern that such a request was essentially asking him to make a judgment on the merits of the case prematurely, before a full trial.

The protective order issued in Nazemian v. NVIDIA illustrates how courts are working to safeguard confidential information in AI-related litigation. Judge Jon S. Tigar established different levels of confidentiality for various types of information, including "CONFIDENTIAL," "HIGHLY CONFIDENTIAL — ATTORNEYS' EYES ONLY," and "HIGHLY CONFIDENTIAL — SOURCE CODE." Notably, "Training Data" was defined broadly as any information used to develop and refine AI models.

This protective order included specific rules for inspecting the AI's underlying code, requiring it to be done on a secure computer in a protected environment and without internet access. It also put in place a restriction preventing anyone who accessed highly confidential information from being involved in the patenting of related AI technologies for a period after the case concludes.

Such orders acknowledge the competitive nature of AI development while still allowing those bringing claims to access the evidence necessary to pursue their cases. The detailed provisions for examining the source code recognize that while the use of others' copyrighted material is being litigated, the AI companies' own intellectual property also deserves protection.

Across these various legal proceedings, the courts have not yet delivered a definitive answer on whether the use of copyrighted materials to train LLMs constitutes fair use. As Judge Lee observed in the Concord v. Anthropic case, the music publishers were essentially asking the court to define the rules of a licensing market for AI training, even though the fundamental question of whether this training is legally permissible as fair use remains unresolved.

This central question also persists in the Kadrey v. Meta case, where Meta has asserted that they did not require licenses to use the training data because their use was fair. While courts have directed AI companies to disclose information about their training processes, a final determination on the legality of this use of copyrighted material is still pending.

It appears that the courts are making a distinction between the copyrighted material used to train the LLM (the input) and the content that the LLM subsequently generates (the output). In the Concord v. Anthropic case, the parties reached an agreement regarding the output, with Anthropic committing to maintain existing safeguards to prevent their AI from reproducing copyrighted lyrics.

Similarly, in Tremblay v. OpenAI, the court ordered OpenAI to produce documents related to their efforts to prevent the LLM from simply regurgitating the training material. This emphasis on preventing the models from directly copying the training data suggests that the courts may view addressing output-based copyright concerns as more straightforward than resolving the complexities of input-based claims.

These judicial decisions are beginning to outline the legal framework for addressing copyright claims related to the training of LLMs. While a conclusive ruling on fair use is still awaited, the courts are establishing procedures for managing the exchange of information and protecting sensitive data. Judges are proceeding cautiously, seemingly reluctant to issue broad orders that could hinder technological progress before the underlying legal principles are more clearly established.

For those of us working with the languages of the Pacific Islands, these developments raise important considerations about how copyright law will shape the creation of language models for languages with smaller amounts of available text. Our organization also uses specialized AI models trained on our own data for these languages. However, the ultimate legal outcomes of these cases could influence how we are able to expand and refine these vital tools. If the courts ultimately determine that training LLMs on copyrighted material is fair use, it could accelerate the development of models for less-resourced languages like Hawaiian, Samoan, CHamorro, Palauan, Marshallese, Chuukese, Yapese, Pohnpeian, Kosraean and Carolinian. Conversely, if licensing becomes a requirement, it might create additional obstacles for languages with fewer commercial resources.

The approach taken by the US judiciary thus far indicates a thoughtful and deliberate pace, allowing legal understanding to evolve alongside the technology itself. As Judge Lee wisely noted, emerging technologies often push the boundaries of existing copyright law. How these boundaries are ultimately defined will have lasting implications not only for major AI corporations but also for specialized language service providers like Huri Translations.
contact

For you:

Pacific Island languages
Domains of Expertise
Leqo

▤ View Complete Article Archive

Cities

  • Las Vegas, NV
  • San Diego, CA
  • West Valley City, UT
  • Sacramento, CA
  • Tokyo
  • Long Beach, CA
  • Euless, TX
  • Nukuʻalofa
  • Santa Ana, CA
  • Portland, OR
  • Little Rock, AR
  • Nouméa
  • Los Angeles CA
  • Tarawa
  • Papeʻete
  • Tacoma, WA
  • Reno, NV
  • Port Moresby PNG
  • San Francisco, CA
  • Oakland, CA
  • Seattle, WA
  • Pago Pago
  • Sydney
  • Independence, MO
  • Salem, OR
  • Killeen, TX
  • Vancouver
  • Suva
  • Honiara
  • Salt Lake City, UT
  • Palikir
  • Houston, TX
  • Anaheim, CA
  • London
  • Hagåtña
  • Manila
  • Santiago
  • Apia
  • Phoenix, AZ
  • Anchorage, AK
  • Port Vila
  • Springdale, AR
  • Oceanside, CA
  • Auckland
  • Paris
  • Mesa, AZ
  • Majuro
  • Honolulu
  • Dallas, TX
  • San Antonio, TX
  • Brisbane
  • Dubai

Domains of expertise

  • Our Insurance Expertise
  • Insurance Products
  • Claims Processing
  • Our Blue Economy Expertise
  • Fisheries Management
  • Marine Conservation
  • Our Aviation Expertise
  • Flight Operations
  • Aircraft Maintenance
  • Our Cultural Expertise
  • Music & Performance
  • Dance & Traditional Arts
  • Our Educational Expertise
  • Academic Institutions
  • Educational Materials
  • Our Energy & Climate Expertise
  • Renewable Energy
  • Power Systems & Grids
  • Our Business Expertise
  • International Trade
  • Corporate Communications
  • Our Space Technology Expertise
  • Satellite Communications
  • Earth Observation
  • Our Agricultural Expertise
  • Agricultural Development
  • Fishing & Marine Resources
  • Precision in Official Communication
  • Legal Documentation
  • Government Communications
  • Our Mining & Resources Expertise
  • Mining Operations
  • Environmental Impact & Compliance
  • Our Maritime Expertise
  • Naval Architecture & Shipbuilding
  • Maritime Operations
  • Our Hospitality Expertise
  • Accommodation Services
  • Travel Services
  • Our Healthcare Expertise
  • Medical Documentation
  • Public Health
  • Our Emergency Management Expertise
  • Disaster Preparedness
  • Emergency Response
  • Empowering Pacific Voices
  • Indigenous Rights Frameworks
  • Land & Resource Rights
  • Our Technology Expertise
  • Technical Documentation
  • Telecoms
  • Our Expertise
  • Content Creation & Production
  • Distribution & Publishing

Our languages

  • Tahitian Translation Services
  • Tokelauan Translation Services
  • Pohnpeian Translation Services
  • Chuukese Localization Services
  • Hawaiʻi Creole English Localization Services
  • Reo Māori Localization Services
  • Carolinian Translation Services
  • Refaluwasch Localization Services
  • Wallisian Translation Services
  • Dorerin Naoero Localization Services
  • CHamorro Localization Services
  • Solomon Islands Pijin Translation Services
  • Gagana Sāmoa Localization Services
  • Na Vosa Vaka Viti Localization Services
  • Samoan Translation Services
  • ʻŌlelo Hawaiʻi Localization Services
  • Hawaiian Pidgin Translation Services
  • Faka ʻUvea Localization Services
  • Yapese Translation Services
  • Vagahau Niuē Localization Services
  • Gana Tuvalu Localization Services
  • Tok Pisin Localization Services
  • Reo Rarotonga Localization Services
  • Fijian Hindi Localization Services
  • Nauruan Translation Services
  • Marshallese Translation Services
  • Rarotongan Translation Services
  • Kusaie Localization Services
  • Bislama Localization Services
  • Palauan Translation Services
  • Fakafutuna Localization Services
  • Gagana Tokelau Localization Services
  • Gilbertese Translation Services
  • Rapanui Translation Services
  • Kajin Ṃajeḷ Localization Services
  • Yapese Localization Services
  • Reo Tahiti Localization Services
  • Lea Fakatonga Localization Services
  • Reko Pakumotu Localization Services
  • Māori Translation Services
  • Solomon Islands Pijin Localization Services
  • Bislama Translation Services
  • Fijian Translation Services
  • Futunan Translation Services
  • Tongan Translation Services
  • Niuean Translation Services
  • Kosraean Translation Services
  • Taetae Ni Kiribati Localization Services
  • Reʻo Rapa Nui Localization Services
  • Chuukese Translation Services
  • Tuvaluan Translation Services
  • Belau Localization Services
  • Fiji Hindi Translation Services
  • Tuamotuan Translation Services
  • Hawaiian Translation Services
  • Pohnpei Localization Services
  • CHamorro Translation Services
  • Tok Pisin Translation Services
View All

Consult now

First name

Last Name

E-mail

Message

  • Terms of use
  • Cookie Policy
  • Privacy Policy
  • Consent Policy
  • Article Index

Huri Translations
Tel. +689 89 205 483
[email protected]
PO BOX 365 Maharepa
98728 Mo'orea
French Polynesia
N°TAHITI 876649

Secure Site Seal

Social