惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
G
Google Developers Blog
Hugging Face - Blog
Hugging Face - Blog
博客园 - 【当耐特】
S
SegmentFault 最新的问题
宝玉的分享
宝玉的分享
博客园 - Franky
博客园_首页
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
WordPress大学
WordPress大学
有赞技术团队
有赞技术团队
月光博客
月光博客
博客园 - 聂微东
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
小众软件
小众软件
Microsoft Security Blog
Microsoft Security Blog
Last Week in AI
Last Week in AI
Vercel News
Vercel News
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
爱范儿
爱范儿
J
Java Code Geeks
博客园 - 叶小钗
Engineering at Meta
Engineering at Meta
阮一峰的网络日志
阮一峰的网络日志

The Register - Offbeat: Legal

Noyb cries foul on LinkedIn withholding profile visitor data China makes it illegal to fire humans if AI takes their jobs Cloudera allegedly overlooked US job candidates: DoJ Australia threatens tech companies with 2.25 percent tax China blocks Meta's acquisition of AI outfit Manus Scotland Yard can keep using live facial recognition on Londoners, say judges UK tribunal sends £2B claim accusing Microsoft of overcharging for licensing to trial Yet another ex-ransomware negotiator admits turning rogue after payoff from crimelords Americans behind Nork IT fraud sentenced to 200 months Indian government investigating TCS after police sting French cops free mother and son after crypto kidnapping EFF: California 3D printer bill threatens digital freedoms IBM pays up under Trump administration's diversity blitz OpenAI CEO Sam Altman home attack suspect charged AI vs the cold hard reality of the legal profession Big Tech has not enforced Australia’s social media ban Big Tech has not enforced Australia’s social media ban China's not thrilled AI experts want to leave the country China's not thrilled AI experts want to leave the country JLR cyber bailout risks dangerous precedent, watchdog warns Patel dodges question about FBI buying location data Patel dodges question about FBI buying location data ChatGPT advised exec on firing Subnautica founders: court Japan to allow ‘proactive cyber-defense’ from October 1st FSF urges AI vendors to liberate LLMs Age verification isn't sage verification when it's inside operating systems India tests whether AI can stop trains hitting elephants Perplexity Comet hurtling toward Amazon ban Lenovo, Nintendo sue US government seeking tariff refunds Google embraces third party app stores and payments
Databricks fails to shake authors' copyright claim
O'Ryan Johnson O'Ryan Johnson · 2026-04-30 · via The Register - Offbeat: Legal

Legal

Databricks can't seem to shake authors' copyright claim that could result in 'extraordinary' damages

Authors say it acquired an LLM that was trained on their copyrighted data, and judge keeps asking for more info

Databricks cannot shake a class action lawsuit targeting its LLM, which several book authors contend was created with a database that contained pirated versions of some of their copyrighted books – and about 196,000 titles in all.

Databricks’ motion to dismiss the case was denied last week by Judge Charles Breyer in U.S. District Court in Northern California, who said the plaintiffs, a group of writers that includes bestsellers and a Pulitzer Prize finalist, had grounds to continue their suit against the data analytics platform.

Databricks LLM, called DBRX, was cobbled together with parts from MosaicLM, which Databricks acquired in 2023. Early versions of that model used a database called RedPajama – which contained Book3 and has since been pulled from Hugging Face for copyright infringement. Databricks is essentially arguing that the authors can't prove that DBRX was trained with the Book3 data, and has testified to that effect.

Databricks closed its acquisition of MosaicLM in July 2023. In a statement at the time, Databricks called Mosaic “a leading generative AI platform known for its state-of-the-art MPT large language models.” MosaicLM released its first MPT model in May 2023 and in a blog announced it had used the RedPajama dataset in training.

Then when Databricks released its DBRX model in March 2024, it said “The development of DBRX was led by the Mosaic team that previously built the MPT model family.” The case hinges on how closely those two steps were tied.

Speaking of the authors, Judge Breyer wrote in his ruling, “They directly tie their infringed works to DBRX, and the employee statements provide supporting inferences when read in context, particularly when viewed alongside other more direct statements."

While Databricks has provided fourteen depositions, thousands of pages of documents, and terabytes of discovery information in its bid to show the court it did nothing wrong, Breyer wants to see more, said Brandon Butler, a copyright lawyer and executive director of Re:Create, a coalition of groups that advocates for balanced copyright laws.

“Judge Breyer basically says, ‘We need to know more before we can say that you didn't actually engage in any infringing copying,’ ” Butler told The Register. “We don't know enough yet, about what happened. Step by step, what did they physically do?”

Butler said potential damages against Databricks are massive if the authors can convince the court that the infringements were willful.

“The damages provisions in copyright law are draconian with a capital D. I mean, they are extraordinary. They are six figures per work infringed up to $150,000,” he said. “This is bet-the-company litigation. If they win, they could get enough damages they just liquidate every asset that belongs to some of these companies, and probably especially a smaller player like Databricks.”

So far several authors have joined the suit, among them young adult best selling author Jason Reynolds, Stuart O’Nan, Brian Keene, and Rebeccas Makkai, whose book The Great Believers was a finalist for the Pulitzer Prize.

Meta won a similar lawsuit last year against book authors who sued for copyright infringement during the creation of its LLAMA models by arguing that its actions were covered by fair use provisions of copyright law. Anthropic also won on a similar fair use claim in a separate case (but had ingested pirated books and agreed to establish a $1.5 billion fund to compensate authors.)

But Databricks has not yet made that argument.

Instead, Databricks' unsuccessful motion said the authors’ complaint was “nonsensical” and encompass actions that predate the training of DBRX.

“By Plaintiffs' strained logic, if a car company experimented on emissions technology with and without a patented component, and later manufactured a car without that component, the patent owner could still assert infringement claims as to the non-infringing car based solely on the earlier experimentation that led to the decision not to include the component,” lawyers for Databricks wrote.

The authors argue they only need to show the court that their works were copyrighted and that those works were then copied by Databricks.

“Databricks copied Books3 multiple times in the process of developing its DBRX models and by so doing, directly infringed Plaintiffs’ copyrights in the asserted works,” the authors who brought the suit stated. “Under Defendants’ logic, as long as an AI company does not incorporate copyrighted books into the final training dataset of a model, it is free to download, store, reproduce, and indefinitely use pirated works for its own benefit. That argument gets it backwards.”

Butler said there are a couple of paths Databricks could take to succeed. First they could argue fair use, which has been a winning argument in the same federal court that is hearing this case. The second is that they could claim the authors cannot show damages and thus have no claim to file suit.

“That may be an argument that would be useful here, which is to say, ‘Whatever happened with all those books back then, none of that ever saw the light of day. It had no impact on our model. It was a mistake, and we undid it, and it had literally no impact in the world. So, why are we here? Why are we wasting the court's time? But I think that's a thing they have to prove, and they haven't proven it yet,” he said. ®