惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
MongoDB | Blog
MongoDB | Blog
博客园_首页
博客园 - 三生石上(FineUI控件)
博客园 - 聂微东
B
Blog RSS Feed
D
Docker
IT之家
IT之家
大猫的无限游戏
大猫的无限游戏
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
阮一峰的网络日志
阮一峰的网络日志
罗磊的独立博客
Recent Announcements
Recent Announcements
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
A
About on SuperTechFans
The GitHub Blog
The GitHub Blog
G
Google Developers Blog
V
V2EX
量子位
雷峰网
雷峰网
月光博客
月光博客
云风的 BLOG
云风的 BLOG
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
T
Tailwind CSS Blog

Hacker News

GitHub - SeanFDZ/macmind: Single-layer transformer in HyperTalk for the classic Macintosh Show HN: Agent-cache – Multi-tier LLM/tool/session caching for Valkey and Redis Bonsai 1-bit WebGPU - a Hugging Face Space by webml-community Moving a large-scale metrics pipeline from StatsD to OpenTelemetry / Prometheus GitHub - Nightmare-Eclipse/RedSun: The Red Sun vulnerability repository GitHub - SethPyle376/hiraeth: Local AWS emulator focused on fast integration testing, with SQS support, SQLite-backed state, and a debug-friendly web UI. GitHub - macOS26/Agent: Any AI, replaces Claude Code, Cursor, OpenClaw. Over 18 LLM providers (Claude, OpenAI, Gemini, Ollama, Zai, HF, Qwen) wired into a native Mac app that writes code, builds Xcode projects, bumps versions, manages git, automates Safari, use AppleScript, JS or Accessibility, extend Agent! w/ MCP Servers, run tasks from your iPhone via Messages. YouTube now lets you turn off Shorts I Made a Terminal Pager Burgers | マクドナルド公式 Commands — HackerNews CLI documentation ChatGPT for Excel PiCore - Raspberry Pi Port of Tiny Core Linux Live Nation illegally monopolized ticketing market, jury finds Google Broke Its Promise to Me. Now ICE Has My Data. Founding Engineer at Adaptional | Y Combinator CRISPR takes important step toward silencing Down syndrome’s extra chromosome GitHub - saffron-health/libretto: The AI toolkit for building reliable browser automations US v. Heppner (S.D.N.Y. 2026) no attorney-client privilege for AI chats [pdf] Retrofitting JIT Compilers into C Interpreters IPv6 – Google The Accursèd Alphabetical Clock Cybersecurity Looks Like Proof of Work Now Fragments: April 14 Cal.com Goes Closed Source: Why AI Security Is Forcing Our Decision | Cal.com - Scheduling Software for Online Bookings Laravel raised money and now injects ads directly into your agent When moving fast, talking is the first thing to break Too much Discussion of the XOR swap trick – Heather Cafe Introduction to Spherical Harmonics for Graphics Programmers The Grand Line
Meta, Zuckerberg Sued Over Alleged Copyright Infringement...
spankibalt · 2026-05-06 · via Hacker News

In a new legal battle in the AI space, Meta and CEO Mark Zuckerberg have been sued by five publishers and author Scott Turow, who allege the tech company illegally copied millions of books, articles and other works to train Meta’s artificial-intelligence systems.

“In their effort to win the AI ‘arms race’ and build a functional generative AI model, Defendants Meta and Zuckerberg followed their well-known motto: ‘move fast and break things,’” the plaintiffs say in their lawsuit. “They first illegally torrented millions of copyrighted books and journal articles from notorious pirate sites and downloaded unauthorized web scrapes of virtually the entire internet. They then copied those stolen fruits many times over to train Meta’s multibillion-dollar generative AI system called Llama. In doing so, Defendants engaged in one of the most massive infringements of copyrighted materials in history.”

The suit was filed Tuesday (May 5) in the U.S. District Court for the Southern District of New York by five publishers (Hachette, Macmillan, McGraw Hill, Elsevier and Cengage) and Turow individually. The proposed class-action suit seeks unspecific monetary damages for the alleged copyright infringement. A copy of the lawsuit is available at this link.

Asked for comment, a Meta spokesperson said, “AI is powering transformative innovations, productivity and creativity for individuals and companies, and courts have rightly found that training AI on copyrighted material can qualify as fair use. We will fight this lawsuit aggressively.”

Authors have sued AI companies for copyright infringement before — and lost.

For example, in June 2025, a federal judge rejected a claim brought by 13 authors, including Sarah Silverman and Junot Díaz, that Meta violated their copyrights by training its AI model on their books. Judge Vincent Chhabria ruled that Meta had engaged in “fair use” when it used a data set of nearly 200,000 books to train its Llama language model for generative AI.

But the latest lawsuit alleges that Meta and Zuckerberg deliberately circumvented copyright-protection mechanisms — and had considered paying to license the works before abandoning that strategy at “Zuckerberg’s personal instruction.” The suit essentially argues that the conduct described falls outside protections afforded by fair-use provisions of the U.S. copyright code.

“Meta — at Zuckerberg’s direction — copied millions of books, journal articles, and other written works without authorization, including those owned or controlled by Plaintiffs and the Class, and then made additional copies of those works to train Llama,” the suit says. “Zuckerberg himself personally authorized and actively encouraged the infringement. Meta also stripped [copyright management information] from the copyrighted works it stole. It did this to conceal its training sources and facilitate their unauthorized use.”

According to the lawsuit, after the release of Llama 1, Meta briefly considered entering into licensing deals with major publishers. Meta discussed increasing the company’s “dataset licensing” budget to as much as $200 million from January to April 2023, per the complaint.

But then in early April 2023, “Meta abruptly stopped its licensing strategy,” according to the lawsuit. “The question of whether to license or pirate [copyrighted material] moving forward was ‘escalated’ to Zuckerberg. After this escalation to Zuckerberg, Meta’s business development team received verbal instructions to stop licensing efforts. One Meta employee presciently described the rationale: ‘if we license once [sic] single book, we won’t be able to lean into the fair use strategy.'”

According to the lawsuit, Meta and Zuckerberg “are well aware of the market for licensing AI training materials.” Meta signed four licenses in 2022 with African-language book publishers for “a limited training set, and it subsequently reached licensing agreements with major news publishers including Fox News, CNN and USA Today,” the suit says.

On Dec. 13, 2023, Meta employees internally circulated a memo concerning the legal risks of using LibGen, a repository of copyrighted material that the Meta memo described as “a dataset we know to be pirated” and added that “we would not disclose use of Libgen datasets used to train,” per the suit. “Ultimately, however, those concerns went unheeded. Zuckerberg and other Meta executives authorized and directed the torrenting of over 267 TB of pirated material — equivalent to hundreds of millions of publications and many times the size of the entire print collection of the Library of Congress,” according to the lawsuit.

As a result of the alleged infringement, Meta’s AI system “readily generates, at speed and scale, substitutes for Plaintiffs’ and the Class’s works on which it was trained,” the lawsuit states. “Those substitutes take multiple forms, including verbatim and near-verbatim copies, replacement chapters of academic textbooks, summaries and alternative versions of famous novels and journal articles, inferior knockoffs that copy creative elements of original works, and derivative works exclusively reserved to rights holders. Llama even tailors outputs to mimic the expressive elements and creative choices of specific authors.”