惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
爱范儿
爱范儿
博客园 - 三生石上(FineUI控件)
Vercel News
Vercel News
M
MIT News - Artificial intelligence
L
LangChain Blog
大猫的无限游戏
大猫的无限游戏
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
Microsoft Azure Blog
Microsoft Azure Blog
J
Java Code Geeks
Recent Announcements
Recent Announcements
Stack Overflow Blog
Stack Overflow Blog
人人都是产品经理
人人都是产品经理
IT之家
IT之家
F
Fortinet All Blogs
博客园 - 聂微东
U
Unit 42
Martin Fowler
Martin Fowler
腾讯CDC
博客园_首页
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
量子位
阮一峰的网络日志
阮一峰的网络日志
博客园 - Franky

OpenSource.net

Building MCP servers the easy way with Apache OpenServerless – OpenSource.net Orchestrating AI in clinical settings with jBPM – OpenSource.net jBPM as an AI orchestration platform – Part 2 – OpenSource.net jBPM as AI Orchestration Platform – Part 1 – OpenSource.net A beginner’s guide to Open Source contribution – OpenSource.net Apache OpenServerless is the easiest way to build your cloud native AI application – OpenSource.net Getting started with Open Source – OpenSource.net A capitalistic value engine – OpenSource.net Open Source and your tech stack – OpenSource.net Welcome to a new era at OpenSource.net – OpenSource.net TLS and networking – OpenSource.net Security and cryptography algorithms: A guide – OpenSource.net Essential Python web security – OpenSource.net Accelerating environmental Open Source – OpenSource.net
The copyright conundrum in artificial intelligence age – ...
opensource.net · 2025-01-10 · via OpenSource.net

Generative AI pits innovation against intellectual property, but practical solutions remain elusive.

The rise of generative AI tools has reignited longstanding debates about copyright law, ownership, and innovation. In a recent podcast, Pamela Samuelson, Richard M. Sherman Distinguished Professor of Law at UC Berkeley, delved into the intricate challenges posed by AI systems to existing intellectual property regimes. Samuelson, a pioneer in digital copyright and co-founder of the Authors Alliance, laid bare the practical difficulties facing regulators, creators, and AI developers alike.

At the heart of the issue lies the question of data provenance and transparency. Generative AI models are typically trained on vast datasets, often comprising billions of works scraped from the internet. Many policymakers, particularly in Europe under the proposed AI Act, are pushing for mandatory disclosure of copyrighted works used in training datasets. Yet, as Samuelson argues, such measures assume an overly simplistic view of the AI landscape.

The problem of scale and feasibility

AI training datasets are colossal, often incorporating publicly available internet data. Major corporations like Google and Meta may comply with stringent transparency rules, but Samuelson highlights that AI development extends far beyond Silicon Valley giants. Small startups, non-profits, and even independent researchers depend on Open Source datasets, such as Common Crawl, to build their models. Requiring them to retain and disclose precise records of every data source is impractical and stifles competition and innovation.

Furthermore, the training process itself complicates matters. AI models do not reproduce copyrighted works; they tokenize data into abstract numerical representations—akin to disassembling a LEGO battleship and using the bricks to build an Eiffel Tower. The input data ceases to exist in a recognizable form, rendering claims of direct copyright infringement tenuous at best. As Samuelson explains, generative AI “learns” patterns from datasets rather than replicating the underlying content, drawing parallels to Renaissance artists studying hands to improve their craft.

man selling fruits
Photo by Tim Mossholder on Unsplash

Licensing: A difficult sell

Collective licensing has been touted as a potential solution to compensate authors whose works are used in AI training. Europe, with its robust history of collective licensing for music and publishing, views this mechanism as viable. However, Samuelson outlines why this approach falters in the AI context.

The sheer volume of data—billions of works, many with negligible commercial value—makes calibrating payments nearly impossible. Imagine a collecting society attempting to distribute fractions of cents to millions of authors; the administrative costs would likely outweigh the actual payouts. Additionally, licensing regimes presuppose a clear distinction between inputs and outputs, but AI models often discard the training datasets after learning, further complicating compensation claims.

More fundamentally, mandating licenses for internet-crawled data risks setting a dangerous precedent. For years, web crawling has operated within legal boundaries, underpinning innovations like search engines. A sudden shift to mandatory licensing could retroactively criminalize commonplace practices, creating uncertainty for developers and chilling technological progress.

The Question of authorship

On the output side, AI-generated works raise questions about authorship. Can AI be recognized as the author of a creative work? Samuelson unequivocally dismisses this notion. U.S. copyright law, she explains, requires human creativity as a prerequisite for protection—a principle reaffirmed by the Supreme Court. However, she acknowledges edge cases: if a human provides detailed prompts and iteratively refines an AI-generated work, the resulting output might meet the threshold for authorship.

This distinction is particularly salient for industries like film and music, where computer-generated content has long coexisted with human creativity. Hollywood studios, for instance, leverage CGI to enhance visual storytelling, but still claim copyright over the final product. As Samuelson notes, rigid policies that disqualify AI-assisted works risk undermining industries that have seamlessly integrated technology into the creative process.

person on balance board
Photo by Gustavo Torres on Unsplash

Towards a balanced framework

Samuelson’s insights underscore the need for nuanced, practical regulations that reflect the realities of AI development. While transparency and compensation are valid concerns, solutions must balance the interests of creators with the imperative to foster innovation. Excessive regulation risks entrenching incumbents and marginalizing new entrants, stifling the very competition that drives technological advancement.

Europe’s AI Act may offer a glimpse of what’s to come: a blanket requirement for transparency without imposing crippling compliance burdens. Yet, as Samuelson cautions, policymakers must resist the temptation to anthropomorphize AI or impose solutions better suited to traditional industries.

Generative AI represents a transformative leap in technology—a tool that, much like the printing press or photography, will reshape creative industries. Rather than seeing AI as a threat, Samuelson advocates for recognizing its potential to empower human creators. The task for regulators, then, is to craft policies that encourage innovation while ensuring creators are fairly valued in this new digital era.

Catch the whole podcast episode or check out the transcript.