惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Security Latest
Security Latest
Apple Machine Learning Research
Apple Machine Learning Research
D
Docker
美团技术团队
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
宝玉的分享
宝玉的分享
月光博客
月光博客
J
Java Code Geeks
V
V2EX
IT之家
IT之家
T
Troy Hunt's Blog
D
DataBreaches.Net
Cloudbric
Cloudbric
Blog — PlanetScale
Blog — PlanetScale
H
Hackread – Cybersecurity News, Data Breaches, AI and More
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
G
Google Developers Blog
MongoDB | Blog
MongoDB | Blog
The GitHub Blog
The GitHub Blog
Jina AI
Jina AI
Google DeepMind News
Google DeepMind News
Engineering at Meta
Engineering at Meta
博客园 - Franky
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
H
Help Net Security
Application and Cybersecurity Blog
Application and Cybersecurity Blog
S
Security @ Cisco Blogs
N
News and Events Feed by Topic
aimingoo的专栏
aimingoo的专栏
S
Security Affairs
Hugging Face - Blog
Hugging Face - Blog
Forbes - Security
Forbes - Security
AI
AI
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
腾讯CDC
H
Heimdal Security Blog
The Cloudflare Blog
S
SegmentFault 最新的问题
Google Online Security Blog
Google Online Security Blog
Webroot Blog
Webroot Blog
有赞技术团队
有赞技术团队
The Hacker News
The Hacker News
Microsoft Security Blog
Microsoft Security Blog
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
罗磊的独立博客
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
博客园 - 聂微东
Help Net Security
Help Net Security
T
The Exploit Database - CXSecurity.com

IEEE Spectrum

NASA Puts Google’s Gemma Large Language Model in Orbit Why AI Needs a "Genie Coefficient" Why Some Coders Now Reach for GLM 5.2 Before Frontier AI Models How a Spinning Drone Exploits Your Eyes to Become Nearly Invisible Why Indonesia’s Fisheries Future Hinges On Data Integrity and Trust Inside ELIZA’s Source Code and Its Multiple Personalities AI Turns DNA Into Tiny Dogs and Mona Lisa Nanostructures How Darth Vader Taught Me Card Counting and AI Security Got Weird The AI Arms Race in Technical Interviews Is Escalating Inside X Square Robot’s Bold Plan for Real Home Robots Large Tabular Models Excel Where LLMs Fail The Hidden Overthinking Flaw That Could Drag AI Services Down There Why Small AI Models Could Power Health Care Where Big Tech Cannot AI’s Wild Power Demands Are Quietly Rewriting Grid Rules Is Melbourne the Place Where AI and Clean Energy Finally Align? The Orbital Data Center Hype Machine Is Already in Orbit What Emily Bender Really Meant by "Stochastic Parrots" Poetry for Engineers: Nine Lives of Nikola Tesla How a Forgotten Wire Turned a Cheap Chip Into a Brainlike Neuron AI Model ConlangCrafter Dreams up Entire New Languages Why Does a Bank Need a Chief Scientist? What it Means to Be a Mathematician When AI Does the Math AI Learns the "Dark Art" of RF Chip Design Can AI Learn to Read the Room? Commemorating 70 Years of Artificial Intelligence IEEE Rolls Out Large Language Models Virtual Training Course Can Sound-Driven Synapses Make AI Both Faster and Greener? How AI Attribution Could Finally Pay Musicians for Training Data Inside GM’s AI Push to Speed Up the Design of Cars and Moon Rovers Are Emotion Reading Robots Still Missing What Matters Most? The Google DeepMind Spinoff Chasing Hidden Drug Targets Save 14 Percent of Energy Used in LLM Training With This Trick AI Can Help Track the World’s Shrinking Glaciers Nvidia’s AI Hardware Comes to Windows in RTX Spark PCs Why Quantum Computers Need a ‘Healthy Chunk’ Of Classical Power How Young Engineers Can Turn AI Into Career Leverage Why Aren’t We Measuring How AI Affects Humans? Majestic’s 128TB AI Server Aims to Smash the LLM Memory Wall Finding Success in Industry as a Chip Designer Why South Africa’s AI Policy Leverage Is Slipping Away Unused AI and Thermal Cameras Help Ships Steer Clear Of Gray Whales Why Reclaiming ‘Social Engineering’ Could Protect Your Autonomy AI with Model-Based Design: Virtual Sensor Modeling - Wiley Science and Engineering Content Hub Millimeter Waves Turn Tiny Insects Into Trackable Data Open-Source AI Could Make It Easier to Build Smart Robots The Future of Physical AI Isn’t Smarter Robots, It’s Smarter Interfaces Agentic AI for Robot Teams How Melbourne’s AI and Data Center Flywheel Is Accelerating Research Innovation Hidden Voice Glitches Could Hijack Audio AI Tools AI Rings Turn Sign Language Into Text In Real Time Graphene Leaf Tattoos Turn Plants Into Living Moisture Meters Accelerating Chipmaking Innovation for the Energy-Efficient AI Era Can AI Chatbots Reason Like Doctors? General AI Outruns Specialized Tools at Transcribing Handwriting Neutralizing the Gigascale Problem: How to Solve the Physical Power Paradox of Extreme AI Training Loads Tiny Data Centers at Substations Aim to Keep AI Power Usage In Check Orbital Bets On a Mesh Of GPU Satellites for AI Inference Can AI Really Build Better AI? AI Chatbot Safety Guardrails for Mental Health Ten Key Enablers for 6G Wireless Communications - Wiley Science and Engineering Content Hub
Māori AI Voice Puts Language Ownership Back In Community Hands
https://www.facebook.com/48576411181 · 2026-05-21 · via IEEE Spectrum

New Zealand is a country famed for its dramatic landscapes, but its linguistic landscape is arguably just as interesting. Of its three official languages, only te reo Māori (the Māori language) could be described as indigenous. Though spoken fluently by just 4.3 percent of the population, national statistics show that about 30 percent of New Zealanders can speak more than a few words or phrases of the language.

But ask ChatGPT to write te reo Māori and it will oblige, fluently answering your questions in the standardized form of the language taught in schools and broadcast on national television. Claude and Perplexity can do the same. This impressive language performance is built on text and audio produced by Māori communities and academics, which was scraped and ingested without their permission, processed outside New Zealand, and returned to users through interfaces owned by large technology companies. For Māori, that is a problem.

“These companies overseas have the resources to produce AI models that work well,” says Te Taka Keegan, a professor at the University of Waikato and codirector of its Artificial Intelligence Institute. “But they scraped all of that data with no input from us, and we don’t own the output. Our language is the most important conveyor we have for our knowledge.…yet we see technology developed outside of Aotearoa [New Zealand] get more and more control over the transfer of that knowledge.”

Motivated by this need for “sovereign digital systems,” as Keegan calls it, he and Kingsley Eng, Keegan’s master’s student at the time, set out to develop a high-fidelity synthetic voice—a text-to-speech system, in other words—for a specific dialect of te reo Māori. Every technical decision Keegan and Eng made along the way was shaped by a foundational constraint typically ignored by the AI sector—that this synthetic voice, and everything used to build it, must remain owned by the people who speak that dialect. What they produced, they hope, offers a replicable blueprint for other minority language communities around the world.

Challenges in Māori AI Voice Models

AI voice models are predominantly built in English, so applying those models to other languages can lead to errors. Te reo Māori has some specific linguistic features, such as the importance of vowel length, that lead to additional challenges for AI voice systems.

As an example, the words for “cake” (keke), “armpit” (kēkē) and “to creak” (kekē) differ only by how long the vowel sounds are. Digraphs—two letters making one sound— are also common, and are pronounced differently than they are in English; “wh” is usually pronounced “f.” In the Māori language, inaccurate pronunciation changes the meanings of words.

In addition, te reo Māori is considered a low-resource language, because, compared with a language like English or Chinese, there’s relatively little potential training data in the form of text, datasets, or recorded speech available in digital formats. To address this problem, Keegan recruited Ngaringi Katipa—a translator, educator, and language mentor—to be the consenting human voice behind the tool.

“Our language is the most important conveyor we have for our knowledge.…yet we see technology developed outside of Aotearoa get more and more control over the transfer of that knowledge.” —Te Taka Keegan, University of Waikato

“We focused on our local dialect, Waikato-Maniapoto, because it’s in the dialects that you see the real beauty of language. They tie it to a specific place and sense of identity,” says Keegan.

“We initially just recorded Ngaringi reading passages from books, which gave us 4.5 hours of data,” says Eng, now a machine learning engineer at the precision toolmaker Extec. “Later, we expanded the dataset by recording from a comprehensive list of sentences and words—including very rare words—given to us by Te Taka’s brother Peter, who is a Māori linguistics expert.” Once cleaned and processed, the final tally was 7 hours and 45 minutes of recordings.

A Māori Text-to-Speech AI Model

Building a text-to-speech system generally takes one of two approaches to data input. The first is character-based, where raw letters are passed directly to the model. The second is phoneme-based, where text is first converted into a phonetic representation, or a description of how each word sounds, before training begins.

“We tried both, but the phoneme approach was far better,” says Eng. “Giving the model phoneme rules off the bat was like a head start.” Phonemes effectively tell the model what certain groups of letters sound like, “which lets you skip some of the learning,” he says. To provide the model with phoneme rules, the researchers used an open-source tool called eSpeak NG, which includes a beta Māori rule set that they adapted further.

Eng tested three open-source neural architectures—Matcha-TTS, Tacotron2, and Piper—to train and transform the recordings into a synthetic voice. Piper, which can run offline on a local machine, had the best results and was chosen for the final build.

Despite using under eight hours of good quality recordings—considerably less than the hundreds of hours typically suggested for training a text-to-speech model—the final AI voice was effective. The primary metric used in text-to-speech research is word error rate, in which a lower percentage indicates higher accuracy. Keegan and Eng’s AI voice achieved an error rate of 6.78 percent, considered “good” by current industry standards.

Throughout the development process, a professional Māori language evaluator assessed the voice, rating it in terms of its naturalness, pronunciation accuracy, and expressiveness.

The researchers also invited 68 fluent speakers of te reo Māori to listen to both human and synthesized audio, and asked them to identify which was which. The listeners correctly identified the voices 65 percent of the time. “We were happy with that because some of the listeners were family members of the speaker—they know her voice really well, but a few still got it wrong,” says Keegan.

While Google provided some funding to the Waikato team, Keegan says it came with no conditions attached and no ownership stake claimed. “They said, we’ve heard about your work with preserving languages, and we wanted to support you. Use the grant whatever way you want.” Ultimately, he says, it allowed them to fairly compensate Katipa for her work.

With the tool now ready for use, the question of ownership remains front of mind for Keegan. From a standard intellectual property perspective, the voice belongs to Katipa. From a Māori perspective, Keegan says, it belongs to the collective: “It’s a treasure that’s been handed down through her ancestors; and her role is to protect it for her children and her grandchildren.”

So rather than release the voice model publicly, Keegan is in discussion with the three iwi (tribes) that Katipa affiliates with—Waikato, Maniapoto, and Raukawa. “Guardianship of this needs to sit with them,” Keegan says, “rather than the university.”

To that end, Keegan found a Wellington-based company, Catalyst IT, that gifted website hosting and the computing power needed to run the voice model for a year.

Data sovereignty is a rapidly growing focus in indigenous AI communities. Te Hiku Media, a Māori media organization in New Zealand’s far north, developed an automatic speech-recognition system that achieves 92 percent accuracy for te reo Māori and 82 percent accuracy for bilingual speech. The organization released the model under a Kaitiakitanga license—a legal instrument stipulating that data can only be used for the benefit of the Māori people.

Elsewhere in the world, the Aina project at the Barcelona Supercomputing Center released Matxa, a multidialect Catalan text-to-speech system also built on open-source architectures. In Quebec, Michael Running Wolf leads the First Languages AI Reality (FLAIR) initiative, which is working to build speech-recognition models for Indigenous languages across North America.

Voice-driven technologies, such as virtual assistants, screen readers, navigation systems, and smart devices, are ubiquitous. For Keegan, these tools can either be a way to “sanitize and colonize our language” or a means to “empower my moko [grandchildren] with their traditional knowledge.” The difference, he says, comes from who develops and owns the technology. “I want my grandchildren and my great-grandchildren to access our knowledge through our own systems. This voice is the first step in achieving that.”

Longer term, his ambition is to use the same open-source, community-owned methodology to build full language models. “It won’t be a te reo Māori large language model,” he says. “It’ll be a Maniapoto large language model, a Tūhoe large language model, et cetera.” Each model would be owned by, and trained on the speech of, the people whose language it speaks.

While that’s a more significant engineering challenge than a text-to-speech system, the Waikato project demonstrates that the necessary infrastructure already exists—efficient training on minimal data, phoneme-based input, open-source tools, and a legal and governance framework for community ownership. “We’ve laid a template so that other iwi throughout the country can do the same thing,” says Keegan. “I am happy to help them do it.”

This story was updated on 21 May, 2026 to correct Te Taka Keegan’s position: He is a professor at the University of Waikato, not an associate professor.