










Le Minh, an audiobook narrator, says he occasionally receives offers to work as an “AI voice trainer” for a few hundred thousand to more than one million Vietnamese dong per hour (VND100,000 = US$3.80).
The 28-year-old Hanoian, who has been in the industry for years, has recorded audiobooks, commercials, and narration videos, but has over the last year been getting more and more work as an AI voice trainer. The pay is high though opportunities are irregular.
Ha Khanh, a linguistics student in Hanoi, works regularly as a contributor for an audiobook application for VND5,000-7,000 ($0.19-$0.27) per minute of narration.
Since last year, the platform has hired her for a project to build voice datasets for AI for higher wages. Instead of reading entire books as she did previously, she now records short scripted texts and uploads the audio files into the system.
For the student living by herself, the income helps cover some of her rent and other expenses. Job advertisements for AI voice trainers have become increasingly common on voice-over and audio-recording forums.
People have to read scripts and record it to create training data for AI systems, and compensation is typically based on hours worked, the number of sentences recorded, or the volume of completed data. Recruiters often seek speakers from specific regions, genders, or age groups.
![]() |
|
Voice data being recorded using a mobile phone. Photo by Trong Dat |
To enable an AI system to read text as naturally as a human, developers must collect large amounts of speech data from real people. For commercial AI voice systems, these datasets can amount to dozens or even hundreds of hours of recordings.
Speaking to VnExpress, Ho Minh Duc, CEO of Vbee, a company specializing in text-to-speech services, says its datasets are built from a variety of sources, including hired contributors, MCs, and professional voice actors.
Some data also comes from publicly available copyrighted sources on the internet, he says. "We look for people with suitable voices. At the same time, some individuals approach our platform because they want to digitize their voices."
A similar process is taking place in academia. While developing Vietnamese-language learning software for foreigners, a research team led by Vu Van Thuong, a postgraduate from the Posts and Telecommunications Institute of Technology, hired teachers to help build a speech database.
Initially, many participants volunteered their time. As the workload increased, the team began paying contributors. They read individual words and sentences displayed on a screen, which were then recorded, processed, and converted into data that AI systems could use to learn Vietnamese pronunciation.
According to Duc, there are currently two commercial models for monetizing voice data in Vietnam. Under the traditional model, audiobook narrators, commercial voice actors, and automated call-center speakers are paid for each recording project, typically by the minute or by assignment.
The second model has emerged alongside advances in AI. Rather than paying solely for specific recordings, companies compensate individuals for digitizing their voices to create AI-generated voice models that can be used for multiple purposes.
This represents a shift from hiring voice talent for individual recordings to leveraging voices through AI technology.
Copyright and ownership of AI voices
This shift has sparked new debates over voice ownership rights. There have been cases where people claimed that an AI-generated voice sounded remarkably like their own, Duc says.
In such situations, developers must trace the origins of the data used to train the system, he says. "It is necessary to determine where the data came from and whether it was used legally. If inappropriate use is identified, the data must be removed and discussions held with the rights holder to resolve copyright issues."
Duc says voice biometrics, the technology capable of identifying individuals through their voices, similar to fingerprint or iris recognition, already exists, but determining whether an AI system has copied an individual’s voice remains difficult to identify in Vietnam due to a lack of unified standards.
![]() |
|
A person using an AI tool to convert text into speech. Photo by Trong Dat |
From a research perspective, Thuong says many AI systems can reproduce the vocal characteristics and intonation of a specific individual using a recording as short as 15 seconds. For this reason, people contributing voice data should be clearly informed about how and where their voices will be used.
Tran Le Hong, Deputy Director General of the Intellectual Property Office of Vietnam - Ministry of Science and Technology, says voice protection was rarely discussed in the past because technology was not advanced enough to imitate voices and apply them to different forms of content. Now that AI can learn and recreate an individual’s voice, the issue deserves serious attention.
He says one of the key challenges is deciding when an AI-generated voice becomes similar enough to be considered a copy of the original. While AI-generated voices may differ in aspects such as pitch, accent, or speaking style, they can still preserve recognizable traits that make listeners associate them with a specific person.
"What level of imitation should be considered acceptable and what level should not? These are questions that require further research and clear answers."
Since April 1, amendments have been made to Vietnam’s Intellectual Property Law that allow data to be used for research, testing, and AI training provided that such use does not "unreasonably prejudice" the legitimate rights and interests of rights holders.
Nevertheless, balancing AI development with the rights of data owners remains a challenge for many countries, including Vietnam.
Meanwhile, both Le Minh and Ha Khanh have stopped "selling their voices to AI" though they continue working as audiobook narrators.
"Who knows? A few years from now, AI may end up taking jobs away from people like me," Minh adds.
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。