Selling Voices to Train AI
After many years in the profession, 28-year-old Lê Minh from Hanoi has become accustomed to recording and providing voiceovers for audiobooks, advertisements, and narration videos. However, in the past year, he has received an increasing number of requests to become an “AI voice trainer”—a provider of voice data for training AI. The compensation is quite high, but it is not regular and depends on each project.
Hà Khánh, a language student in Hanoi, is a regular contributor to an audiobook app. She is paid 5,000-7,000 VND per minute of reading, excluding taxes. Since the end of last year, this organization has offered her a chance to participate in a project to create voice data for AI with a higher compensation rate.
Instead of reading entire books as before, Khánh now reads short texts according to a template and sends the audio files back to the system. For a student living away from home, the money from this job is enough to help her share some of the costs of rent and living expenses.
On forums for the voice recording and dubbing community, many job postings for “AI voice trainers” have also appeared. According to the descriptions, participants are required to read scripts and record audio to create training data for AI. Compensation is calculated hourly, based on the number of sentences or the volume of completed data. Readers must meet criteria related to region, gender, or age.

Recording voice data using a mobile phone. Photo: Trọng Đạt
To enable an AI system to read text fluently like a human, developers must collect a large amount of voice data from real people. For commercial AI voice systems, this data pool can reach dozens or even hundreds of hours of recordings.
In an interview with VnExpress, Hồ Minh Đức, CEO of Vbee—a company specializing in text-to-speech services—stated that their data is built from various sources, such as hiring collaborators, hosts, or professional voice actors.
Additionally, some data comes from publicly available copyrighted sources on the Internet. “We actively seek individuals with suitable reading voices. Conversely, there are also those who want to digitize their voices and thus approach our platform,” Đức explained.
A similar situation occurs in the research environment. When creating software to teach Vietnamese to foreigners, the team led by master’s student Vũ Văn Thương at the Post and Telecommunications Institute had to hire a team of teachers to build a voice data repository.
Initially, many participants supported the project for free. However, as the workload increased, the team began offering compensation. They read each word and sentence displayed on the screen, and the system recorded, processed, and converted this into data for AI to learn how to pronounce Vietnamese.
According to Đức, there are two commercial voice exploitation models in Vietnam. In the traditional model, voice readers for audiobooks, advertisements, or automated call centers are paid per recording, calculated by minute or by task.
In contrast, the second model has emerged with the development of AI. Instead of only exploiting specific recordings, companies pay for the digitization of voices to create AI reading voices, which can then be used for various purposes. Essentially, this represents a shift from hiring traditional voice actors to utilizing AI voice technology.
AI Voice Reading Copyright
This transition has also sparked new debates about ownership rights concerning voice. According to Hồ Minh Đức, there have been cases where users reported that an AI voice sounded very similar to theirs. In such cases, developers must trace back the source of the data used to train the system.
“It is essential to determine where the data comes from and whether it has copyright. If it is found that the data used is inappropriate, it must be removed, and the developers must work with the owners to resolve copyright issues,” he stated.
He noted that there is now voice biometric technology worldwide, similar to fingerprint or iris recognition. However, in Vietnam, determining whether AI has copied from a specific individual remains a challenging issue due to the lack of standardized criteria or regulations.

A person using AI tools to convert text into speech. Photo: Trọng Đạt
From a research perspective, master’s student Vũ Văn Thương believes that with just a short recording—about 15 seconds—many AI systems can replicate the voice characteristics and intonation of a specific individual. Therefore, data contributors must be clearly informed about how their voices will be used and within what scope.
Trần Lê Hồng, Deputy Director of the Intellectual Property Department at the Ministry of Science and Technology, mentioned that in the past, voice protection was rarely considered due to technology’s limitations in mimicking and applying voices to other content. However, as AI can now learn and reproduce an individual’s voice, this issue requires serious attention.
To be protected, he stated, the challenge lies in identifying what factors shape a voice. Furthermore, AI does not always use the original voice of a person but can modify it. “The extent to which this is acceptable and the extent to which it is not need to be researched and answered,” he raised.
As of April 1, the amended Intellectual Property Law officially takes effect, allowing the exploitation of data for research, experimentation, and AI training, provided it does not “unreasonably affect” the rights and legitimate interests of the owners. However, balancing the need for AI development with the rights of data owners remains a problem for many countries, including Vietnam.
Meanwhile, both Lê Minh and Hà Khánh have stopped “selling their voices to AI,” although they still maintain their work in simple audiobook reading. “Who knows, maybe in a few years, AI will take away jobs from people like me,” Minh said.