FPT and Nvidia Release Vietnamese Dataset with 900,000 ‘Personas’
Nemotron-Personas-Vietnam consists of 900,000 Vietnamese “personas,” serving to train and fine-tune artificial intelligence models. Each persona represents a hypothetical Vietnamese individual, with information about name, age, place of residence, occupation, income, marital status, and more. The personas are not based on real people but are synthetic data created by an AI system, reflecting the statistical distribution and verification methods to accurately represent the social realities of Vietnam.
Nemotron-Personas-Vietnam was released on Hugging Face, the world’s leading platform for sharing open-source AI models and datasets, last week. It allows for commercial and non-commercial use as long as the source is acknowledged.

One classification in the Nemotron-Personas-Vietnam dataset is based on occupational groups. Photo: Hugging Face
According to FPT, the publicly released version of Nemotron-Personas-Vietnam has a total capacity of 118 million tokens (the units that AI models use to read and process language) – a scale large enough to support developers in creating training data, fine-tuning, or evaluating Vietnamese AI models.
Among the 900,000 persona records, each is described through various fields of information, including occupation, skills, career goals, sports interests, arts, travel, culinary preferences, age, gender, education level, marital status, residency area, and locality. Describing personas in multiple dimensions allows developers to filter, group, and create data scenarios tailored to specific user groups, professions, or application needs.
Most popular AI models today are trained on English data and Western contexts. When applied in Vietnam, these models may not fully understand the differences in language, culture, professions, regional variations, communication styles, and the real needs of users. With more local data, AI developers can reduce biases in the training process and build AI models that better meet the needs of Vietnamese users and society.

“AI sovereignty must be built from the ground up to reflect the local language, culture, and economic realities. The Nemotron-Personas-Vietnam dataset provides AI developers with the necessary resources to create AI solutions tailored to Vietnamese people and potentially expand to the region,” said Associate Professor Dr. Ngo Xuan Bach, Director of the AI Product Division at FPT Smart Cloud.
Nemotron-Personas is Nvidia’s method for constructing realistic statistical profiles of simulated individuals. The dataset from the United States contains approximately 6 million records, while South Korea has 7 million. FPT stated that it developed the dataset based on Nvidia’s methodology, local expertise, data verification capabilities, data infrastructure, and AI research capabilities.