To the central content area
Toggle Dark/Light Mode Dark Mode
:::

MODA’s Taiwan Sovereign AI Training Corpus Has Added 20 Million Tokens of Hakka Language Data, Enriching the Basic Foundations of Local AI Language Data

To enhance Taiwan’s AI capabilities in local culture and languages, the Taiwan Sovereign AI Training Corpus, an initiative developed by the Ministry of Digital Affairs (MODA), recently collaborated with the Hakka Affairs Council (HAC) to officially add valuable content such as Hakka language resources, publications, historical and cultural materials, and research reports. Approximately 20 million tokens of high-quality training data have been added for the first time, injecting richer cultural and linguistic vitality into Taiwan's sovereign AI corpus. This will help improve the AI’s understanding of Hakka language and culture, thereby supporting diverse applications such as Hakka-language smart customer service and education. 

Hakka is one of Taiwan’s national languages, and carries a profound historical and cultural heritage. The corpus provided by HAC for this endeavor is pivotal, and mainly includes official government publications such as poetry collections, Hakka customs, and books for Hakka language education, as well as vocabulary teaching materials and question banks for all levels of Hakka language proficiency certifications, oral transcripts, cultural materials, and academic and policy research materials for initiatives such as Hakka language revitalization, policy development, and cultural governance. 

MODA pointed out that the corpus not only showcases the fruitful results of the preservation of Hakka culture in Taiwan, but also helps improve AI models’ ability to understand and generate different Hakka accents, semantics, and cultural contexts. In the future, it can be widely used in innovative AI applications such as Hakka smart customer service, local cultural knowledge services, and public policy. 

MODA noted that, since it was launched, the Taiwan Sovereign AI Training Corpus has accumulated more than 1.5 billion tokens. It will continue to work with government agencies to collect high-quality language data, while also deepening public-private partnerships to collect corpora from non-governmental sectors, thereby gradually building a database that can fully represent Taiwan’s sovereignty while also featuring professional knowledge. The Hakka language corpus is now officially available. AI model development teams, academic research institutions, and related organizations who wish to use this technology may submit an application on the Taiwan Sovereign AI Training Corpus website (https://taic.moda.gov.tw). This will enable collaborative efforts to drive the vigorous development and diverse applications of Hakka AI technology. MODA also hopes that research results, real-world applications, and other high-quality data with Taiwanese characteristics will be fed back into the corpus, thereby enabling more innovative ideas to take root and collectively strengthening Taiwan's sovereign AI.

Go Top