Rostislav Konyashkin, First Vice Minister of Artificial Intelligence and Digital Development of Kazakhstan, explained how state digital services work with the Kazakh language and spoke about the shortage of high-quality Kazakh-language data for training artificial intelligence... reports infohub.kz.
During a briefing at the government, journalists asked Konyashkin what language eGov "thinks" in. The Vice Minister replied that the eGov portal is available in three languages, and artificial intelligence models, including eGov GPT, also understand and "think" in Kazakh, Russian, and English.
One journalist drew attention to the lack of quality datasets for training AI in the Kazakh language. According to him, the largest open dataset was created by Nazarbayev University – about 1,200-1,300 hours of recordings based on parliamentary speeches, without taking into account regional pronunciation features.
Konyashkin said that Kazakhstan has already launched the development of its own language models – KAZ-LLM and ALEM-LLM – so that they take into account national specifics and the historical heritage of the country. In the framework of training and development of the models, the maximum possible data from education, culture, and archival materials in the Kazakh language were used.
According to Konyashkin, ideally, models should draw data from the internet, but there is still significantly more content in English than in Kazakh. He stressed that the ministry is interested in increasing the volume of Kazakh-language content online – this will allow better training of language models. Currently, developers rely mainly on archival materials, textbooks, and content from the cultural sphere. In this work, the ministry is assisted by Nazarbayev University and the ISSAI research center. Konyashkin added that the ministry is open to other organizations ready to join in collecting and preparing Kazakh-language data.


