On September 30, the National Institute of the Korean Language released 16 new corpora—including Korean-foreign language parallel corpora and Korean Sign Language corpora—on the “Corpora for All” website. With this release, the total number of Korean corpora distributed by the National Institute of the Korean Language has increased to 174.
The National Institute of the Korean Language has been building and making available to industry and academia high-quality language resources necessary for the development and research of Korean-specialized artificial intelligence (AI). These new corpora are designed to enhance AI’s capabilities in Korean language and culture, facilitate expanded communication through the Korean language, and improve the public’s proficiency in the Korean language.
The largest portion consists of eight Korean-foreign language parallel corpora. Written Korean materials, such as newspaper articles, and spoken materials from the National Institute of the Korean Language’s “Everyday Conversation Corpus 2024” have been paired with translations into eight languages—Vietnamese, Indonesian, Thai, Hindi, Cambodian Khmer, Filipino Tagalog, Russian, and Uzbek. These serve as essential data for improving the quality of AI-based foreign language interpretation and translation models, with one corpus released per language.
There are three types of Korean Sign Language corpora: a raw corpus consisting of videos of two deaf individuals conversing in sign language; an annotated corpus in which these videos were translated into Korean and the sign language words were segmented and annotated; and a Korean Sign Language–Korean parallel corpus in which the conversation videos were transcribed into Korean. These corpora can be used to develop sign language interpretation technologies that enhance communication convenience for sign language users.
Data for training AI to evaluate people’s proficiency in the Korean language has also been released. These include a raw writing corpus comprising argumentative essays of approximately 1,000 characters written by university students at national and public universities and general adults across nine regions nationwide, and a writing grading corpus containing the results of these essays graded by two expert graders.
In addition, there is the “Instruction-Based Writing Correction Support Corpus,” in which user question types were subdivided into 872 categories based on Korean language counseling cases, question-and-answer materials were created for each category, and a language model generated Q&A responses based on this data; the “Korean Language Historical Corpus,” containing early 20th-century documents such as new-style novels, the *Jeuguk Shinmun* newspaper, and the academic journal *Hangul*; and the “Basic Korean Language Knowledge Corpus,” which organizes examples of foreign word transcription, standard pronunciation, and Romanization rules into a common format, have been released.
The corpora are available to any researcher, developer, or business operator seeking to utilize them for research and technology development in the fields of Korean language studies and language information processing. Users can download the resources by completing an online agreement form on the “Corpus for All” website and obtaining approval. However, the data may not be used for purposes other than those approved; it may not be transferred or lent to third parties; and prior approval must be obtained before publicly releasing any results derived from the data.
An official from the National Institute of the Korean Language stated, “To support the development of autonomous AI systems that are proficient in the Korean language and well-versed in Korean culture, we plan to continuously build and release a cumulative total of 340 Korean language and cultural corpora by 2030.”




![[인터뷰] 심리 상담사가 직접 말하는 심리 상담사의 하루](https://en.swn.kr/wp-content/uploads/sites/24/2015/07/KakaoTalk_20141219_092335502.jpg)