Morphologically rich and low-resource languages present challenges for tokenization and other natural language processing (NLP) tasks. Standard tokenization algorithms often fail to segment words into meaningful morphological units. At the same time, datasets for many lowresource languages remain limited. We introduce a new 1dataset of child-directed language in Kazakh and Russian compiled from multiple sources including child–adult dialogue transcripts, fairy tales, cartoons and children’s literature texts. The dataset is manually collected and contains 1M tokens in Kazakh and 2.3M tokens in Russian. Dialogue data includes metadata such as child age and speaker identity, enabling research on language acquisition and linguistic development. We describe the process of data collection, corpus organization, and basic statistical properties of the dataset. The corpus provides a new resource for research on low-resource NLP, language acquisition, and morphologically-aware tokenization methods.
A Kazakh–Russian Corpus of Child-Directed Language for Low-Resource Languages
Albina Mukusheva
Writing – Original Draft Preparation
;Achille FuscoWriting – Review & Editing
;Cristiano ChesiSupervision
2026-01-01
Abstract
Morphologically rich and low-resource languages present challenges for tokenization and other natural language processing (NLP) tasks. Standard tokenization algorithms often fail to segment words into meaningful morphological units. At the same time, datasets for many lowresource languages remain limited. We introduce a new 1dataset of child-directed language in Kazakh and Russian compiled from multiple sources including child–adult dialogue transcripts, fairy tales, cartoons and children’s literature texts. The dataset is manually collected and contains 1M tokens in Kazakh and 2.3M tokens in Russian. Dialogue data includes metadata such as child age and speaker identity, enabling research on language acquisition and linguistic development. We describe the process of data collection, corpus organization, and basic statistical properties of the dataset. The corpus provides a new resource for research on low-resource NLP, language acquisition, and morphologically-aware tokenization methods.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


