普通视图

Received before yesterday

KannadaLit4NLP: A comprehensive classical kannada literary dataset of Vachanas, Tripadis, and Kagga with scholarly interpretations for natural language processing

2026年7月2日 18:00

Data Brief. 2026 Jun 19;67:112983. doi: 10.1016/j.dib.2026.112983. eCollection 2026 Aug.

ABSTRACT

This article presents KannadaLit4NLP, a large-scale, machine-readable corpus of Kannada literary texts designed to support natural language processing (NLP) research for a low-resource language. The dataset comprises 24,746 literary verses from three major Kannada literary traditions-Vachanas (11th-19th century), Tripadis (16th century), and Kagga (20th century)-along with 22,369 corresponding interpretations curated from scholarly sources. The corpus captures linguistic, stylistic, and semantic variations across historical periods and literary forms. The dataset was developed through a systematic pipeline that included source identification, digitisation via optical character recognition (OCR), manual verification, and structured annotation. Each entry is organised in a structured format that includes the original verse, metadata (literary form, author, and source), and associated interpretation(s), enabling its use in tasks such as semantic textual similarity, textual entailment, information retrieval, and generative modelling. KannadaLit4NLP addresses the limited availability of culturally grounded Kannada datasets by providing a resource that integrates classical and modern literary content with interpretative annotations. The dataset can facilitate the development and evaluation of NLP models in areas such as semantic understanding, translation, and knowledge representation, while also supporting computational studies of literary and cultural texts. The dataset is made publicly available to encourage further research and reproducibility in Kannada NLP.

PMID:42389175 | PMC:PMC13320459 | DOI:10.1016/j.dib.2026.112983

❌