BrajText-Saar: A structured cultural dataset for cultural text mining in indian heritage texts
Data Brief. 2026 Jul 16;68:113086. doi: 10.1016/j.dib.2026.113086. eCollection 2026 Aug.
ABSTRACT
India is one of the countries where many regional languages are spoken in different parts of the country. Each part of the country has its unique cultural traditions; therefore, an abundance of cultural texts is available in Indian Regional languages. The cultural texts are mostly ancient scriptures that represent the community's heritage, values, and ancient knowledge. Braj is one of the Indian regional languages that represents the cultural heritage of the Braj region. The BrajText-Saar dataset contains texts from the Braj language, which features a variety of prehistoric scripts. The data was collected from the Maan Mandir trust's portal, which operates to maintain Braj culture and heritage. The Braj text primarily explains the divine actions and the life of Lord Krishna in the form of poetry and devotional songs. Offline Braj literature from renowned writers is available in manuscript form. This represents an opportunity for cultural text mining in the Braj Language. The developed dataset opens new paths in various research areas, including language studies, sentiment analysis, and emotion analysis, and is also suitable for digital humanities research.
PMID:42564907 | PMC:PMC13446094 | DOI:10.1016/j.dib.2026.113086