普通视图

Received before yesterday学术期刊(海外)

KannadaLit4NLP: A comprehensive classical kannada literary dataset of Vachanas, Tripadis, and Kagga with scholarly interpretations for natural language processing

2026年7月2日 18:00

Data Brief. 2026 Jun 19;67:112983. doi: 10.1016/j.dib.2026.112983. eCollection 2026 Aug.

ABSTRACT

This article presents KannadaLit4NLP, a large-scale, machine-readable corpus of Kannada literary texts designed to support natural language processing (NLP) research for a low-resource language. The dataset comprises 24,746 literary verses from three major Kannada literary traditions-Vachanas (11th-19th century), Tripadis (16th century), and Kagga (20th century)-along with 22,369 corresponding interpretations curated from scholarly sources. The corpus captures linguistic, stylistic, and semantic variations across historical periods and literary forms. The dataset was developed through a systematic pipeline that included source identification, digitisation via optical character recognition (OCR), manual verification, and structured annotation. Each entry is organised in a structured format that includes the original verse, metadata (literary form, author, and source), and associated interpretation(s), enabling its use in tasks such as semantic textual similarity, textual entailment, information retrieval, and generative modelling. KannadaLit4NLP addresses the limited availability of culturally grounded Kannada datasets by providing a resource that integrates classical and modern literary content with interpretative annotations. The dataset can facilitate the development and evaluation of NLP models in areas such as semantic understanding, translation, and knowledge representation, while also supporting computational studies of literary and cultural texts. The dataset is made publicly available to encourage further research and reproducibility in Kannada NLP.

PMID:42389175 | PMC:PMC13320459 | DOI:10.1016/j.dib.2026.112983

CHINTEXDB-PERU28: A unique dataset of traditional textile iconographies from Chinchero, Peru for cultural preservation and image recognition

Data Brief. 2026 May 10;66:112835. doi: 10.1016/j.dib.2026.112835. eCollection 2026 Jun.

ABSTRACT

This dataset was collected during on-site fieldwork conducted in the district of Chinchero, located in the province of Urubamba, Cusco, Peru, a region internationally recognized for its rich Andean textile tradition rooted in Inca Culture heritage. The dataset comprises high-quality Photographic images of traditional handwoven Andean textile iconographies produced by local artisan communities. These images were captured directly at textile centers where the fabrics are woven, dyed and finished using ancestral techniques measuring authentic representation of colors, textures, and symbolic patterns under natural and controlled conditions. The dataset consists of 1358 images organized into 28 distinct classes, each corresponding to a specific textile iconography characteristic of the Chinchero tradition. The images are provided in a processed and curated format, facilitating organization enables systematic analysis of visual motifs that are often challenging to distinguish due to their intricate geometric patterns and cultural symbolism. The primary reuse potential of this dataset lies in its application to Artificial Intelligence (AI) and Machine Learning (ML) research focused on image classification, pattern recognition, and cultural heritage preservation. Researchers can leverage the dataset to develop and evaluate models capable of identifying and differentiating traditional Andean textile iconographies, addressing the growing difficulty faced by younger generations, local communities, and visitors in recognizing the cultural expressions. Additionally, the dataset supports interdisciplinary research in digital humanities, ethnography, textile studies, and cultural informatics contributing to the documentation and preservation of intangible cultural heritage. By making this dataset publicly available, this work aims to support the development of AI-driven tools for cultural preservation, educational applications, and heritage awareness, while fostering collaboration between researchers, technologists, and local artisan communities to safeguard ancestral knowledge for future generations.

PMID:42220648 | PMC:PMC13217882 | DOI:10.1016/j.dib.2026.112835

Stacks and Intersections: Feminist Thinking in Digital Humanities, a view from these islands

This discussion is about feminisms, Digital Humanities (DH), stacks, and archives. We argue for Full Stack Feminism, as a methodological approach, informed by the successes of previous feminist interventions in DH.

How digital are the Digital Humanities? An analysis of two scholarly blogging platforms

2015年2月13日 19:00

PLoS One. 2015 Feb 12;10(2):e0115035. doi: 10.1371/journal.pone.0115035. eCollection 2015.

ABSTRACT

In this paper we compare two academic networking platforms, HASTAC and Hypotheses, to show the distinct ways in which they serve specific communities in the Digital Humanities (DH) in different national and disciplinary contexts. After providing background information on both platforms, we apply co-word analysis and topic modeling to show thematic similarities and differences between the two sites, focusing particularly on how they frame DH as a new paradigm in humanities research. We encounter a much higher ratio of posts using humanities-related terms compared to their digital counterparts, suggesting a one-way dependency of digital humanities-related terms on the corresponding unprefixed labels. The results also show that the terms digital archive, digital literacy, and digital pedagogy are relatively independent from the respective unprefixed terms, and that digital publishing, digital libraries, and digital media show considerable cross-pollination between the specialization and the general noun. The topic modeling reproduces these findings and reveals further differences between the two platforms. Our findings also indicate local differences in how the emerging field of DH is conceptualized and show dynamic topical shifts inside these respective contexts.

PMID:25675441 | PMC:PMC4326279 | DOI:10.1371/journal.pone.0115035

The Possibility of Using African Languages as Media of Teaching and Learning in South Africa

This study sets out to examine the possibility of using African languages as media of teaching and learning in South African schools. Literature is consistent that (a) language is a crucial means of communication and gaining access to important knowledge and skills, and (b) mother tongue is the only language that promotes effective teaching and learning and that any language, which is not a mother tongue, is a barrier to teaching and learning. In South Africa, there are nine African official languages, but English is the media of instruction used by South African learners, which is a barrier to teaching and learning. This study revealed that using one or two African languages may improve teaching, learning, and the academic performance of the learners, but the problem is how to implement because it will be difficult to use many African languages as media of instruction. The use of nine African languages as media of instruction in South Africa will promote tribalism, which was dominant during the apartheid era, and it will be costly to the government. Therefore, this study supports the use of English as a media of instruction because it will promote unity in South Africa, it will not be costly, and it is an international language.

❌