普通视图

Received before yesterday学术期刊(海外)

KannadaLit4NLP: A comprehensive classical kannada literary dataset of Vachanas, Tripadis, and Kagga with scholarly interpretations for natural language processing

2026年7月2日 18:00

Data Brief. 2026 Jun 19;67:112983. doi: 10.1016/j.dib.2026.112983. eCollection 2026 Aug.

ABSTRACT

This article presents KannadaLit4NLP, a large-scale, machine-readable corpus of Kannada literary texts designed to support natural language processing (NLP) research for a low-resource language. The dataset comprises 24,746 literary verses from three major Kannada literary traditions-Vachanas (11th-19th century), Tripadis (16th century), and Kagga (20th century)-along with 22,369 corresponding interpretations curated from scholarly sources. The corpus captures linguistic, stylistic, and semantic variations across historical periods and literary forms. The dataset was developed through a systematic pipeline that included source identification, digitisation via optical character recognition (OCR), manual verification, and structured annotation. Each entry is organised in a structured format that includes the original verse, metadata (literary form, author, and source), and associated interpretation(s), enabling its use in tasks such as semantic textual similarity, textual entailment, information retrieval, and generative modelling. KannadaLit4NLP addresses the limited availability of culturally grounded Kannada datasets by providing a resource that integrates classical and modern literary content with interpretative annotations. The dataset can facilitate the development and evaluation of NLP models in areas such as semantic understanding, translation, and knowledge representation, while also supporting computational studies of literary and cultural texts. The dataset is made publicly available to encourage further research and reproducibility in Kannada NLP.

PMID:42389175 | PMC:PMC13320459 | DOI:10.1016/j.dib.2026.112983

The Naxi Dongba MOOC: A Test Case for Digital Revitalisation of Endangered Writing Systems

2026年6月28日 08:00
Can digital humanities help revitalise endangered scripts? Drawing on the case of the Naxi Dongba script, this article shows how a MOOC can support the sustainable digital revitalisation of endangered writing systems, moving beyond documentation toward active transmission and engagement.

An Alchemical <em>Prima Materia</em> for the Digital Age: Making the Early Modern Latin Alchemical Prints (EMLAP) Dataset

2026年6月25日 18:00

Ambix. 2026 Jun 25:1-30. doi: 10.1080/00026980.2026.2668735. Online ahead of print.

ABSTRACT

This article documents the creation of EMLAP (Early Modern Latin Alchemical Prints), a machine-readable corpus of one hundred Latin alchemical printed works produced within the TOME (The Origins of Modern Encyclopaedism, 2023-2025) project. It first situates the corpus within both the digital humanities landscape and the historiography of alchemy, where the availability of reliable machine-readable texts remains limited. It then addresses the challenges of converting early modern Latin printed text into machine-readable format with a high standard of quality. The article argues that producing a high-quality transcribed corpus at scale still requires human scholarly intervention, and that a transcription project must balance the ideal of digital edition standards against the practical constraints of time and resources. The article describes the practical experience of building EMLAP: the selection of the Transkribus platform for AI-powered automatic text recognition, the choice of transcription models, the development of human quality standards to complement automated metrics, and the construction of a computational pipeline to process and enrich the transcriptions. The EMLAP corpus has been made publicly available in open access (Zenodo repository) as well as in the form of a website that offers different search opportunities for researchers.

PMID:42345208 | DOI:10.1080/00026980.2026.2668735

Visions in the Machine: Automated Tagging of the William Blake Archive

What can multimodal AI actually see in William Blake's visionary art? This pilot study finds that AI reliably retrieves Blake's objects and motifs but falters, measurably, at his personal iconography, mapping precisely where machine assistance ends, and scholarly interpretation begins.

CHINTEXDB-PERU28: A unique dataset of traditional textile iconographies from Chinchero, Peru for cultural preservation and image recognition

Data Brief. 2026 May 10;66:112835. doi: 10.1016/j.dib.2026.112835. eCollection 2026 Jun.

ABSTRACT

This dataset was collected during on-site fieldwork conducted in the district of Chinchero, located in the province of Urubamba, Cusco, Peru, a region internationally recognized for its rich Andean textile tradition rooted in Inca Culture heritage. The dataset comprises high-quality Photographic images of traditional handwoven Andean textile iconographies produced by local artisan communities. These images were captured directly at textile centers where the fabrics are woven, dyed and finished using ancestral techniques measuring authentic representation of colors, textures, and symbolic patterns under natural and controlled conditions. The dataset consists of 1358 images organized into 28 distinct classes, each corresponding to a specific textile iconography characteristic of the Chinchero tradition. The images are provided in a processed and curated format, facilitating organization enables systematic analysis of visual motifs that are often challenging to distinguish due to their intricate geometric patterns and cultural symbolism. The primary reuse potential of this dataset lies in its application to Artificial Intelligence (AI) and Machine Learning (ML) research focused on image classification, pattern recognition, and cultural heritage preservation. Researchers can leverage the dataset to develop and evaluate models capable of identifying and differentiating traditional Andean textile iconographies, addressing the growing difficulty faced by younger generations, local communities, and visitors in recognizing the cultural expressions. Additionally, the dataset supports interdisciplinary research in digital humanities, ethnography, textile studies, and cultural informatics contributing to the documentation and preservation of intangible cultural heritage. By making this dataset publicly available, this work aims to support the development of AI-driven tools for cultural preservation, educational applications, and heritage awareness, while fostering collaboration between researchers, technologists, and local artisan communities to safeguard ancestral knowledge for future generations.

PMID:42220648 | PMC:PMC13217882 | DOI:10.1016/j.dib.2026.112835

Introduction to Community, Activism, and Innovation in Latin American and Latinx Public Digital HumanitiesEntwodiksyon sou Kominote, Aktivis, ak Inovasyon nan Syans Imanitè Nimerik Piblik an Amerik Latin ak LatinoIntroduction à La Communauté, l’Activisme et l’Innovation dans les Humanités Numériques Publiques Latino-américaines et LatinxIntroducción a la comunidad, el activism y la inovacción en las humanidades digitales públicas latinoamericanas y latinxIntrodução à Comunidade, ativismo e inovação nas humanidades digitais públicas latino-Americanas e latinx

The use of digital tools in Latin American and Latinx digital humanities is often about building connections beyond academic spaces and engaging with real problems faced by actual people. Itilizasyon zouti dijital nan syans imanitè dijital Amerik Latin nan ak Latino yo souvan gen pou wè ak bati koneksyon ki depase espas akademik yo epi angaje yo ak pwoblèm reyèl moun reyèl ap fè fas. L'utilisation des outils numériques dans les humanités numériques latino-américaines et latinx vise souvent à tisser des liens au-delà des espaces académiques et à s'attaquer aux problèmes concrets auxquels sont confrontées des personnes réelles. El uso de herramientas digitales en las humanidades digitales latinoamericanas y latinas suele centrarse en establecer conexiones más allá de los espacios académicos y en abordar problemas reales que enfrentan personas reales. O uso de ferramentas digitais nas humanidades digitais latino-americanas e latinx frequentemente diz respeito a estabelecer conexões para além dos espaços acadêmicos e a engajar-se com problemas reais enfrentados por pessoas reais.

Fostering Transborder Thinking at the Intersection of Digital-Public Humanities and Border Epistemologies with United Fronteras

This article explores the intersection of digital humanities and border epistemologies, providing pedagogical frameworks that can help other researchers, teachers and community members analyze tools, practices, and ethical knowledge production to challenge dominant narratives of the Mexico-U.S. border.

Documenting the Movement for Mexican American Studies (MAS) in Texas Through Critical Latinx Public Digital Humanities (DH) and Chicanx Feminisms

Sylvia Mendoza Aviña, assistant professor of Mexican American Studies (MAS) at the University of Texas, San Antonio (UTSA), introduces the MAS Muxeres Oral History Project, a digital storytelling project that uses ArcGIS StoryMaps to share the oral histories and intimate archival materials of the muxeres who build and sustain MAS programs in Yanawana/San Antonio, Texas, and also explores the transformational possibilities of fusing critical Latinx digital humanities with MAS and Ethnic Studies research.

Good Women, Mediocre Men: Hierarchy in Narrative and Digital Prosopography

Why is information about mediocre eighteenth-century actor Alexander Pope (1762-1835) so readily available? And why does his information dictate what we know about his more talented wives, Elizabeth Younge (1740-1797) and Maria Ann Campion (1777-1803)? Good Women, Mediocre Men: Hierarchy in Narrative and Gendered Prosopography explores how the gendered assumptions embedded in historical biographies continue to condition 21st century digital scholarship.

Infrastructures of Listening: The ManoWhisper Podcast Analysis Pipeline

ManoWhisper is an end-to-end research pipeline for collecting, transcribing, and analyzing hateful and misogynistic podcast content, built to support peer-reviewed and policy-facing research on gender-based extremism. This paper argues the tool reframes harmful media as a site of feminist methodological inquiry, with implications for understanding how such content spreads across platforms and into AI training data
❌