❌

普通视图

Received yesterday — 2026年10月5日学术期刊(海外)
Received before yesterday学术期刊(海外)

Contextual Word Embeddings for Paracelsian Lexicography: Tangled Terminologies and their Origins in Ruland's Alchemical Dictionary

2026年9月4日 18:00

Ambix. 2026 May-Aug;73(2-3):205-238. doi: 10.1080/00026980.2026.2692190. Epub 2026 Sep 4.

ABSTRACT

Martin Ruland the Younger's Lexicon Alchemiae (1612) is one of the most influential alchemical dictionaries of the early modern period, yet its sources and compilation methods remain poorly understood. This study applies computational approaches to investigate the vocabulary underlying Ruland's lexicon and to identify potential textual influences. Using a TEI-XML encoded version of the Lexicon Alchemiae, a standard data format for encoding textual data in the digital humanities, we extract its headwords and compare them against large-scale digital corpora of Latin literature, including the Early Modern Latin Alchemical Prints (EMLAP) dataset and the broader GreLa database. By combining frequency-based lexical comparison with contextual word embeddings generated through a Latin BERT (Bidirectional Encoder Representations from Transformers) model, the analysis traces both the distribution and semantic behaviour of terms across earlier alchemical and scientific texts. The article is accompanied by an interactive web application allowing readers to explore additional case studies. Our analysis indicates that Ruland drew not only on earlier Paracelsian word lists, but also on large-scale contemporary compilations such as Andreas Libavius's Alchemia (1597), suggesting that the Lexicon Alchemiae should be understood within a broader movement to systematise and professionalise alchemical knowledge in the late sixteenth and early seventeenth centuries.

PMID:42695817 | DOI:10.1080/00026980.2026.2692190

Computational tools in the history of physics: Distant reading Franklin's reception in 18th-century France

2026年8月22日 18:00

Stud Hist Philos Sci. 2026 Aug 22;119:102196. doi: 10.1016/j.shpsa.2026.102196. Online ahead of print.

ABSTRACT

This article examines the use of computational tools in the history of science. For this reason, a historical case study is chosen: the interaction between Benjamin Franklin and Jean-Antoine Nollet over their competing theories of electricity in 18th-century France, focusing on the reception of Franklin's one-fluid theory in a Nollet-dominated Parisian Academy of Sciences. We built a database to study the reception and eventual acceptance of Franklin's theory in France. The historian of science John L. Heilbron portrays the Academy as divided and paralyzed, while Roderick W. Home argues that Nollet's dominance led the Academy to conduct the debate in Nollet's shadow. Using a database of publications on electricity related to the Parisian Academy between 1745 and 1785, we examined Franklin's theory in France quantitatively, revealing a pattern of influence and intellectual dominance that changed abruptly after Nollet's death. We also produced data comparing Nollet and another important figure in 18th-century French electricity, Jean-Baptiste Le Roy, and reconstructed this debate using statistics. Our findings demonstrate that databases and scientometric methods can be fruitfully applied in the history and philosophy of science. In particular, we can confirm part of the dynamics of this historical event, especially related to Nollet's influence within the Parisian Academy, an extra-scientific factor. For further exploration of this case study and to test these tools more thoroughly, a larger database of electricity-related publications in France is envisioned as the next step.

PMID:42632168 | DOI:10.1016/j.shpsa.2026.102196

Uncertainty-aware joint modeling for Sanskrit compound splitting and segmentation

2026年8月12日 18:00

Front Artif Intell. 2026 Jul 28;9:1873672. doi: 10.3389/frai.2026.1873672. eCollection 2026.

ABSTRACT

Sanskrit compound word splitting is challenging because Sandhi-driven phonological transformations obscure word boundaries, making splitting non-deterministic, especially in multi-split cases where multiple hidden boundaries must be identified and constituent segments reconstructed. In this study, a joint end-to-end Transformer-based multi-task architecture is proposed to address this problem by integrating boundary detection and segmented sequence generation within a unified framework. The model employs a shared character-level Transformer encoder that feeds a BiLSTM-CRF boundary prediction branch, which produces sequence-consistent boundary locations and posterior marginals, and a Transformer decoder that generates the segmented output autoregressively. To couple the two tasks, boundary probabilities and entropy-based uncertainty are derived from the CRF marginals. These signals are used to gate encoder representations and to construct a global boundary-aware context that conditions each decoding step, thereby enabling more robust decoding under boundary ambiguity. Experimental results demonstrate consistent improvements over state-of-the-art methods. The model achieves exact-match boundary location accuracy of 84.39%, exact segmentation accuracy of 79.34%, and character-level accuracy of 87.79%. It also achieves boundary-level Precision, Recall, and F1 scores of 92.84%, 91.66%, and 92.25%, respectively. These results indicate that uncertainty-aware coupling and structured BiLSTM-CRF supervision improve segmentation performance while maintaining strong boundary detection, thereby enabling more accurate morphological analysis for NLP and digital humanities applications.

PMID:42582249 | PMC:PMC13457376 | DOI:10.3389/frai.2026.1873672

KannadaLit4NLP: A comprehensive classical kannada literary dataset of Vachanas, Tripadis, and Kagga with scholarly interpretations for natural language processing

2026年7月2日 18:00

Data Brief. 2026 Jun 19;67:112983. doi: 10.1016/j.dib.2026.112983. eCollection 2026 Aug.

ABSTRACT

This article presents KannadaLit4NLP, a large-scale, machine-readable corpus of Kannada literary texts designed to support natural language processing (NLP) research for a low-resource language. The dataset comprises 24,746 literary verses from three major Kannada literary traditions-Vachanas (11th-19th century), Tripadis (16th century), and Kagga (20th century)-along with 22,369 corresponding interpretations curated from scholarly sources. The corpus captures linguistic, stylistic, and semantic variations across historical periods and literary forms. The dataset was developed through a systematic pipeline that included source identification, digitisation via optical character recognition (OCR), manual verification, and structured annotation. Each entry is organised in a structured format that includes the original verse, metadata (literary form, author, and source), and associated interpretation(s), enabling its use in tasks such as semantic textual similarity, textual entailment, information retrieval, and generative modelling. KannadaLit4NLP addresses the limited availability of culturally grounded Kannada datasets by providing a resource that integrates classical and modern literary content with interpretative annotations. The dataset can facilitate the development and evaluation of NLP models in areas such as semantic understanding, translation, and knowledge representation, while also supporting computational studies of literary and cultural texts. The dataset is made publicly available to encourage further research and reproducibility in Kannada NLP.

PMID:42389175 | PMC:PMC13320459 | DOI:10.1016/j.dib.2026.112983

The Naxi Dongba MOOC: A Test Case for Digital Revitalisation of Endangered Writing Systems

2026年6月28日 08:00
Can digital humanities help revitalise endangered scripts? Drawing on the case of the Naxi Dongba script, this article shows how a MOOC can support the sustainable digital revitalisation of endangered writing systems, moving beyond documentation toward active transmission and engagement.

An Alchemical <em>Prima Materia</em> for the Digital Age: Making the Early Modern Latin Alchemical Prints (EMLAP) Dataset

2026年6月25日 18:00

Ambix. 2026 May-Aug;73(2-3):145-174. doi: 10.1080/00026980.2026.2668735. Epub 2026 Jun 25.

ABSTRACT

This article documents the creation of EMLAP (Early Modern Latin Alchemical Prints), a machine-readable corpus of one hundred Latin alchemical printed works produced within the TOME (The Origins of Modern Encyclopaedism, 2023-2025) project. It first situates the corpus within both the digital humanities landscape and the historiography of alchemy, where the availability of reliable machine-readable texts remains limited. It then addresses the challenges of converting early modern Latin printed text into machine-readable format with a high standard of quality. The article argues that producing a high-quality transcribed corpus at scale still requires human scholarly intervention, and that a transcription project must balance the ideal of digital edition standards against the practical constraints of time and resources. The article describes the practical experience of building EMLAP: the selection of the Transkribus platform for AI-powered automatic text recognition, the choice of transcription models, the development of human quality standards to complement automated metrics, and the construction of a computational pipeline to process and enrich the transcriptions. The EMLAP corpus has been made publicly available in open access (Zenodo repository) as well as in the form of a website that offers different search opportunities for researchers.

PMID:42345208 | DOI:10.1080/00026980.2026.2668735

Visions in the Machine: Automated Tagging of the William Blake Archive

What can multimodal AI actually see in William Blake's visionary art? This pilot study finds that AI reliably retrieves Blake's objects and motifs but falters, measurably, at his personal iconography, mapping precisely where machine assistance ends, and scholarly interpretation begins.

CHINTEXDB-PERU28: A unique dataset of traditional textile iconographies from Chinchero, Peru for cultural preservation and image recognition

Data Brief. 2026 May 10;66:112835. doi: 10.1016/j.dib.2026.112835. eCollection 2026 Jun.

ABSTRACT

This dataset was collected during on-site fieldwork conducted in the district of Chinchero, located in the province of Urubamba, Cusco, Peru, a region internationally recognized for its rich Andean textile tradition rooted in Inca Culture heritage. The dataset comprises high-quality Photographic images of traditional handwoven Andean textile iconographies produced by local artisan communities. These images were captured directly at textile centers where the fabrics are woven, dyed and finished using ancestral techniques measuring authentic representation of colors, textures, and symbolic patterns under natural and controlled conditions. The dataset consists of 1358 images organized into 28 distinct classes, each corresponding to a specific textile iconography characteristic of the Chinchero tradition. The images are provided in a processed and curated format, facilitating organization enables systematic analysis of visual motifs that are often challenging to distinguish due to their intricate geometric patterns and cultural symbolism. The primary reuse potential of this dataset lies in its application to Artificial Intelligence (AI) and Machine Learning (ML) research focused on image classification, pattern recognition, and cultural heritage preservation. Researchers can leverage the dataset to develop and evaluate models capable of identifying and differentiating traditional Andean textile iconographies, addressing the growing difficulty faced by younger generations, local communities, and visitors in recognizing the cultural expressions. Additionally, the dataset supports interdisciplinary research in digital humanities, ethnography, textile studies, and cultural informatics contributing to the documentation and preservation of intangible cultural heritage. By making this dataset publicly available, this work aims to support the development of AI-driven tools for cultural preservation, educational applications, and heritage awareness, while fostering collaboration between researchers, technologists, and local artisan communities to safeguard ancestral knowledge for future generations.

PMID:42220648 | PMC:PMC13217882 | DOI:10.1016/j.dib.2026.112835

Introduction to Community, Activism, and Innovation in Latin American and Latinx Public Digital HumanitiesEntwodiksyon sou Kominote, Aktivis, ak Inovasyon nan Syans Imanitè Nimerik Piblik an Amerik Latin ak LatinoIntroduction à La Communauté, l’Activisme et l’Innovation dans les Humanités Numériques Publiques Latino-américaines et LatinxIntroducción a la comunidad, el activism y la inovacción en las humanidades digitales públicas latinoamericanas y latinxIntrodução à Comunidade, ativismo e inovação nas humanidades digitais públicas latino-Americanas e latinx

The use of digital tools in Latin American and Latinx digital humanities is often about building connections beyond academic spaces and engaging with real problems faced by actual people. Itilizasyon zouti dijital nan syans imanitè dijital Amerik Latin nan ak Latino yo souvan gen pou wè ak bati koneksyon ki depase espas akademik yo epi angaje yo ak pwoblèm reyèl moun reyèl ap fè fas. L'utilisation des outils numériques dans les humanités numériques latino-américaines et latinx vise souvent à tisser des liens au-delà des espaces académiques et à s'attaquer aux problèmes concrets auxquels sont confrontées des personnes réelles. El uso de herramientas digitales en las humanidades digitales latinoamericanas y latinas suele centrarse en establecer conexiones más allá de los espacios académicos y en abordar problemas reales que enfrentan personas reales. O uso de ferramentas digitais nas humanidades digitais latino-americanas e latinx frequentemente diz respeito a estabelecer conexões para além dos espaços acadêmicos e a engajar-se com problemas reais enfrentados por pessoas reais.

Fostering Transborder Thinking at the Intersection of Digital-Public Humanities and Border Epistemologies with United Fronteras

This article explores the intersection of digital humanities and border epistemologies, providing pedagogical frameworks that can help other researchers, teachers and community members analyze tools, practices, and ethical knowledge production to challenge dominant narratives of the Mexico-U.S. border.

Documenting the Movement for Mexican American Studies (MAS) in Texas Through Critical Latinx Public Digital Humanities (DH) and Chicanx Feminisms

Sylvia Mendoza Aviña, assistant professor of Mexican American Studies (MAS) at the University of Texas, San Antonio (UTSA), introduces the MAS Muxeres Oral History Project, a digital storytelling project that uses ArcGIS StoryMaps to share the oral histories and intimate archival materials of the muxeres who build and sustain MAS programs in Yanawana/San Antonio, Texas, and also explores the transformational possibilities of fusing critical Latinx digital humanities with MAS and Ethnic Studies research.

Good Women, Mediocre Men: Hierarchy in Narrative and Digital Prosopography

Why is information about mediocre eighteenth-century actor Alexander Pope (1762-1835) so readily available? And why does his information dictate what we know about his more talented wives, Elizabeth Younge (1740-1797) and Maria Ann Campion (1777-1803)? Good Women, Mediocre Men: Hierarchy in Narrative and Gendered Prosopography explores how the gendered assumptions embedded in historical biographies continue to condition 21st century digital scholarship.
❌