普通视图
-
2 - DHQ(Digital Humanities Quarterly)
- Building open access, calibrated, syntactically annotated corpora over the history of French: A Case study of the application and adaptation of annotation tools for historical syntax
-
7 - PubMed
- Contextual Word Embeddings for Paracelsian Lexicography: Tangled Terminologies and their Origins in Ruland's Alchemical Dictionary
Contextual Word Embeddings for Paracelsian Lexicography: Tangled Terminologies and their Origins in Ruland's Alchemical Dictionary
Ambix. 2026 May-Aug;73(2-3):205-238. doi: 10.1080/00026980.2026.2692190. Epub 2026 Sep 4.
ABSTRACT
Martin Ruland the Younger's Lexicon Alchemiae (1612) is one of the most influential alchemical dictionaries of the early modern period, yet its sources and compilation methods remain poorly understood. This study applies computational approaches to investigate the vocabulary underlying Ruland's lexicon and to identify potential textual influences. Using a TEI-XML encoded version of the Lexicon Alchemiae, a standard data format for encoding textual data in the digital humanities, we extract its headwords and compare them against large-scale digital corpora of Latin literature, including the Early Modern Latin Alchemical Prints (EMLAP) dataset and the broader GreLa database. By combining frequency-based lexical comparison with contextual word embeddings generated through a Latin BERT (Bidirectional Encoder Representations from Transformers) model, the analysis traces both the distribution and semantic behaviour of terms across earlier alchemical and scientific texts. The article is accompanied by an interactive web application allowing readers to explore additional case studies. Our analysis indicates that Ruland drew not only on earlier Paracelsian word lists, but also on large-scale contemporary compilations such as Andreas Libavius's Alchemia (1597), suggesting that the Lexicon Alchemiae should be understood within a broader movement to systematise and professionalise alchemical knowledge in the late sixteenth and early seventeenth centuries.
PMID:42695817 | DOI:10.1080/00026980.2026.2692190
-
7 - PubMed
- Computational tools in the history of physics: Distant reading Franklin's reception in 18th-century France
Computational tools in the history of physics: Distant reading Franklin's reception in 18th-century France
Stud Hist Philos Sci. 2026 Aug 22;119:102196. doi: 10.1016/j.shpsa.2026.102196. Online ahead of print.
ABSTRACT
This article examines the use of computational tools in the history of science. For this reason, a historical case study is chosen: the interaction between Benjamin Franklin and Jean-Antoine Nollet over their competing theories of electricity in 18th-century France, focusing on the reception of Franklin's one-fluid theory in a Nollet-dominated Parisian Academy of Sciences. We built a database to study the reception and eventual acceptance of Franklin's theory in France. The historian of science John L. Heilbron portrays the Academy as divided and paralyzed, while Roderick W. Home argues that Nollet's dominance led the Academy to conduct the debate in Nollet's shadow. Using a database of publications on electricity related to the Parisian Academy between 1745 and 1785, we examined Franklin's theory in France quantitatively, revealing a pattern of influence and intellectual dominance that changed abruptly after Nollet's death. We also produced data comparing Nollet and another important figure in 18th-century French electricity, Jean-Baptiste Le Roy, and reconstructed this debate using statistics. Our findings demonstrate that databases and scientometric methods can be fruitfully applied in the history and philosophy of science. In particular, we can confirm part of the dynamics of this historical event, especially related to Nollet's influence within the Parisian Academy, an extra-scientific factor. For further exploration of this case study and to test these tools more thoroughly, a larger database of electricity-related publications in France is envisioned as the next step.
PMID:42632168 | DOI:10.1016/j.shpsa.2026.102196
Uncertainty-aware joint modeling for Sanskrit compound splitting and segmentation
Front Artif Intell. 2026 Jul 28;9:1873672. doi: 10.3389/frai.2026.1873672. eCollection 2026.
ABSTRACT
Sanskrit compound word splitting is challenging because Sandhi-driven phonological transformations obscure word boundaries, making splitting non-deterministic, especially in multi-split cases where multiple hidden boundaries must be identified and constituent segments reconstructed. In this study, a joint end-to-end Transformer-based multi-task architecture is proposed to address this problem by integrating boundary detection and segmented sequence generation within a unified framework. The model employs a shared character-level Transformer encoder that feeds a BiLSTM-CRF boundary prediction branch, which produces sequence-consistent boundary locations and posterior marginals, and a Transformer decoder that generates the segmented output autoregressively. To couple the two tasks, boundary probabilities and entropy-based uncertainty are derived from the CRF marginals. These signals are used to gate encoder representations and to construct a global boundary-aware context that conditions each decoding step, thereby enabling more robust decoding under boundary ambiguity. Experimental results demonstrate consistent improvements over state-of-the-art methods. The model achieves exact-match boundary location accuracy of 84.39%, exact segmentation accuracy of 79.34%, and character-level accuracy of 87.79%. It also achieves boundary-level Precision, Recall, and F1 scores of 92.84%, 91.66%, and 92.25%, respectively. These results indicate that uncertainty-aware coupling and structured BiLSTM-CRF supervision improve segmentation performance while maintaining strong boundary detection, thereby enabling more accurate morphological analysis for NLP and digital humanities applications.
PMID:42582249 | PMC:PMC13457376 | DOI:10.3389/frai.2026.1873672
-
2 - DHQ(Digital Humanities Quarterly)
- True Digital Poiesis: A Conceptual Framework and Interface Design for Generative AI in the Humanities
True Digital Poiesis: A Conceptual Framework and Interface Design for Generative AI in the Humanities
-
2 - DHQ(Digital Humanities Quarterly)
- Responsible AI and the Middle Ages: Detecting Historical Toxicity in Medieval Datasets
Responsible AI and the Middle Ages: Detecting Historical Toxicity in Medieval Datasets
-
7 - PubMed
- KannadaLit4NLP: A comprehensive classical kannada literary dataset of Vachanas, Tripadis, and Kagga with scholarly interpretations for natural language processing
KannadaLit4NLP: A comprehensive classical kannada literary dataset of Vachanas, Tripadis, and Kagga with scholarly interpretations for natural language processing
Data Brief. 2026 Jun 19;67:112983. doi: 10.1016/j.dib.2026.112983. eCollection 2026 Aug.
ABSTRACT
This article presents KannadaLit4NLP, a large-scale, machine-readable corpus of Kannada literary texts designed to support natural language processing (NLP) research for a low-resource language. The dataset comprises 24,746 literary verses from three major Kannada literary traditions-Vachanas (11th-19th century), Tripadis (16th century), and Kagga (20th century)-along with 22,369 corresponding interpretations curated from scholarly sources. The corpus captures linguistic, stylistic, and semantic variations across historical periods and literary forms. The dataset was developed through a systematic pipeline that included source identification, digitisation via optical character recognition (OCR), manual verification, and structured annotation. Each entry is organised in a structured format that includes the original verse, metadata (literary form, author, and source), and associated interpretation(s), enabling its use in tasks such as semantic textual similarity, textual entailment, information retrieval, and generative modelling. KannadaLit4NLP addresses the limited availability of culturally grounded Kannada datasets by providing a resource that integrates classical and modern literary content with interpretative annotations. The dataset can facilitate the development and evaluation of NLP models in areas such as semantic understanding, translation, and knowledge representation, while also supporting computational studies of literary and cultural texts. The dataset is made publicly available to encourage further research and reproducibility in Kannada NLP.
PMID:42389175 | PMC:PMC13320459 | DOI:10.1016/j.dib.2026.112983
-
5 - ZfdG (Zeitschrift für digitale Geisteswissenschaften)
- Aiding Provenance Research. A Computer-Assisted Image Retrieval in Auction Catalogs
Aiding Provenance Research. A Computer-Assisted Image Retrieval in Auction Catalogs
-
2 - DHQ(Digital Humanities Quarterly)
- The Naxi Dongba MOOC: A Test Case for Digital Revitalisation of Endangered Writing Systems
The Naxi Dongba MOOC: A Test Case for Digital Revitalisation of Endangered Writing Systems
-
7 - PubMed
- An Alchemical <em>Prima Materia</em> for the Digital Age: Making the Early Modern Latin Alchemical Prints (EMLAP) Dataset
An Alchemical <em>Prima Materia</em> for the Digital Age: Making the Early Modern Latin Alchemical Prints (EMLAP) Dataset
Ambix. 2026 May-Aug;73(2-3):145-174. doi: 10.1080/00026980.2026.2668735. Epub 2026 Jun 25.
ABSTRACT
This article documents the creation of EMLAP (Early Modern Latin Alchemical Prints), a machine-readable corpus of one hundred Latin alchemical printed works produced within the TOME (The Origins of Modern Encyclopaedism, 2023-2025) project. It first situates the corpus within both the digital humanities landscape and the historiography of alchemy, where the availability of reliable machine-readable texts remains limited. It then addresses the challenges of converting early modern Latin printed text into machine-readable format with a high standard of quality. The article argues that producing a high-quality transcribed corpus at scale still requires human scholarly intervention, and that a transcription project must balance the ideal of digital edition standards against the practical constraints of time and resources. The article describes the practical experience of building EMLAP: the selection of the Transkribus platform for AI-powered automatic text recognition, the choice of transcription models, the development of human quality standards to complement automated metrics, and the construction of a computational pipeline to process and enrich the transcriptions. The EMLAP corpus has been made publicly available in open access (Zenodo repository) as well as in the form of a website that offers different search opportunities for researchers.
PMID:42345208 | DOI:10.1080/00026980.2026.2668735
-
5 - ZfdG (Zeitschrift für digitale Geisteswissenschaften)
- Topic Modeling für die Geschichtswissenschaft
Topic Modeling für die Geschichtswissenschaft
-
2 - DHQ(Digital Humanities Quarterly)
- Visions in the Machine: Automated Tagging of the William Blake Archive
Visions in the Machine: Automated Tagging of the William Blake Archive
-
5 - ZfdG (Zeitschrift für digitale Geisteswissenschaften)
- Das konzeptuelle Modell des Zaubermärchens und seine digitale Umsetzung
Das konzeptuelle Modell des Zaubermärchens und seine digitale Umsetzung
-
7 - PubMed
- CHINTEXDB-PERU28: A unique dataset of traditional textile iconographies from Chinchero, Peru for cultural preservation and image recognition
CHINTEXDB-PERU28: A unique dataset of traditional textile iconographies from Chinchero, Peru for cultural preservation and image recognition
Data Brief. 2026 May 10;66:112835. doi: 10.1016/j.dib.2026.112835. eCollection 2026 Jun.
ABSTRACT
This dataset was collected during on-site fieldwork conducted in the district of Chinchero, located in the province of Urubamba, Cusco, Peru, a region internationally recognized for its rich Andean textile tradition rooted in Inca Culture heritage. The dataset comprises high-quality Photographic images of traditional handwoven Andean textile iconographies produced by local artisan communities. These images were captured directly at textile centers where the fabrics are woven, dyed and finished using ancestral techniques measuring authentic representation of colors, textures, and symbolic patterns under natural and controlled conditions. The dataset consists of 1358 images organized into 28 distinct classes, each corresponding to a specific textile iconography characteristic of the Chinchero tradition. The images are provided in a processed and curated format, facilitating organization enables systematic analysis of visual motifs that are often challenging to distinguish due to their intricate geometric patterns and cultural symbolism. The primary reuse potential of this dataset lies in its application to Artificial Intelligence (AI) and Machine Learning (ML) research focused on image classification, pattern recognition, and cultural heritage preservation. Researchers can leverage the dataset to develop and evaluate models capable of identifying and differentiating traditional Andean textile iconographies, addressing the growing difficulty faced by younger generations, local communities, and visitors in recognizing the cultural expressions. Additionally, the dataset supports interdisciplinary research in digital humanities, ethnography, textile studies, and cultural informatics contributing to the documentation and preservation of intangible cultural heritage. By making this dataset publicly available, this work aims to support the development of AI-driven tools for cultural preservation, educational applications, and heritage awareness, while fostering collaboration between researchers, technologists, and local artisan communities to safeguard ancestral knowledge for future generations.
PMID:42220648 | PMC:PMC13217882 | DOI:10.1016/j.dib.2026.112835
-
2 - DHQ(Digital Humanities Quarterly)
- Introduction to Community, Activism, and Innovation in Latin American and Latinx Public Digital HumanitiesEntwodiksyon sou Kominote, Aktivis, ak Inovasyon nan Syans Imanitè Nimerik Piblik an Amerik Latin ak LatinoIntroduction à La Communauté, l’Activisme et l’Innovation dans les Humanités Numériques Publiques Latino-américaines et LatinxIntroducción a la comunidad, el activism y la inovacción en las humanidades digitales públicas latinoamericanas y latinxIntrodução à Comunidade, ativismo e inovação nas humanidades digitais públicas latino-Americanas e latinx
Introduction to Community, Activism, and Innovation in Latin American and Latinx Public Digital HumanitiesEntwodiksyon sou Kominote, Aktivis, ak Inovasyon nan Syans Imanitè Nimerik Piblik an Amerik Latin ak LatinoIntroduction à La Communauté, l’Activisme et l’Innovation dans les Humanités Numériques Publiques Latino-américaines et LatinxIntroducción a la comunidad, el activism y la inovacción en las humanidades digitales públicas latinoamericanas y latinxIntrodução à Comunidade, ativismo e inovação nas humanidades digitais públicas latino-Americanas e latinx
-
2 - DHQ(Digital Humanities Quarterly)
- Fostering Transborder Thinking at the Intersection of Digital-Public Humanities and Border Epistemologies with United Fronteras
Fostering Transborder Thinking at the Intersection of Digital-Public Humanities and Border Epistemologies with United Fronteras
-
2 - DHQ(Digital Humanities Quarterly)
- Social Media as a Digital Humanities Platform? A conversation between Rendering Revolution co-founders, Jonathan Michael Square and Siobhan Meï
Social Media as a Digital Humanities Platform? A conversation between Rendering Revolution co-founders, Jonathan Michael Square and Siobhan Meï
-
2 - DHQ(Digital Humanities Quarterly)
- Baobabs, Networks and Digital Sovereignty: Afro-Brazilian and Indigenous Community Digital Territories as Communitarian DH
Baobabs, Networks and Digital Sovereignty: Afro-Brazilian and Indigenous Community Digital Territories as Communitarian DH
-
2 - DHQ(Digital Humanities Quarterly)
- Documenting the Movement for Mexican American Studies (MAS) in Texas Through Critical Latinx Public Digital Humanities (DH) and Chicanx Feminisms
Documenting the Movement for Mexican American Studies (MAS) in Texas Through Critical Latinx Public Digital Humanities (DH) and Chicanx Feminisms
-
2 - DHQ(Digital Humanities Quarterly)
- Good Women, Mediocre Men: Hierarchy in Narrative and Digital Prosopography