❌

阅读视图

A knowledge graph dataset of medieval and renaissance geographical works

Data Brief. 2026 Sep 3;69:113224. doi: 10.1016/j.dib.2026.113224. eCollection 2026 Aug.

ABSTRACT

Medieval and Renaissance Latin geographical works constitute a major source for understanding how space, places, and territories were described and conceptualised in pre-modern Europe. However, information about these works, their manuscript transmission, and the places they mention remains dispersed across catalogues, archives, and specialised scholarship. Here we present the IMAGO knowledge graph, a semantically structured dataset representing 343 Latin geographical works written between the 6th and the 15th centuries. The dataset integrates curated information provided by domain experts, including authors, works, manuscripts, printed editions, libraries, literary genres, and mentioned places. Data were initially collected in tabular form and subsequently enriched through semi-automatic reconciliation with external authority sources such as Wikidata and the MIRABILE digital archive. Domain experts further expanded the dataset using a dedicated annotation tool. The curated data were transformed into an OWL 2 DL knowledge graph aligned with the IMAGO ontology and published following FAIR and Linked Open Data principles. The knowledge graph was validated through automated reasoning, expert review, and query-based evaluation. The resulting dataset enables systematic exploration of textual, bibliographic, and spatial relationships within medieval and Renaissance geographical literature and supports reuse in historical, philological, and digital humanities research.

PMID:42757014 | PMC:PMC13583948 | DOI:10.1016/j.dib.2026.113224

  •  

Sustainability, Sound, and Sites: Case Studies in Reconciling Minimal Computing Practices with the Preservation and Presentation of Audio Media

How can we design sustainable web projects involving audio? While minimal computing workflows and the Endings Project provide an excellent general foundation for developing sustainable sites, most minimal projects are geared towards images and text with little regard to the complexities of audio. Additionally, minimal computing’s insistence on pruning dependencies can clash with complex [...]

  •  

The style of scorn: computationally operationalizing the relationship between anger and complexity in “Ik Ga Leven” by Lale Gül

Abstract
We present a computational analysis of the relationship between stylistic complexity and angry sentiment in Lale Gül’s novel Ik Ga Leven. We begin with the close-reading observation that, in this novel, the level of textual complexity increases when the narrator becomes angrier. This leads to the research questions: “What is the relationship between the level of complexity and the level of anger in this novel? Can this relationship be measured computationally?” To operationalize anger, we use a sentiment dictionary and a machine learning model fine-tuned on angry paragraphs. To operationalize textual complexity, we use Kolmogorov Complexity, a Dutch-language readability tool, and selected textual features. We contextualize the novel’s complexity by comparing it to the LitRiddle corpus of Dutch literature. The analysis indicates a relation between levels of anger and complexity in the novel. Our article thus illustrates the unusual position of Gül’s work in the contemporary Dutch literary landscape and illuminates the relationship between literary style and intradiegetic emotion, particularly anger. We also reflect on the uses and limitations of the tools, and on the challenge of combining sociocultural, stylistic, and content-level analysis of literature in computational literary studies. Finally, we identify some areas for future research.
  •  

Swearing online: exploring dimensions of variability in born-digital data

Abstract
Swearing involves the use of words interpreted as obscene, derogatory, or offensive, and so can have a significant societal impact in the digital world. While it is widely acknowledged that the functions of swearing vary according to who is doing it and in what context, other dimensions of variability in online swearing remain to be fully explored. This article argues that we need to move beyond purely distributional text analytics of inferred social variables in analyzing online swearing if we are to systematically examine the social impacts of swearing in born-digital data. Using computational text analysis and network analysis methods, three large datasets representing snapshots of English Twitter/X communication in Australia, the United Kingdom, and the United States are examined through four interrelated research questions: (1) Do rates of vulgarity differ significantly across these three English-speaking regions and across different times of day? (2) To what extent do users employ non-standard orthographic and typographic variations to obscure vulgar expressions? (3) How does vulgarity correlate with users’ positions within social networks, specifically their network integration and follower counts? (4) What patterns emerge when these dimensions of variability are examined together? Findings demonstrate that while it remains important to examine who swears online, it is also critical to examine other dimensions of variability, including where, when, how, and with whom swear words are used. The implications of extending our understanding of these different dimensions of variability in born-digital data for studies of online swearing and the data-intensive humanities more generally are also discussed.
  •  

Topic mining and evolution analysis of digital humanities research based on the BERTopic model

Abstract
This study employs a deep learning-based topic modeling approach to analyze research topics and evolution trends in digital humanities, offering valuable insights for pertinent research and practices. Literature data were collected from the Web of Science database, with a total of 5,901 valid records analyzed using the BERTopic model and dynamic topic modeling techniques. The study identifies seven major research topics in digital humanities, including digital transformation and interdisciplinary innovation; literary and philosophical studies in digital humanities; cultural heritage, archaeology and semantic technologies; digital pedagogy and artificial intelligence (AI)-enhanced education; social science methods and digital research practice; digital scholarship, libraries, and research infrastructure; and computational literary and linguistic analysis. These topics’ evolution reflects the field’s maturation and the growing prominence of technological advancements, particularly AI and digital pedagogy. The findings underscore how integrating AI, digital tools, and interdisciplinary methods drives the future of digital humanities research. This study enhances the BERTopic model by integrating Sentence-Bidirectional Encoder Representations from Transformers-based embeddings, MultiDimensional Scaling for dimensionality reduction, KMeans clustering, and Log-Likelihood Ratio weighting to enhance topic coherence and accuracy. It reveals the growth of key research areas such as AI-driven pedagogy and digital transformation in digital humanities. The innovation involves utilizing this adjusted model to gain profound insights into the dynamic evolution of research topics, elucidating the influence of digital tools and AI on the field.
  •  

Digitized archives, content providers, and slow scholarship: why archival researchers should care about digital provenance

Abstract
Recent research has stressed the importance of recognizing the political nature of the digitization of primary sources, but how much do we really know about the choices, labour, and lacunae embedded within and across our preferred databases? Following Tom Nesmith’s redefinition of provenance as encompassing a much fuller view of records’ histories and contexts, this article argues that articulating the digital provenance of our sources can unlock and contextualize data ethics and informed use, for both ourselves and our students. It begins by defining digital provenance and its relationship to digital literacy, before undertaking a close reading of the UK digitized archive sector with two brief case studies, centring on the growth of commercial content providers. It explores the consequences of this shift in primary source provision and associated barriers to developing a full understanding of the technical histories and contingent present underlying our digital archival landscape. The second part takes a reflective view on the difficulty of doing this slow, careful work in a sector under intense pressure. The article concludes with recommendations for future study and a proposal that we seek inspiration from the critical turn to the archives seen in the 1990s and 2000s to naturalize questions of digital provenance, both to better understand our sources and as a way of learning about the wider materialities and contingencies of digitized data in today’s world.
  •  

ChatGPT’s ability to imitate writing styles: an analysis guided by forensic text comparison

Abstract
This study examines ChatGPT’s ability to learn and replicate individuals’ unique writing styles. Using a one-shot training method, short texts from 300 authors of product reviews were used to train ChatGPT to generate new texts mimicking each author’s style. Likelihood ratio-based forensic text comparison experiments were conducted on same-author and different-author text pairs, involving human-written, machine-generated, and mixed texts. The findings indicate that replicating writing styles remains a significant challenge for ChatGPT. Linguistic analysis revealed that while both humans and machines exhibited preferred words and expressions, these preferences were more strongly associated with the machine.
  •  

Counting costs: the categorization of eighteenth-century theatrical expenditure for the Theatronomics database

Abstract
The design and implementation of categories is an important part of digital humanities research, for which theoretical discussion is still relatively underdeveloped. This article presents an instructive case study from the Theatronomics project: the categorization of some 159,000 items of theatrical expenditure recorded 1732–1809. It explores the key considerations involved in categorizing financial data for a database; explains why the team adopted a pragmatic–explorative method, highly manual in nature; illustrates the major problems confronted during the process; and offers suggestions for further thinking on, and enacting of, digital humanities categorization. It advances three main arguments. First, categorization is not a ‘natural’ or ‘common-sense’ process. Scholars working on databases should consider the multiple different methods of categorization available, and the various and sometimes contradictory principles involved in categorization, before designing a method that aligns with their project’s aims, sources, and resources. Second, scholars should assess the uncertainty in their dataset, and consider how to represent it in their categorization model. Third, a model that involves multiple layers of categories can be effective in solving some of the key problems involved in categorization. This article also explains how Theatronomics’s expenses data, categorized in accordance with those three arguments, contributes to scholarship on Georgian theatre, Georgian business, and Georgian London.
  •