Feminist Digital Humanities: Intersections in Practice RhodyL. M.SchreibmanS. (eds), Urbana-Champaign, IL, University of Illinois Press, 2024, pp. 290. ISBN 9780252047732, £19.95, (eBook).
The Routledge introduction to English Canadian literature and digital humanitiesBarrettP. and RogerS. (eds), New York, Routledge, 2025, pp. 346, £41.99 (P/B). ISBN 9781032331256.
We present a computational analysis of the relationship between stylistic complexity and angry sentiment in Lale Gül’s novel Ik Ga Leven. We begin with the close-reading observation that, in this novel, the level of textual complexity increases when the narrator becomes angrier. This leads to the research questions: “What is the relationship between the level of complexity and the level of anger in this novel? Can this relationship be measured computationally?” To operationalize anger, we use a sentiment dictionary and a machine learning model fine-tuned on angry paragraphs. To operationalize textual complexity, we use Kolmogorov Complexity, a Dutch-language readability tool, and selected textual features. We contextualize the novel’s complexity by comparing it to the LitRiddle corpus of Dutch literature. The analysis indicates a relation between levels of anger and complexity in the novel. Our article thus illustrates the unusual position of Gül’s work in the contemporary Dutch literary landscape and illuminates the relationship between literary style and intradiegetic emotion, particularly anger. We also reflect on the uses and limitations of the tools, and on the challenge of combining sociocultural, stylistic, and content-level analysis of literature in computational literary studies. Finally, we identify some areas for future research.
Swearing involves the use of words interpreted as obscene, derogatory, or offensive, and so can have a significant societal impact in the digital world. While it is widely acknowledged that the functions of swearing vary according to who is doing it and in what context, other dimensions of variability in online swearing remain to be fully explored. This article argues that we need to move beyond purely distributional text analytics of inferred social variables in analyzing online swearing if we are to systematically examine the social impacts of swearing in born-digital data. Using computational text analysis and network analysis methods, three large datasets representing snapshots of English Twitter/X communication in Australia, the United Kingdom, and the United States are examined through four interrelated research questions: (1) Do rates of vulgarity differ significantly across these three English-speaking regions and across different times of day? (2) To what extent do users employ non-standard orthographic and typographic variations to obscure vulgar expressions? (3) How does vulgarity correlate with users’ positions within social networks, specifically their network integration and follower counts? (4) What patterns emerge when these dimensions of variability are examined together? Findings demonstrate that while it remains important to examine who swears online, it is also critical to examine other dimensions of variability, including where, when, how, and with whom swear words are used. The implications of extending our understanding of these different dimensions of variability in born-digital data for studies of online swearing and the data-intensive humanities more generally are also discussed.
This study employs a deep learning-based topic modeling approach to analyze research topics and evolution trends in digital humanities, offering valuable insights for pertinent research and practices. Literature data were collected from the Web of Science database, with a total of 5,901 valid records analyzed using the BERTopic model and dynamic topic modeling techniques. The study identifies seven major research topics in digital humanities, including digital transformation and interdisciplinary innovation; literary and philosophical studies in digital humanities; cultural heritage, archaeology and semantic technologies; digital pedagogy and artificial intelligence (AI)-enhanced education; social science methods and digital research practice; digital scholarship, libraries, and research infrastructure; and computational literary and linguistic analysis. These topics’ evolution reflects the field’s maturation and the growing prominence of technological advancements, particularly AI and digital pedagogy. The findings underscore how integrating AI, digital tools, and interdisciplinary methods drives the future of digital humanities research. This study enhances the BERTopic model by integrating Sentence-Bidirectional Encoder Representations from Transformers-based embeddings, MultiDimensional Scaling for dimensionality reduction, KMeans clustering, and Log-Likelihood Ratio weighting to enhance topic coherence and accuracy. It reveals the growth of key research areas such as AI-driven pedagogy and digital transformation in digital humanities. The innovation involves utilizing this adjusted model to gain profound insights into the dynamic evolution of research topics, elucidating the influence of digital tools and AI on the field.
Recent research has stressed the importance of recognizing the political nature of the digitization of primary sources, but how much do we really know about the choices, labour, and lacunae embedded within and across our preferred databases? Following Tom Nesmith’s redefinition of provenance as encompassing a much fuller view of records’ histories and contexts, this article argues that articulating the digital provenance of our sources can unlock and contextualize data ethics and informed use, for both ourselves and our students. It begins by defining digital provenance and its relationship to digital literacy, before undertaking a close reading of the UK digitized archive sector with two brief case studies, centring on the growth of commercial content providers. It explores the consequences of this shift in primary source provision and associated barriers to developing a full understanding of the technical histories and contingent present underlying our digital archival landscape. The second part takes a reflective view on the difficulty of doing this slow, careful work in a sector under intense pressure. The article concludes with recommendations for future study and a proposal that we seek inspiration from the critical turn to the archives seen in the 1990s and 2000s to naturalize questions of digital provenance, both to better understand our sources and as a way of learning about the wider materialities and contingencies of digitized data in today’s world.
This study examines ChatGPT’s ability to learn and replicate individuals’ unique writing styles. Using a one-shot training method, short texts from 300 authors of product reviews were used to train ChatGPT to generate new texts mimicking each author’s style. Likelihood ratio-based forensic text comparison experiments were conducted on same-author and different-author text pairs, involving human-written, machine-generated, and mixed texts. The findings indicate that replicating writing styles remains a significant challenge for ChatGPT. Linguistic analysis revealed that while both humans and machines exhibited preferred words and expressions, these preferences were more strongly associated with the machine.
The design and implementation of categories is an important part of digital humanities research, for which theoretical discussion is still relatively underdeveloped. This article presents an instructive case study from the Theatronomics project: the categorization of some 159,000 items of theatrical expenditure recorded 1732–1809. It explores the key considerations involved in categorizing financial data for a database; explains why the team adopted a pragmatic–explorative method, highly manual in nature; illustrates the major problems confronted during the process; and offers suggestions for further thinking on, and enacting of, digital humanities categorization. It advances three main arguments. First, categorization is not a ‘natural’ or ‘common-sense’ process. Scholars working on databases should consider the multiple different methods of categorization available, and the various and sometimes contradictory principles involved in categorization, before designing a method that aligns with their project’s aims, sources, and resources. Second, scholars should assess the uncertainty in their dataset, and consider how to represent it in their categorization model. Third, a model that involves multiple layers of categories can be effective in solving some of the key problems involved in categorization. This article also explains how Theatronomics’s expenses data, categorized in accordance with those three arguments, contributes to scholarship on Georgian theatre, Georgian business, and Georgian London.
Traditionally, the comparison of textual witnesses is achieved through manual collation. This study introduces a computational approach adapting methods from bioinformatics: pairwise sequence alignment and dimensionality reduction, to measure and visualize textual relationships across a corpus. We apply global (Needleman–Wunsch) and local (Smith–Waterman) alignment algorithms directly to character strings, generating quantitative similarity scores which are then represented through t-distributed stochastic neighbour embedding. We also test a language-specific modification of the Needleman–Wunsch algorithm on two manuscripts in our corpus. Unlike automated collation methods that aim for semantic accuracy, this approach focuses on corpus-wide similarity patterns. The test corpus contains twenty-four medieval manuscripts of the Liturgical Targum, preserved in Jewish festival prayer books. Previous (manual) philological analysis had already identified two textual families among the Targum units within these prayer books. Our computational method successfully and independently replicates these families and reflects the overall coherence of the corpus. Crucially, it enabled new insights overlooked in the manual study: the new identification of a textual subgroup and the discovery that two manuscripts were written by the same scribe. Local alignment proves effective for identifying the closest textual parallels of a fragmentary manuscript. The language-specific alignment modification test on two manuscripts indicates improved alignment algorithm performance for Hebrew script. This article demonstrates that combining pairwise sequence alignment with dimensionality reduction is a powerful exploratory tool for engaging with a text corpus. The method requires only accurate transcriptions to produce maps of textual relationships that can guide subsequent detailed collation and interpretation.
This study investigates the interrelationship between syntactic complexity [as measured by dependency distance (DD)] and clause length (specifically those less than ten characters) in Chinese Hua’er folk songs. The findings demonstrate that the right truncated modified Zipf–Alekseev model effectively describes the inverse relationship between DDs and clause lengths. An examination of the internal self-organization mechanism of the Hua’er system reveals that as sentence length increases, the mean dependency distance (MDD) decreases. This observation provides novel evidence for MDD particularly in short clauses less than ten characters. Through the application of writer’s view, it is observed that the syntactic frequency structure exhibits self-organization, tending towards the golden section. The results yield that longer sentences are not necessarily difficult in Chinese Hua’er folk songs.
This study presents the design, implementation, and evaluation of the Greater Mekong Subregion (GMS) Ethnic Groups Knowledge Graph (KG) and its accompanying web application–an innovative, culturally inclusive platform for modelling the ethnographic diversity of 375 ethnic groups across Thailand, Laos, Myanmar, Cambodia, Vietnam, and China. Addressing the limitations of traditional Knowledge Organization Systems (KOS), the project integrates semantic technologies, local ethnographic data, and community-informed classification frameworks to represent complex relationships involving language, religion, cultural practices, geography, and historical migration of ethnic groups in the GMS. This study employs structured data transformation, entity–relationship modelling, Neo4j-based graph construction, and interactive visualization using React.js and D3.js. Evaluation results, based on extrinsic performance benchmarks and domain-specific expert validation, demonstrate substantial improvements in usability, task efficiency, and data interpretability compared with conventional databases. The findings support using knowledge graphs for ethically grounded, context-sensitive knowledge infrastructures that foster epistemic justice, cultural sustainability, and inclusive access to ethnographic data. In addition, this work contributes a replicable digital humanities model that blends KOS, knowledge management, and semantic web principles to empower interdisciplinary research and community engagement.
Existing research on the literary generation capabilities of artificial intelligence (AI) has mostly centered on the English-language context. While this has certainly provided an important reference, it may also obscure the differences across language systems. This article takes short Chinese online science fiction as the evaluation object and, based on the perspective of prompt engineering and multi-model comparison, constructs a five-dimensional quantitative evaluation system to systematically test and analyze the story generation abilities of four general Chinese large language models (LLMs): DeepSeek, Doubao, KIMI, and Wenxin Yiyan. By adjusting structured prompts and temperature parameters, combined with a human-AI collaborative scoring mechanism, it evaluates their performance in aspects such as creative setting, character construction, structural coherence, stylistic fluency, and philosophical depth. The results show that DeepSeek exhibits significant advantages. Further cyberpunk rewriting experiment on The Legend of the White Snake shows that DeepSeek can, to a certain extent, recode specific cultural imagery. This article argues that the literary generation capability of AI should not be evaluated merely by plot completeness or linguistic fluency; rather, it should also be examined in terms of whether it can enter the narrative paradigms and cultural expectations of a specific linguistic community. The evaluation of Chinese local LLMs therefore has independent methodological significance.
We evaluate six flagship Large Language Models (LLMs) (spanning prior- and current-generation systems across the ChatGPT, Claude, and Gemini families) on Pali-to-English translation for selected suttas from the Majjhima Nikaya. We introduce an automated pipeline that segments scripture into context-linked JSON, enforces format-faithful outputs via prompt engineering, and closes the loop with artificial intelligence (AI)-based quality assessment using the Generative Pretrained Transformer (GPT) Estimation Metric Based Assessment (GEMBA) in both no-reference (NR) and human-reference modes, alongside METEOR, TER, and BERTScore. Results show that modern LLMs can produce high-quality translations of ancient Pali: GPT-5 is the most consistent on NR-GEMBA (fewest sub-80 rated lines), while Claude 4 Sonnet aligns most closely with the human reference on traditional metrics; a compact head-to-head matrix further reveals a line-wise edge for Claude 3.5 Sonnet in pair-wise wins. Inter-evaluator agreement is moderate-to-high. Overall, the workflow can broaden access to Buddhist scripture while leaving final interpretive authority with human scholars, and it provides a scalable template for evaluating AI translations of ancient texts.
Large language models are increasingly capable of producing creative texts, yet most studies on AI-generated poetry focus on English—a language that dominates training data. In this article, we examine the perception of AI- and human-written Czech poetry. We ask if Czech native speakers are able to identify it and how they aesthetically judge it. Participants performed at chance level when guessing authorship (45.8 per cent correct on average), indicating that Czech AI-generated poems were largely indistinguishable from human-written ones. Aesthetic evaluations revealed a strong authorship bias: when participants believed a poem was AI-generated, they rated it as less favorably, even though AI poems were in fact rated equally or more favorably than human ones on average. The logistic regression model uncovered that the more the people liked a poem, the less probable was that they accurately assign the authorship. Familiarity with poetry or literary background had no effect on recognition accuracy. Our findings show that AI can convincingly produce poetry even in a morphologically complex, low-resource (with respect to the training data of AI models) Slavic language such as Czech. The results suggest that readers’ beliefs about authorship and the aesthetic evaluation of the poem are interconnected.
There has been extensive research on the phenomenon of code switching, meaning the use of two or more languages or language varieties, within texts. Until recently, most code switching studies in the digital humanities have tagged the mixed languages manually, as automatic language identification methods have so far performed too unreliably to be useful, although research on automatic methods is growing. This paper aims to improve methods for identifying snippets of a second language in historical and literary texts by evaluating three automatic methods of increasing complexity: dictionary method, dedicated language identification packages, and fine-tuned large language models (LLMs). We evaluate the methods on the test case of manually tagged French snippets in English literary texts, and report that fine-tuned LLMs performed with the highest overall detection rates in our experiment with different models, although language pair-specific methodological questions remain. We have published our code and fine-tuned LLMs to assist research on this language pair, and these methods may be expanded to more language pairs and broader applications in the future. Finally, we report observations of French usage in the English literary texts in our corpus, dating 1814–1920.
This article repositions Burrows’s Delta as a flexible family of distance measures for exploratory and unsupervised stylometry, where interpretability and stability are as important as predictive accuracy. We introduce two probabilistic extensions, Rank-Turbulence Delta and Jensen–Shannon Delta, by reinterpreting uncentred standardized word-frequency vectors as non-negative representations that can be normalized into probability distributions and compared using information and rank-based divergences. Building on this representation, we derive token-level decompositions for classical Delta, Cosine Delta, Jensen–Shannon Delta, and Rank-Turbulence Delta, enabling each document distance to be expressed as a sum of interpretable lexical contributions. We evaluate the proposed framework in clustering and nearest-neighbour attribution experiments on four literary corpora in English, German, French, and Russian, including an extended Russian benchmark drawn from SOCIOLIT (639 works by 89 authors, XVIII–XXI centuries). To substantiate interpretability, we additionally verify that top-contributing tokens are stable under small perturbations (mfw variation and bootstrap resampling) and that removing these tokens reduces inter-author separation in the expected direction. Overall, the results show that probabilistic and rank-based geometries extend Delta without abandoning its core intuition, while providing a reproducible bridge between aggregate distances and philologically meaningful lexical signals.
This study investigates the relationships among emotional arcs, thematic topics, and narrative arcs in a corpus of 100 English fairy tales drawn from Project Gutenberg. Sentence-level sentiment scores were computed using Flair and then aggregated to construct emotional arcs, which captured consistent patterns of affective change throughout each narrative. Thematic topics were extracted by applying Latent Dirichlet Allocation to TF-IDF vectors and were subsequently validated using perplexity, coherence metrics, and expert review. Narrative arcs were annotated based on Freytag’s five-stage model, which provided a structural framework to trace story progression from exposition to resolution. The analysis demonstrates that: (1) fairy tales exhibit six distinct shapes of emotional arcs; (2) a reciprocal relationship exists between these emotional arcs and the thematic topics of “Fate” and “Growth”; (3) emotional arcs function as a covert driving force that underlies the development of narrative arcs.
Translation invisibility is a common issue when working with bibliographic translation data from national libraries (Teichmann and Roman 2024) and tools to visualize the network of transfer still need to be developed for the Germanophone context. Johan Heilbron proposed a network model for translated fiction which maps the relationships between central and peripheral languages. However, it has not been explored which authors connect linguistic communities. In this paper, we build on Heilbron’s center-periphery language model to investigate the following questions: Do authors form a distinct group that connect languages? Which authors connect peripheral languages to central languages and which authors are unique to language groups? By applying a network model to the bibliographic data of 31,519 translated titles originally published in German, representing 3,986 authors that have been translated into eighty-one languages extracted from the German National Library (DNB), this paper presents first-time network visualizations of the DNB’s collection of translations. By applying a community detection algorithm to the network graph, we found two distinct groups of authors: representatives of the canon that connect the central to the peripheral languages, and groups of authors that stay within central and peripheral languages, whose translations are highly specialized in certain target languages and genres. Hence, we argue that we can use Heilbron’s center-periphery language network model to identify authors of German fiction who played a crucial role in pushing the canon beyond its national boundaries by means of their translations, and in turn, make visible the network of translational transfer between languages.
This study analyses visitor sentiment at the Chinese Language Museum using SnowNLP sentiment analysis and LDA topic modelling applied to online reviews. Findings reveal overwhelmingly positive evaluations (82.3 per cent), driven by cultural identity, heritage value, and interactive experiences. Neutral feedback concerns regional tourism appeal and functional utility, while negative critiques focus on service gaps and mismatched expectations. By critically examining the insights and limitations of such digital traces, the research proposes a four-dimensional influence framework (attraction, transformation, guarantee, and regulation), offering a theoretically informed, visitor-centred approach for museums to enhance experiential design and cultural communication through perceptual analysis.