阅读视图

A digital humanities approach to Chinese export watercolours: a case study on the Victoria and Albert Museum Collection

Abstract
Chinese Export Watercolours (CEW) is a type of painting created for export to European and North American markets during the eighteenth and nineteenth centuries, which catered to Western consumers’ aesthetic tastes. These paintings blend Chinese and European painting techniques, depict detailed traditional Chinese occupations, customs, plants, animals, architectures, and landscapes of the late Qing Dynasty, and they form a unique artistic style. This article introduces collaborative research between the Victoria and Albert Museum (V&A) and University College London. By focusing on the V&A CEW collection, this article presents a new digital humanities research model for studying Chinese export art, which comprises different parts. First, based on the acquisition archives and digitized images, the study analyses its catalogue metadata and acquisition records to trace the collecting history and documented provenance. Secondly, it classifies 2,350 paintings by themes using deep learning methods. It then conducts semi-automatic image annotation and in-depth analysis to explore the iconographies and their evolution in history. The findings reveal a changing iconography distribution before and after 1840 within the V&A holdings, offering new, collection-based insights into Sino-British cultural interactions and global trade history during this period. The model not only contributes to in-depth analysis of the V&A CEW collection but also provides a new way for future research on Chinese export art and beyond.
  •  

Incidence and evidence: early modern stress patterns in stylometry classification

Abstract
Stylometry—the science of measuring writing styles—typically relies on the counting of word or letter frequencies to offer judgements on the authorship of anonymous texts. However, a number of 20th-century tables remain prominently in use by authorship scholars. Most significant among these are Philip Timberlake’s 1931 The Feminine Ending in English Blank Verse, and Ants Oras’s 1960 Pause Patterns in Elizabethan and Jacobean Drama. Timberlake’s study counted how frequently early modern verse lines ended on an additional, unstressed syllable, while Oras’s study counted the pause positions in lines of verse. The works of Timberlake and Oras occupy a contentious place in the study of authorship, in which they are occasionally framed as a safer alternative to modern methods. Richard Proudfoot and Nicola Bennett, for example, cited Timberlake’s study as one that could help them avoid the ‘controversy about the relative value and reliability of different ‘non-traditional’ methods’ in authorship studies (Proudfoot and Bennett). This article examines the evidentiary value of Oras and Timberlake’s data when applied to machine-learning stylometry tests. Results from this research suggest that Oras and Timberlake’s data lead to an extremely marginal increase in accuracy for stylometry experiments, and do not justify their use over modern approaches, such as the counts of the most frequently occurring words.
  •  

TEI-encoded image-based editions of Middle English religious poetry: the case study of Prick of Conscience

Abstract
The past century has seen the publication of numerous editions of various mediaeval texts. These are typically graphemic critical editions, which attempt to establish the text’s archetype by synthesizing content drawn from different manuscript sources. This proliferation has been accelerated by the advancement of digital humanities and the subsequent increase in the digitization of library archives. Given this context, this article presents conventions for generating a distinct type of edition: the graphetic and diplomatic image-based edition. This format ensures fidelity to the original source and is suitable for detailed study either of the language or the document itself. While XML-TEI is a common format for producing digital editions, its standard conventions may not always adequately represent the specific features encountered by mediaevalists in their manuscript source. Consequently, this article aims to provide guidance in the production of an XML-TEI edition that addressed challenges related to tagging different manuscript characteristics. These challenges include encoding features such as Middle English specific letters and symbols, manuscript illumination, and elements related to mediaeval scribal practices (e.g. additions or deletions). These conventions were developed following the TEI-P5 guidelines for editions (version 4.10.2) and were applied specifically to edit the Prologue of the poem Prick of Conscience. The study draws upon five different manuscripts: MS D.5, English MS 90, MS. e Mus. 76, CCA-DCc/LitMs/D/13, and MS Dd.12.69.
  •  

Multimodal RAG for cultural heritage: a technical exploration based on Jino traditional knowledge

Abstract
This article presents a Multimodal Retrieval-Augmented Generation (RAG) system for the digital preservation of traditional knowledge (TK) from the Jino ethnic group in China, a small indigenous community whose knowledge is primarily transmitted through oral narratives, ritual practices, and place-based ecological experience. The system integrates text, audio, and image data, using the m3e-base model for embedding generation and Facebook AI Similarity Search for semantic search. A Flask backend supports cross-modal queries, with OpenAI models handling keyword extraction and generation. Deployed via Docker and Cloudflare Tunnel, the system is embedded in a WordPress interface for public access. Beyond implementation, the study addresses challenges in data collection, intellectual property, and cultural authenticity. Results indicate that AI-driven multimodal retrieval can support sustainable TK transmission and inform future digital heritage infrastructures.
  •  

Climax chapter recognition of a Chinese novel based on plot fluctuation of a chapter and sentiment changing of main character

Abstract
The climax chapter of a Chinese novel is the part where the conflict develops to be tense and sharp. Such a chapter determines the fate of the main character, the development of the plot, and the end of the story, which makes it crucial to a Chinese novel. How to quickly and accurately recognize the climax chapter is important to attract readers’ interest. The climax chapter of a Chinese novel has two distinctive characteristics: dramatic development of plot fluctuation and sentiment changing of the main character. Therefore, this paper tries to propose a model to recognize the climax chapter, which both considers the plot fluctuation of a chapter and the sentiment changing of the main character on the corpus The Legend of the Condor Heroes (射雕英雄传), based on a series of deep learning models. The plot fluctuation of a chapter and the sentiment changing of the main character can be, respectively, depicted as plot fluctuation curve of a chapter and sentiment changing curve of the main character. The chapter difference curve can be drawn by combining the above two curves and based on which the novel climax chapter can be determined. Comparative experimental results on the experimental corpus show that, compared to existing models, the performance of the model proposed in this paper is much better, its F1 value arriving at 87.50 per cent. Ablation experiment results show that the F1 value of the model considering sentiment changing of main character is 9.41 per cent higher than that of the model considering the plot fluctuation of a chapter.
  •  

Decoding AI authorship: can LLMs truly mimic human style across literature and politics?

Abstract
Amidst the rising capabilities of generative AI to mimic specific human styles, this study investigates the ability of state-of-the-art large language models (LLMs), including GPT-4o, Gemini 1.5 Pro, and Claude Sonnet 3.5, to emulate the authorial signatures of prominent literary and political figures: Walt Whitman, William Wordsworth, Donald Trump, and Barack Obama. Utilizing a zero-shot prompting framework with strict thematic alignment, we generated synthetic corpora evaluated through a complementary framework combining transformer-based classification (BERT) and feature-based machine learning (XGBoost). Our methodology integrates Linguistic Inquiry and Word Count (LIWC) markers, perplexity, and readability indices to assess divergence between AI-generated and human-authored text. Results demonstrate that AI-generated mimicry remains highly detectable, with XGBoost models trained on a restricted set of eight stylometric features achieving accuracy comparable to high-dimensional neural classifiers. Post-hoc feature importance analysis indicates that perplexity is the most influential discriminative metric, suggesting systematic differences in the distributional regularity of AI outputs relative to the greater variability observed in human writing. While LLMs exhibit distributional convergence with human authors on low-dimensional heuristic features, such as syntactic complexity and readability, they do not yet fully replicate the nuanced affective density and stylistic variance inherent in the human-authored corpus. By isolating measurable statistical divergences in current generative mimicry, this study provides a structured benchmark for LLM stylistic behavior and offers insights for authorship attribution in digital humanities and social media contexts.
  •  

Lemmatization of long unit words of historical Japanese

Abstract
Lemmatization plays a crucial role in digital humanities research, as it is essential for identifying the canonical forms of words, especially for morphologically rich languages such as Japanese. This article examines methods for estimating lemmas of long unit words (LUWs) in historical Japanese documents, which cannot be estimated by a dictionary-based method. After validating the estimation accuracy using an annotated corpus from the Heian to Muromachi periods (C.E. 794–1573), the results show that lemmas can be estimated with an accuracy of over 90 per cent for different surface forms. This suggests that the method can support and streamline manual annotation tasks. Additionally, it demonstrates that the estimation can be directly applied to digital humanities research. Specifically, by performing hierarchical clustering on unannotated Edo period (C.E. 1603–1868) documents, it became possible to focus more on stylistic features in the clustering process using the estimated LUW lemmas. As a result, a valid clustering outcome was achieved, reflecting reasonable classifications based on factors such as the author and the writing period.
  •  

Computational curation of traditional martial arts: from archives to systems

Abstract
Intangible Cultural Heritage (ICH) poses distinct challenges for documentation and digital representation, as it is expressed through embodied practices and ideologies shaped by complex histories. Recent advances in digital tools have enabled richer modes of recording and opened new possibilities for studying how ICH knowledge may be preserved, experienced, and transmitted. Addressing these opportunities requires methodologies that reconcile the dynamic nature of living practices with the constraints of traditional archival systems, particularly in relation to accessing and analysing knowledge embedded in archival data. This work responds by proposing a systems thinking approach applied to traditional martial arts (TMA), implemented through computational experimentation that synthesizes three critical aspects: encoding embodied practice within its contextual meanings; translating humanities scholarship into viable representational models; and devising interaction strategies to facilitate engagement with the encoded knowledge. Drawing on the strengths and gaps of the Hong Kong Martial Arts Living Archive with respect to knowledge transmission, this research combines theoretical inquiry with empirical studies to examine the multifaceted episteme of TMA, which informs the development of the Martial Arts Ontology (MAon), formalizing key TMA concepts. The coupling of MAon-based annotation with deep-learned motion embeddings encodes TMA knowledge from multimodal materials into operable data that carry the semantics of texts, acts, and minds. These encodings support the development of an interactive knowledge system for informative and exploratory learning, while the process informs a computational curation pipeline in which interdisciplinary researchers act as translators and operators, mediating the culturally meaningful datafication of embodied knowledge archives.
  •  

Ethical dilemmas in artificial intelligence-generated art: authorship, ownership, and the blurring of creative boundaries

Abstract
The aim of this theoretical study was to examine the ethical dilemmas related to the use of generative artificial intelligence (AI) in visual art, with a focus on authorship, ownership, and the shifting boundaries of creativity. The research was conducted in 2024–5 using methods of conceptual analysis, legal comparison, and interdisciplinary synthesis. A total of forty-three peer-reviewed sources were analysed across the fields of art theory, cultural philosophy, intellectual property law, and cognitive aesthetics. The results show that authorship in AI-generated art is increasingly viewed as distributed. While primary responsibility is attributed to human users—through prompt design, selection, and interpretation—there is growing recognition of algorithmic systems as co-contributors. Legal frameworks in six jurisdictions (Canada, the UK, the USA, Germany, Japan, and China) remain fragmented: most require human input for copyright protection, yet differ in how they define ‘authorship’. Analysis of empirical findings from existing studies confirms that audience perception is shaped by authorship attribution: artworks known to be created by humans received significantly higher aesthetic ratings (mean = 5.72, AI = 4.99, P < .001). At the same time, the perceived originality of AI outputs led to divided judgements about their legal and cultural legitimacy. In conclusion, the study highlights that AI enhances formal production but lacks intentionality, emotion, and cultural awareness. These findings are relevant for updating legal definitions, designing transparent AI tools, and creating educational programs on human–AI collaboration in creative fields.
  •  

Deep learning-based classification of Yinshan rock art and future research directions

Abstract
Rock art is universally recognized as a significant cultural heritage. However, resources dedicated to its research remain insufficient, leading to a lack of sustained attention and conservation efforts. As immovable cultural relics, numerous rock art sites across China and the globe have yet to be discovered, documented, studied, or made publicly accessible, with fundamental classification work still pending. And a substantial amount of rock art is located in mountainous areas, gullies, or on cliffs, making on-site manual identification and classification challenging. Furthermore, even when working with collected image datasets, the task of applying objective and consistent classification standards across large volumes of materials is notoriously difficult. Emerging digital humanities methods facilitate the automation and intellectualization of certain research steps, providing relatively objective, convenient, and efficient basic classification measures. This study focuses on the Yinshan rock art, employing deep learning algorithms to conduct a classification experiment. The results show a test set classification accuracy of approximately 80% and an overall F1-score exceeding 78%, validating the feasibility of digital humanities approaches in rock art studies. Future applications of digital humanities in this field promise to be even more diverse, encompassing tasks such as object detection and recognition in rock art, and the construction of knowledge graphs. These methods can offer new perspectives for the preservation and cultural transmission of rock art, thereby enriching research on Chinese rock art.
  •  

Link visions together: visualizing geographies of late Qing and Republican China

Abstract
The integration of spatial methodologies into historical research has reshaped traditional narratives by foregrounding geography as both an analytical framework and an epistemological lens. This paper combines Historical GIS (HGIS) with a spatial epistemological approach to investigate how heterogeneous historical sources can be aligned and interpreted through space. Using the CHMap and LoGaRT platforms, this study spatializes diverse materials, including local gazetteers, land survey maps and IIIF-based open map resources, and examines two contrasting case studies: the evolution of riverine sandbanks and flood processes along the Jingjiang section of the Yangtze River, and the spatiotemporal distribution of the material embodiments of Confucian learning in Jiaxing, Zhejiang Province. These cases demonstrate that spatial alignment operates bidirectionally: it both restores the spatial attributes embedded in historical sources and generates new relational knowledge across domains. By establishing a shared spatial coordinate system for disparate data, the study validates the methodological value of space as a mechanism of knowledge alignment and positions spatial reconstruction as an epistemological framework for historical knowledge production.
  •  

Human evaluation of large language models in Old English translation: a qualitative analysis

Abstract
This study evaluates the performance of large language models (LLMs) in translating Old English (OE) texts into Contemporary English. We analyze translations generated by LLaMA2 and GPT-4 through their respective interfaces, Meta Llama 2 Chat and ChatGPT. Our methodology employs a human evaluation approach for three Old English excerpts: ‘The life of St. Æthelthryth’ from Ælfric′s Lives of Saints, ‘Cynewulf and Cyneheard’ from The Anglo-Saxon Chronicle, and ‘Ohthere’s voyage’ from the Old English translation of Orosius’ Historiarum adversum paganos libri septem. Through qualitative analysis focusing on morphology, syntax, and lexicon and comparison against a golden corpus of human translations, we examine the coherence, adequacy, and precision of LLM-generated translations. Our findings reveal significant variation in translation quality across different LLMs and source texts. While GPT-4 demonstrates remarkable competence in translating Old English, particularly in morphological and lexical accuracy, both models show inconsistencies with complex syntactic arrangements. The study highlights the potential of LLMs for historical language translation and emphasizes the necessity of human revision and the impact of source text complexity on translation quality. This research provides insights into the capabilities and limitations of AI in processing low-resource historical languages.
  •  

Comparative computational textual criticism of the Gospel of Mark and the paradigm of textual clusters

Abstract
In New Testament textual criticism, the paradigm of textual clusters has become a controversial subject. Within this paradigm, the traditional practice of classifying textual witnesses (including manuscripts, translations, and church fathers who quoted the text) into families and larger groups facilitates the task of weighing their support for competing variant readings. But recently, textual critics have raised objections to this paradigm. One point of contention in particular has been how much agreement and disagreement between witnesses is required to define a group and distinguish between groups. To investigate the matter, we apply five digital approaches to a subset of the Editio Critica Maior (ECM) collation of the Gospel according to Mark. Three approaches, classical multidimensional scaling (CMDS), partitioning around medoids (PAM), and non-negative matrix factorization (NMF), are explicitly clustering-based approaches. Two others, network analysis and reticulating cladistics, reconstruct models of the textual tradition whose branches correspond to clusters. The responses of these approaches to contamination and coincidental agreement provide another axis of comparison. We find that all five consistently replicate traditional textual clusters, although contamination and coincidental agreement result in misclassifications of individual witnesses to varying degrees. Notably, all five approaches identify the ‘Caesarean’ group, whose existence has long been debated in the literature, as well as the ‘Western’ group, whose existence has more recently been challenged based on its meager evidence among Greek manuscripts. The convergence of the different approaches on similar results speaks to the continuing relevance of the paradigm of textual clusters.
  •  

Can algorithms detect lost scholarly communities? A network-based study of the Lower Yangtze in Late Imperial China (1644–1912)

Abstract
This article proposes a network-based framework for analyzing scholarly communities in Qing China (1644–1912), introducing the concept of “lost scholarly communities” to identify structurally cohesive but historiographically overlooked intellectual groups. Drawing on Qingru Xue’an (清儒學案, Cases of Qing Confucian Learning) as the primary source, supplemented by China Biographical Database and China Historical Geographic Information System, the study reconstructs over 3,000 academic relationships among approximately 1,300 scholars in the Jiangnan region. These relationships—spanning mentorship, social interaction, and correspondence—are extracted, structurally coded, and transformed into a dynamic temporal network. The methodological framework combines dynamic clustering to track the longitudinal aggregation of communities with static Louvain-based detection to reveal latent groupings across space and time. The study demonstrates that Qing intellectual communities did not always align with named schools or regional factions; instead, relational density and interactional structure better account for the emergence of scholarly cohesion. By identifying “lost communities” within the network, the research challenges conventional ideological or genealogical classifications and offers a relational redefinition of academic affiliation. More broadly, it provides a transferable model for historical knowledge network analysis, applicable to other periods and corpora. This approach complements textual and conceptual methods in intellectual history while expanding the digital humanities tool kit for mapping knowledge circulation in historical contexts.
  •  

Machine learning techniques for exploring influence, commonalities, and shared origin of scripts: cases of Ethiopic, Armenian, Georgian, and Caucasian Albanian scripts

Abstract
The morphological similarities between the Armenian, Georgian, and Caucasian Albanian scripts and the Ethiopic script have long intrigued both casual observers and scholars. However, prior studies have relied primarily on qualitative or historical analysis, often lacking objective or computational rigor. This study addresses that gap by applying machine learning and deep learning methods to explore potential structural relationships among these scripts. Using over 28,000 images of Ethiopic characters, we trained a deep convolutional neural network and augmented the dataset to enhance generalization. The resulting model, FeedelLigence, analyzes cross-script similarities through transformation-invariant distance measures, cosine distance (CD), and mutual information (MI). Our findings indicate notable structural and symbolic proximity between Ethiopic and the three comparison scripts. Armenian showed the strongest similarity, with the highest MI (0.7428 bits) and the lowest CD (0.0774). Georgian and Caucasian Albanian followed, with MI scores of 0.6843 and 0.6561 bits, and CDs of 0.1558 and 0.2498, respectively. These results provide computational evidence of significant structural overlap, suggesting possible historical connections or shared influences. In a broader cultural context, such affinities align with historical patterns of script evolution and cross-civilizational exchange. By combining artificial intelligence with comparative script analysis, this study offers a novel, quantitative perspective on the relationships among ancient writing systems—advancing our understanding beyond traditional human-centered approaches.
  •