❌

普通视图

Received before yesterday美 - 普林斯顿大学(Princeton)

Welcome to the new academic year!

2026年9月15日 02:53

Welcome to the new academic year! I write to you from this, my 20th year (!!!) at Princeton and my 12th (!!!!) at the helm of the Center for Digital Humanities. Thank you for your ongoing support of our unique interdisciplinary community. It’s been fantastic to have Associate Director Paul Verthaler join me this past year as we continue our mission of centering the humanities in conversations about the technologies that are transforming our research, culture, and society.

As always, we remain at the cutting edge of developments in AI, while continuing to foster thoughtful, timely, and urgent discussions about a future where humanistic approaches and values play a central role in our increasingly technical world.

Over the past year, we took up this core question — how AI and the humanities intersect — with Princeton faculty and graduate students in our Modeling Culture seminar, and by the campus-wide audiences who attended our workshops, events, and talks. Across CDH-affiliated labs, Princeton scholars join their domain expertise with testing the benefits and limits of AI applications. As always, we partner and advise on a wide range of projects — from curating data from ancient Middle Eastern texts, to developing technologies for African languages, to learning the possibilities of how AI might sustain long-form narrative. Our research software engineers are building custom tools for scholars who want to analyze complex, often multilingual or historical texts that conventional AI systems weren't built for, while also exploring open models and experimenting with new ways of accessing and exploring research materials.

A major milestone for the CDH in 2026 was taking on the role of publisher for the long-running Journal of Cultural Analytics (JCA), a leading journal for scholarship on the computational study of culture. JCA is just one piece of our involvement in a broader "Cultural AI" initiative. We’re also launching a working group, a collection of pedagogical experiments and discussions, and new graduate student and faculty opportunities. In collaboration with the newly launched Princeton Societal AI group at the newly launched Data and Intelligent Systems Initiative, we’re happy to support and participate in new cross-campus efforts to ensure that culture and society are at the forefront of how Princeton approaches data science, computational science, and AI.

We invite you to join these important and timely conversations. CDH is always looking for new collaborators, and we welcome inquiries from scholars at any level or from any division. Sign up for a consultation to discuss research or teaching ideas. Students can find us through our growing slate of undergraduate courses under the CDH course code. Subscribe to our monthly newsletter to hear about what's new and what's happening. Please join us at our Open House on Wednesday, September 16, and in October for a series of events with Leif Weatherby from NYU, media theorist and one of the most important voices in our understanding of how AI is affecting culture.

Last but not least, we're delighted to share news about the CDH family: Seyi Olojo joins us as a Postdoctoral Research Associate working on epistemological questions about African language technologies, and Christine Roughan has been promoted from postdoc to the staff position of CDH Project Manager.

Meet Senior Thesis Prize Co-Winner Alison Fortenberry ’26

2026年9月9日 09:19

Congratulations to Alison Fortenberry, co-winner of the 2026 CDH Senior Thesis Prize! Alison recently graduated with an A.B. in African American Studies with minors in Religion, American Studies, and English. Her thesis, entitled “Until Change Comes: The Hidden History of Denominational Formation and Racial Exclusion in Twentieth Century American Pentecostalism,” was advised by Wallace Best (African American Studies).

We asked Alison about her award-winning work.

Tell us about your project. How would you summarize it in two to three sentences?

My project explored the racialized formation of American Pentecostal denominations in the early twentieth century. American Pentecostalism began as an interracial revival movement largely led by African Americans, but it soon split into racially segregated denominations. My thesis argued that, for White Pentecostals, denominationalism was a project of racial exclusion and racial differentiation. 

How did you get interested in your topic?

I grew up in a Pentecostal church, and three years ago, one of my mentors in the congregation passed away. At her funeral, her family insisted that her church membership card be added to a church archive, explaining that members of their family had been historically denied access to the church because they were Black. This surprised me, because the church I grew up in was predominantly Black, and our existing church histories obscured the church’s exclusionary past. I wanted to correct those church histories. At first, the project began just as an effort to make a church archive (to date, I have cataloged and digitized tens of thousands of church records), but as I read through church records and began to explore broader literature about the denomination’s history, I realized that I wanted to make this story the focus of my thesis. [Check out this link to explore some of the items Alison digitized from the church archive.]

Fortenberry 4

Alison Fortenberry ’26

What role do digital humanities methods play in your analysis?

I tried to let the archive and broader corpus of literature I was drawing on guide my use of digital humanities methods in my analysis. A major component of my thesis was a historiography of a predominantly White American Pentecostal denomination, the Assemblies of God (AG). As I got into American Pentecostal literature, I noticed that historical authors who were affiliated with the AG tended to write the denomination’s history much differently than authors who were not, often downplaying the AG’s role in segregating an interracial religious movement. While I was able to demonstrate the differences in AG-authored histories through close readings, I wanted a way to quantify these trends. I worked with Andy Janco in Digital Scholarship to develop a Scattertext project, which tracked how frequently authors of different denominational backgrounds used specific words, then plotted them on a graph. With Scattertext, I was able to demonstrate that AG-affiliated authors use the names of Pentecostalism’s early Black leaders, for example, at disproportionately low rates. 

ScattertextExample
What was a challenge you overcame as you worked on your project?

A major goal of my project was to reconstruct the lives of Black Pentecostals in White churches who had been excluded from White church archives. One of my subjects, Albertina Cooper, was a Black woman who attended a White church for over thirty years, yet in an archive of tens of thousands of objects, her name was only listed four times. Trying to flesh out the motivations and feelings behind her decisions was challenging. For example, I wondered why Albertina chose to attend this discriminatory White church when there were other Pentecostal churches in the city, and I assumed that maybe this was just the closest Pentecostal church to her home. I was able to use archived church minutes and a Philadelphia city directory to map Albertina Cooper’s home in relation to all of the city’s Pentecostal churches in the 1930s. I was surprised to find that the closest Pentecostal church to Albertina’s home was actually a Black Pentecostal church and that Albertina had to travel over three miles to get to her White congregation. Albertina’s choice to attend her church was not just a matter of convenience, but an active decision to be the only Black member of a White congregation. Digital mapping helped me reconstruct Albertina’s life and motivations in the face of archival silences.

MapExample
What was one surprising thing you learned in the process—either about yourself or about your topic?

Through this process, I was surprised to learn that, even as someone with more traditional humanistic training and interests, digital tools are accessible to me. I never thought that I had the ability to conduct more quantitative or computational research, but with the help of Digital Scholarship, the Maps & Geospatial Information Center, and the Center for Digital Humanities, I was able to incorporate digital humanities methods to deepen and extend my qualitative research. 

How might your work on your thesis inform your future?

This fall, I’m returning to Princeton to begin a PhD program in the Department of Religion. I plan to continue using methods from the digital humanities in my doctoral work and beyond.

Read about co-winner Christina Li ’26 here.

Meet Senior Thesis Prize Co-Winner Christina Li ’26

2026年9月7日 10:07

Congratulations to Christina Li, co-winner of the 2026 CDH Senior Thesis Prize! Christina recently graduated with a B.S.E. in Operations Research and Financial Engineering and minors in Art History and French Language and Culture. Her thesis, entitled “The Value of Art,” was advised by Alain Kornhauser (ORFE).

We asked Christina about her award-winning project.

Let’s start with a quick summary. How would you describe your work in a few sentences?

My thesis investigates how monetary value is formed in the fine art auction market through social, political, economic, and artistic forces—or a lack thereof. Using a dataset of fine art auction sales, I analyze the factors that determine whether an artwork sells and, if it does, how much it sells for. I find that social prestige and institutional recognition are the strongest predictors of success at auction. Ultimately, however, my thesis asks what art is worth to the individual and to society, beyond its monetary value.

kirchner

Ernst Ludwig Kirchner, Street, Berlin (1913), in the Degenerate Art exhibition (upper level, room 4). Courtesy of The Museum of Modern Art (MoMA).

kirchner_christies

Ernst Ludwig Kirchner, Berliner Strassenszene (Berlin Street Scene), 1913–14. Sold at Christie’s.

Christina explains: “I use the story of the Degenerate Art exhibition, and these Kirchner paintings in particular, as an introduction to my thesis, framing the discussion about the relationship between art, value, and institutions.” The top painting included a label that "noted it was ‘purchased with the taxes of working German people’ by the Nationalgalerie in 1920, for 12,000 German marks (or $100,000 USD today). This seemingly high figure was intended to provoke visitors.” In contrast, the “very similar Kirchner [below] was sold at auction at Christie’s in 2006 for over $38 million USD.”

What drew you to your topic?

I am deeply interested in art history, and working with auction data allowed me to combine this passion with my technical background in Operations Research and Financial Engineering. At Princeton, I have taken several courses in modern and contemporary art history where we discussed the commercialization of art.

Tell us about the computational aspects of your project.

To analyze auction results, I estimated hedonic regressions with fixed effects, difference-in-diffferences, and event study models around major exhibitions, allowing me to examine how institutional recognition affected artists’ market values over time. I also implemented neural network models to predict auction outcomes based on artwork, artist, and auction characteristics, providing a complementary approach to understanding the factors that drive prices. 

What is one takeaway—either personal or intellectual—from your senior thesis experience?

I am inspired and encouraged by the potential for the humanities to be deeply interrogated and enriched through the application of engineering tools and principles. 

Any connections between your thesis and your future plans?

My dream is to become a curator. My thesis has reinforced my interest in the important role that art plays in our institutions and societies. 

Read about co-winner Alison Fortenberry ’26 here.

Check Out Fall 2026 Graduate DH Course Offerings in NJ and NY

2026年9月4日 10:46

Looking to explore DH-relevant courses–both at Princeton and beyond?

We’ve made a list of graduate courses for fall 2026 offered throughout the Inter-University Doctoral Consortium (IUDC).

With the approval of the certificate director, these courses can be used to satisfy the elective requirement for the CDH’s Graduate Certificate in Digital Humanities.

To enroll in a course, submit a completed registration form to the IUDC coordinator at both your host and home institutions.

PRINCETON UNIVERSITY

IUDC Details

Note: for a list of grad seminars offered or cross-listed by the CDH that automatically count toward the Graduate Certificate in DH, please see Graduate Courses in DH.

Architecture

Art & Archaeology

Computer Science

East Asian Studies

Electrical & Computer Engineering

Politics

Psychology

Sociology

Public & International Affairs

Princeton Writing Program

COLUMBIA UNIVERSITY

IUDC Details

Comparative Literature: Russian

English

History

Music

Political Science

Psychology

Quantitative Methods: Social Sciences

Religion

CUNY GRADUATE CENTER

IUDC Details

Data Analysis and Visualization (see courses)

  • DATA 70500: Working with Data: Fundamentals (Tim Shortell)
  • DATA 71000: Data Analysis Methods (Howard Everson)

Digital Humanities (see courses)

  • DHUM 70000: Introduction to Digital Humanities (Krystyna Michael)
  • DHUM 70600 - 01: Special Topics in Computational Fundamentals: AI-Assisted Programming (Stephen Zweibel)
  • DHUM 70600 - 02: Special Topics in Computational Fundamentals: Introduction to R (Sam O’Hana)
  • DHUM 73000: Visualization and Design (Julia Bloom)
  • DHUM 78000 - 01: Special Topics: AI and Machine Learning for Artists and Humanists (Michael Mandiberg)
  • DHUM 78000 - 02: Special Topics: Labor, Literature, and the Digital (Jeff Allred)

Environmental Sciences (see courses)

  • EES 79903: Inhabiting Digital Worlds (Setha Low)

English (see courses)

  • ENGL 89000: Mining the Archives (David Reynolds)

Linguistics (see courses)

  • LING 79400: Seminar in Speech Science: Hybrid Society (Human/AI) and the Pursuit of (Speech) Science (Douglas H. Whalen and Ofer Tchernichovski)

Music (see courses)

  • MUS 72900: Computer-Assisted Composition (Steve Everett)

Philosophy (see courses)

  • PHIL 78700: AI and the Arts (Jesse Prinz)

Sociology (see courses)

  • SOC 82301: Computer Mapping for NY and LA, and Global Cities: GIS with Manifo: Basic and Advanced Techniques (David Halle)

FORDHAM UNIVERSITY

IUDC Details

(Fordham courses can be accessed here)

Center for Ethics Education

  • Center for Ethics Education 5375: Technology & Ethics (Zein Murib)
  • Center for Ethics Education 5475: Algorithmic Ethics (Mathias Klang)

NEW SCHOOL FOR SOCIAL RESEARCH

IUDC Details

Liberal Studies

New York University

IUDC Details

(NYU course listings can be accessed here)

Animal Studies

  • ANST-GA 2500: Animals, Art, and Technology (Gal Nissim)

Center for Experimental Humanities

  • CEH-GA 1089: Ethnography and Technology of Media (Natasha Schüll)

Cinema Studies

  • CINE-GT 1025: Creative Artificial Intelligence: Ethical Co-Creation in the Arts and Humanities (Marina Hassapopoulou)

Digital Humanities and Social Science

  • DHSS-GA 1120: Introduction to Programming

History

  • HIST-GA 2033: Creating Digital History (Leah Potter)

Institute for the Study of the Ancient World

  • ISAW-GA 3023: Special Topics in Digital Humanities for the Ancient World: Databases & Network Analysis for the Ancient World (Sebastian Heath)

Media, Culture, and Communication

  • MCC-GE 2100 Seminar in Media Criticism
  • MCC-GE 2112 Politics of The Gaze: Sensory Formations Mod
  • MCC-GE 2113 Fame: Social Theories of Charisma, Recognition, and Renown
  • MCC-GE 2236 Queer and Trans Game Studies
  • MCC-GE 3010 Special Topics in Critical Theory: The Grounds of Media: Energy, Infrastructure and Land in Materialist Media Theory
  • MCC-GE 3302: NYLON New York: Technology, Culture, Politics (Sophie Gonick and Natasha Schüll)

Near Eastern Studies

  • NEST-GA 3002: Topics in Anthropology: Critical Thinking with AI: Media & Power in MENA (Jared McCormick)

Politics

  • POL-GA 3300: Seminar in American Government and Politics: Seminar on Social Media and Politics

Sociology

  • SOC-GA 3000: Interdisciplinary Seminar: Generative AI in Sociology (Christopher Barrie)

RUTGERS UNIVERSITY

IUDC Details

(Rutgers course listings can be accessed here)

German, Russian, and East European Languages (see courses)

  • 16:470:670:01: Habit and Habitation: On Walter Benjamin’s Media Aesthetics and Philosophy of Technology (Astrid Deuber-Mankowsky)

Italian (see courses)

  • 16:560:680: Ethics: AI and Digital Humanism in Language Acquisition (Carmela Scala)

Music

  • 16:700:515: Computer Composition

Political Science

  • 16:790:567: Global Politics of Internet Security (Ihab Darwish)

Sociology (see courses)

  • 16:920:574: Topics in Sociology: AI and Social Science Research (Di Zhou) (half-semester course)
  • 16:920:575: Topics in Sociology: AI and Society (Di Zhou) (half-semester course)

School of Communication and Information (see courses)

  • 17:610:562: Problem Solving with Data (Kiran Garimella)
  • 17:610:579: Ethics, Values, and Change in Information Practices (Britt Paris)

STONY BROOK UNIVERSITY

IUDC Details

(Stony Brook course listings can be accessed here)

Art History

  • ARH 520: Media Aesthetics (Zabet Patterson)

English

  • EGL 608: Relationship of Literature and Other Disciplines: Literature and Critical Data Studies (Katherine Johnston)

Princeton's Meredith Martin Co-Organized First-Ever ICML Workshop on Culture and AI

2026年9月2日 23:17

Meredith Martin, Professor of English and Director for Center for Digital Humanities at Princeton, was a co-organizer of the inaugural "Culture × AI: Evaluating AI as a Cultural Technology," a workshop held at the International Conference on Machine Learning (ICML) on Friday, July 10, 2026, in Seoul, South Korea.

The ICML is the leading annual conference for machine learning research, publishing foundational work underlying much of modern AI research. This year, ICML’s workshop program was particularly selective: “Culture × AI” was one of only 44 workshops chosen out of 247 proposals, and marks the first time ICML has hosted a workshop dedicated to this topic.

"Culture × AI" examines how generative AI systems function as cultural technologies — producing text, images, and video shaped by vast amounts of social and cultural data — and asks how the humanities, arts, and qualitative social sciences can inform AI development from the outset, rather than being applied only after deployment to mitigate harm. The workshop's central theme, "Interpretive Technologies," explored how humanistic traditions of meaning-making and contextual and aesthetic judgment can be built into AI design.

Speakers included prominent voices in Cultural AI: Lauren Klein (Emory University), Ted Underwood (University of Illinois Urbana-Champaign), Maria Antoniak (University of Colorado, Boulder), and Joel Z. Leibo (Google DeepMind), along with a panel discussion, a lightning talk session featuring accepted papers from researchers at McGill University, KAIST, Duke University, and the University of Groningen, and a closing poster session.

Martin's involvement reflects Princeton's growing engagement with the intersection of the humanities and artificial intelligence, extending the work of the university's Center for Digital Humanities into international, cross-institutional conversations about how cultural knowledge can shape the future of AI systems.

More information about the workshop is available at doingaidifferently.org/culturexaiworkshop.

Journal of Cultural Analytics publishes data essay on Cairo Geniza, bridging medieval Middle East with data science and AI

2026年9月2日 06:23

The Center for Digital Humanities (CDH) at Princeton University announces the publication of "Princeton Geniza Project Datasets," a new data essay by Rebecca Sutton Koeser (CDH) and Marina Rustow (Near Eastern Studies/History), published in the Journal of Cultural Analytics (Koeser and Rustow 2026, 11(3): 1–35, DOI: 10.22148/jca.1084). The essay presents, for the first time, a full account of the data behind the Princeton Geniza Project (PGP) — one of the richest digital resources for the study of the medieval Islamic and Jewish Mediterranean world.

The PGP dataset opens rich and complex historical material to data science approaches

The Cairo Geniza — a repository of discarded largely Hebrew-script writings in the Ben Ezra Synagogue in Fustat, Old Cairo — preserved hundreds of thousands of documents from everyday medieval life c. 950–1250, from personal letters and legal contracts to trade accounts. These fragments offer scholars an unusually candid window onto medieval culture stretching from Spain to Sumatra. Since the 1980s, Princeton’s Geniza Project has been collecting, documenting and annotating these fragments, which are now scattered across some seventy libraries worldwide. Today, the PGP contains 35,855 Geniza documents — with descriptions, and information about document types, languages and scripts, dates across multiple calendars, transcriptions, translations, and bibliographic citations — along with information on 1,802 people and 486 places.

The publication of the dataset – a complement to the public-facing PGP website – makes the complex material accessible for data science and AI researchers and historians alike. Project PI Marina Rustow’s collaboration with the Center for Digital Humanities has made the PGP one of the leading efforts bringing medieval manuscripts into the world of data science and AI — alongside projects such as the multi-institutional OpenITI project for Arabic texts and the National Library of Israel's KTIV for Hebrew manuscripts.

The essay documents the PGP’s 40-year history — from an IBM-funded 1986 pilot to the five-and-a-half-year PGP–CDH partnership (2020–2025) that rebuilt the database from the ground up — and explains the modeling decisions behind the data, including the many-to-many relationship between physical fragments and the documents written on them.

"Providing both a well-structured dataset, as well as the contextual information of its creation, enables a broader community of researchers to engage with the Geniza material," states the authors. "The study of the medieval Mediterranean world can now be explored, interrogated, and modeled — inviting the tools of data science and AI to reveal patterns across hundreds of thousands of fragments that no single scholar could trace by hand."

About the Dataset

The Princeton Geniza Project dataset is archived on Zenodo (DOI: 10.5281/zenodo.18716581) and updated quarterly. The essay itself is open access under a Creative Commons Attribution 4.0 license. The Princeton Geniza Project is the flagship project of the Princeton Geniza Lab (PGL), directed by Marina Rustow; PGL is an affiliate lab of the Center for Digital Humanities.

Read the essay: https://doi.org/10.22148/jca.1084

Two CDH-Affiliated Scholars Win Major NEH Grants for Cairo Geniza Research

2026年8月26日 04:45

Eve Krakowski and Marina Rustow, co-director and director of Princeton's Geniza Lab — an affiliate lab of the Center for Digital Humanities — have each been awarded $300,000 from the National Endowment for the Humanities to produce new scholarly editions and translations from the Cairo Geniza, a cache of over 400,000 medieval manuscripts. Krakowski's project traces the social history of death in medieval Egypt, while Rustow's brings to light Jewish traders' letters, legal documents, and ledgers from the medieval Indian Ocean trade.

The news follows a related CDH milestone: Rustow and CDH's Rebecca Sutton Koeser recently published a data essay opening the Princeton Geniza Project's rich and complex historical material to data science approaches.

Read more at Princeton University →

Related posts

Princeton Geniza Lab

At the forefront of DH scholarship on the Cairo Geniza, preserving and providing access to this vast and invaluable collection of historical texts. Director: Marina Rustow

genizafragments.jpg

Princeton Geniza Project

Accessing the medieval Islamic world through digital tools

Built by CDH
twitter3.jpg

New Publication Details Cooperative Broadband and Community-Supported Infrastructure

2026年6月5日 03:55

DH Strategist and head of CDH graduate programs Grant Wythoff has published a new chapter in the latest Debates in DH volume, Critical Infrastructure Studies and Digital Humanities. The chapter — Alternative Infrastructures for Digital Equity: Community-Based Internet Access —  details the creation of Philly Community Wireless (PCW), a community-controlled network. PCW offers free public Wi-Fi across the Philadelphia neighborhoods of Norris Square, Fairhill, and Kensington. Today, dozens of homes and businesses host PCW rooftop antennas. The network provides access to 50+ city blocks and 7,600+ monthly users. It is maintained by dozens of volunteers, community partners, and a dedicated staff.

PCW was first launched in the summer of 2020 with the support of a Rapid Response Grant from the Princeton Humanities Council and student interns from the Pace Center for Civic Engagement’s RISE Program. During the Covid pandemic’s earliest days, a group of organizers, researchers, librarians, and neighbors — including Wythoff and the chapter’s co-authors Alex Wermer-Colan, Devren Washington, and Allan Gomez — gathered to address the problem of internet accessibility at a time when most of daily life had moved online. Their goals were to expand internet access, grow tech literacy, and build community autonomy.

The chapter reflects on the lessons learned in building PCW’s technical and social infrastructures, puts concepts from digital humanities and community technology into conversation, and finally, gestures toward the future of community-supported infrastructure. The authors write:

Our question is how infrastructure cooperatives that provide community mesh networks for urban environments can facilitate collaboration between neighbors while fostering community consensus on network architecture, organizational structure, and the geographic redistribution of digital resources. Ultimately, our theory of change is that access leads to adoption in a fuller social sense: if one empowers the communities most affected by the biases and harms of the tech industry to own and control last-mile internet infrastructure in their neighborhood, such communities will be equipped not just to use broadband technology, but to advocate for better outcomes in the way that such technologies are utilized, governed, and regulated.

You can read the open access Manifold edition of the chapter now: Alternative Infrastructures for Digital Equity: Community-Based Internet Access

Graduate Fellows Scan, Model, and Map: New Discoveries from Sermons to Ballet

2026年6月4日 00:47

This spring, six CDH Graduate Fellows arrived with their research in progress, asking six different questions across multiple disciplines ranging from History and Comparative Literature to Music and English. Their findings? Working with data and computational methods rarely unfolded the way they expected, and while oftentimes arduous, the labor uncovered “strange and beautiful” discoveries.

"We are receiving more competitive applications to the Graduate Fellowship than ever before," said Grant Wythoff, who directs CDH graduate student programs. "Some of these emerging scholars bring knowledge of data curation standards and machine learning methods. Others are tuned into the latest debates on AI's political and epistemological impacts. The mix of voices makes for an incredibly exciting group dynamic."

Cecelia Ramsey (French and Italian) came to the CDH with a question about literary afterlives: what makes a book experience a revival many years after its initial release? To study this at scale, she worked with BiblioBase, a database of the nineteenth-century Bibliographie de la France, tracking gaps between editions and reeditions and looking for patterns in a book's reintroduction.

The data was messy—inconsistent titles, variable spellings of authors' names—and rather than cleaning the inconsistencies away, Cecelia explored them as an opportunity to learn more about the nature of reeditions and the format of the Bibliographie itself. "Interacting with the messy data taught me how slippery the very name of a work can be," she reflected.

It was also her first substantial engagement with DH methods, and she described the fellowship's atmosphere as essential. When Grant Wythoff opened the semester by telling the cohort it was normal not to know things—that the DH world is so interdisciplinary that all scholars often feel that way—it changed what was possible. "This introduction made it a space where it's normal to ask questions, to learn, and to just be openly curious," Cecelia said. "What a gift."

Hand-drawn diagram on aged paper showing connected circles with handwritten names such as “Ignez de Guinea,” “Maria,” “Luiza,” “Antonio,” and others, arranged like a branching family tree.

A pedigree chart tracing the lineage of Ignez de Guiné, the matriarch of several prominent Portuguese families.

Amanda Pinheiro (History) has been working with a database developed over the last ten years, containing 115,545 baptismal, notarial, and judicial documents from ten villages in colonial Brazil. These records detail the eighteenth- and nineteenth-century lives of roughly 6,000 individuals who inhabited the south of Brazil and may have migrated through the frontier zone between the Spanish and Portuguese empires. The issue: that number is inflated by duplicate names, with varying spellings and characteristics across documents. Her fellowship project used Splink, a Python library for probabilistic record linkage, to calculate the statistical likelihood that two records refer to the same individual—generating a unique identifier for each person so she could cross-reference this database with her new archival findings.

I realized that automation necessarily requires diligent and continuous manual labor.

Amanda Pinheiro

What surprised her was how much human judgment, or as she describes it, laborious decision-making, the automation required. "I realized that automation necessarily requires diligent and continuous manual labor," Amanda reflected. "The two are interconnected and walk hand-in-hand in the digital realm." With Wouter Haverals’ (Associate Research Scholar, CDH; Perkins Fellow, Humanities Council) guidance, she completed Splink tests and generated an analytical report on the quality of her datasets—one she hopes to publish and add to her metadata in the future.

Amy Weng (English) asked whether a seventeenth-century English preacher's confessional affiliation—Anglican, Nonconformist, or Catholic—leaves a detectable fingerprint in their printed sermons. Her project, Godly and Learned Divines (GoLD), represented each of 809 preachers across 2,877 books as a 1,000-dimensional vector built from scripture citations, named references, topic modeling, and entity types, then trained a Random Forest classifier to predict denomination. The model achieved an F1-score (a machine learning metric used to evaluate the performance of a classification model) of 0.80 for Anglicans, 0.51 for Nonconformists, and 0.12 for Catholics.

Two side-by-side scatter plots of data points labeled by legend categories “Anglican,” “Catholic,” and “Nonconformist,” with the left in 2D and the right a 3D plot labeled axes PC1, PC2, and PC3.

Wikidata-Linkable Preachers in EEBO-TCP

Clustered based on the distribution of topics, named entities, and scriptural references in their sermons

The results were striking. Place of education turned out to be the least important feature—far less predictive than the types of sources a preacher reached for. "Bible versions matter more than the proportions of Bible divisions," Amy concluded, "and ancient entities once again outrank medieval and contemporary references. Generally, godly learnedness—patterns of referencing scripture—distinguishes preachers across confessional divides more than overall learnedness."

Amy credited Jacob Murel (Research Software Engineer, Classics) for sustained mentorship on using large language models for orthographic standardization, and Wouter Haverals for introducing her to Wikidata reconciliation.

Pierre Azou (French and Italian) examines the relationship between literature and political violence in his doctoral research and found himself drawn to the “digital sphere” as the space where the questions he studies in published books are being reconfigured. For his fellowship project investigating the link between "manliness" and insecurity in contemporary French public discourse, he turned to two foundational texts in the French debate on masculinity: Élisabeth Badinter's XY, de l'identité masculine (1992) and Éric Zemmour's Le Premier Sexe (2006). These works contain opposing premises, one theorizing a fragile masculinity, the other insisting it is strong but under siege, yet both binding manliness tightly to a language of threat and crisis.

Using keyness analysis and topic modeling in Python allowed Pierre to compare the density of each author's clusters of insecurity-related words (fear, violence, war, crisis, domination). He identified the author's statistically distinctive vocabulary and examined the semantic neighborhoods of shared terms. The most productive approach was contextual: looking at what words appear near a key shared term like virilité in each text. "It turns out the same word lives in completely different semantic environments in Badinter and Zemmour," Pierre noted.

A technical challenge gave him pause early on—preprocessing French text that contained English-language citations required combining stopword lists (words like “a,” “the,” “and” or “un,” “le,” “et”) and filtering bibliographic noise—but his more substantive reflection was on what data cleaning actually does. "It reminded me that cleaning decisions in DH are more than purely technical, as they also shape the findings."

Cleaning decisions in DH are more than purely technical, as they also shape the findings.

Pierre Azou

Nathaniel Gallant (Comparative Literature) studies the relationship between Buddhism and the history of dramatic and poetic theory across Japanese and Tibetan literary traditions. In his daily research, he relies on well-developed digital tools and databases built for pre-modern Japanese sources—resources that reflect years of philological groundwork by scholars who came before him. For Tibetan studies, that infrastructure is still being built. DH projects in the field are scattered across academic, non-profit, and private spheres, with no centralized view of what exists or where the gaps are.

Nathaniel’s fellowship project addressed that directly: he created a database cataloging existing DH projects in Tibetan studies, with visualizations mapping networks of funding sources, text archives, OCR and LLM development projects, and institutional stakeholders. The goal was to understand current patterns in project development and identify potential directions for future text-digitization projects, particularly in the history of Tibetan literature and poetry.

The hours of scanning documents, mental grappling...crystallized into something coherent, beautiful.

Rachel Glodo

Black-and-white illustration of multiple ballet dancers in tutus descending a staircase, with Russian text on the right including “Балетная труппа” and names and dates printed below.

Each issue of the Yearbook of the Imperial Theaters includes detailed lists (spiski) of creators and artists.

Rachel Glodo (Music) is reconstructing the world of the Imperial Ballet in the Russian Silver Age through eighteen volumes of the Yearbook of the Imperial Theaters (1890–1908)—elaborate annual retrospectives documenting productions, performers, choreographers, musicians, designers, and administrators across St. Petersburg and Moscow. The challenge was getting that data out of the page and into a form that a researcher could query. Rachel used optical text recognition (OTR) to convert nineteenth-century printed Cyrillic into machine-readable text, while Andy Janco (Digital Scholarship Specialist) developed custom Python scripts, based on her project design, to convert images of lists and tables into structured spreadsheets.

What Rachel hadn't anticipated was how much the project would begin with physical, analog labor. "The most challenging part of my project wasn't the implementation of DH methodologies," she said, "but the quotidian task of scanning and saving thousands of images spanning 18 volumes." She described it as "the strange and beautiful juxtaposition of 'distant' and 'close' readings that characterizes DH." And she was surprised by how much the technology itself shifted between her original proposal and the start of the fellowship—she ended up using an entirely different processing strategy than she had planned, with Christine Roughan (Postdoctoral Research Associate, CDH/MARBAS) and Andy as crucial partners in identifying her priorities and methods.

A eureka moment was had when they ran the Python scripts together for the first time. "All the hours of scanning documents, mental grappling, design, and redesign suddenly crystallized into something coherent, beautiful, and—almost miraculously—exactly what I needed," she recalled. "It was a glorious moment."

Open spread of an aged printed table in Russian with dates across the top including “25 октября 1896 г.” and “13 ноября,” filled with dense columns of text and numbers arranged in a grid.

A record of all productions on the Imperial stages, including ballets and operas.

Throughout the semester, the cohort's monthly sessions became as important as the technical work itself. "The regular meetings provided me with more productive time and space to learn about digital tools than scheduling different consultations could have," Amanda said. For Cecelia, the cross-disciplinary exchange was its own kind of finding: "It's exciting to step outside your discipline and be invited into someone else's world while it's still in the making—while they're still experimenting and puzzling through the challenges."

Interested in applying for a Graduate Fellowship? Visit here, or head to the CDH Graduate Program page to see more opportunities for graduate students.

Related posts

Graduate Fellowships

A one-semester studio for workshopping research in progress.

jcasey_just_data_sms_0597.original

Announcing the Fall 2026 CDH Graduate Fellows

2026年6月3日 21:11

We are delighted to announce the recipients of the Fall 2026 CDH Graduate Fellowship. Grad Fellows are mentored by CDH staff to employ computational and data-driven methods in their research. As a cohort, fellows explore tools, methods, and best practices that will benefit them throughout their careers.

Please join us in congratulating this cohort:

Maximilian Diemer (History) is researching the origins of meritocratic practice in eighteenth-century France and Britain, with a focus on professionalisation in the army and royal service.

Aneka Kazlyna (History) studies the history of astronomy and the broader mathematical sciences, with particular interests in the histories of observation and the circulation of knowledge across cultures.

Shiqi Pan (East Asian Studies) studies the medieval history of the Huai River region in China, exploring it as an internal frontier shaped by environmental, political, and cultural change.

Benjamin Price (Art and Archaeology) studies art, anarchism, and histories of science in late-nineteenth century France.

Sun Shen (Politics) is studying presidential influence in the making of U.S. foreign policy.

Filippo Ugolini (East Asian Studies) examines the discourse of romance in mid-to-late Tang China (8th–9th century) at the intersection of literary analysis, gender studies, and socio-economic history.

Tirzah Anderson (History) is researching Afro-Indigenous worldmaking in Indian Territory and Oklahoma from the 1890s through the 1930s.

We look forward to working with this remarkable cohort and sharing more about their projects and fellowship outcomes as the year unfolds.

Graduate Fellowships

A one-semester studio for workshopping research in progress.

jcasey_just_data_sms_0597.original

Applications open for AI Postdoctoral Research Fellows

2026年5月28日 21:28

Location

AI Lab-41 William Street, Princeton, NJ 08540

Open Date

May 25, 2026

Salary Range or Pay Grade

$100,000

Deadline

Jul 01, 2026 at 11:59 PM Eastern Time

Description

Princeton University is seeking to hire AI Postdoctoral Research Fellows to join the Princeton Laboratory for Artificial Intelligence with affiliation to the Center for Digital Humanities to conduct research focused on societal AI. 

Postdoctoral Research Fellows will be part of a dedicated research community that includes leading faculty, research fellows, scientists, software engineers, postdocs, and graduate students. Fellows will have access to a state-of-the-art AI Lab GPU cluster to support large-scale experimentation and evaluation.

Postdoctoral Research Fellows will be appointed through the Princeton Laboratory for Artificial Intelligence (the AI Lab) and will receive a competitive salary, benefits, office space and equipment, dedicated compute resources, and annual research support.

Appointments will be made at the Postoctoral Research Associate rank. 

Appointments are for one year with the possibility of renewal pending satisfactory performance and continued funding.

Princeton University places high value on in-person collaborations and interactions, and we expect selected candidates to participate in-person. The work location for this position is in-person on campus at Princeton University.

This position is subject to the University's background check policy

These positions are not eligible for sponsorship of an H-1B visa requiring consular processing; other visa sponsorships (including H-1B visas not requiring consular processing) may be available, as appropriate.

Qualifications

Ideal candidates will have demonstrated experience in “societal AI” especially as related to culture. We seek scholars working on AI as a "cultural technology" — including its role in disciplinary practice across the arts and humanities and in knowledge infrastructures — as well as those examining AI's broader societal effects – including on creative and cultural industries (literature, music, film, visual arts). A strong foundation in computational humanities or cultural analytics, combined with humanistic and interpretive approaches to AI systems, is expected.

Applicants must have a PhD.

Application Instructions

Applicant must provide the following:

  • Cover letter
  • Current curriculum vitae
  • Research statement
  • Contact information for three references is required

Application Process

This institution is using Interfolio's Faculty Search to conduct this search. Applicants to this position receive a free Dossier account and can send all application materials, including confidential letters of recommendation, free of charge.

Equal Employment Opportunity Statement

Princeton University is an Equal Opportunity Employer and all qualified applicants will receive consideration for employment without regard to age, race, color, religion, sex, sexual orientation, gender identity or expression, national origin, disability status, protected veteran status, or any other characteristic protected by law.

Pay Transparency Disclosure

The University considers factors such as (but not limited to) the scope and responsibilities of the position, candidate's qualifications, work experience, education/training, key skills, market, collective bargaining agreements as applicable, and organizational considerations when extending an offer. The posted salary range represents the University's good faith and reasonable estimate for a full-time position; salaries for part-time positions are pro-rated accordingly.

The University also offers a comprehensive benefits program to eligible employees. Please see this link for more information.

CDH announces inaugural Affiliate Labs

2026年5月15日 22:13

The Center for Digital Humanities is excited to announce its Affiliate Labs designation – a formal recognition of our partnerships with Princeton research groups whose work sits at the intersection of humanistic inquiry and digital and computational methods. Affiliation reflects the depth and ongoing nature of these collaborations, and signals CDH’s commitment to fostering a broad and vibrant ecosystem of digital humanities across campus. Affiliate Labs represent a range of approaches — from open-access archival preservation and computational text analysis to the development of language technologies and digital reconstructions of historical places.

“The Affiliate Labs program recognizes what is already happening at Princeton — rigorous, innovative work at the intersection of humanities and computation — and makes space for those partnerships to grow," states Jeri Wieringa (Assistant Director, CDH). "We’re excited to see where these collaborations take us.”

We are proud to introduce the following inaugural Affiliate Labs:

Princeton Geniza Lab, directed by Marina Rustow (Khedouri A. Zilkha Professor of Jewish Civilization in the Near East), is at the forefront of digital humanities scholarship on the Cairo Geniza, preserving and providing access to this vast and invaluable collection of historical texts. CDH has had the privilege of collaborating with the PGL on the Princeton Geniza Project since 2020.

African Language Technologies Lab, co-directed by Christiane Fellbaum (Professor of Linguistics) and Happy Buzaaba (Associate Research Scholar, AI Lab with affiliations in CDH and Africa World Initiative), works to increase the representation of African languages in rapidly advancing language technologies driven by large language models, while foregrounding the values and cultures those languages carry. Through a series of projects, courses, and speaker events, the lab generates campus-wide conversation about this critical and underserved area of language technology research.

REACH² Lab, led by Paul Vierthaler (Assistant Professor of East Asian Studies, Associate Faculty Director of CDH), centers on the digital analysis of the literature, culture, and history of East Asia. Vierthaler and his collaborators work on projects ranging from quantitative analyses of large-scale text and image corpora to digital reconstructions of historical places, welcoming students, faculty, librarians, and researchers interested in any dimension of computational East Asian studies.

Poetry's Data Lab, directed by Meredith Martin (Professor of English, Faculty Director of CDH), draws on the Princeton Prosody Archive to analyze patterns in anglophone poetry teaching over time, asking what poetry used in teaching texts — at scale — can reveal about canonicity, book history, and the development of English as a discipline. The lab's work is linked to the broader Ends of Prosody project and forthcoming special issues of the Journal of Cultural Analytics and Victorian Poetry.

We look forward to sharing more about each of these labs and the work emerging from these partnerships in the months ahead. To learn more or inquire about affiliation, visit the CDH Affiliate Labs page.

Affiliate Labs

Princeton Geniza Lab

At the forefront of DH scholarship on the Cairo Geniza, preserving and providing access to this vast and invaluable collection of historical texts. Director: Marina Rustow

genizafragments.jpg

African Language Technologies Lab

Focused on increasing the representation of African languages in NLP, LLMs, and AI. Directors: Christiane Fellbaum, Happy Buzaaba

Infrastructure for African Languages

REACH² Lab

Centered on digital analysis of the literature, culture, and history of East Asia. Director: Paul Vierthaler

Composite image combining an old map drawing with modern data visualizations, colored lines, and Chinese text on a dark background

Poetry's Data Lab

Using the poetry found in the Princeton Prosody Archive datasets to analyze patterns of anglophone poetry teaching over time. Director: Meredith Martin

Multiple stylized versions of the opening line of Paradise Lost displayed in colored quadrants with varied typography and notation

Graduate research transcends disciplines at annual joint colloquium

2026年5月7日 03:22

May 4, 2026

By Alaina O'Regan

Last Friday, Princeton graduate students gathered to present research spanning the evolution of social behavior through natural selection, the quantum nature of atomic nuclei, and the effect of tropical cyclones on global climate.

The annual data and computation joint graduate certificate colloquium, held on April 24, brought together students across the humanities, social sciences, natural sciences, and engineering to share their work and exchange ideas.

This year, the Center for Digital Humanities (CDH) joined the Center for Statistics and Machine Learning (CSML) and the Princeton Institute for Computational Science and Engineering (PICSciE) for the first time. The event was held in the new Commons Visualization Lab.

“Few, if any, graduate student events focused on research have participation across all areas of scholarship at the University,” said Michael E. Mueller, interim director of PICSciE, director of the graduate certificate in computational science and engineering, and acting chair of the department of mechanical and aerospace engineering. 

“Some of the most interesting research questions came from students driven by sheer curiosity about domains completely removed from their own.”

CDH grad certificate students included:

  • Laura Nelson, History 
    Adviser: Laura Edwards
    “Unseen fragments, seen lives: From data to biography in digital public history”
  • April Gilbert, Comparative Literature 
    Adviser: Claudia J. Brodsky
    “Narrating narrative’s lifespan: Exploring data from the conference programs of the International Society for the Study of Narrative (ISSN)”
  • Sharifa Lookman, Art & Archaeology 
    Adviser: Carolina Mangone
    “From pixel to foundry: Recuperating sixteenth-century bronze techniques and technicians through 3D-imaging and historical reconstruction.”

Read the original story: Graduate research transcends disciplines at annual joint colloquium (Source: MAE)

Check Out Our Curated Fall 2026 Course List!

2026年4月14日 01:08

Fall course selection starts this month! Wondering what to register for? We've compiled a list of courses on media studies, technology, data and culture, and more.

This list includes both undergraduate and graduate courses. Grad courses are marked with an asterisk (*).

ANTHROPOLOGY

ANT 422 / EAS 422: Digital China: Technology and Society (Jamie Wong)

ARCHITECTURE

ARC 311 / STC 311: Building Science and Technology: Building Systems (Peter Pelsinksi)

*ARC 573: Pro Seminar: Computation, Energy, Technology in Architecture (Forrest M. Meggers)

ART & ARCHAEOLOGY

ART 248: Photography and the Making of the Modern World (Monica C. Bravo)

ART 320: Ink, Paper, Wood, Metal, Stone: The Printed Image in the Western World (Jun P. Nakamura)

*ART 577 / HUM 577: Seminar in Modern Art: Science and Its Fictions in the Long Nineteenth Century (Rachael Z. DeLue)

COMPARATIVE LITERATURE

COM 200: Who Are You Really? Authenticity in Life, Literature, and the Internet (Avram C. Alpert and Wendy Laura Belcher)

COMPUTER SCIENCE

COS 126 / EGR 126: Computer Science: An Interdisciplinary Approach (Adam Finkelstein, Maryam Hedayati, and Alan Kaplan)

COS 209: Algorithms in the Wild (Matt Weinberg)

COS 217: Introduction to Programming Systems (Christopher M. Moretti, Kevin Alarcón Negy, and Jaswinder P. Singh)

COS 226: Algorithms and Data Structures (Marcel Dall’Agnol and Kevin Wayne)

COS 240: Reasoning about Computation (Zeev Dvir, Iasonas Petras)

COS 302 / SML 305 / ECE 305: Mathematics for Numerical Computing and Machine Learning (Ryan P. Adams)

COS 324: Introduction to Machine Learning (Adji Bousso Dieng and Vikram Ramaswamy)

COS 326: Functional Programming and Formal Methods (David P. Walker)

COS 330: Great Ideas in Theoretical Computer Science (Pravesh K. Kothari and Pedro Paredes)

COS 333: Advanced Programming Techniques (Robert M. Dondero)

COS 424: Reasoning with Data (Manoel Horta Ribeiro and Xiaoyan Li)

COS 429: Computer Vision (Olga Russakovsky)

COS 436: Human-Computer Interaction (Parastoo Abtahi)

*COS 516 / ECE 516: Automated Reasoning About Software (Zachary Kincaid)

*COS 522 / MAT 578: Computational Complexity (Gillat Kol)

*COS 529: Advanced Computer Vision (Felix Heide)

*COS 597C: Advanced Topics in Computer Science: AI Agents (Karthik Narasimhan)

CREATIVE WRITING

CWR 213: Writing Speculative Fiction (Ed Park)

EAST ASIAN STUDIES

EAS 212 / COM 235: Embodied Visions: The Body in East Asian Film and Media (staff)

EAS 313: Japanese Horror: Cognitive and Embodied Approaches (staff)

EAS 328 / CDH 328: Building Qianlong’s World: Digital Worldbuilding and Cultural Heritage (Paul A. Vierthaler)

EAS 407 / CDH 407: Hacking Chinese Studies: An Introduction to Text Mining for Chinese Literature and Culture (Paul A. Vierthaler)

*EAS 529: Readings in East Asian Film and Media (Steven Chung)

ELECTRICAL & COMPUTER ENGINEERING

ECE 364: Machine Learning for Predictive Data Analytics (Niraj K. Jha)

ECE 435: Machine Learning and Pattern Recognition (Mengdi Wang)

ECE 488: Fundamental Image Processing: From Mars to Hollywood with a Stop at the Hospital (Guillermo Sapiro)

*ECE 535: Machine Learning and Pattern Recognition (Hossein Valavi)

ECONOMICS

ECO 202: Statistics and Data Analysis for Economics (staff)

ECO 326 / COS 206 / ECE 326: Economics of Digital Connectivity and Artificial Intelligence (Swati Bhatt)

ENGINEERING

EGR 200 / ENT 200: Creativity, Innovation, and Design (Alice Kogan)

EGR 361 / ENT 361 / URB 361 / AAS 348: The Reclamation Studio: Humanistic Design applied to Systemic Bias (Majora J. Carter)

ENVIRONMENTAL STUDIES

ENV 221: AI for Global Good (Jamie M. Caldwell and Jeffrey R. Smith)

GERMAN

GER 525: Studies in German Film: Media in Film (Joseph W. Vogl)

HISTORY

HIS 291: The Scientific Revolution (Matthew L. Jones)

JOURNALISM

JRN 414: Data Journalism: Investigative Journalism and Open Source Investigations (staff)

NEAR EASTERN STUDIES

NES 369 / HIS 251 / JDS 351 / CDH 369: The World of the Cairo Geniza (Marina Rustow)

PHILOSOPHY

PHI 312: Computability and Logic (John P. Burgess)

POLITICS

POL 345 / SOC 305 / SPI 211: Introduction to Quantitative Social Science (Nicolas Idrobo)

*POL 571: Empirical Research Methods for Political Science (Jonathan F. Mummolo)

*POL 573: Quantitative Analysis II (Rocío Titiunik)

PSYCHOLOGY

PSY 360 / COS 360: Computational Models of Cognition (Brenden M. Lake)

PSY 362 / CHV 362: Can Machines Be Moral? (Molly J. Crockett)

*PSY 503: Foundations in Statistical Methods for Psychological Science (Suyog Chandramouli)

*PSY 505: Current Issues in Statistical Methods and Research Methods for Psychological Science (Suyog Chandramouli)

STATISTICS & MACHINE LEARNING

SML 201: Introduction to Data Science (Daisy Yan Huang) (section 1 and section 2)

SML 301 / COS 301: Data Intelligence: Modern Data Science Methods (Derek E. Sollberger)

SOCIOLOGY

SOC 301: Statistical Methods in Sociology (Tod G. Hamilton)

*SOC 500: Applied Social Statistics (Brandon W. Stewart)

*SOC 505: Research Seminar in Empirical Investigation (Yu Xie)

*SOC 557: Technology Studies (Half-Term): Artificial Intelligence, Computation and Society (Zeynep Tufekci)

SPANISH

SPA 368 / TRA 368: Spanish into English Translation in the Age of AI (Catalina Arango)

PUBLIC & INTERNATIONAL AFFAIRS

SPI 352 / COS 352: Artificial Intelligence, Law, & Public Policy (Peter Henderson)

SPI 365: Tech/Ethics (Steven A. Kelts)

SPI 374: Challenges in Regulating AI: A Sociological Approach (Zeynep Tufekci)

*SPI 507B: Quantitative Analysis for Policymakers (staff)

*SPI 507C: Quantitative Analysis for Policymakers (Advanced) (Eduardo Morales)

SCIENCE AND TECHNOLOGY CENTER

STC 349 / ENV 349 / JRN 349: Writing about Science (Michael D. Lemonick)

URBAN STUDIES

URB 385 / SOC 385 / HUM 385 / ARC 385: Mapping Gentrification (Aaron P. Shkuda)

VISUAL ARTS

VIS 215 / CWR 215: Graphic Design: Typography (David W. Reinfurt)

VIS 216: Graphic Design: Visual Form (David W. Reinfurt)

VIS 218: Graphic Design: Image (Laura Coombs)

VIS 220: Animation I (Tim Szetela)

PRINCETON WRITING PROGRAM

WRI 139 and 140: How to Raise a Machine (Allen Durgin)

*WRI 501: Reading and Writing About the Scientific Literature (Andrea L. DiGiorgio)

*WRI 502: Writing a Grant Proposal in Quantitative Disciplines (Andrea L. DiGiorgio)

Introducing MuSE: Linking Music-Theoretical Concepts Across Languages

2026年4月13日 23:37

This past January, the CDH kicked off a new Collaborative Research Partnership with Anna Yu Wang (Assistant Professor of Music) and Jürgen Hackl (Assistant Professor of Civil and Environmental Engineering), the researchers behind the larger Music Theory in the Plural project. MuSE—short for Multilingual Semantic Embeddings—asks: when scholars write about music theory across languages, including Chinese, Japanese, Spanish, and Portuguese, are they talking about the same things? The project sets out to evaluate whether multilingual LLMs can translate the domain-specific discourse of music theory without flattening its nuance, and to test computational methods for discovering related concepts across those language traditions.

Led by RSE Laure Thompson, the CDH team is working with the scholar-translated articles in Music Theory Online volume 30, number 4, which presents articles written in Chinese, Japanese, Portuguese, and Spanish (among others) alongside their English translations. This resource allows the team to assess automated translation against expert scholarly ones. From there, the team will experiment with embedding models as tools for surfacing cross-linguistic connections across a broader corpus. The work aims not just to answer these questions for music theory, but to contribute broader, critically needed comparative research on how LLMs perform with specialized humanistic content.

Like all CDH Research Partnerships, MuSE begins with a project charter that defines the project's scope, deliverables, team roles, and terms of collaboration. The MuSE collaboration is one small part of the larger research agenda of the Music Theory in the Plural project, and it is through chartering that the project team and the RSE team define how the pieces will fit together. You can download the MuSE project charter here and follow the project's progress on its project page.

The Center for Digital Humanities is seeking proposals for innovative and computationally-engaged research partnerships from Princeton faculty. The next application cycle closes April 17, 2026. Apply now.

MuSE (Multilingual Semantic Embeddings)

Linking concepts in music-theoretical texts across languages

Built by CDH
music

Collaborative Research Partnerships

Faculty are welcome to apply to work with the CDH Research Software Engineering team!

collaborate_AdobeStock_429621181

AI and the humanities: Across the Princeton campus, an era of collaboration is underway

2026年2月26日 04:28

By Allison Gasparini, Center for Statistics and Machine Learning, and AI Lab

Originally published on the Princeton homepage.

With the launch of the Princeton Laboratory for Artificial Intelligence and the New Jersey AI Hub, over the last few years Princeton University has firmly established its presence at the forefront of artificial intelligence research — including transformative work in humanities scholarship. 

From piecing together fragments of ancient texts with language models to exploring the future of human-robot interactions, Princeton scholars aren’t just exploring what AI can do for the humanities. They’re uncovering what the humanities can do for AI.

Already, AI tools are appearing in all facets of our society and culture. “It’s the world that our kids are going to inherit,” said Meredith Martin, professor of English, faculty director of Princeton’s Center for Digital Humanities (CDH) and a grant recipient from the international Schmidt Sciences Humanities and AI Virtual Institute. “We should try our hardest to put the humanities into every aspect of AI development, not only in the input data and the interpretation of the results,” she said.

Princeton humanities scholars had already been using machine learning in their research — largely by partnering with the robust community of humanities research software engineers and digital humanities experts at CDH, established more than a decade ago. And now, with the AI Lab, they are proving to be some of the best positioned collaborators for the future of humanities and AI. 

“Our goal in the AI Lab is to support the transformative impact of AI on research across the Princeton campus, and we’re excited about the many opportunities to do so in the humanities,” said AI Lab Director Tom Griffiths.

“The most valuable and productive relationship between artificial intelligence and humanistic research is collaboration,” said Rachael DeLue, director of the University’s sweeping new Humanities Initiative. Here, a sample of the scholarship now underway.

Insights into ancient texts

Split image showing a seated person indoors in a jacket and collared shirt on the left, and a close-up of a printed page with large Chinese characters and vertical columns of text on the right.

Photo by Matthew Raspanti, Office of Communications

Paul Vierthaler uses machine learning to study premodern Chinese books such as “Xianqing ouji 閒情偶寄 (Leisure Notes),” a collection of essays published in 1671, pictured at right.

Fifteen years ago, Paul Vierthaler became fascinated by a particular figure common in Ming and Qing dynasty literature. Wei Zhongxian, a late Ming dynasty eunuch, nearly took over the imperial government in the 1620s. His infamy became such that within a year of his death, half a dozen novels and unofficial histories on his exploits had already been published. “I became really interested in how people talk about historical events within so-called ‘unreliable genres,’” said Vierthaler.

Studying the historical figure of Wei and the stories he’d inspired raised new questions for Vierthaler: How common were these types of narratives in imperial China? What more could be learned from studying bibliographic information, and how could he even begin to study that information at scale? “I realized, if I wanted to try to get a grip on how these kinds of narratives existed in the Chinese literary tradition, I needed to start thinking more broadly,” he said.

The realization led Vierthaler, now an assistant professor of Chinese literature and interdisciplinary data science, to the global catalog WorldCat, which holds digitized bibliographic records from tens of thousands of libraries around the world. To grapple with the vast amounts of data, Vierthaler turned to computational methods and machine learning/AI — which has altered the scope of his work.   

Using machine learning to analyze and extract data from written descriptions of premodern Chinese books — which may include information on author, illustrations, content and more — Vierthaler initially set out to understand whether genres containing suspect stories about historical events increased or decreased in popularity over time. But soon his focus expanded. “It’s exploded into a much larger project when I began to apply these same tools to digitized versions of the books themselves,” he said.

More recently he has been using machine learning methods to study the likely authorship of anonymously published works and to detect historical documents inserted into novels. With the advent of transformer-based language models, Vierthaler is now training specialized language models on premodern Chinese corpora, hoping to pick up minute nuances in the texts he studies. “There’s a movement now in the humanities aimed at training much more targeted, bespoke smaller language models,” Vierthaler said.

Instead of training bigger and bigger models, Vierthaler’s work has wider implications for using custom training sets to capture and retain the nuance necessary for the study of culture. “Humanities scholars can bring an understanding of the historical background and composition of training materials, which can help identify blind spots that might have otherwise been missed,” he said.

That same movement for highly specialized language models drives the work of Marina Rustow, the Khedouri A. Zilkha Professor of Jewish Civilization in the Near East and professor of Near Eastern studies and history. Rustow is training a model designed to transcribe fragments of medieval texts — and save humans time-consuming, painstaking labor.

Split image showing a fragment of a handwritten manuscript with dense script on the left and a person seated on a wooden floor by a window on the right, holding a small object beside a patterned surface.

Geniza fragment courtesy of Cambridge University Library. Photo: Sameer A. Khan.

Left: A government decree from the Fatimid period in Egypt (969–1171) shows the original Arabic inscription with wide line spacing. Right: Marina Rustow.

Rustow runs the Geniza Lab, a group dedicated to studying an enormous cache of paper and parchment recovered from a medieval synagogue in Cairo. The documents are unique because, unlike most ancient texts preserved today, they’re not the work of society’s elites and philosophers. They’re everyday records from the masses, things like complaints about business travel, heated personal letters and descriptions of stomachaches.

The fragments offer a broad perspective of society in that period, which is invaluable to Rustow as a social historian of the medieval Middle East. But they’re written in dialects of Arabic, Hebrew and Aramaic no one speaks today and in handwriting that can border on illegible.

Rustow said it can take her two full days to transcribe just one document. She hopes the machine learning model she’s training will turn that around for the 36,000 Geniza fragments that she and her colleagues have uploaded to the Princeton Geniza Project database for public access. 

Since the Geniza’s discovery in 1896, “it has taken researchers 130 years to transcribe 7,000 documents, and we have another 29,000 documents to transcribe,” said Rustow. With machine learning, she hopes to save scholars 530 years of transcription drudgery.

Like Rustow, Barbara Graziosi is on a mission to make premodern texts free and accessible for all. “Ideally, I’d like to see everything that we have from before the invention of printing preserved, made accessible, translated, well edited, well understood and well studied,” said Graziosi, the Ewing Professor of Greek Language and Literature and professor of classics.

Split image of a person reading a large historic book in a library setting and a close-up of hands pointing to lines of text in an open manuscript.

Photos by Denise Applewhite, Office of Communications

Barbara Graziosi uses AI as a "collaborator" to study ancient texts, including this 13th-century Byzantine manuscript of Aristotle's "Organon" from Princeton's Special Collections.

Graziosi is contributing to that mission by filling in the gaps of fragmented ancient Greek text. Over millennia, words and phrases written on documents are lost, chewed away by mice, eroded by moisture or obscured by stains. When a student approached her and suggested — before the advent of ChatGPT — that language modeling could generate suggestions to fill the gaps in these papyri, Graziosi began work on a machine learning tool attuned to the nuances of ancient Greek. The result of that work is the Logion Project, which Graziosi leads.

Graziosi said AI works best as a collaborator for humanists, not a replacement for highly trained scholars who dedicate years to studying these difficult texts. The tool Graziosi developed provides several suggested words to fill a given gap in the text. Seeing multiple suggestions can jog the thinking of a scholar who might face a block after spending hours reading and rereading the same passage.

The humanities, as Graziosi sees it, can help shape a future where the strengths of humans and the strengths of machines work in harmonious collaboration. “It’s very important that we keep the conversation going and that we respect human expertise as well as machine confidence,” she said. “I hope more humanists will get involved with AI because their perspective is exactly what’s needed now.”

Modern-language applications

Person gesturing toward a large display showing a multilingual sentence comparison and named-entity recognition examples for African languages, with a whiteboard in the background.

When Happy Buzaaba moved to Japan in 2015 to study for his Ph.D. at the University of Tsukuba, he couldn’t speak any Japanese.

“I had to rely on translation apps for my everyday life,” said Buzaaba, who is now an associate research scholar at Princeton Language and Intelligence with affiliations to the Center for Digital Humanities, the African Humanities Colloquium, the Princeton Institute for International and Regional Studies (PIIRS), and the Africa World Initiative.

Japanese-to-English translation is widely available, as both are well-studied languages with vast digital footprints. But some languages — in particular, many of those spoken in Africa — have a scant internet presence on which to train AI models. “I started thinking, imagine you went to a country where they speak a language that you don’t understand, and it’s also not supported by any existing technology,” said Buzaaba.

With computational linguist Christiane Fellbaum, Buzaaba is now introducing African languages to LLMs by creating large collections of syntactically annotated text, called treebanks, for 11 African languages.

The idea is that the rich annotations, filled with linguistic knowledge, can be used to train LLMs on the African languages, even though there’s not as much text as what’s available for Japanese or English. “We can actually create models that perform well on these languages, even with less amount of data,” Buzaaba said.

He and his colleagues have already released three African language models, which benchmarks show to be the best performing models of their kind. “The main goal here is not just creating tools, but accessibility,” he said. He has also brought this work into the classroom, on campus and in a PIIRS Global Seminar in Kenya.

AI in arts and architecture

Split image showing a seated person in a library or gallery space with books displayed on shelves, and another person seated at a computer workstation in a dimly lit room.

Photo by Matthew Raspanti, Office of Communications

Elizabeth Margulis uses AI to study how people describe their musical experiences in her Music Cognition lab (pictured at right: Itamar Jalon, a postdoc in psychology and music).

Humanities faculty in music, creative writing and architecture, among other disciplines, are using AI to understand the very essence of human creativity and to inform new work.

Elizabeth Margulis, a professor of music and acting department chair, is trying to understand how music shapes our emotions, imaginings and the thoughts that arise when we let our minds wander. Machine learning, she said, has been instrumental in advancing the studies she conducts for her Music Cognition Lab.

Margulis uses AI to study how people describe their musical experiences. “Where machine learning has been really helpful is giving us a way into unconstrained, free-response descriptions of what music evokes,” she said.

Researchers at the Music Cognition Lab collect these responses from volunteers, who enter a booth, listen to a musical excerpt, and then describe in writing the imaginative scenarios and emotions that arise. The lab has also worked in collaboration with Princeton University Concerts. At a Takács Quartet performance this past spring, the researchers gathered free-response descriptions from hundreds of concertgoers.

With the help of large language models, Margulis and her team analyze the descriptions, looking for patterns they might not have elucidated without the help of AI tools. “What’s so cool about machine learning is it helps us see structure in what seem like singular, subjective experiences,” said Margulis.

What she’s found so far is that people from different cultures frequently have remarkably different emotional and imaginative reactions after listening to the same piece of music. For example, one atonal excerpt by Anton Webern often conjured up a sense of impending doom for English speakers from the American Midwest. However, Dong speakers from the Guizhou province in China tended to imagine joyfully playing outside with friends.

Margulis hopes this work opens new avenues for understanding spontaneous thought in a way that could be applied to clinical settings somewhere down the road. “Think about ADHD or anxiety — both have these components that reside in patterns of spontaneous thought,” said Margulis. “Music gives us a powerful way to study the susceptibility of those thoughts to perceptual influence.”

A.M. Homes, professor of the practice in creative writing and the Lewis Center for the Arts, has been doing a lot thinking herself lately about AI. “I’m one of the writers whose books have been fed to AI to train on,” Homes said. “I sit on the Writers Guild of America’s Council on AI, and we are very concerned about how AI is being used in the entertainment world.”

Seated person in a dark suit on a white chair against a blue wall, with a narrow shelf of colorful books mounted behind them.

Photo by Matthew Raspanti, Office of Communications

A.M. Homes in creative writing is working on a novel that explores themes of grief and what could happen when people turn to artificial intelligence for comfort.

To navigate this complicated moment of murky boundaries surrounding AI use, Homes is doing what she does best — writing about it. AI isn’t a tool she uses in her creative life; instead she’s working on a novel that interweaves themes of grief and explores what could happen when people turn to artificial intelligence for comfort.

Homes is a fiction writer who taps into the ideas percolating through society and culture at large. She sees the author’s role as being an artist who conceptualizes worlds and futures that don’t exist, inviting readers to think critically about the one they live in. “Whenever I’m writing something, what I really want to inspire is discussion and conversation,” she said.

In Arash Adel’s ideal future, humans and AI aren’t at creative odds but work together. While pursuing his Ph.D. at ETH Zurich, Adel focused on computational design and robotic integration into architecture construction. “But I wondered about the role of humans,” he said. 

Split image of a person holding a controller indoors and an aerial view of a curved wooden pavilion with one person standing inside on grass.

From L to R: Photos by Daniel Ruan and Bob Berg; courtesy of Arash Adel

Arash Adel and his team have recently built Timbrelyn, a robotically fabricated structure on the historic grounds of the 1969 Woodstock Festival in Bethel, N.Y.

Now an assistant professor in the School of Architecture with an affiliation at Princeton Robotics, Adel investigates human-robot collaboration where people supervise and instruct while robots perform some of the physically demanding and potentially dangerous construction tasks. This type of human-robot partnership, he said, is driven forward with AI models.

In 2024, Adel’s research group put this approach into practice for their Timbrelyn installation on the grounds of the 1969 Woodstock Festival in Bethel, N.Y., a raised wooden platform created from intricate layers of lumber. Using AI vision, robots scanned inventories of reclaimed and new lumber to identify wood elements that met design specifications while minimizing waste. After selection, the robots grasped and processed the elements using a saw before assembling them with a human collaborator.

Adel and his group are now working on a project that involves AI assisting in the design process as well. The goal is to ultimately develop a pipeline where humans and robots collaborate from inception to final construction.

“Humans are very intuitive, but we struggle to process large amounts of information at once,” said Adel. “The role of the AI is to augment human creativity.”

Connecting engineers and humanists: the Center for Digital Humanities

Split image showing annotated pages from a classic literary text on the left and a seated person in a studio portrait on the right, resting an arm on a knee against a dark background.

Photo by Kristopher Johnson

The Princeton Prosody Archive (left) is a searchable database of thousands of English-language digitized works published between 1559 and 1928, directed by Meredith Martin (right).

By the time ChatGPT exploded onto the scene in 2022, the Center for Digital Humanities had already been situated at the cutting edge of humanities-technology collaboration for the better part of a decade. 

The center equips Princeton humanities faculty to thrive in a tech-dominated landscape, connecting them with software engineers who build the bespoke software that underlies projects (like Rustow’s and Graziosi’s) and teaching humanists and software engineers how to successfully collaborate.   

“We at CDH had already built the necessary collaborative infrastructure for projects involving both software engineers and humanists,” Martin said. With the rapid proliferation of generative AI tools, she has noticed a surge of humanities scholars approaching CDH with questions about how the new technology might transform their work.

There are obvious advantages to AI: faster processing of larger datasets, high performance computing, quicker pattern recognition. But these technological leaps aren’t enough on their own, Martin said. “Humanists have to bring a lot of knowledge to that interaction for it to work out.”

For that reason, the staff at CDH think carefully about how machine learning might fit into a research project and whether a particular approach would be the right fit for their question. At the same time, Martin sees an opportunity for humanists themselves to shape the AI tools.

“There’s no reason why humanists can’t feel empowered to build better models, to participate in model architecture, to think about the kinds of data on which various models are trained and why,” she said.

This desire to bring humanists into the AI fold helped inspire a three-part project developed by CDH that spans the 2025-26 academic year and beyond. The project, Modeling Culture: New Humanities Practices in the Age of AI, brings together Princeton faculty and researchers from other universities for a seminar series to think critically about AI.

Martin ran one of the seminars this fall with Matthew Jones from the Department of History and Andrew Janco, a digital scholarship specialist at Firestone Library, on the problems and questions of modeling. “The main feeling has been one of real empowerment and excitement,” said Martin. “In the room, you can feel people leaning forward.”

CDH becomes first U.S. institution to participate in major European research infrastructure project

2026年2月24日 03:40

The Center for Digital Humanities (CDH) is proud to announce its collaboration with ATRIUM (Advancing fronTier Research In the arts and hUManities), a major European research infrastructure project funded by the European Commission and coordinated by DARIAH-EU. Princeton CDH is the first U.S.-based institution to work with ATRIUM partners, aiming to generate new avenues of transatlantic collaboration for digital humanities research.

As part of this partnership, CDH, the UNESCO Chair on Digital Methods for the Humanities and Social Sciences, and the Athena Research Centre will co-organize the ATRIUM Summer School titled “From Maps to Data and Data to Maps: Exploring Spatial Histories.”

The four-day workshop, which will take place at the Athens University of Economics and Business on June 29 – July 2, 2026, will bring together graduate students and early-career scholars from the U.S. and Europe to explore cutting-edge methods for analyzing and visualizing spatial data drawn from historical maps and geographic sources. Participants will gain hands-on experience with tools and approaches that are shaping the future of spatial humanities research.

“This collaboration opens an important channel between U.S. and European digital humanities communities,” said Meredith Martin, Faculty Director of the Center for Digital Humanities. “Affiliation with ATRIUM will allow us to connect American students and scholars with the innovative research trends, tools, and networks being developed across Europe — and contribute Princeton’s own expertise to that exchange.”

ATRIUM unites leading humanities research infrastructures across Europe, creating shared access to advanced digital tools, datasets, and expertise. The project also supports intensive training opportunities that foster cross-border scholarly exchange and build capacity in emerging digital methods.

“ATRIUM is about consolidating the European research infrastructure landscape, and meaningful international partnerships are an integral part of that effort,” said Toma Tasovac, Principal Investigator of ATRIUM. “We are delighted to build on our existing relationship with Princeton CDH as a DARIAH Cooperating Partner, and to explore new avenues of collaboration within the ATRIUM framework.”

The summer school reflects CDH’s ongoing commitment to advancing interdisciplinary, computationally engaged humanities research and to developing international partnerships that expand opportunities for scholars at all career stages.

For more information about the summer school, visit: ATRIUM Summer School 2026 Call for Participation. To learn more about the ATRIUM project, visit: atrium-research.eu.

ATRIUM is funded by the European Union under Grant Agreement n. 101132163.

Revisiting the 2025 Vienna HTR Winter School for Medievalists

2026年2月18日 22:34

Earlier this winter, CDH / MARBAS Postdoctoral Research Associate Christine Roughan returned to Vienna for the second year in a row to share her experience using Handwritten Text Recognition (HTR) technology for medieval texts.

The workshops were part of HTR Winter School 2025, hosted by the Institute for Medieval Research of the Austrian Academy of Sciences in collaboration with MARBAS and the Institute for Habsburg and Balkan Studies. The in-person sessions followed three virtual workshops that brought together scholars of Carolingian Latin, Byzantine Greek, and Syriac, among other languages.

“This year I reprised my role as a group leader for the Syriac HTR group alongside Ephrem Aboud Ishac (Austrian Academy of Sciences),” Christine explained. “In addition, I provided instruction to the cohort as a whole on how to apply their HTR training in different contexts, so that their new skills were not tethered to only a single tool.”

As she noted in an interview last year, Christine became involved with Winter School after giving a talk at the Institute for Medieval Research, where she met several of the organizers. At the time, she said that Winter School offered an opportunity for her to hone her skills in teaching methods that play an important role in her own work.

Even more important this year: learning about how this year’s participants will use workshop content to advance their own work.

“My favorite part of the experience was definitely hearing about the variety of research topics the participants were engaged in,” Christine explained. “Seeing their enthusiasm for how the Winter School experience would equip them to dive into those projects was really great, especially in the final days when everyone was now practiced with the methodologies and ready to take off on their own.”

For more information on Winter School, visit https://www.oeaw.ac.at/en/imafo/the-institute/detail/htr-of-historical-sources.

From Philadelphia to Luxembourg: RSE Team at Fall Conferences

2026年2月18日 22:19

Last Fall, the CDH Research Software Engineering team traveled near and far to share their cutting-edge work and learn from others about new developments in the fields of research software engineering and digital humanities.

🔔 Philadelphia

From October 6–8, the RSE team participated in the third annual conference of the United States Research Software Engineer Association. Held in Philadelphia, US-RSE’25 brought together RSEs from across universities, laboratories, industry, and other institutions, as well as their managers and allies, to discuss “Code, Practices, and People.” CDH represented the small — yet mighty! — Humanities contingent in a field of mostly scientists and social scientists.

Wide shot of a conference ballroom with chandeliers and an audience seated facing a large screen reading Welcome to Philadelphia! for USRSE 2025.

US-RSE'25

Lead RSE Rebecca Sutton Koeser presented a notebook on undate, a Python package for computing with uncertain and partially-unknown dates, such as those in Sylvia Beach’s lending library records and in the fragmentary Hebrew texts in the Princeton Geniza Project. Koeser also collaborated on two posters: “Community Code Review in the Digital Humanities,” which detailed the history, process, and future work of the DHTech Code Review Working Group; and “Surveying the Digital Humanities Research Software Engineering Landscape,” which reported on survey results about the backgrounds and career paths of DH developers.

Photo of two people standing indoors on either side of a large conference poster covered with charts, diagrams, and text, displayed on a stand against a wall.

Rebecca Koeser (left); Julia Damerow (right)

Postdoctoral Researcher Christine Roughan presented a paper, also co-written with Koeser, titled “Integrating ATR Software with University HPC Infrastructure: balancing diverse compute needs.” The paper and corresponding presentation described the methods and outcomes of Bringing HTR to the HPC: A Pilot to Customize eScriptorium for Princeton, a Research Partnership with the CDH under the umbrella of the Princeton Open HTR Initiative (funded by a 2024–25 Princeton Language + Intelligence Seed Grant). Conference attendees were fascinated to hear about how Koeser and Roughan implemented an instance of eScriptorium — the current leader in open-source handwritten text recognition software — on Princeton’s high-performance computing hardware, which enabled professors and students without advanced technical skills to train large text-recognition models customized to their documents’ needs.

Person speaking at a podium with a microphone at the Marriott Philadelphia Old City, with a patterned wall behind them and an audience seated in the foreground.

Christine Roughan presents at US-RSE'25

Assistant Director Jeri Wieringa and Project Manager Mary Naydan presented on a panel about supporting and managing RSE projects. Their presentation “Creating Research Software with Humanities Faculty” highlighted the CDH’s chartering process, which helps transition humanities faculty from the individual, expansive mode of traditional humanities scholarship to the collaborative, modular mode of computational research. Their illustrative opening skit, which set the stage for the talk, drew lots of laughter and resonated with audience members. The rest of the panel was just as engaging, sharing lessons learned from many different types of organizations and fields, from the multi-institutional development of medical technologies, to a large laboratory focused on national security, to a lone RSE’s personal project management workflow at a research university (Naydan’s favorite presentation of the conference!).

Wide shot of a conference room with an audience seated facing a projected slide about faculty training, while two presenters stand at the front near a podium.

Mary Naydan and Jeri Wieringa present at US-RSE’25 in Philadelphia.

The technical talks affirmed that the CDH RSE team is ahead of the curve on best practices in Python development, such as using uv for installing packages and choosing Marimo over Jupyter for notebooks. Koeser noted the field-wide shift in starting to think about notebooks as a form of publication, and CDH Research Software Engineer Hao Tan was inspired by Reed Maxwell’s keynote on creating groundwater simulations using physics-informed machine learning. Tan reflects, “Explainability is still a real issue, but rather than rejecting AI outright, we should learn to leverage its strengths and mitigate its weaknesses — through comparative evaluation, transparent step-by-step reasoning, and other methods we develop.”

Many of the conference’s presentations — from Maxwell’s keynote to the Birds of a Feather workshop “AI in Practice” — showed the field of research software engineering grappling with, adapting to, and incorporating AI. While our team entered the conference thinking our challenges were somewhat unique to the humanities, we were surprised to see RSEs from across disciplines encountering similar challenges around this topic: from defining research questions, to gathering sufficient data, to disabusing researchers about what AI can actually do.

🇱🇺 Luxembourg

Unsurprisingly, AI was also a popular topic at the 2025 Computational Humanities Research Conference, held at the Luxembourg Centre for Contemporary and Digital History (C²DH) at the University of Luxembourg from December 9–12, 2025. Many of the presentations focused on benchmarking various models for automatic transcription tasks, or using chatbots to scale up annotation data, from identifying “acts of God” in contemporary Christian fiction to assessing a popular song’s “narrativity.” Miguel Escobar Varela’s keynote, “‘A watch by Kran Kamu’: Exploratory fine tuning for cultural reliability,” discussed using supervised fine-tuning on large open-weight models to yield reliable results within highly specific cultural contexts, such as Southeast Asian historical newspapers — a common problem facing computational humanities researchers given the specialized nature of our data and the scarcity of it for fine-tuning.

Photo of two people standing indoors beside a research poster titled Unstable Data, with charts and text displayed on a wall behind them.

Rebecca Koeser (left); Mary Naydan (right)

Rebecca Koeser and Mary Naydan presented a poster based on their short paper “Unstable Data and the Unusual Case of the Prosody Excerpt in the Digital Library” (co-authored with Meredith Martin). Using the HathiTrust materials contained in the Princeton Prosody Archive as a case study, Koeser and Naydan cautioned researchers that the page-level data provided by cultural heritage aggregators is not as stable as we might assume. This instability can lead to erroneous data, flawed conclusions, and difficulties building on previous scholarship.

Wide shot of a lecture hall with a projected slide titled Word Segmentation – What is the Problem Here?, showing Japanese text examples, while a presenter stands at a podium pointing toward the screen.

Hao Tan presents at 2025 CHR Conference in Luxembourg.

Hao Tan delivered a lightning talk, “When Larger LLMs Aren’t Enough: Word Segmentation in Historical Chinese Texts,” which used word segmentation in historical Chinese texts as a case study to highlight how large language models, while powerful, can quietly introduce risks when applied to humanities research. The talk sparked conversations with researchers working on East Asian materials across Europe, the US, and Singapore, especially around the tricky parts of historical text processing, with projects ranging from power relations in historical fiction to poetic imagery and stylistic change in epitaphs.

In addition to Tan’s lightning talk, Koeser and Naydan found two others particularly interesting: Katarina Mohar’s on “Speculative Reconstruction and the Ethics of the Fragment: Early Experiments with Generative AI in Art History,” and Antonina Martynenko, Artjoms Šeļa and Petr Plecháč’s on "Where Empires End: Tracing the Geography of a ‘Soaring Spirit’ in Poetry.” Mohar discussed the possibilities and limitations of using Generative AI to fill in gaps in medieval paintings, and provided practical recommendations for how to use it responsibly. Martynenko et al. examined the spatial imagination of European poets by mapping the distance and directionality of place mentions.

For Koeser, one of the most interesting presentations was “Cluster Ambiguity in Networks as Substantive Knowledge,” which describes a method for running a clustering algorithm multiple times to measure how often edge nodes connect nodes in the same community, allowing researchers to identify ambiguous data. Koeser is interested in the interpretive power and potential applications of this method, such as identifying ambiguous characters in novels. Another highlight was Taylor Arnold and Lauren Tilton’s presentation on “Sitcom Form and Function: Pacing and Production in a Collection of Thirty U.S. Series,” which examined how trends in visual and aural pacing changed over time using a combination of large-scale computational analysis and close reading: an example of truly multimodal research and scalable reading.

The RSE team was energized by these shifts in the field: leaning into ambiguity; using audio, video, and visual data rather than defaulting to text; and combining different scales of reading (close and distant) to draw more responsible conclusions. We are excited to carry what we learned into our work this year on multilingual machine translation and term clustering in music theoretical texts and using vision-based LLMs to aid historical document transcription and data extraction.

View from an elevated platform with a hand on a railing overlooking a large industrial complex courtyard with a tall metal structure, surrounding buildings, and scattered tables below.

Photo by Hao Tan in Luxembourg.

Journal of Cultural Analytics Enters New Chapter with CDH, Joins Open Journals Collective

2026年2月18日 20:41

In January 2026, Princeton University's Center for Digital Humanities (CDH) began serving as publisher of the Journal of Cultural Analytics (JCA), a leading open-access publication in computational approaches to culture. Today, CDH announces JCA’s vision for expanding cultural analytics scholarship amid rapid technological change and the launch of a new website, supported by Schmidt Sciences’ Humanities and AI Virtual Institute (HAVI).

"The Journal of Cultural Analytics has been instrumental in advancing computational methods in the humanities," said Meredith Martin, faculty director of the CDH and professor of English at Princeton, who serves as one of the journal's three editors alongside Amelia Acker (Rutgers University) and Tanya Clement (University of Texas at Austin). "We are honored to lead JCA's continued evolution and grateful to Andrew Piper for his pioneering work in establishing this field-changing, scholarly venue."

Building on a Strong Foundation

Founded by Piper at McGill University's Department of Languages, Literatures, and Cultures, JCA has published groundbreaking data-driven research about culture since 2016. The journal encourages transparent research practices, including open sharing of data and code. It has become a cornerstone publication for scholars working at the intersection of digital humanities, computational social sciences, and computational approaches to culture.

"The idea for the journal was born in 2015 as a response to a shared sense that our field needed a venue dedicated to the critical use of computation to study culture," said Piper. "After a decade of growth, the journal has far exceeded my hopes. I'm extremely happy to see it continue under the leadership of the new editors and its new institutional home at Princeton's Center for Digital Humanities."

Looking Ahead: Expanding Scope and Impact

JCA is broadening its vision to serve an ever-evolving interdisciplinary and international scholarly community invested in cultural study and the methods by which we interrogate the digital in culture – especially in the age of AI. Central to this vision is the commitment to publishing work that goes beyond method for method's sake, asking instead how computational approaches to culture at scale can reshape what we know and how we know it.

The editorial board has expanded to 43 scholars representing institutions across North America, Europe, Asia, and Australia, reflecting JCA’s commitment to international perspectives and increasing representation from junior scholars. This expanded scope will support the journal’s growing focus on multi-lingual and multi-modal approaches to culture.

A new Special Features section, edited by Laura McGrath (Temple University), will highlight shorter, timely essays on computational cultural analysis written in an accessible style for non-specialist audiences, designed to spark discussion on new methodologies, datasets, or research.

JCA will deepen its focus on critical engagement with data, which is increasingly significant for AI researchers returning to smaller, human-curated cultural models. Welcoming a new data editor, Sarah Reiff Conell (Princeton University Library), JCA will revise its data-essay and dataset-review format in collaboration with the scholarly “data collectives” (such as Post45 and 19thC Data Collective), and provide a directory of datasets for cultural studies.

Upcoming Special Issues will explore topics ranging from computational humanities in the Global South to data-driven approaches to poetry and a retrospective on ten years of the JCA. The journal is currently accepting Special Issue proposals for 2027.

New Infrastructure for Open Access, Community-Led Publishing

The transformative support from Schmidt Sciences’ HAVI program has enabled JCA's growth and modernization, expanding the editorial team with new roles for graduate students—providing both recognition and compensation for the labor required to run an academic journal and an opportunity to train the next generation of computational humanities scholars. The grant has also enabled the journal to migrate to Janeway, an open-source publishing platform developed by the Open Library of Humanities, featuring a redesigned user interface and customizable workflow management system.

In this new phase, JCA maintains its commitment to diamond open access—free to read and free to publish, with no article processing charges (APCs) or publishing fees for authors or universities. JCA has also joined the Open Journals Collective, a coalition of libraries and university-based publishers that launched in March 2025, providing journals with technological support, financial sustainability, and community governance through a library-funded model that keeps research freely accessible and journals editorially independent.

"I'm thrilled to have such a prestigious Princeton journal carrying the banner for diamond open access as part of the launch collection. We're excited for JCA, and for the promise of the new, sustainable funding model OJC is delivering," said Matthew Kopel, Princeton's Open Access & Intellectual Property Librarian, who also sits on the Open Journals Collective Library Board.

More information about the journal's new direction, upcoming issues, and submission guidelines can be found on JCA's newly launched platform at https://culturalanalytics.org.

Editorial Team

Editors

  • Meredith Martin, Princeton University
  • Tanya Clement, University of Texas at Austin
  • Amelia Acker, Rutgers University

Special Features Editor

  • Laura McGrath, Temple University

Data Editor

  • Sarah Reiff Conell, Princeton University Library

Graduate Editorial Assistants

  • Cecelia Ramsey, Princeton University (Managing Editor)
  • Odalis Garcia Gorra, University of Texas at Austin
  • Haiqi Zhou, McGill University
  • Emilien Arnaud, Princeton University

Former Editorial Assistant

  • Katrin Rohrbacher
Three logos on a white background for the Center for Digital Humanities at Princeton, Schmidt Sciences, and JCA, displayed side by side.

About the Center for Digital Humanities

Princeton's Center for Digital Humanities, founded in 2014, advances computational and data-intensive humanities scholarship through collaborative research, innovative pedagogy, and community building to create a more just future. The center develops better practices in technological development and research while bringing humanistic perspectives to data science applications.

About the Open Journals Collective

The Open Journals Collective is a growing coalition of libraries and university-based publishers providing sustainable, community-led alternatives to commercial academic publishing. Through diamond open access and collective funding models, OJC supports hundreds of journals while ensuring research remains freely accessible to all.

About Schmidt Sciences

Schmidt Sciences is a nonprofit organization founded in 2024 by Eric and Wendy Schmidt that works to accelerate scientific knowledge and breakthroughs with the most promising tools to support a thriving planet. The organization prioritizes research in areas poised for impact, including AI and advanced computing, astrophysics, biosciences, climate, and space—as well as supporting researchers in a variety of disciplines through its science systems program. The Humanities and Artificial Intelligence Virtual Institute (HAVI) intends to spur innovative, domain-specific research outcomes from humanities scholars through the integral application of AI-inspired tools and techniques, as well as produce insights from the humanities that will advance the development of AI.

❌