Welcome to the new academic year! I write to you from this, my 20th year (!!!) at Princeton and my 12th (!!!!) at the helm of the Center for Digital Humanities. Thank you for your ongoing support of our unique interdisciplinary community. It’s been fantastic to have Associate Director Paul Verthaler join me this past year as we continue our mission of centering the humanities in conversations about the technologies that are transforming our research, culture, and society.
As always, we remain at the cutting edge of developments in AI, while continuing to foster thoughtful, timely, and urgent discussions about a future where humanistic approaches and values play a central role in our increasingly technical world.
Over the past year, we took up this core question — how AI and the humanities intersect — with Princeton faculty and graduate students in our Modeling Culture seminar, and by the campus-wide audiences who attended our workshops, events, and talks. Across CDH-affiliated labs, Princeton scholars join their domain expertise with testing the benefits and limits of AI applications. As always, we partner and advise on a wide range of projects — from curating data from ancient Middle Eastern texts, to developing technologies for African languages, to learning the possibilities of how AI might sustain long-form narrative. Our research software engineers are building custom tools for scholars who want to analyze complex, often multilingual or historical texts that conventional AI systems weren't built for, while also exploring open models and experimenting with new ways of accessing and exploring research materials.
A major milestone for the CDH in 2026 was taking on the role of publisher for the long-running Journal of Cultural Analytics (JCA), a leading journal for scholarship on the computational study of culture. JCA is just one piece of our involvement in a broader "Cultural AI" initiative. We’re also launching a working group, a collection of pedagogical experiments and discussions, and new graduate student and faculty opportunities. In collaboration with the newly launched Princeton Societal AI group at the newly launched Data and Intelligent Systems Initiative, we’re happy to support and participate in new cross-campus efforts to ensure that culture and society are at the forefront of how Princeton approaches data science, computational science, and AI.
We invite you to join these important and timely conversations. CDH is always looking for new collaborators, and we welcome inquiries from scholars at any level or from any division. Sign up for a consultation to discuss research or teaching ideas. Students can find us through our growing slate of undergraduate courses under the CDH course code. Subscribe to our monthly newsletter to hear about what's new and what's happening. Please join us at our Open House on Wednesday, September 16, and in October for a series of events with Leif Weatherby from NYU, media theorist and one of the most important voices in our understanding of how AI is affecting culture.
Last but not least, we're delighted to share news about the CDH family: Seyi Olojo joins us as a Postdoctoral Research Associate working on epistemological questions about African language technologies, and Christine Roughan has been promoted from postdoc to the staff position of CDH Project Manager.
Congratulations to Alison Fortenberry, co-winner of the 2026 CDH Senior Thesis Prize! Alison recently graduated with an A.B. in African American Studies with minors in Religion, American Studies, and English. Her thesis, entitled “Until Change Comes: The Hidden History of Denominational Formation and Racial Exclusion in Twentieth Century American Pentecostalism,” was advised by Wallace Best (African American Studies).
We asked Alison about her award-winning work.
Tell us about your project. How would you summarize it in two to three sentences?
My project explored the racialized formation of American Pentecostal denominations in the early twentieth century. American Pentecostalism began as an interracial revival movement largely led by African Americans, but it soon split into racially segregated denominations. My thesis argued that, for White Pentecostals, denominationalism was a project of racial exclusion and racial differentiation.
How did you get interested in your topic?
I grew up in a Pentecostal church, and three years ago, one of my mentors in the congregation passed away. At her funeral, her family insisted that her church membership card be added to a church archive, explaining that members of their family had been historically denied access to the church because they were Black. This surprised me, because the church I grew up in was predominantly Black, and our existing church histories obscured the church’s exclusionary past. I wanted to correct those church histories. At first, the project began just as an effort to make a church archive (to date, I have cataloged and digitized tens of thousands of church records), but as I read through church records and began to explore broader literature about the denomination’s history, I realized that I wanted to make this story the focus of my thesis. [Check out this link to explore some of the items Alison digitized from the church archive.]
Alison Fortenberry ’26
What role do digital humanities methods play in your analysis?
I tried to let the archive and broader corpus of literature I was drawing on guide my use of digital humanities methods in my analysis. A major component of my thesis was a historiography of a predominantly White American Pentecostal denomination, the Assemblies of God (AG). As I got into American Pentecostal literature, I noticed that historical authors who were affiliated with the AG tended to write the denomination’s history much differently than authors who were not, often downplaying the AG’s role in segregating an interracial religious movement. While I was able to demonstrate the differences in AG-authored histories through close readings, I wanted a way to quantify these trends. I worked with Andy Janco in Digital Scholarship to develop a Scattertext project, which tracked how frequently authors of different denominational backgrounds used specific words, then plotted them on a graph. With Scattertext, I was able to demonstrate that AG-affiliated authors use the names of Pentecostalism’s early Black leaders, for example, at disproportionately low rates.
What was a challenge you overcame as you worked on your project?
A major goal of my project was to reconstruct the lives of Black Pentecostals in White churches who had been excluded from White church archives. One of my subjects, Albertina Cooper, was a Black woman who attended a White church for over thirty years, yet in an archive of tens of thousands of objects, her name was only listed four times. Trying to flesh out the motivations and feelings behind her decisions was challenging. For example, I wondered why Albertina chose to attend this discriminatory White church when there were other Pentecostal churches in the city, and I assumed that maybe this was just the closest Pentecostal church to her home. I was able to use archived church minutes and a Philadelphia city directory to map Albertina Cooper’s home in relation to all of the city’s Pentecostal churches in the 1930s. I was surprised to find that the closest Pentecostal church to Albertina’s home was actually a Black Pentecostal church and that Albertina had to travel over three miles to get to her White congregation. Albertina’s choice to attend her church was not just a matter of convenience, but an active decision to be the only Black member of a White congregation. Digital mapping helped me reconstruct Albertina’s life and motivations in the face of archival silences.
What was one surprising thing you learned in the process—either about yourself or about your topic?
Through this process, I was surprised to learn that, even as someone with more traditional humanistic training and interests, digital tools are accessible to me. I never thought that I had the ability to conduct more quantitative or computational research, but with the help of Digital Scholarship, the Maps & Geospatial Information Center, and the Center for Digital Humanities, I was able to incorporate digital humanities methods to deepen and extend my qualitative research.
How might your work on your thesis inform your future?
This fall, I’m returning to Princeton to begin a PhD program in the Department of Religion. I plan to continue using methods from the digital humanities in my doctoral work and beyond.
Congratulations to Christina Li, co-winner of the 2026CDH Senior Thesis Prize! Christina recently graduated with a B.S.E. in Operations Research and Financial Engineering and minors in Art History and French Language and Culture. Her thesis, entitled “The Value of Art,” was advised by Alain Kornhauser (ORFE).
We asked Christina about her award-winning project.
Let’s start with a quick summary. How would you describe your work in a few sentences?
My thesis investigates how monetary value is formed in the fine art auction market through social, political, economic, and artistic forces—or a lack thereof. Using a dataset of fine art auction sales, I analyze the factors that determine whether an artwork sells and, if it does, how much it sells for. I find that social prestige and institutional recognition are the strongest predictors of success at auction. Ultimately, however, my thesis asks what art is worth to the individual and to society, beyond its monetary value.
Ernst Ludwig Kirchner, Street, Berlin (1913), in the Degenerate Art exhibition (upper level, room 4). Courtesy of The Museum of Modern Art (MoMA).
Ernst Ludwig Kirchner, Berliner Strassenszene (Berlin Street Scene), 1913–14. Sold at Christie’s.
Christina explains: “I use the story of the Degenerate Art exhibition, and these Kirchner paintings in particular, as an introduction to my thesis, framing the discussion about the relationship between art, value, and institutions.” The top painting included a label that "noted it was ‘purchased with the taxes of working German people’ by the Nationalgalerie in 1920, for 12,000 German marks (or $100,000 USD today). This seemingly high figure was intended to provoke visitors.” In contrast, the “very similar Kirchner [below] was sold at auction at Christie’s in 2006 for over $38 million USD.”
What drew you to your topic?
I am deeply interested in art history, and working with auction data allowed me to combine this passion with my technical background in Operations Research and Financial Engineering. At Princeton, I have taken several courses in modern and contemporary art history where we discussed the commercialization of art.
Tell us about the computational aspects of your project.
To analyze auction results, I estimated hedonic regressions with fixed effects, difference-in-diffferences, and event study models around major exhibitions, allowing me to examine how institutional recognition affected artists’ market values over time. I also implemented neural network models to predict auction outcomes based on artwork, artist, and auction characteristics, providing a complementary approach to understanding the factors that drive prices.
What is one takeaway—either personal or intellectual—from your senior thesis experience?
I am inspired and encouraged by the potential for the humanities to be deeply interrogated and enriched through the application of engineering tools and principles.
Any connections between your thesis and your future plans?
My dream is to become a curator. My thesis has reinforced my interest in the important role that art plays in our institutions and societies.
Note: for a list of grad seminars offered or cross-listed by the CDH that automatically count toward the Graduate Certificate in DH, please seeGraduate Courses in DH.
Meredith Martin, Professor of English and Director for Center for Digital Humanities at Princeton, was a co-organizer of the inaugural "Culture × AI: Evaluating AI as a Cultural Technology," a workshop held at the International Conference on Machine Learning (ICML) on Friday, July 10, 2026, in Seoul, South Korea.
The ICML is the leading annual conference for machine learning research, publishing foundational work underlying much of modern AI research. This year, ICML’s workshop program was particularly selective: “Culture × AI” was one of only 44 workshops chosen out of 247 proposals, and marks the first time ICML has hosted a workshop dedicated to this topic.
"Culture × AI" examines how generative AI systems function as cultural technologies — producing text, images, and video shaped by vast amounts of social and cultural data — and asks how the humanities, arts, and qualitative social sciences can inform AI development from the outset, rather than being applied only after deployment to mitigate harm. The workshop's central theme, "Interpretive Technologies," explored how humanistic traditions of meaning-making and contextual and aesthetic judgment can be built into AI design.
Speakers included prominent voices in Cultural AI: Lauren Klein (Emory University), Ted Underwood (University of Illinois Urbana-Champaign), Maria Antoniak (University of Colorado, Boulder), and Joel Z. Leibo (Google DeepMind), along with a panel discussion, a lightning talk session featuring accepted papers from researchers at McGill University, KAIST, Duke University, and the University of Groningen, and a closing poster session.
Martin's involvement reflects Princeton's growing engagement with the intersection of the humanities and artificial intelligence, extending the work of the university's Center for Digital Humanities into international, cross-institutional conversations about how cultural knowledge can shape the future of AI systems.
The Center for Digital Humanities (CDH) at Princeton University announces the publication of "Princeton Geniza Project Datasets," a new data essay by Rebecca Sutton Koeser (CDH) and Marina Rustow (Near Eastern Studies/History), published in the Journal of Cultural Analytics (Koeser and Rustow 2026, 11(3): 1–35, DOI: 10.22148/jca.1084). The essay presents, for the first time, a full account of the data behind the Princeton Geniza Project (PGP) — one of the richest digital resources for the study of the medieval Islamic and Jewish Mediterranean world.
The PGP dataset opens rich and complex historical material to data science approaches
The Cairo Geniza — a repository of discarded largely Hebrew-script writings in the Ben Ezra Synagogue in Fustat, Old Cairo — preserved hundreds of thousands of documents from everyday medieval life c. 950–1250, from personal letters and legal contracts to trade accounts. These fragments offer scholars an unusually candid window onto medieval culture stretching from Spain to Sumatra. Since the 1980s, Princeton’s Geniza Project has been collecting, documenting and annotating these fragments, which are now scattered across some seventy libraries worldwide. Today, the PGP contains 35,855 Geniza documents — with descriptions, and information about document types, languages and scripts, dates across multiple calendars, transcriptions, translations, and bibliographic citations — along with information on 1,802 people and 486 places.
The publication of the dataset – a complement to the public-facing PGP website – makes the complex material accessible for data science and AI researchers and historians alike. Project PI Marina Rustow’s collaboration with the Center for Digital Humanities has made the PGP one of the leading efforts bringing medieval manuscripts into the world of data science and AI — alongside projects such as the multi-institutional OpenITI project for Arabic texts and the National Library of Israel's KTIV for Hebrew manuscripts.
The essay documents the PGP’s 40-year history — from an IBM-funded 1986 pilot to the five-and-a-half-year PGP–CDH partnership (2020–2025) that rebuilt the database from the ground up — and explains the modeling decisions behind the data, including the many-to-many relationship between physical fragments and the documents written on them.
"Providing both a well-structured dataset, as well as the contextual information of its creation, enables a broader community of researchers to engage with the Geniza material," states the authors. "The study of the medieval Mediterranean world can now be explored, interrogated, and modeled — inviting the tools of data science and AI to reveal patterns across hundreds of thousands of fragments that no single scholar could trace by hand."
About the Dataset
The Princeton Geniza Project dataset is archived on Zenodo (DOI: 10.5281/zenodo.18716581) and updated quarterly. The essay itself is open access under a Creative Commons Attribution 4.0 license. The Princeton Geniza Project is the flagship project of the Princeton Geniza Lab (PGL), directed by Marina Rustow; PGL is an affiliate lab of the Center for Digital Humanities.
Eve Krakowski and Marina Rustow, co-director and director of Princeton's Geniza Lab — an affiliate lab of the Center for Digital Humanities — have each been awarded $300,000 from the National Endowment for the Humanities to produce new scholarly editions and translations from the Cairo Geniza, a cache of over 400,000 medieval manuscripts. Krakowski's project traces the social history of death in medieval Egypt, while Rustow's brings to light Jewish traders' letters, legal documents, and ledgers from the medieval Indian Ocean trade.
The news follows a related CDH milestone: Rustow and CDH's Rebecca Sutton Koeser recently published a data essay opening the Princeton Geniza Project's rich and complex historical material to data science approaches.
At the forefront of DH scholarship on the Cairo Geniza, preserving and providing access to this vast and invaluable collection of historical texts. Director: Marina Rustow
DH Strategist and head of CDH graduate programs Grant Wythoff has published a new chapter in the latest Debates in DH volume, Critical Infrastructure Studies and Digital Humanities. The chapter — Alternative Infrastructures for Digital Equity: Community-Based Internet Access — details the creation of Philly Community Wireless (PCW), a community-controlled network. PCW offers free public Wi-Fi across the Philadelphia neighborhoods of Norris Square, Fairhill, and Kensington. Today, dozens of homes and businesses host PCW rooftop antennas. The network provides access to 50+ city blocks and 7,600+ monthly users. It is maintained by dozens of volunteers, community partners, and a dedicated staff.
PCW was first launched in the summer of 2020 with the support of a Rapid Response Grant from the Princeton Humanities Council and student interns from the Pace Center for Civic Engagement’s RISE Program. During the Covid pandemic’s earliest days, a group of organizers, researchers, librarians, and neighbors — including Wythoff and the chapter’s co-authors Alex Wermer-Colan, Devren Washington, and Allan Gomez — gathered to address the problem of internet accessibility at a time when most of daily life had moved online. Their goals were to expand internet access, grow tech literacy, and build community autonomy.
The chapter reflects on the lessons learned in building PCW’s technical and social infrastructures, puts concepts from digital humanities and community technology into conversation, and finally, gestures toward the future of community-supported infrastructure. The authors write:
Our question is how infrastructure cooperatives that provide community mesh networks for urban environments can facilitate collaboration between neighbors while fostering community consensus on network architecture, organizational structure, and the geographic redistribution of digital resources. Ultimately, our theory of change is that access leads to adoption in a fuller social sense: if one empowers the communities most affected by the biases and harms of the tech industry to own and control last-mile internet infrastructure in their neighborhood, such communities will be equipped not just to use broadband technology, but to advocate for better outcomes in the way that such technologies are utilized, governed, and regulated.
This spring, six CDH Graduate Fellows arrived with their research in progress, asking six different questions across multiple disciplines ranging from History and Comparative Literature to Music and English. Their findings? Working with data and computational methods rarely unfolded the way they expected, and while oftentimes arduous, the labor uncovered “strange and beautiful” discoveries.
"We are receiving more competitive applications to the Graduate Fellowship than ever before," said Grant Wythoff, who directs CDH graduate student programs. "Some of these emerging scholars bring knowledge of data curation standards and machine learning methods. Others are tuned into the latest debates on AI's political and epistemological impacts. The mix of voices makes for an incredibly exciting group dynamic."
Cecelia Ramsey (French and Italian) came to the CDH with a question about literary afterlives: what makes a book experience a revival many years after its initial release? To study this at scale, she worked with BiblioBase, a database of the nineteenth-century Bibliographie de la France, tracking gaps between editions and reeditions and looking for patterns in a book's reintroduction.
The data was messy—inconsistent titles, variable spellings of authors' names—and rather than cleaning the inconsistencies away, Cecelia explored them as an opportunity to learn more about the nature of reeditions and the format of the Bibliographie itself. "Interacting with the messy data taught me how slippery the very name of a work can be," she reflected.
It was also her first substantial engagement with DH methods, and she described the fellowship's atmosphere as essential. When Grant Wythoff opened the semester by telling the cohort it was normal not to know things—that the DH world is so interdisciplinary that all scholars often feel that way—it changed what was possible. "This introduction made it a space where it's normal to ask questions, to learn, and to just be openly curious," Cecelia said. "What a gift."
A pedigree chart tracing the lineage of Ignez de Guiné, the matriarch of several prominent Portuguese families.
Amanda Pinheiro (History) has been working with a database developed over the last ten years, containing 115,545 baptismal, notarial, and judicial documents from ten villages in colonial Brazil. These records detail the eighteenth- and nineteenth-century lives of roughly 6,000 individuals who inhabited the south of Brazil and may have migrated through the frontier zone between the Spanish and Portuguese empires. The issue: that number is inflated by duplicate names, with varying spellings and characteristics across documents. Her fellowship project used Splink, a Python library for probabilistic record linkage, to calculate the statistical likelihood that two records refer to the same individual—generating a unique identifier for each person so she could cross-reference this database with her new archival findings.
I realized that automation necessarily requires diligent and continuous manual labor.
Amanda Pinheiro
What surprised her was how much human judgment, or as she describes it, laborious decision-making, the automation required. "I realized that automation necessarily requires diligent and continuous manual labor," Amanda reflected. "The two are interconnected and walk hand-in-hand in the digital realm." With Wouter Haverals’ (Associate Research Scholar, CDH; Perkins Fellow, Humanities Council) guidance, she completed Splink tests and generated an analytical report on the quality of her datasets—one she hopes to publish and add to her metadata in the future.
Amy Weng (English) asked whether a seventeenth-century English preacher's confessional affiliation—Anglican, Nonconformist, or Catholic—leaves a detectable fingerprint in their printed sermons. Her project, Godly and Learned Divines (GoLD), represented each of 809 preachers across 2,877 books as a 1,000-dimensional vector built from scripture citations, named references, topic modeling, and entity types, then trained a Random Forest classifier to predict denomination. The model achieved an F1-score (a machine learning metric used to evaluate the performance of a classification model) of 0.80 for Anglicans, 0.51 for Nonconformists, and 0.12 for Catholics.
Wikidata-Linkable Preachers in EEBO-TCP
Clustered based on the distribution of topics, named entities, and scriptural references in their sermons
The results were striking. Place of education turned out to be the least important feature—far less predictive than the types of sources a preacher reached for. "Bible versions matter more than the proportions of Bible divisions," Amy concluded, "and ancient entities once again outrank medieval and contemporary references. Generally, godly learnedness—patterns of referencing scripture—distinguishes preachers across confessional divides more than overall learnedness."
Amy credited Jacob Murel (Research Software Engineer, Classics) for sustained mentorship on using large language models for orthographic standardization, and Wouter Haverals for introducing her to Wikidata reconciliation.
Pierre Azou (French and Italian) examines the relationship between literature and political violence in his doctoral research and found himself drawn to the “digital sphere” as the space where the questions he studies in published books are being reconfigured. For his fellowship project investigating the link between "manliness" and insecurity in contemporary French public discourse, he turned to two foundational texts in the French debate on masculinity: Élisabeth Badinter's XY, de l'identité masculine (1992) and Éric Zemmour's Le Premier Sexe (2006). These works contain opposing premises, one theorizing a fragile masculinity, the other insisting it is strong but under siege, yet both binding manliness tightly to a language of threat and crisis.
Using keyness analysis and topic modeling in Python allowed Pierre to compare the density of each author's clusters of insecurity-related words (fear, violence, war, crisis, domination). He identified the author's statistically distinctive vocabulary and examined the semantic neighborhoods of shared terms. The most productive approach was contextual: looking at what words appear near a key shared term like virilité in each text. "It turns out the same word lives in completely different semantic environments in Badinter and Zemmour," Pierre noted.
A technical challenge gave him pause early on—preprocessing French text that contained English-language citations required combining stopword lists (words like “a,” “the,” “and” or “un,” “le,” “et”) and filtering bibliographic noise—but his more substantive reflection was on what data cleaning actually does. "It reminded me that cleaning decisions in DH are more than purely technical, as they also shape the findings."
Cleaning decisions in DH are more than purely technical, as they also shape the findings.
Pierre Azou
Nathaniel Gallant (Comparative Literature) studies the relationship between Buddhism and the history of dramatic and poetic theory across Japanese and Tibetan literary traditions. In his daily research, he relies on well-developed digital tools and databases built for pre-modern Japanese sources—resources that reflect years of philological groundwork by scholars who came before him. For Tibetan studies, that infrastructure is still being built. DH projects in the field are scattered across academic, non-profit, and private spheres, with no centralized view of what exists or where the gaps are.
Nathaniel’s fellowship project addressed that directly: he created a database cataloging existing DH projects in Tibetan studies, with visualizations mapping networks of funding sources, text archives, OCR and LLM development projects, and institutional stakeholders. The goal was to understand current patterns in project development and identify potential directions for future text-digitization projects, particularly in the history of Tibetan literature and poetry.
The hours of scanning documents, mental grappling...crystallized into something coherent, beautiful.
Rachel Glodo
Each issue of the Yearbook of the Imperial Theaters includes detailed lists (spiski) of creators and artists.
Rachel Glodo (Music) is reconstructing the world of the Imperial Ballet in the Russian Silver Age through eighteen volumes of the Yearbook of the Imperial Theaters (1890–1908)—elaborate annual retrospectives documenting productions, performers, choreographers, musicians, designers, and administrators across St. Petersburg and Moscow. The challenge was getting that data out of the page and into a form that a researcher could query. Rachel used optical text recognition (OTR) to convert nineteenth-century printed Cyrillic into machine-readable text, while Andy Janco (Digital Scholarship Specialist) developed custom Python scripts, based on her project design, to convert images of lists and tables into structured spreadsheets.
What Rachel hadn't anticipated was how much the project would begin with physical, analog labor. "The most challenging part of my project wasn't the implementation of DH methodologies," she said, "but the quotidian task of scanning and saving thousands of images spanning 18 volumes." She described it as "the strange and beautiful juxtaposition of 'distant' and 'close' readings that characterizes DH." And she was surprised by how much the technology itself shifted between her original proposal and the start of the fellowship—she ended up using an entirely different processing strategy than she had planned, with Christine Roughan (Postdoctoral Research Associate, CDH/MARBAS) and Andy as crucial partners in identifying her priorities and methods.
A eureka moment was had when they ran the Python scripts together for the first time. "All the hours of scanning documents, mental grappling, design, and redesign suddenly crystallized into something coherent, beautiful, and—almost miraculously—exactly what I needed," she recalled. "It was a glorious moment."
A record of all productions on the Imperial stages, including ballets and operas.
Throughout the semester, the cohort's monthly sessions became as important as the technical work itself. "The regular meetings provided me with more productive time and space to learn about digital tools than scheduling different consultations could have," Amanda said. For Cecelia, the cross-disciplinary exchange was its own kind of finding: "It's exciting to step outside your discipline and be invited into someone else's world while it's still in the making—while they're still experimenting and puzzling through the challenges."
Interested in applying for a Graduate Fellowship?Visit here, or head to theCDH Graduate Program page to see more opportunities for graduate students.
We are delighted to announce the recipients of the Fall 2026 CDH Graduate Fellowship. Grad Fellows are mentored by CDH staff to employ computational and data-driven methods in their research. As a cohort, fellows explore tools, methods, and best practices that will benefit them throughout their careers.
Please join us in congratulating this cohort:
Maximilian Diemer (History) is researching the origins of meritocratic practice in eighteenth-century France and Britain, with a focus on professionalisation in the army and royal service.
Aneka Kazlyna (History) studies the history of astronomy and the broader mathematical sciences, with particular interests in the histories of observation and the circulation of knowledge across cultures.
Shiqi Pan (East Asian Studies) studies the medieval history of the Huai River region in China, exploring it as an internal frontier shaped by environmental, political, and cultural change.
Benjamin Price (Art and Archaeology) studies art, anarchism, and histories of science in late-nineteenth century France.
Sun Shen (Politics) is studying presidential influence in the making of U.S. foreign policy.
Filippo Ugolini (East Asian Studies) examines the discourse of romance in mid-to-late Tang China (8th–9th century) at the intersection of literary analysis, gender studies, and socio-economic history.
Tirzah Anderson (History) is researching Afro-Indigenous worldmaking in Indian Territory and Oklahoma from the 1890s through the 1930s.
We look forward to working with this remarkable cohort and sharing more about their projects and fellowship outcomes as the year unfolds.
Princeton University is seeking to hire AI Postdoctoral Research Fellows to join the Princeton Laboratory for Artificial Intelligence with affiliation to the Center for Digital Humanities to conduct research focused on societal AI.
Postdoctoral Research Fellows will be part of a dedicated research community that includes leading faculty, research fellows, scientists, software engineers, postdocs, and graduate students. Fellows will have access to a state-of-the-art AI Lab GPU cluster to support large-scale experimentation and evaluation.
Postdoctoral Research Fellows will be appointed through the Princeton Laboratory for Artificial Intelligence (the AI Lab) and will receive a competitive salary, benefits, office space and equipment, dedicated compute resources, and annual research support.
Appointments will be made at the Postoctoral Research Associate rank.
Appointments are for one year with the possibility of renewal pending satisfactory performance and continued funding.
Princeton University places high value on in-person collaborations and interactions, and we expect selected candidates to participate in-person. The work location for this position is in-person on campus at Princeton University.
This position is subject to the University's background check policy
These positions are not eligible for sponsorship of an H-1B visa requiring consular processing; other visa sponsorships (including H-1B visas not requiring consular processing) may be available, as appropriate.
Qualifications
Ideal candidates will have demonstrated experience in “societal AI” especially as related to culture. We seek scholars working on AI as a "cultural technology" — including its role in disciplinary practice across the arts and humanities and in knowledge infrastructures — as well as those examining AI's broader societal effects – including on creative and cultural industries (literature, music, film, visual arts). A strong foundation in computational humanities or cultural analytics, combined with humanistic and interpretive approaches to AI systems, is expected.
Applicants must have a PhD.
Application Instructions
Applicant must provide the following:
Cover letter
Current curriculum vitae
Research statement
Contact information for three references is required
Application Process
This institution is using Interfolio's Faculty Search to conduct this search. Applicants to this position receive a free Dossier account and can send all application materials, including confidential letters of recommendation, free of charge.
Equal Employment Opportunity Statement
Princeton University is an Equal Opportunity Employer and all qualified applicants will receive consideration for employment without regard to age, race, color, religion, sex, sexual orientation, gender identity or expression, national origin, disability status, protected veteran status, or any other characteristic protected by law.
Pay Transparency Disclosure
The University considers factors such as (but not limited to) the scope and responsibilities of the position, candidate's qualifications, work experience, education/training, key skills, market, collective bargaining agreements as applicable, and organizational considerations when extending an offer. The posted salary range represents the University's good faith and reasonable estimate for a full-time position; salaries for part-time positions are pro-rated accordingly.
The University also offers a comprehensive benefits program to eligible employees. Please see this link for more information.
The Center for Digital Humanities is excited to announce its Affiliate Labs designation – a formal recognition of our partnerships with Princeton research groups whose work sits at the intersection of humanistic inquiry and digital and computational methods. Affiliation reflects the depth and ongoing nature of these collaborations, and signals CDH’s commitment to fostering a broad and vibrant ecosystem of digital humanities across campus. Affiliate Labs represent a range of approaches — from open-access archival preservation and computational text analysis to the development of language technologies and digital reconstructions of historical places.
“The Affiliate Labs program recognizes what is already happening at Princeton — rigorous, innovative work at the intersection of humanities and computation — and makes space for those partnerships to grow," states Jeri Wieringa (Assistant Director, CDH). "We’re excited to see where these collaborations take us.”
We are proud to introduce the following inaugural Affiliate Labs:
Princeton Geniza Lab, directed by Marina Rustow (Khedouri A. Zilkha Professor of Jewish Civilization in the Near East), is at the forefront of digital humanities scholarship on the Cairo Geniza, preserving and providing access to this vast and invaluable collection of historical texts. CDH has had the privilege of collaborating with the PGL on the Princeton Geniza Project since 2020.
African Language Technologies Lab, co-directed by Christiane Fellbaum (Professor of Linguistics) and Happy Buzaaba (Associate Research Scholar, AI Lab with affiliations in CDH and Africa World Initiative), works to increase the representation of African languages in rapidly advancing language technologies driven by large language models, while foregrounding the values and cultures those languages carry. Through a series of projects, courses, and speaker events, the lab generates campus-wide conversation about this critical and underserved area of language technology research.
REACH² Lab, led by Paul Vierthaler (Assistant Professor of East Asian Studies, Associate Faculty Director of CDH), centers on the digital analysis of the literature, culture, and history of East Asia. Vierthaler and his collaborators work on projects ranging from quantitative analyses of large-scale text and image corpora to digital reconstructions of historical places, welcoming students, faculty, librarians, and researchers interested in any dimension of computational East Asian studies.
Poetry's Data Lab, directed by Meredith Martin (Professor of English, Faculty Director of CDH), draws on the Princeton Prosody Archive to analyze patterns in anglophone poetry teaching over time, asking what poetry used in teaching texts — at scale — can reveal about canonicity, book history, and the development of English as a discipline. The lab's work is linked to the broader Ends of Prosody project and forthcoming special issues of theJournal of Cultural Analytics and Victorian Poetry.
We look forward to sharing more about each of these labs and the work emerging from these partnerships in the months ahead. To learn more or inquire about affiliation, visit the CDH Affiliate Labs page.
At the forefront of DH scholarship on the Cairo Geniza, preserving and providing access to this vast and invaluable collection of historical texts. Director: Marina Rustow
Using the poetry found in the Princeton Prosody Archive datasets to analyze patterns of anglophone poetry teaching over time. Director: Meredith Martin
Last Friday, Princeton graduate students gathered to present research spanning the evolution of social behavior through natural selection, the quantum nature of atomic nuclei, and the effect of tropical cyclones on global climate.
The annual data and computation joint graduate certificate colloquium, held on April 24, brought together students across the humanities, social sciences, natural sciences, and engineering to share their work and exchange ideas.
“Few, if any, graduate student events focused on research have participation across all areas of scholarship at the University,” said Michael E. Mueller, interim director of PICSciE, director of the graduate certificate in computational science and engineering, and acting chair of the department of mechanical and aerospace engineering.
“Some of the most interesting research questions came from students driven by sheer curiosity about domains completely removed from their own.”
CDH grad certificate students included:
Laura Nelson, History Adviser: Laura Edwards “Unseen fragments, seen lives: From data to biography in digital public history”
April Gilbert, Comparative Literature Adviser: Claudia J. Brodsky “Narrating narrative’s lifespan: Exploring data from the conference programs of the International Society for the Study of Narrative (ISSN)”
Sharifa Lookman, Art & Archaeology Adviser: Carolina Mangone “From pixel to foundry: Recuperating sixteenth-century bronze techniques and technicians through 3D-imaging and historical reconstruction.”
Fall course selection starts this month! Wondering what to register for? We've compiled a list of courses on media studies, technology, data and culture, and more.
This list includes both undergraduate and graduate courses. Grad courses are marked with an asterisk (*).
This past January, the CDH kicked off a new Collaborative Research Partnership with Anna Yu Wang (Assistant Professor of Music) and Jürgen Hackl (Assistant Professor of Civil and Environmental Engineering), the researchers behind the larger Music Theory in the Plural project. MuSE—short for Multilingual Semantic Embeddings—asks: when scholars write about music theory across languages, including Chinese, Japanese, Spanish, and Portuguese, are they talking about the same things? The project sets out to evaluate whether multilingual LLMs can translate the domain-specific discourse of music theory without flattening its nuance, and to test computational methods for discovering related concepts across those language traditions.
Led by RSE Laure Thompson, the CDH team is working with the scholar-translated articles in Music Theory Online volume 30, number 4, which presents articles written in Chinese, Japanese, Portuguese, and Spanish (among others) alongside their English translations. This resource allows the team to assess automated translation against expert scholarly ones. From there, the team will experiment with embedding models as tools for surfacing cross-linguistic connections across a broader corpus. The work aims not just to answer these questions for music theory, but to contribute broader, critically needed comparative research on how LLMs perform with specialized humanistic content.
Like all CDH Research Partnerships, MuSE begins with a project charter that defines the project's scope, deliverables, team roles, and terms of collaboration. The MuSE collaboration is one small part of the larger research agenda of the Music Theory in the Plural project, and it is through chartering that the project team and the RSE team define how the pieces will fit together. You can download the MuSE project charter here and follow the project's progress on its project page.
The Center for Digital Humanities is seeking proposals for innovative and computationally-engaged research partnerships from Princeton faculty. The next application cycle closes April 17, 2026. Apply now.
With the launch of the Princeton Laboratory for Artificial Intelligence and the New Jersey AI Hub, over the last few years Princeton University has firmly established its presence at the forefront of artificial intelligence research — including transformative work in humanities scholarship.
From piecing together fragments of ancient texts with language models to exploring the future of human-robot interactions, Princeton scholars aren’t just exploring what AI can do for the humanities. They’re uncovering what the humanities can do for AI.
Already, AI tools are appearing in all facets of our society and culture. “It’s the world that our kids are going to inherit,” said Meredith Martin, professor of English, faculty director of Princeton’s Center for Digital Humanities (CDH) and a grant recipient from the international Schmidt Sciences Humanities and AI Virtual Institute. “We should try our hardest to put the humanities into every aspect of AI development, not only in the input data and the interpretation of the results,” she said.
Princeton humanities scholars had already been using machine learning in their research — largely by partnering with the robust community of humanities research software engineers and digital humanities experts at CDH, established more than a decade ago. And now, with the AI Lab, they are proving to be some of the best positioned collaborators for the future of humanities and AI.
“Our goal in the AI Lab is to support the transformative impact of AI on research across the Princeton campus, and we’re excited about the many opportunities to do so in the humanities,” said AI Lab Director Tom Griffiths.
“The most valuable and productive relationship between artificial intelligence and humanistic research is collaboration,” said Rachael DeLue, director of the University’s sweeping new Humanities Initiative. Here, a sample of the scholarship now underway.
Insights into ancient texts
Photo by Matthew Raspanti, Office of Communications
Paul Vierthaler uses machine learning to study premodern Chinese books such as “Xianqing ouji 閒情偶寄 (Leisure Notes),” a collection of essays published in 1671, pictured at right.
Fifteen years ago, Paul Vierthaler became fascinated by a particular figure common in Ming and Qing dynasty literature. Wei Zhongxian, a late Ming dynasty eunuch, nearly took over the imperial government in the 1620s. His infamy became such that within a year of his death, half a dozen novels and unofficial histories on his exploits had already been published. “I became really interested in how people talk about historical events within so-called ‘unreliable genres,’” said Vierthaler.
Studying the historical figure of Wei and the stories he’d inspired raised new questions for Vierthaler: How common were these types of narratives in imperial China? What more could be learned from studying bibliographic information, and how could he even begin to study that information at scale? “I realized, if I wanted to try to get a grip on how these kinds of narratives existed in the Chinese literary tradition, I needed to start thinking more broadly,” he said.
The realization led Vierthaler, now an assistant professor of Chinese literature and interdisciplinary data science, to the global catalog WorldCat, which holds digitized bibliographic records from tens of thousands of libraries around the world. To grapple with the vast amounts of data, Vierthaler turned to computational methods and machine learning/AI — which has altered the scope of his work.
Using machine learning to analyze and extract data from written descriptions of premodern Chinese books — which may include information on author, illustrations, content and more — Vierthaler initially set out to understand whether genres containing suspect stories about historical events increased or decreased in popularity over time. But soon his focus expanded. “It’s exploded into a much larger project when I began to apply these same tools to digitized versions of the books themselves,” he said.
More recently he has been using machine learning methods to study the likely authorship of anonymously published works and to detect historical documents inserted into novels. With the advent of transformer-based language models, Vierthaler is now training specialized language models on premodern Chinese corpora, hoping to pick up minute nuances in the texts he studies. “There’s a movement now in the humanities aimed at training much more targeted, bespoke smaller language models,” Vierthaler said.
Instead of training bigger and bigger models, Vierthaler’s work has wider implications for using custom training sets to capture and retain the nuance necessary for the study of culture. “Humanities scholars can bring an understanding of the historical background and composition of training materials, which can help identify blind spots that might have otherwise been missed,” he said.
That same movement for highly specialized language models drives the work of MarinaRustow, the Khedouri A. Zilkha Professor of Jewish Civilization in the Near East and professor of Near Eastern studies and history. Rustow is training a model designed to transcribe fragments of medieval texts — and save humans time-consuming, painstaking labor.
Geniza fragment courtesy of Cambridge University Library. Photo: Sameer A. Khan.
Left: A government decree from the Fatimid period in Egypt (969–1171) shows the original Arabic inscription with wide line spacing. Right: Marina Rustow.
Rustow runs the Geniza Lab, a group dedicated to studying an enormous cache of paper and parchment recovered from a medieval synagogue in Cairo. The documents are unique because, unlike most ancient texts preserved today, they’re not the work of society’s elites and philosophers. They’re everyday records from the masses, things like complaints about business travel, heated personal letters and descriptions of stomachaches.
The fragments offer a broad perspective of society in that period, which is invaluable to Rustow as a social historian of the medieval Middle East. But they’re written in dialects of Arabic, Hebrew and Aramaic no one speaks today and in handwriting that can border on illegible.
Rustow said it can take her two full days to transcribe just one document. She hopes the machine learning model she’s training will turn that around for the 36,000 Geniza fragments that she and her colleagues have uploaded to the Princeton Geniza Project database for public access.
Since the Geniza’s discovery in 1896, “it has taken researchers 130 years to transcribe 7,000 documents, and we have another 29,000 documents to transcribe,” said Rustow. With machine learning, she hopes to save scholars 530 years of transcription drudgery.
Like Rustow, Barbara Graziosi is on a mission to make premodern texts free and accessible for all. “Ideally, I’d like to see everything that we have from before the invention of printing preserved, made accessible, translated, well edited, well understood and well studied,” said Graziosi, the Ewing Professor of Greek Language and Literature and professor of classics.
Photos by Denise Applewhite, Office of Communications
Barbara Graziosi uses AI as a "collaborator" to study ancient texts, including this 13th-century Byzantine manuscript of Aristotle's "Organon" from Princeton's Special Collections.
Graziosi is contributing to that mission by filling in the gaps of fragmented ancient Greek text. Over millennia, words and phrases written on documents are lost, chewed away by mice, eroded by moisture or obscured by stains. When a student approached her and suggested — before the advent of ChatGPT — that language modeling could generate suggestions to fill the gaps in these papyri, Graziosi began work on a machine learning tool attuned to the nuances of ancient Greek. The result of that work is the Logion Project, which Graziosi leads.
Graziosi said AI works best as a collaborator for humanists, not a replacement for highly trained scholars who dedicate years to studying these difficult texts. The tool Graziosi developed provides several suggested words to fill a given gap in the text. Seeing multiple suggestions can jog the thinking of a scholar who might face a block after spending hours reading and rereading the same passage.
The humanities, as Graziosi sees it, can help shape a future where the strengths of humans and the strengths of machines work in harmonious collaboration. “It’s very important that we keep the conversation going and that we respect human expertise as well as machine confidence,” she said. “I hope more humanists will get involved with AI because their perspective is exactly what’s needed now.”
Modern-language applications
When Happy Buzaaba moved to Japan in 2015 to study for his Ph.D. at the University of Tsukuba, he couldn’t speak any Japanese.
Japanese-to-English translation is widely available, as both are well-studied languages with vast digital footprints. But some languages — in particular, many of those spoken in Africa — have a scant internet presence on which to train AI models. “I started thinking, imagine you went to a country where they speak a language that you don’t understand, and it’s also not supported by any existing technology,” said Buzaaba.
With computational linguist Christiane Fellbaum, Buzaaba is now introducing African languages to LLMs by creating large collections of syntactically annotated text, called treebanks, for 11 African languages.
The idea is that the rich annotations, filled with linguistic knowledge, can be used to train LLMs on the African languages, even though there’s not as much text as what’s available for Japanese or English. “We can actually create models that perform well on these languages, even with less amount of data,” Buzaaba said.
He and his colleagues have already released three African language models, which benchmarks show to be the best performing models of their kind. “The main goal here is not just creating tools, but accessibility,” he said. He has also brought this work into the classroom, on campus and in a PIIRS Global Seminar in Kenya.
AI in arts and architecture
Photo by Matthew Raspanti, Office of Communications
Elizabeth Margulis uses AI to study how people describe their musical experiences in her Music Cognition lab (pictured at right: Itamar Jalon, a postdoc in psychology and music).
Humanities faculty in music, creative writing and architecture, among other disciplines, are using AI to understand the very essence of human creativity and to inform new work.
Elizabeth Margulis, a professor of music and acting department chair, is trying to understand how music shapes our emotions, imaginings and the thoughts that arise when we let our minds wander. Machine learning, she said, has been instrumental in advancing the studies she conducts for her Music Cognition Lab.
Margulis uses AI to study how people describe their musical experiences. “Where machine learning has been really helpful is giving us a way into unconstrained, free-response descriptions of what music evokes,” she said.
Researchers at the Music Cognition Lab collect these responses from volunteers, who enter a booth, listen to a musical excerpt, and then describe in writing the imaginative scenarios and emotions that arise. The lab has also worked in collaboration with Princeton University Concerts. At a Takács Quartet performance this past spring, the researchers gathered free-response descriptions from hundreds of concertgoers.
With the help of large language models, Margulis and her team analyze the descriptions, looking for patterns they might not have elucidated without the help of AI tools. “What’s so cool about machine learning is it helps us see structure in what seem like singular, subjective experiences,” said Margulis.
What she’s found so far is that people from different cultures frequently have remarkably different emotional and imaginative reactions after listening to the same piece of music. For example, one atonal excerpt by Anton Webern often conjured up a sense of impending doom for English speakers from the American Midwest. However, Dong speakers from the Guizhou province in China tended to imagine joyfully playing outside with friends.
Margulis hopes this work opens new avenues for understanding spontaneous thought in a way that could be applied to clinical settings somewhere down the road. “Think about ADHD or anxiety — both have these components that reside in patterns of spontaneous thought,” said Margulis. “Music gives us a powerful way to study the susceptibility of those thoughts to perceptual influence.”
A.M. Homes, professor of the practice in creative writing and the Lewis Center for the Arts, has been doing a lot thinking herself lately about AI. “I’m one of the writers whose books have been fed to AI to train on,” Homes said. “I sit on the Writers Guild of America’s Council on AI, and we are very concerned about how AI is being used in the entertainment world.”
Photo by Matthew Raspanti, Office of Communications
A.M. Homes in creative writing is working on a novel that explores themes of grief and what could happen when people turn to artificial intelligence for comfort.
To navigate this complicated moment of murky boundaries surrounding AI use, Homes is doing what she does best — writing about it. AI isn’t a tool she uses in her creative life; instead she’s working on a novel that interweaves themes of grief and explores what could happen when people turn to artificial intelligence for comfort.
Homes is a fiction writer who taps into the ideas percolating through society and culture at large. She sees the author’s role as being an artist who conceptualizes worlds and futures that don’t exist, inviting readers to think critically about the one they live in. “Whenever I’m writing something, what I really want to inspire is discussion and conversation,” she said.
In Arash Adel’s ideal future, humans and AI aren’t at creative odds but work together. While pursuing his Ph.D. at ETH Zurich, Adel focused on computational design and robotic integration into architecture construction. “But I wondered about the role of humans,” he said.
From L to R: Photos by Daniel Ruan and Bob Berg; courtesy of Arash Adel
Arash Adel and his team have recently built Timbrelyn, a robotically fabricated structure on the historic grounds of the 1969 Woodstock Festival in Bethel, N.Y.
Now an assistant professor in the School of Architecture with an affiliation at Princeton Robotics, Adel investigates human-robot collaboration where people supervise and instruct while robots perform some of the physically demanding and potentially dangerous construction tasks. This type of human-robot partnership, he said, is driven forward with AI models.
In 2024, Adel’s research group put this approach into practice for their Timbrelyn installation on the grounds of the 1969 Woodstock Festival in Bethel, N.Y., a raised wooden platform created from intricate layers of lumber. Using AI vision, robots scanned inventories of reclaimed and new lumber to identify wood elements that met design specifications while minimizing waste. After selection, the robots grasped and processed the elements using a saw before assembling them with a human collaborator.
Adel and his group are now working on a project that involves AI assisting in the design process as well. The goal is to ultimately develop a pipeline where humans and robots collaborate from inception to final construction.
“Humans are very intuitive, but we struggle to process large amounts of information at once,” said Adel. “The role of the AI is to augment human creativity.”
Connecting engineers and humanists: the Center for Digital Humanities
Photo by Kristopher Johnson
The Princeton Prosody Archive (left) is a searchable database of thousands of English-language digitized works published between 1559 and 1928, directed by Meredith Martin (right).
By the time ChatGPT exploded onto the scene in 2022, the Center for Digital Humanities had already been situated at the cutting edge of humanities-technology collaboration for the better part of a decade.
The center equips Princeton humanities faculty to thrive in a tech-dominated landscape, connecting them with software engineers who build the bespoke software that underlies projects (like Rustow’s and Graziosi’s) and teaching humanists and software engineers how to successfully collaborate.
“We at CDH had already built the necessary collaborative infrastructure for projects involving both software engineers and humanists,” Martin said. With the rapid proliferation of generative AI tools, she has noticed a surge of humanities scholars approaching CDH with questions about how the new technology might transform their work.
There are obvious advantages to AI: faster processing of larger datasets, high performance computing, quicker pattern recognition. But these technological leaps aren’t enough on their own, Martin said. “Humanists have to bring a lot of knowledge to that interaction for it to work out.”
For that reason, the staff at CDH think carefully about how machine learning might fit into a research project and whether a particular approach would be the right fit for their question. At the same time, Martin sees an opportunity for humanists themselves to shape the AI tools.
“There’s no reason why humanists can’t feel empowered to build better models, to participate in model architecture, to think about the kinds of data on which various models are trained and why,” she said.
This desire to bring humanists into the AI fold helped inspire a three-part project developed by CDH that spans the 2025-26 academic year and beyond. The project, Modeling Culture: New Humanities Practices in the Age of AI, brings together Princeton faculty and researchers from other universities for a seminar series to think critically about AI.
Martin ran one of the seminars this fall with Matthew Jones from the Department of History and Andrew Janco, a digital scholarship specialist at Firestone Library, on the problems and questions of modeling. “The main feeling has been one of real empowerment and excitement,” said Martin. “In the room, you can feel people leaning forward.”
The Center for Digital Humanities (CDH) is proud to announce its collaboration with ATRIUM (Advancing fronTier Research In the arts and hUManities), a major European research infrastructure project funded by the European Commission and coordinated by DARIAH-EU. Princeton CDH is the first U.S.-based institution to work with ATRIUM partners, aiming to generate new avenues of transatlantic collaboration for digital humanities research.
As part of this partnership, CDH, the UNESCO Chair on Digital Methods for the Humanities and Social Sciences, and the Athena Research Centre will co-organize the ATRIUM Summer School titled “From Maps to Data and Data to Maps: Exploring Spatial Histories.”
The four-day workshop, which will take place at the Athens University of Economics and Business on June 29 – July 2, 2026, will bring together graduate students and early-career scholars from the U.S. and Europe to explore cutting-edge methods for analyzing and visualizing spatial data drawn from historical maps and geographic sources. Participants will gain hands-on experience with tools and approaches that are shaping the future of spatial humanities research.
“This collaboration opens an important channel between U.S. and European digital humanities communities,” said Meredith Martin, Faculty Director of the Center for Digital Humanities. “Affiliation with ATRIUM will allow us to connect American students and scholars with the innovative research trends, tools, and networks being developed across Europe — and contribute Princeton’s own expertise to that exchange.”
ATRIUM unites leading humanities research infrastructures across Europe, creating shared access to advanced digital tools, datasets, and expertise. The project also supports intensive training opportunities that foster cross-border scholarly exchange and build capacity in emerging digital methods.
“ATRIUM is about consolidating the European research infrastructure landscape, and meaningful international partnerships are an integral part of that effort,” said Toma Tasovac, Principal Investigator of ATRIUM. “We are delighted to build on our existing relationship with Princeton CDH as a DARIAH Cooperating Partner, and to explore new avenues of collaboration within the ATRIUM framework.”
The summer school reflects CDH’s ongoing commitment to advancing interdisciplinary, computationally engaged humanities research and to developing international partnerships that expand opportunities for scholars at all career stages.
Earlier this winter, CDH / MARBAS Postdoctoral Research Associate Christine Roughan returned to Vienna for the second year in a row to share her experience using Handwritten Text Recognition (HTR) technology for medieval texts.
The workshops were part of HTR Winter School 2025, hosted by the Institute for Medieval Research of the Austrian Academy of Sciences in collaboration with MARBAS and the Institute for Habsburg and Balkan Studies. The in-person sessions followed three virtual workshops that brought together scholars of Carolingian Latin, Byzantine Greek, and Syriac, among other languages.
“This year I reprised my role as a group leader for the Syriac HTR group alongside Ephrem Aboud Ishac (Austrian Academy of Sciences),” Christine explained. “In addition, I provided instruction to the cohort as a whole on how to apply their HTR training in different contexts, so that their new skills were not tethered to only a single tool.”
As she noted in an interview last year, Christine became involved with Winter School after giving a talk at the Institute for Medieval Research, where she met several of the organizers. At the time, she said that Winter School offered an opportunity for her to hone her skills in teaching methods that play an important role in her own work.
Even more important this year: learning about how this year’s participants will use workshop content to advance their own work.
“My favorite part of the experience was definitely hearing about the variety of research topics the participants were engaged in,” Christine explained. “Seeing their enthusiasm for how the Winter School experience would equip them to dive into those projects was really great, especially in the final days when everyone was now practiced with the methodologies and ready to take off on their own.”
Last Fall, the CDH Research Software Engineering team traveled near and far to share their cutting-edge work and learn from others about new developments in the fields of research software engineering and digital humanities.
🔔 Philadelphia
From October 6–8, the RSE team participated in the third annual conference of the United States Research Software Engineer Association. Held in Philadelphia, US-RSE’25 brought together RSEs from across universities, laboratories, industry, and other institutions, as well as their managers and allies, to discuss “Code, Practices, and People.” CDH represented the small — yet mighty! — Humanities contingent in a field of mostly scientists and social scientists.
Postdoctoral Researcher Christine Roughan presented a paper, also co-written with Koeser, titled “Integrating ATR Software with University HPC Infrastructure: balancing diverse compute needs.” The paper and corresponding presentation described the methods and outcomes of Bringing HTR to the HPC: A Pilot to Customize eScriptorium for Princeton, a Research Partnership with the CDH under the umbrella of the Princeton Open HTR Initiative (funded by a 2024–25 Princeton Language + Intelligence Seed Grant). Conference attendees were fascinated to hear about how Koeser and Roughan implemented an instance of eScriptorium — the current leader in open-source handwritten text recognition software — on Princeton’s high-performance computing hardware, which enabled professors and students without advanced technical skills to train large text-recognition models customized to their documents’ needs.
Christine Roughan presents at US-RSE'25
Assistant Director Jeri Wieringa and Project Manager Mary Naydan presented on a panel about supporting and managing RSE projects. Their presentation “Creating Research Software with Humanities Faculty” highlighted the CDH’s chartering process, which helps transition humanities faculty from the individual, expansive mode of traditional humanities scholarship to the collaborative, modular mode of computational research. Their illustrative opening skit, which set the stage for the talk, drew lots of laughter and resonated with audience members. The rest of the panel was just as engaging, sharing lessons learned from many different types of organizations and fields, from the multi-institutional development of medical technologies, to a large laboratory focused on national security, to a lone RSE’s personal project management workflow at a research university (Naydan’s favorite presentation of the conference!).
Mary Naydan and Jeri Wieringa present at US-RSE’25 in Philadelphia.
The technical talks affirmed that the CDH RSE team is ahead of the curve on best practices in Python development, such as using uv for installing packages and choosing Marimo over Jupyter for notebooks. Koeser noted the field-wide shift in starting to think about notebooks as a form of publication, and CDH Research Software Engineer Hao Tan was inspired by Reed Maxwell’skeynote on creating groundwater simulations using physics-informed machine learning. Tan reflects, “Explainability is still a real issue, but rather than rejecting AI outright, we should learn to leverage its strengths and mitigate its weaknesses — through comparative evaluation, transparent step-by-step reasoning, and other methods we develop.”
Many of the conference’s presentations — from Maxwell’s keynote to the Birds of a Feather workshop “AI in Practice” — showed the field of research software engineering grappling with, adapting to, and incorporating AI. While our team entered the conference thinking our challenges were somewhat unique to the humanities, we were surprised to see RSEs from across disciplines encountering similar challenges around this topic: from defining research questions, to gathering sufficient data, to disabusing researchers about what AI can actually do.
🇱🇺 Luxembourg
Unsurprisingly, AI was also a popular topic at the 2025 Computational Humanities Research Conference, held at the Luxembourg Centre for Contemporary and Digital History (C²DH) at the University of Luxembourg from December 9–12, 2025. Many of the presentations focused on benchmarking various models for automatic transcription tasks, or using chatbots to scale up annotation data, from identifying “acts of God” in contemporary Christian fiction to assessing a popular song’s “narrativity.” Miguel Escobar Varela’s keynote, “‘A watch by Kran Kamu’: Exploratory fine tuning for cultural reliability,” discussed using supervised fine-tuning on large open-weight models to yield reliable results within highly specific cultural contexts, such as Southeast Asian historical newspapers — a common problem facing computational humanities researchers given the specialized nature of our data and the scarcity of it for fine-tuning.
Rebecca Koeser (left); Mary Naydan (right)
Rebecca Koeser and Mary Naydan presented a poster based on their short paper “Unstable Data and the Unusual Case of the Prosody Excerpt in the Digital Library” (co-authored with Meredith Martin). Using the HathiTrust materials contained in the Princeton Prosody Archive as a case study, Koeser and Naydan cautioned researchers that the page-level data provided by cultural heritage aggregators is not as stable as we might assume. This instability can lead to erroneous data, flawed conclusions, and difficulties building on previous scholarship.
Hao Tan presents at 2025 CHR Conference in Luxembourg.
Hao Tan delivered a lightning talk, “When Larger LLMs Aren’t Enough: Word Segmentation in Historical Chinese Texts,” which used word segmentation in historical Chinese texts as a case study to highlight how large language models, while powerful, can quietly introduce risks when applied to humanities research. The talk sparked conversations with researchers working on East Asian materials across Europe, the US, and Singapore, especially around the tricky parts of historical text processing, with projects ranging from power relations in historical fiction to poetic imagery and stylistic change in epitaphs.
In addition to Tan’s lightning talk, Koeser and Naydan found two others particularly interesting: Katarina Mohar’s on “Speculative Reconstruction and the Ethics of the Fragment: Early Experiments with Generative AI in Art History,” and Antonina Martynenko, Artjoms Šeļa and Petr Plecháč’s on "Where Empires End: Tracing the Geography of a ‘Soaring Spirit’ in Poetry.” Mohar discussed the possibilities and limitations of using Generative AI to fill in gaps in medieval paintings, and provided practical recommendations for how to use it responsibly. Martynenko et al. examined the spatial imagination of European poets by mapping the distance and directionality of place mentions.
For Koeser, one of the most interesting presentations was “Cluster Ambiguity in Networks as Substantive Knowledge,” which describes a method for running a clustering algorithm multiple times to measure how often edge nodes connect nodes in the same community, allowing researchers to identify ambiguous data. Koeser is interested in the interpretive power and potential applications of this method, such as identifying ambiguous characters in novels. Another highlight was Taylor Arnold and Lauren Tilton’s presentation on “Sitcom Form and Function: Pacing and Production in a Collection of Thirty U.S. Series,” which examined how trends in visual and aural pacing changed over time using a combination of large-scale computational analysis and close reading: an example of truly multimodal research and scalable reading.
The RSE team was energized by these shifts in the field: leaning into ambiguity; using audio, video, and visual data rather than defaulting to text; and combining different scales of reading (close and distant) to draw more responsible conclusions. We are excited to carry what we learned into our work this year on multilingual machine translation and term clustering in music theoretical texts and using vision-based LLMs to aid historical document transcription and data extraction.
In January 2026, Princeton University's Center for Digital Humanities (CDH) began serving as publisher of the Journal of Cultural Analytics (JCA), a leading open-access publication in computational approaches to culture. Today, CDH announces JCA’s vision for expanding cultural analytics scholarship amid rapid technological change and the launch of a new website, supported by Schmidt Sciences’ Humanities and AI Virtual Institute (HAVI).
"The Journal of Cultural Analytics has been instrumental in advancing computational methods in the humanities," said Meredith Martin, faculty director of the CDH and professor of English at Princeton, who serves as one of the journal's three editors alongside Amelia Acker (Rutgers University) and Tanya Clement (University of Texas at Austin). "We are honored to lead JCA's continued evolution and grateful to Andrew Piper for his pioneering work in establishing this field-changing, scholarly venue."
Building on a Strong Foundation
Founded by Piper at McGill University's Department of Languages, Literatures, and Cultures, JCA has published groundbreaking data-driven research about culture since 2016. The journal encourages transparent research practices, including open sharing of data and code. It has become a cornerstone publication for scholars working at the intersection of digital humanities, computational social sciences, and computational approaches to culture.
"The idea for the journal was born in 2015 as a response to a shared sense that our field needed a venue dedicated to the critical use of computation to study culture," said Piper. "After a decade of growth, the journal has far exceeded my hopes. I'm extremely happy to see it continue under the leadership of the new editors and its new institutional home at Princeton's Center for Digital Humanities."
Looking Ahead: Expanding Scope and Impact
JCA is broadening its vision to serve an ever-evolving interdisciplinary and international scholarly community invested in cultural study and the methods by which we interrogate the digital in culture – especially in the age of AI. Central to this vision is the commitment to publishing work that goes beyond method for method's sake, asking instead how computational approaches to culture at scale can reshape what we know and how we know it.
The editorial board has expanded to 43 scholars representing institutions across North America, Europe, Asia, and Australia, reflecting JCA’s commitment to international perspectives and increasing representation from junior scholars. This expanded scope will support the journal’s growing focus on multi-lingual and multi-modal approaches to culture.
A new Special Features section, edited by Laura McGrath (Temple University), will highlight shorter, timely essays on computational cultural analysis written in an accessible style for non-specialist audiences, designed to spark discussion on new methodologies, datasets, or research.
JCA will deepen its focus on critical engagement with data, which is increasingly significant for AI researchers returning to smaller, human-curated cultural models. Welcoming a new data editor, Sarah Reiff Conell (Princeton University Library), JCA will revise its data-essay and dataset-review format in collaboration with the scholarly “data collectives” (such as Post45 and 19thC Data Collective), and provide a directory of datasets for cultural studies.
Upcoming Special Issues will explore topics ranging from computational humanities in the Global South to data-driven approaches to poetry and a retrospective on ten years of the JCA. The journal is currently accepting Special Issue proposals for 2027.
New Infrastructure for Open Access, Community-Led Publishing
The transformative support from Schmidt Sciences’ HAVI program has enabled JCA's growth and modernization, expanding the editorial team with new roles for graduate students—providing both recognition and compensation for the labor required to run an academic journal and an opportunity to train the next generation of computational humanities scholars. The grant has also enabled the journal to migrate to Janeway, an open-source publishing platform developed by the Open Library of Humanities, featuring a redesigned user interface and customizable workflow management system.
In this new phase, JCA maintains its commitment to diamond open access—free to read and free to publish, with no article processing charges (APCs) or publishing fees for authors or universities. JCA has also joined the Open Journals Collective, a coalition of libraries and university-based publishers that launched in March 2025, providing journals with technological support, financial sustainability, and community governance through a library-funded model that keeps research freely accessible and journals editorially independent.
"I'm thrilled to have such a prestigious Princeton journal carrying the banner for diamond open access as part of the launch collection. We're excited for JCA, and for the promise of the new, sustainable funding model OJC is delivering," said Matthew Kopel, Princeton's Open Access & Intellectual Property Librarian, who also sits on the Open Journals Collective Library Board.
More information about the journal's new direction, upcoming issues, and submission guidelines can be found on JCA's newly launched platform at https://culturalanalytics.org.
Editorial Team
Editors
Meredith Martin, Princeton University
Tanya Clement, University of Texas at Austin
Amelia Acker, Rutgers University
Special Features Editor
Laura McGrath, Temple University
Data Editor
Sarah Reiff Conell, Princeton University Library
Graduate Editorial Assistants
Cecelia Ramsey, Princeton University (Managing Editor)
Odalis Garcia Gorra, University of Texas at Austin
Haiqi Zhou, McGill University
Emilien Arnaud, Princeton University
Former Editorial Assistant
Katrin Rohrbacher
About the Center for Digital Humanities
Princeton's Center for Digital Humanities, founded in 2014, advances computational and data-intensive humanities scholarship through collaborative research, innovative pedagogy, and community building to create a more just future. The center develops better practices in technological development and research while bringing humanistic perspectives to data science applications.
About the Open Journals Collective
The Open Journals Collective is a growing coalition of libraries and university-based publishers providing sustainable, community-led alternatives to commercial academic publishing. Through diamond open access and collective funding models, OJC supports hundreds of journals while ensuring research remains freely accessible to all.
About Schmidt Sciences
Schmidt Sciences is a nonprofit organization founded in 2024 by Eric and Wendy Schmidt that works to accelerate scientific knowledge and breakthroughs with the most promising tools to support a thriving planet. The organization prioritizes research in areas poised for impact, including AI and advanced computing, astrophysics, biosciences, climate, and space—as well as supporting researchers in a variety of disciplines through its science systems program. The Humanities and Artificial Intelligence Virtual Institute (HAVI) intends to spur innovative, domain-specific research outcomes from humanities scholars through the integral application of AI-inspired tools and techniques, as well as produce insights from the humanities that will advance the development of AI.