Open-Source Python Tools for Russian Text Processing in Digital Humanities

Open-Source Python Tools for Russian Text Processing in Digital Humanities

●1 ●3 ●26
calendar_today • schedule4 min read

Practical tools for OCR, accentuation, and text analysis that every developer should know about


The Problem

If you're working with Russian texts in the humanities, you've probably faced this situation: you need to process a pre-1917 text, but every NLP tool is trained on modern Russian. Or you want to analyze verse, but stress marks are missing. Or you have a TEI XML document that needs conversion, but existing tools are either proprietary or dead.

While large language models are everywhere, they're surprisingly bad at these specific tasks. As Orekhov (2025) points out in his review, "modern versions of artificial intelligence cannot handle subtle nuances tied to specific subject domains."

Worse, even when good solutions exist, nobody knows about them. Orekhov mentions at least two separate attempts to rebuild functionality that was already available in open-source Python packages — a clear sign that the real problem isn't technical, but informational.


What's Actually Available

Here are the tools Orekhov covers. All are open-source, Python-based, and created by researchers at HSE University's School of Linguistics.

1. Pre-1917 Russian Orthography Support

OCR Model for Old Orthography

Repository: Serovvans/trocr-preferform-orthography

The model uses Tesseract-OCR infrastructure and handles the pre-reform Russian alphabet (including Ѣ, Ѳ, Ѵ, and the final Ъ). It's actually trained on modern Russian text in old orthography (not Old Church Slavonic — the model's description incorrectly labels it as "Old Russian").

Quality metrics:

  • CER (Character Error Rate): 0.095
  • WER (Word Error Rate): 0.298

This means only about 70% of words are correctly recognized. Not great, but it's free and works well enough when combined with GPT-4 for correction, as Orekhov demonstrates with an example where GPT-4o successfully cleans up OCR errors while preserving the old orthography.

Example usage:

from recognize_page import recognize_page

page_path = "page.png"
text = recognize_page(page_path, text_output_path="output/file.txt")
print(f"Text from page:\n{text}")

Dependencies: tesseract-ocr, poppler, pytesseract, Pillow, transformers, torch, numpy, tqdm, pdf2image

⚠️ Note: For 18th-century texts with that weird 'm'-shaped 'т', this model doesn't work well. You'll need proprietary Transkribus models (which Orekhov warns against due to their dependence on a private company with an uncertain future).

Old to New Orthography Converter

Repository: preferom2modern on PyPI

Simple, dependency-free string conversion. Can output plain text or TEI XML with <choice> tags preserving both versions:

<choice>
  <reg>пример</reg>
  <orig>примѣръ</orig>
</choice>

Works from both command line and Python. Orekhov notes it's been unmaintained for years, but that's actually fine — it worked correctly from day one.


2. Literary and Poetic Text Analysis

Accentuation Module for Russian Verse

Repository: ru-accent-poet on PyPI

This is a hybrid approach: dictionary-based + recurrent neural network. It's designed specifically for poetic texts where stress placement is crucial for meter analysis.

Performance: The main downside is speed. A pre-annotated dataset of classical Russian prose was published in 2024 to let researchers skip the processing step.

Citation: Короткова Ю. О. (2022). Комбинированный словарно-нейросетевой акцентуатор для разметки русского поэтического текста. Труды Института русского языка им. В. В. Виноградова, 3(33), 181–190.

Direct Speech Detector

Repository: neiriz/direct-speech-detector

This isn't just a rule-based approach — it's a full pipeline combining regex (for direct speech marks) with morphological parsing, parsing, and text segmentation. It handles cases where direct speech and authorial narration are mixed.

Use cases:

  • Comparing dialogue density across novels (there's previous work by Sobchuk, 2016)
  • Analyzing character speech patterns
  • Extracting dialogue for stylistic analysis
Formulaicity Assessment Module

Repository: evgenilatyshev/formulaicity

This is for folklore studies. It detects formulaic expressions — repeated phrases characteristic of oral traditions. The algorithm combines:

  • Vocabulary variability coefficient (VocD)
  • N-gram coefficients
  • Binomial expressions
  • Fixed phrases
  • Interjections

Example output (from a folk song analyzed by the module's author, V. Sidnenko):

Found n-grams: 
- черный ворон

Formulaicity coefficient: 1.207

Dependencies: spacy, pymorphy2, numpy, sklearn

TEI XML Converter

Repository: TEItransformer on PyPI

TEI is the standard XML format for digital humanities. This module, developed by Kostyanitsyna and Skorinkin (2024), converts TEI XML to:

  • DOCX (editable format)
  • HTML (with search functionality)
  • JSON (for data analysis)

Under the hood, it's actually an XSLT wrapper — which Orekhov candidly calls an "outdated and unreliable technology" for XML processing, but acknowledges that building a pure Python alternative proved too complex.

The generated HTML includes a search interface that lets you query by character, speaker, or any TEI-annotated field.


3. Journal Publication Preprocessing

cgi-processor

Repository: nevmenandr/cgi_processor

This is a niche tool for the journal "Цифровые гуманитарные исследования" (Digital Humanities Research). It handles LaTeX preprocessing:

  • Non-breaking spaces between initials and surnames
  • Em-dashes in number ranges
  • Other typographic rules specific to Russian academic publishing

The journal uses LaTeX for production, and manual typographic cleanup is error-prone. This automates it, saving editors from the tedious process of manually inserting special characters.


Why This Matters

Here's the thing: these tools solve real problems that LLMs can't handle well. They're infrastructure. They're the kind of boring, practical code that enables interesting research.

The bigger issue, as Orekhov emphasizes, is visibility. Researchers often don't even check if their problem has already been solved. There's a mental block: "my problem is so specific, surely nobody else has built this."

But they have. And it's open source.


Reference

Orekhov, B. V. (2025). Otkrytye komp'yuternye instrumenty dlya resheniya zadach otsifrovki i analiza russkoyazychnogo teksta v oblasti Digital Humanities [Open computer tools for digitization and analysis of Russian-language text in Digital Humanities]. Tsifrovye gumanitarnye issledovaniya, (2), 71–83. DOI: 10.31860/cgi-2025-2-71-83

Available (in Russian) at: https://nevmenandr.github.io/portfolio/assets/pdf/cgi2025-2-tools.pdf

2 Comments

1 vote
1
🔥 Join developers growing publicly
Share your knowledge, build in public, and grow your developer presence with a global community.

More Posts

The Audit Trail of Things: Using Hashgraph as a Digital Caliper for Provenance

Ken W. Algerverified - Apr 28

When Stylometry Meets Intellectual History: What "The Bakhtin Circle" Teaches Us About Authorship

nevmenandr - Sep 24

Dashboard Operasional Armada Rental Mobil dengan Python + FastAPI

Masbadar - Mar 12

7 Best Tools for Founders Building in Public (2026 Guide)

Udit060 - Jul 15

Academia.edu's podcast summary of Computer text analysis for Digital Humanities

nevmenandr - Sep 10
chevron_left
1.5k Points • 30 Badges
36Posts
7Comments
12Connections
Digital Humanities researcher

Related Jobs

View all jobs →

Commenters (This Week)

1 comment
1 comment
1 comment

Contribute meaningful comments to climb the leaderboard and earn badges!