Practical tools for OCR, accentuation, and text analysis that every developer should know about
The Problem
If you're working with Russian texts in the humanities, you've probably faced this situation: you need to process a pre-1917 text, but every NLP tool is trained on modern Russian. Or you want to analyze verse, but stress marks are missing. Or you have a TEI XML document that needs conversion, but existing tools are either proprietary or dead.
While large language models are everywhere, they're surprisingly bad at these specific tasks. As Orekhov (2025) points out in his review, "modern versions of artificial intelligence cannot handle subtle nuances tied to specific subject domains."
Worse, even when good solutions exist, nobody knows about them. Orekhov mentions at least two separate attempts to rebuild functionality that was already available in open-source Python packages — a clear sign that the real problem isn't technical, but informational.
What's Actually Available
Here are the tools Orekhov covers. All are open-source, Python-based, and created by researchers at HSE University's School of Linguistics.
1. Pre-1917 Russian Orthography Support
OCR Model for Old Orthography
Repository: Serovvans/trocr-preferform-orthography
The model uses Tesseract-OCR infrastructure and handles the pre-reform Russian alphabet (including Ѣ, Ѳ, Ѵ, and the final Ъ). It's actually trained on modern Russian text in old orthography (not Old Church Slavonic — the model's description incorrectly labels it as "Old Russian").
Quality metrics:
- CER (Character Error Rate): 0.095
- WER (Word Error Rate): 0.298
This means only about 70% of words are correctly recognized. Not great, but it's free and works well enough when combined with GPT-4 for correction, as Orekhov demonstrates with an example where GPT-4o successfully cleans up OCR errors while preserving the old orthography.
Example usage:
from recognize_page import recognize_page
page_path = "page.png"
text = recognize_page(page_path, text_output_path="output/file.txt")
print(f"Text from page:\n{text}")
Dependencies: tesseract-ocr, poppler, pytesseract, Pillow, transformers, torch, numpy, tqdm, pdf2image
⚠️ Note: For 18th-century texts with that weird 'm'-shaped 'т', this model doesn't work well. You'll need proprietary Transkribus models (which Orekhov warns against due to their dependence on a private company with an uncertain future).
Old to New Orthography Converter
Repository: preferom2modern on PyPI
Simple, dependency-free string conversion. Can output plain text or TEI XML with <choice> tags preserving both versions:
<choice>
<reg>пример</reg>
<orig>примѣръ</orig>
</choice>
Works from both command line and Python. Orekhov notes it's been unmaintained for years, but that's actually fine — it worked correctly from day one.
2. Literary and Poetic Text Analysis
Accentuation Module for Russian Verse
Repository: ru-accent-poet on PyPI
This is a hybrid approach: dictionary-based + recurrent neural network. It's designed specifically for poetic texts where stress placement is crucial for meter analysis.
Performance: The main downside is speed. A pre-annotated dataset of classical Russian prose was published in 2024 to let researchers skip the processing step.
Citation: Короткова Ю. О. (2022). Комбинированный словарно-нейросетевой акцентуатор для разметки русского поэтического текста. Труды Института русского языка им. В. В. Виноградова, 3(33), 181–190.
Direct Speech Detector
Repository: neiriz/direct-speech-detector
This isn't just a rule-based approach — it's a full pipeline combining regex (for direct speech marks) with morphological parsing, parsing, and text segmentation. It handles cases where direct speech and authorial narration are mixed.
Use cases:
- Comparing dialogue density across novels (there's previous work by Sobchuk, 2016)
- Analyzing character speech patterns
- Extracting dialogue for stylistic analysis
Repository: evgenilatyshev/formulaicity
This is for folklore studies. It detects formulaic expressions — repeated phrases characteristic of oral traditions. The algorithm combines:
- Vocabulary variability coefficient (VocD)
- N-gram coefficients
- Binomial expressions
- Fixed phrases
- Interjections
Example output (from a folk song analyzed by the module's author, V. Sidnenko):
Found n-grams:
- черный ворон
Formulaicity coefficient: 1.207
Dependencies: spacy, pymorphy2, numpy, sklearn
TEI XML Converter
Repository: TEItransformer on PyPI
TEI is the standard XML format for digital humanities. This module, developed by Kostyanitsyna and Skorinkin (2024), converts TEI XML to:
- DOCX (editable format)
- HTML (with search functionality)
- JSON (for data analysis)
Under the hood, it's actually an XSLT wrapper — which Orekhov candidly calls an "outdated and unreliable technology" for XML processing, but acknowledges that building a pure Python alternative proved too complex.
The generated HTML includes a search interface that lets you query by character, speaker, or any TEI-annotated field.
3. Journal Publication Preprocessing
cgi-processor
Repository: nevmenandr/cgi_processor
This is a niche tool for the journal "Цифровые гуманитарные исследования" (Digital Humanities Research). It handles LaTeX preprocessing:
- Non-breaking spaces between initials and surnames
- Em-dashes in number ranges
- Other typographic rules specific to Russian academic publishing
The journal uses LaTeX for production, and manual typographic cleanup is error-prone. This automates it, saving editors from the tedious process of manually inserting special characters.
Why This Matters
Here's the thing: these tools solve real problems that LLMs can't handle well. They're infrastructure. They're the kind of boring, practical code that enables interesting research.
The bigger issue, as Orekhov emphasizes, is visibility. Researchers often don't even check if their problem has already been solved. There's a mental block: "my problem is so specific, surely nobody else has built this."
But they have. And it's open source.
Reference
Orekhov, B. V. (2025). Otkrytye komp'yuternye instrumenty dlya resheniya zadach otsifrovki i analiza russkoyazychnogo teksta v oblasti Digital Humanities [Open computer tools for digitization and analysis of Russian-language text in Digital Humanities]. Tsifrovye gumanitarnye issledovaniya, (2), 71–83. DOI: 10.31860/cgi-2025-2-71-83
Available (in Russian) at: https://nevmenandr.github.io/portfolio/assets/pdf/cgi2025-2-tools.pdf