When Stylometry Meets Intellectual History: What

When Stylometry Meets Intellectual History: What "The Bakhtin Circle" Teaches Us About Authorship

●1 ●3 ●24
calendar_today • schedule3 min read

Every developer who has worked with stylometry knows the drill: you feed a corpus of texts with known authors into a machine learning pipeline, extract a feature set (usually the most frequent words), apply some dimensionality reduction (PCA, anyone?), and watch as the unknown text confidently clusters with its true author. It's elegant, it's powerful, and for most practical cases—from disputed Federalist Papers to identifying anonymous blog posts—it works remarkably well.

But what happens when you point this finely-tuned machinery at a collective intellectual project? What if the "author" isn't an individual at all, but a circle of thinkers who read each other's drafts, argued over concepts, and published under each other's names? This is exactly the fascinating problem tackled by Orekhov and Vasilenko in their recent paper "Quantitative Authorship Attribution and the Bakhtin Circle: The Problem of Authorship in Intellectual Communities" (2026).

The Core Problem: Attribution in a Collaborative Context

The Bakhtin Circle—that legendary group of Soviet thinkers including Mikhail Bakhtin, Valentin Voloshinov, and Pavel Medvedev—presents a nightmare scenario for traditional authorship attribution. These scholars didn't just influence each other; they actively collaborated, shared ideas, and in some cases, published works that may have been jointly authored or even ghost-written. The question "Who really wrote Marxism and the Philosophy of Language?" has haunted literary scholars for decades.

The paper, which I'll cite properly in a moment, approaches this not as a simple "who wrote what" puzzle, but as a methodological challenge: how do we test authorship attribution methods when the ground truth itself is contested, and when the very concept of "individual authorship" breaks down?

The Technical Approach: Stylo and Delta

The researchers used the excellent R package stylo (Eder et al., 2016)—a toolkit that has become the workhorse of computational stylistics. If you haven't used it, imagine scikit-learn for text analysis but with a specialized focus on authorship problems. It implements the classic "Delta" method (Burrows' distance) along with various machine learning classifiers.

Their pipeline is refreshingly straightforward:

  1. Feature extraction: They used the most frequent words (MFW) as features—a common and effective choice in stylometry
  2. Distance calculation: They computed the Delta distance between texts
  3. Visualization: They used Principal Component Analysis (PCA) and cluster analysis to visualize textual relationships
  4. Validation: They performed cross-validation to test the stability of their attributions

What makes this study particularly interesting for developers is the data strategy. Instead of relying on contested attributions, they built their training set from texts with uncontroversial authorship—works that everyone agrees were written by Bakhtin, Voloshinov, or Medvedev individually. The test set then included the disputed works.

This approach is worth remembering. In real-world attribution problems, we rarely have perfect ground truth. We often have to construct it from a subset of cases where attribution is clear, then test our methods on the ambiguous ones.

The Surprising (and Cautionary) Result

Here's where it gets interesting for practitioners. The stylometric analysis did not produce the clean, unambiguous clusters that we often see in textbook examples. The disputed texts didn't neatly align with a single author. Instead, they showed varying degrees of similarity to different circle members.

Now, a less careful researcher might have concluded "the method doesn't work." But Orekhov and Vasilenko make a much more sophisticated argument: the ambiguous results actually reflect the collaborative reality of the Bakhtin Circle. When Voloshinov and Bakhtin worked closely together, exchanging ideas and editing each other's prose, the resulting texts naturally contain elements of multiple writing styles. The "noise" in our stylometric analysis isn't a bug—it's a feature that tells us about the social dynamics of intellectual production.

This is a crucial insight for developers: our tools are never neutral. They encode assumptions about authorship (e.g., that style is stable and unique to individuals) that may not hold in all contexts. When we see ambiguous results, we should ask: is this a methodological failure, or is the data telling us something about the nature of authorship in this particular case?

Practical Implementation: R Code for Stylometric Analysis

To make this concrete, here's how you can perform a similar analysis using the stylo package in R. The following code demonstrates the key steps used in the paper—from loading the corpus to performing cluster analysis and cross-validation.

Setting Up the Environment

First, install and load the necessary packages:

# Install stylo from CRAN if you haven't already
# install.packages("stylo")

# Load the package
library(stylo)

Basic Stylometric Analysis

The main workhorse function is stylo(). Here's how to run a basic analysis on your own corpus:

# --- Basic stylometric analysis ---
# Place your plain text files (.txt) in a folder, e.g., "bakhtin_corpus"
# Each file should be named with author information, e.g., "Bakhtin_text1.txt"

results <- stylo(
    gui = FALSE,                        # Set to TRUE for interactive mode
    corpus.dir = "bakhtin_corpus",      # Path to your text files
    mfw.min = 100,                      # Minimum number of most frequent words
    mfw.max = 100,                      # Maximum number (fixed set here)
    mfw.incr = 1,                       # Increment step (ignored if min=max)
    analysis.type = "CA",               # "CA" for Cluster Analysis
    distance.measure = "delta",         # Burrows' Delta distance
    write.png = TRUE                    # Save plots as PNG files
)

# View the distance matrix (first 10 rows and columns)
print("Delta distance matrix (first 10 texts):")
print(results$distance.matrix[1:10, 1:10])

Advanced: Manual Feature Selection and Cross-Validation

For more control over the process—mirroring the approach in Orekhov and Vasilenko—you can manually extract features and perform cross-validation:

# --- Manual feature extraction ---
# Load all texts from the corpus directory
corpus <- load.corpus(files = "all", corpus.dir = "bakhtin_corpus")

# Create a table of frequencies for all words
all_freq <- make.frequencies(corpus)

# Select the top N most frequent words as features
n_features <- 150
mfw_list <- frequent.features(all_freq, cutoff = n_features)

# Create a table of relative frequencies for these features
# This is your feature matrix: rows = texts, columns = word frequencies
feature_table <- make.table.of.frequencies(corpus, mfw_list)

# --- Delta distance calculation ---
# Calculate Delta distance matrix manually
delta_matrix <- dist.delta(feature_table)

# Perform PCA for visualization
pca_results <- prcomp(delta_matrix, scale. = TRUE)

# Create a PCA plot with labels
plot(pca_results$x[,1], pca_results$x[,2],
     type = "n",
     xlab = "PC1", ylab = "PC2",
     main = "PCA of Bakhtin Circle Corpus")
text(pca_results$x[,1], pca_results$x[,2],
     labels = rownames(pca_results$x),
     cex = 0.7)

# --- Cross-validation to test classification accuracy ---
# Use the classify() function to perform leave-one-out cross-validation
# This tests how well the model can predict authorship for unknown texts

# First, create a vector of known author labels
# These should match the filenames or be provided separately
# For example: author_labels <- c("Bakhtin", "Bakhtin", "Voloshinov", ...)

# Perform cross-validation using the Delta distance
cv_results <- classify(
    corpus.dir = "bakhtin_corpus",      # Path to corpus
    mfw = n_features,                   # Number of features to use
    distance.measure = "delta",         # Distance measure
    cv.folds = 10,                      # Number of cross-validation folds
    classification.method = "knn"       # k-nearest neighbors classification
)

# Print classification accuracy
print("Cross-validation accuracy:")
print(paste("Overall accuracy:", 
            round(cv_results$accuracy * 100, 2), "%"))

# View the confusion matrix
print("Confusion matrix:")
print(cv_results$confusion.matrix)

Processing the Bakhtin Circle Dataset

If you want to replicate the analysis from the paper, you would modify the code to handle multiple feature sets and visualize the ambiguous cases:

# --- Iterative feature selection to find optimal MFW ---
# Test different numbers of most frequent words
# to find the range that yields the best separation

mfw_range <- seq(from = 50, to = 500, by = 50)
accuracy_results <- data.frame(mfw = mfw_range, accuracy = NA)

for (i in seq_along(mfw_range)) {
    cat("Testing MFW =", mfw_range[i], "...\n")
    cv_result <- classify(
        corpus.dir = "bakhtin_corpus",
        mfw = mfw_range[i],
        distance.measure = "delta",
        cv.folds = 10,
        classification.method = "knn"
    )
    accuracy_results$accuracy[i] <- cv_result$accuracy
}

# Plot accuracy by number of features
plot(accuracy_results$mfw, accuracy_results$accuracy,
     type = "b",
     xlab = "Number of Most Frequent Words",
     ylab = "Cross-Validation Accuracy",
     main = "Feature Selection for Bakhtin Circle Attribution")

# --- Visualizing the ambiguous cases ---
# In the paper, disputed texts showed mixed affiliations
# You can highlight these in your plots

# Identify disputed texts (e.g., those with uncertain authorship)
# For demonstration, let's assume you have a vector of disputed text names
# disputed_texts <- c("Marxism_and_Philosophy_of_Language.txt", ...)

# In PCA plot, color disputed texts differently
is_disputed <- rownames(pca_results$x) %in% disputed_texts
plot_colors <- ifelse(is_disputed, "red", "blue")

plot(pca_results$x[,1], pca_results$x[,2],
     col = plot_colors,
     pch = ifelse(is_disputed, 17, 16),
     xlab = "PC1", ylab = "PC2",
     main = "PCA with Disputed Texts Highlighted")
text(pca_results$x[,1], pca_results$x[,2],
     labels = rownames(pca_results$x),
     pos = 3, cex = 0.6)

The Open Dataset: A Resource for the Community

One of the most valuable contributions of this paper is the release of their data on OSF. The dataset (Orekhov & Vasilenko, 2025) contains:

  • The full corpus of texts used in the analysis
  • Metadata about authorship attributions (both certain and disputed)
  • The feature tables generated during analysis

This is a significant resource for developers working in stylometry. It provides a test case for authorship attribution in a "difficult" context—one where the usual assumptions about individual style don't straightforwardly apply. If your algorithm can handle the Bakhtin Circle, it can probably handle most real-world attribution problems.

Key Takeaways for Developers

1. Always validate your assumptions. Stylometry works best when authorship is individual and style is consistent. Before applying it, ask: does the text corpus meet these assumptions? If not, how might that affect the results?

2. Embrace ambiguity. Not every classification problem yields a clear answer. Sometimes the "noise" contains important information about the data generation process. In the Bakhtin Circle case, the ambiguous classifications revealed patterns of intellectual collaboration that traditional scholarship had debated for decades.

3. Use the right tools for the right problem. stylo is a mature, well-documented package that handles the core tasks of authorship attribution elegantly. Its GUI interface makes it accessible to non-programmers, while its R backend allows for deep customization. The paper demonstrates its effectiveness even on challenging datasets.

4. Share your data. The authors' decision to release their dataset on OSF is a model for reproducible research. If you're working on stylometric problems, consider doing the same—it allows others to validate your findings and build upon your work.

5. Tune your parameters carefully. The choice of the number of most frequent words (MFW) can significantly impact results, as shown in the feature selection plot above. Always test a range of values and consider the stability of your classifications across different feature sets.

The Full Citation

For those who want to dig deeper, here's the proper reference (paper in Russian):

Orekhov, B.V., & Vasilenko, A.G. (2026). Quantitative Authorship Attribution and the Bakhtin Circle: The Problem of Authorship in Intellectual Communities. New Philological Bulletin, (1), 40-55. DOI: 10.54770/20729316-2026-1-40

And the dataset:

Orekhov, B.V., & Vasilenko, A.G. (2025). Quantitative Authorship Attribution and the Bakhtin Circle: the Problem of Authorship in Intellectual Communities (stylometry data). OSF. DOI: 10.17605/OSF.IO/R9WBZ

2 Comments

1 vote
1
🔥 Join developers growing publicly
Share your knowledge, build in public, and grow your developer presence with a global community.

More Posts

What Developers Already Know About Data Center Delays

Tom Smithverified - Sep 28

Code, Poetry & Neural Ghosts: What AI Teaches Us About Verse

nevmenandr - Aug 15

Legacy in the Data: Transforming Family Medical History into a Blueprint for Longevity

Huifer - Jan 29

Unlocking New Horizons: What GitHub Copilot Teaches Us About the Future of Code

Sunny - Oct 3, 2025

When the Memory Gate Met a Real Archive: What 90 Experiments Taught Us About Cheap LLM Slop

Flamehaven - Jun 4
chevron_left
1.4k Points • 28 Badges
36Posts
7Comments
10Connections
Digital Humanities researcher

Related Jobs

View all jobs →

Commenters (This Week)

2 comments
2 comments
1 comment

Contribute meaningful comments to climb the leaderboard and earn badges!