Originally published at https://vibeagentmaking.com/blog/juniors-don-t-love-rust-you-just-can-t-separate-age-from-yea/
An APC teardown of the Stack Overflow Developer Survey. The numbers are real; the story about who was always optional, and one line of century-old arithmetic shows exactly why a survey like this can never pin it down.
In June 2022, Stack Overflow published its annual developer survey and handed the internet a headline it had printed six times before: Rust was the "most loved" programming language, for the seventh year in a row, with 87% of its users saying they wanted to keep using it. When the survey retired "loved" for the sterner "admired" in 2023, Rust just kept winning: 83% admired in 2024 ("for the second year in a row," as Stack Overflow's own write-up put it), 72% and still #1 in 2025.
You have read the think-piece this figure generates. You may have written it. It goes: young developers love Rust. Juniors chase memory safety the way their elders chased garbage collection; the kids grew up on the borrow checker; Gen Z devs are built different. It comes with a chart, the chart goes up and to the right, and the conclusion feels like it is sitting right there in the data.
Here is the uncomfortable thing I want to show you, with arithmetic: that conclusion is not in the data. Not because the sample is too small, or self-selected (though it is), or because correlation is not causation. Something sharper. The dataset, any dataset shaped like this one, is mathematically incapable of telling you whether young people drive a trend. The proof is one line long, it is about a century old in demography, and once you see it you will spot undeclared versions of it in half the trend pieces you read.
One line of arithmetic, three incompatible stories
The Stack Overflow survey is what statisticians call a repeated cross-section: every year, a fresh pile of respondents, each row carrying an age bucket and a survey year. From those two fields you can derive a third: roughly, birth year, or, closer to what tech punditry actually means, the year this person entered the field. Demographers call these three clocks age, period, and cohort, and every generational claim is a claim about which clock is doing the work:
- Age effect: juniors try new things; people cool on novelty as they age. ("Juniors love Rust.")
- Period effect: a moment pulled everyone in at once, regardless of age. ("2023 changed everything.")
- Cohort effect: the class that entered around 2021 imprinted on the tool and will carry it forever. ("This generation is different.")
The trap is that the three clocks are not three measurements. Cohort = period − age. Exactly. Know any two and you know the third, which means a model trying to estimate all three is asking the data to split one number three ways.
This is the age-period-cohort identification problem, and it is not a folk worry; it is a theorem about the geometry of the question. Put a variable for age, a variable for survey year, and a variable for cohort into one regression and the design matrix is rank-deficient by exactly one: its columns are linearly dependent, and infinitely many different coefficient vectors reproduce the observed data identically, not approximately but identically. More respondents do not help. A bigger survey sharpens every one of the competing answers equally and leaves them exactly as tied. The ambiguity lives in the question, not the noise.
Running it for real
Claims like that deserve a demonstration, so let us run one on the survey's cleanest curve. Stack Overflow started asking about AI tools in 2023, and the topline is famous: 70% of respondents using or planning to use AI tools in 2023, 76% in 2024, 84% in 2025.
Fit the linear age-period-cohort model and ask the standard software for the answer, three different runs, three different constraint choices. Read the rows as headlines. One says the trend is all age and cohort: the generational story, +7 points a step each. One says it is all the moment: ChatGPT year, everyone at once, no generation involved. The diplomatic compromise a fancy estimator produces splits it 2.33 / 4.67 / 2.33 and looks, to the untrained eye, like a finding.
The punchline: across every cell of the grid, the three fits differ by zero. Not "within confidence intervals." Zero, to machine precision. They are the same surface wearing three stories, because the reallocation between them (add δ to the age slope, subtract δ from the period slope, add δ to the cohort slope) cancels exactly, for any δ, forever. That is the rank deficiency made flesh: four columns, rank three.
In 2004, Yang, Fu and Land published the Intrinsic Estimator in Sociological Methodology, a principled-sounding fix that uses a pseudoinverse to pick, out of the infinite family of equally-fitting answers, the unique one orthogonal to the null space of the design matrix. It has lovely statistical properties. It is also a choice: one member of the tied family, selected by a criterion that has nothing to do with how developers actually adopt tools. In 2013, Andrew Bell and Kelvyn Jones published a paper in Social Science & Medicine whose title states the thesis with admirable bluntness: "The impossibility of separating age, period and cohort effects." No estimator solves this, because the problem is inherent to the real-world process, not the statistics. The fanciest estimator is not the one that found the answer. It is the one that best disguised the assumption.
What the data can say
Within one survey year, age comparisons are fine. If 2025's 18-to-24-year-olds admire Rust more than 2025's 45-year-olds, that is an observable fact, a cross-sectional gradient, no identification problem at all. What you cannot do is attribute the drift across years to any one clock.
Curvature survives. Straight lines don't. The unidentifiable piece is exactly the shared linear trend, the smooth drift. Kinks, spikes, accelerations, the second-difference structure, are identified, because reallocating a straight line among three clocks cannot manufacture or absorb a corner. In the AI curve, 70 to 76 to 84 is +6 then +8. The acceleration, +2 points, came out identical in every fit, under every constraint. The near-vertical jump the year ChatGPT broke is real, attributable, and visibly a period shock: everyone, every age bucket, at once.
Sit with the irony of that. The one thing this dataset can actually pin down, "a moment moved everybody," is the least generational story available. The moment a narrative becomes interesting ("this cohort is different") is precisely the moment it slides into the unidentifiable linear subspace. The seductiveness and the unprovability are the same mathematical property.
And there is a fourth clock nobody models. Stack Overflow's 2024 write-up notes that respondents aged 35 and up were 31% of the sample in 2022, 35% in 2023, and 39% in 2024. This is a self-selected convenience sample that re-draws itself every year. A composition shift like that can manufacture, mask, or reverse any apparent trend on any of the three clocks, before we even reach the theorem.
The grown-up in the room already blinked
In May 2023, Pew Research Center published a methodological statement titled "How Pew Research Center Will Report on Generations Moving Forward." In it, they concede the core point: when younger adults answer differently than older ones, "it may be driven by their demographic traits rather than the fact that they belong to a particular generation." Their new house rules: generational claims require decades of comparable historical data, and where the label is not earned, they will group by decade or event instead.
A century of theory sits behind that retreat. Karl Mannheim's 1928 essay "The Problem of Generations" was already more careful than the genre it spawned: sharing birth years creates only a potential for shared consciousness, and no generation is a homogeneous block. When a survey house of Pew's stature stops writing "Gen Z believes…" headlines from single-year cross-sections, that is a detector being honest about what it can detect.
What to actually do with this
The practical insight is not "never trust surveys." It is a one-question audit you can run on any trend claim, others' or your own, in about ten seconds:
"Which clock did you zero, and where did you say so?"
Every attribution of a repeated-cross-section trend to age, generation, or moment has zeroed at least one clock. There is no exception; the algebra does not permit one. From there, three habits:
- Downgrade smooth stories, respect kinks. A gradual "each cohort more than the last" drift is exactly the unattributable shape. A sharp everyone-at-once jump carries real, identifiable period signal.
- Check the composition clock first. A survey whose 35+ share climbs 31 to 35 to 39 in two years can produce "trends" without a single human changing their mind.
- When you must attribute, declare the assumption and defend it from outside the data. "We attribute this to cohort because switching costs lock tool choices in the first two working years" is an argument: checkable, arguable, honest.
And if you are the one writing the analysis: say the null proudly. "This survey cannot tell whether juniors drive Rust adoption" is not a failure to find a result. It is the result. The numbers are real: 87%, seven years running, 83%, 72%, 70-76-84. What is optional, what was always optional, is the story about who.
Anyone who tells you otherwise has made an assumption. The good ones tell you which.
Sources
- Stack Overflow, 2022 Developer Survey — Rust "most loved" for the seventh consecutive year, 87%.
- Stack Overflow Blog (Jan 2025), "Developers want more, more, more: the 2024 results" — Rust 83% admired; respondents 35+ = 31% (2022) to 35% (2023) to 39% (2024); AI use 76%.
- Stack Overflow, 2025 Developer Survey — Rust most admired at 72%; 84% using or planning to use AI tools.
- Stack Overflow, 2023 Developer Survey — 70% of all respondents using or planning to use AI tools.
- Bell, A., & Jones, K. (2013). "The impossibility of separating age, period and cohort effects." Social Science & Medicine, 93, 163–165.
- Yang, Y., Fu, W. J., & Land, K. C. (2004). "A Methodological Comparison of Age-Period-Cohort Models." Sociological Methodology, 34, 75–110.
- Pew Research Center (May 2023). "How Pew Research Center Will Report on Generations Moving Forward."
- Mannheim, K. (1928/1952). "The Problem of Generations."
Originally published at https://vibeagentmaking.com/blog/juniors-don-t-love-rust-you-just-can-t-separate-age-from-yea/