Why do some papers keep getting cited, and most don't?

August 20, 2026 - 4 minute read -
research science citations.md matthew-effect.md research-metrics.md blog

Why do some papers keep getting cited, and most don’t?

Years ago, at Automattic, I looked at what made a blog post collect likes. The strongest predictor I found wasn’t the topic, the length, or the time of day. It was whether the same author’s earlier posts had collected likes. Likes bred likes. If you already had them, you got more. If you didn’t, you mostly didn’t.

Sociologists have a name for this: the Matthew effect. The rich get richer. In citation studies it goes by “preferential attachment” or “cumulative advantage,” and it is the standard answer to a question every researcher asks: why do a handful of papers get cited thousands of times while most sink without a trace?

A recent preprint says that answer is incomplete. In “Community-centric modeling of citation dynamics explains collective citation patterns in science, law, and patents” (arXiv, January 2025), Sadamori Kojaku and colleagues argue that whether a paper keeps getting cited is not mainly about the paper. It is about the community the paper sits in, and where that community’s attention happens to drift. Success, in their framing, is less about you and more about the crowd you run with.

What rich-get-richer can’t explain

Pure cumulative advantage has an awkward blind spot. If citations only flow to papers that already have citations, then a paper ignored at birth should stay ignored forever. But that is not what happens. Some papers sit unread for years, sometimes decades, and then suddenly surge. The literature calls them sleeping beauties, and the delay is called delayed recognition. The authors note that models of individual citation trajectories fail to reproduce this. You cannot get a sleeping beauty out of a machine whose only rule is “the popular get more popular.”

The idea: citations are collective decisions

Their move is to stop modelling a paper as a lone object with intrinsic pull, and start modelling the citing side. A citation, in this view, is a decision made by the authors who are writing now, and those authors cluster into communities in a high-dimensional “knowledge space.” Communities drift through that space over time. As a community moves, the work near its current location gets cited, and work it has moved away from goes quiet. That drift produces recency even though the model is never told about time. There is no explicit aging term. Papers age because attention wanders off.

It also explains sleeping beauties in a way rich-get-richer cannot. A paper published in a sparse, out-of-the-way corner of knowledge space waits. Years later a community wanders into that corner, and the old paper wakes up. The model reproduces this across all three corpora it was tested on: science, law, and patents, using a newly available U.S. case-law dataset for the legal side. The three domains share the same heavy-tailed patterns, which is part of why the authors think they are onto a general mechanism rather than a quirk of academia.

The model keeps three ingredients working together, as the authors describe them: relevance, which is how close a paper sits to where a community’s attention is; cumulative advantage, the old rich-get-richer pull, still present, just no longer alone; and fitness, something like intrinsic quality. Preferential attachment is not wrong here. It is one of three, not the whole story.

What this does not tell you

Two cautions before anyone gets excited. First, this is a preprint. It hasn’t been through peer review as I write this, and preprints get revised, sometimes heavily. Read it as a strong hypothesis, not a settled result. Second, and more important, it is a model. It reproduces collective, aggregate patterns: the shape of the distribution, the abundance of sleeping beauties across a whole field. It says almost nothing about your specific paper. Knowing that attention drifts through communities doesn’t tell you when, or whether, a community will drift toward your particular corner. A model that explains the crowd is not a plan for the individual.

Does it tell you anything useful?

Carefully, maybe one thing. If citations track where communities are looking, then being legible to the community you actually belong to isn’t vanity. It’s how you get seen at all. That is not the same as saying “promote harder and you’ll get cited.” The paper makes no such claim, and I won’t put it in the authors’ mouths. It’s closer to a reframing: a paper nobody in your community can find, or place, is a paper sitting in an empty corner of the space, waiting for a drift that may never come.

I find the collective framing convincing, and I might be wrong about that. If the mechanism is real, the uncomfortable implication is that a lot of what we call a paper’s quality is really a statement about its neighbourhood. Would your best-cited work have done as well one community over? I genuinely don’t know. Neither, quite, does the model.