Every study on this subject has to solve one problem: who agreed to be measured

Recruitment decides the answer before anyone picks up a ruler. It is the largest source of disagreement between studies and the least discussed.

By Flavius Cojocaru · 6 August 2026 · 7 min read

Two research teams measure men in the same country, in the same decade, using the same definition of the measurement, and come back with figures that differ by more than a centimetre.

Neither team is careless. Neither is lying. The difference is upstream of the ruler, in the question every study on this subject has to answer first and almost none can answer well: who agreed to take part?

Volunteer bias, stated plainly

Suppose you advertise for participants in a study measuring penis size. Who replies?

Not a random sample of men. Replying requires being comfortable enough with the subject to walk into a room and be measured by a stranger, and the willingness to do that is not evenly distributed. Men confident about the answer are, on the whole, likelier to volunteer than men who dread it.

The direction of the resulting bias is upward, and researchers raise it as the standard limitation of this literature rather than as a courtesy. It is why a study that recruited volunteers by advertisement should be read differently from one that measured consecutive patients attending a clinic for an unrelated reason.

The four questions to ask of any figure

This is the checklist that separates a number worth quoting from one that merely exists, and it applies to every row in this game.

  1. Who was measured, and why were they there? Volunteers respond to advertisements; patients attend for a reason. Both are informative and neither is the general population.
  2. How many? A study of two hundred men gives an estimate with a wide interval around it. A pooled analysis across tens of thousands narrows that interval considerably. A single small study is imprecise rather than wrong, and readers confuse the two constantly.
  3. Measured by whom, under what protocol? Clinician-measured or self-reported changes the answer substantially and predictably. Bone-pressed or not changes it by roughly a centimetre.
  4. What was the age range? A study of men in their twenties and one of men in their sixties are measuring different populations. On how the figure changes with age, the literature is thinner than you would hope.

Why this makes country comparisons so treacherous

Now stack the problem. A cross-country comparison requires that every country's figure was produced under conditions similar enough to be compared.

In practice that condition is almost never met. One country's number comes from a large clinician-measured study of outpatients; another's from a smaller volunteer sample; a third's from a study using a different protocol for the same nominal measurement. Rank those three and you have produced a ranking of study designs wearing the costume of a ranking of countries.

This is the single most important caveat on this site, and the reason each row here carries its sample size and its method rather than just its figure. A gap between two countries that is smaller than the gap between two protocols tells you nothing about the countries and quite a lot about how late you got to the methods section.

What each design gets right and wrong

RecruitmentMain strengthMain bias
Advertised volunteersCooperative, easy to runLikely skews upward
Clinic outpatientsNo self-selection into topicNot the general population
Conscripts or screeningLarge, near-universal in cohortNarrow age band
Self-measurement surveyVery large samples possibleSystematically overstated

Why the pooled reviews are worth more than any single study

This is the reason a systematic review beats even a good individual study, and why the figures used as anchors on this site come from pooled work rather than from someone's favourite paper.

A review sets a protocol before it begins, states which studies qualify and which are excluded and why, and combines what survives with weights reflecting size and quality. Individual biases do not vanish, but they no longer determine the answer on their own, and the review makes its own filtering inspectable in a way a single study never can.

So two independent reviews, conducted years apart across largely different sets of studies, landing within a few millimetres of each other is a far stronger result than either review alone. Agreement between methods that could have disagreed is the closest thing this field offers to a fact.

What follows for reading the map

Nothing here says the country figures are worthless. It says they are estimates with error bars, that the error bars are wider than the differences between many adjacent countries, and that anyone presenting them as a clean league table is either not reading the methods or hoping you won't.

A map that colours in every country on earth has, necessarily, filled in the countries where no primary study exists. The interesting question about such a map is never what it says about a country. It is where the number came from, and the honest answer for a great many countries is nowhere.

Before you go

This piece is published for general interest and education. It is not medical advice, and nothing in it is a substitute for speaking to a doctor. Every figure quoted is a population average from a published study, and the variation between individual men inside any one of those populations is far larger than any difference between them.

If something about your own body is worrying you, please take it to a clinician rather than to an article.

Sources

More guides