Every study on this subject has to solve one problem: who agreed to be measured
Recruitment decides the answer before anyone picks up a ruler. It is the largest source of disagreement between studies and the least discussed.
Two research teams measure men in the same country, in the same decade, using the same definition of the measurement, and come back with figures that differ by more than a centimetre.
Neither team is careless. Neither is lying. The difference is upstream of the ruler, in the question every study on this subject has to answer first and almost none can answer well: who agreed to take part?
Volunteer bias, stated plainly
Suppose you advertise for participants in a study measuring penis size. Who replies?
Not a random sample of men. Replying requires being comfortable enough with the subject to walk into a room and be measured by a stranger, and the willingness to do that is not evenly distributed. Men confident about the answer are, on the whole, likelier to volunteer than men who dread it.
The direction of the resulting bias is upward, and researchers raise it as the standard limitation of this literature rather than as a courtesy. It is why a study that recruited volunteers by advertisement should be read differently from one that measured consecutive patients attending a clinic for an unrelated reason.
The four questions to ask of any figure
This is the checklist that separates a number worth quoting from one that merely exists, and it applies to every row in this game.
- Who was measured, and why were they there? Volunteers respond to advertisements; patients attend for a reason. Both are informative and neither is the general population.
- How many? A study of two hundred men gives an estimate with a wide interval around it. A pooled analysis across tens of thousands narrows that interval considerably. A single small study is imprecise rather than wrong, and readers confuse the two constantly.
- Measured by whom, under what protocol? Clinician-measured or self-reported changes the answer substantially and predictably. Bone-pressed or not changes it by roughly a centimetre.
- What was the age range? A study of men in their twenties and one of men in their sixties are measuring different populations. On how the figure changes with age, the literature is thinner than you would hope.
Why this makes country comparisons so treacherous
Now stack the problem. A cross-country comparison requires that every country's figure was produced under conditions similar enough to be compared.
In practice that condition is almost never met. One country's number comes from a large clinician-measured study of outpatients; another's from a smaller volunteer sample; a third's from a study using a different protocol for the same nominal measurement. Rank those three and you have produced a ranking of study designs wearing the costume of a ranking of countries.
This is the single most important caveat on this site, and the reason each row here carries its sample size and its method rather than just its figure. A gap between two countries that is smaller than the gap between two protocols tells you nothing about the countries and quite a lot about how late you got to the methods section.
What each design gets right and wrong
| Recruitment | Main strength | Main bias |
|---|---|---|
| Advertised volunteers | Cooperative, easy to run | Likely skews upward |
| Clinic outpatients | No self-selection into topic | Not the general population |
| Conscripts or screening | Large, near-universal in cohort | Narrow age band |
| Self-measurement survey | Very large samples possible | Systematically overstated |
Why the pooled reviews are worth more than any single study
This is the reason a systematic review beats even a good individual study, and why the figures used as anchors on this site come from pooled work rather than from someone's favourite paper.
A review sets a protocol before it begins, states which studies qualify and which are excluded and why, and combines what survives with weights reflecting size and quality. Individual biases do not vanish, but they no longer determine the answer on their own, and the review makes its own filtering inspectable in a way a single study never can.
So two independent reviews, conducted years apart across largely different sets of studies, landing within a few millimetres of each other is a far stronger result than either review alone. Agreement between methods that could have disagreed is the closest thing this field offers to a fact.
What follows for reading the map
Nothing here says the country figures are worthless. It says they are estimates with error bars, that the error bars are wider than the differences between many adjacent countries, and that anyone presenting them as a clean league table is either not reading the methods or hoping you won't.
A map that colours in every country on earth has, necessarily, filled in the countries where no primary study exists. The interesting question about such a map is never what it says about a country. It is where the number came from, and the honest answer for a great many countries is nowhere.
Before you go
This piece is published for general interest and education. It is not medical advice, and nothing in it is a substitute for speaking to a doctor. Every figure quoted is a population average from a published study, and the variation between individual men inside any one of those populations is far larger than any difference between them.
If something about your own body is worrying you, please take it to a clinician rather than to an article.
Sources
More guides
- Almost every adolescent comparison is a comparison of clocksPuberty runs on a schedule that varies by years between boys the same age. Nearly all the distress of that period comes from mistaking a difference in timing for a difference in outcome.
- What the evidence actually says about every method of changing the numberAn industry worth a great deal of money rests on a body of evidence that can be summarised honestly in one page. Here it is, category by category.
- How a number with no source becomes a fact everyone knowsThe country maps on this subject are mostly fiction, and they got there by a repeatable process. Once you can see the process, you cannot unsee it.
- Girth is the measurement that has a practical consequence, and nobody discusses itCondom fit affects breakage, slippage and whether one gets used at all. It is the one place where a dimension genuinely matters, and it is the one nobody talks about.