An IQ score is a position, not a quantity
The number on an intelligence test report is a deviation IQ. It is manufactured, not measured: the test publisher administers the battery to a large standardisation sample chosen to match the population on age, sex, region, ethnicity and parental education, then transforms the raw item totals so that the resulting scores have a mean of exactly 100 and a standard deviation of exactly 15. The number tells you where a person sits relative to that sample. It carries no units and it is not a count of anything.
This matters because it makes the percentile the more fundamental figure. Saying "115" is shorthand for "one standard deviation above the standardisation sample mean", which is shorthand for "above about 84% of the sample". The percentile survives a change of scale; the raw numeral does not. A score of 132 on the old Stanford-Binet Form L-M, which used a standard deviation of 16, describes the identical position as a 130 on a Wechsler scale. Anyone quoting an IQ without naming the scale has left out half the information.
The original ratio IQ, from Stern and the early Binet-Simon work, was genuinely a ratio: mental age divided by chronological age, times 100. Wechsler abandoned it in the 1930s because mental age stops growing in adulthood while chronological age does not, so the ratio drifts downward for every adult who ages. Every current battery — the Wechsler scales, Stanford-Binet 5, Woodcock-Johnson, KABC — uses the deviation form instead, with age-specific norms so that a 40-year-old is compared with 40-year-olds.
From score to percentile to rarity
Three steps take you from the reported score to everything else. First, standardise: subtract the mean and divide by the standard deviation, z = (IQ − 100) ÷ 15. The z-score is the distance from the mean expressed in standard deviations, and it is the only quantity that is comparable across tests.
Second, integrate the normal curve. The percentile is Φ(z) × 100, where Φ is the cumulative normal distribution — the area under the bell curve to the left of z. There is no closed-form expression for Φ; every table and every calculator, including this one, evaluates it numerically. The familiar landmarks are worth memorising: z = 1 gives 84.1%, z = 2 gives 97.7%, z = 3 gives 99.87%.
Third, invert the tail to get rarity. The proportion at or above the score is 1 − Φ(z), so the number of people you would expect to sample before finding one scoring that high is the reciprocal, 1 ÷ (1 − Φ(z)). This is the figure that makes the tails intuitive. A score of 115 is roughly 1 in 6. A score of 130 is roughly 1 in 44. A score of 145 is roughly 1 in 741. Each additional 15 points is not a fixed increment in rarity — it multiplies rarity by a factor that itself keeps growing.
Running the chain backwards uses the inverse normal function. Given a percentile, z = Φ⁻¹(p), and then IQ = 100 + 15z. Report writers need this when a manual prints percentile ranks but the referral question is phrased in standard scores, and gifted-programme coordinators need it when a cutoff is written as "the top 2%" rather than as a score.
Worked example: a full-scale score of 124 on a Wechsler scale
A report gives a Full Scale IQ of 124 with a 95% confidence interval of 119 to 129. The scale has a mean of 100 and a standard deviation of 15.
- Standardise. 124 − 100 = 24 points above the mean. Divide by the standard deviation: 24 ÷ 15 = z = 1.60.
- Find the area to the left. The standard normal table gives Φ(1.60) = 0.9452. The percentile rank is 0.9452 × 100 = 94.5. The score is at or above about 94.5% of the standardisation sample.
- Invert the upper tail. The proportion above is 1 − 0.9452 = 0.0548, so the rarity is 1 ÷ 0.0548 = 1 in 18.2 people.
- Translate to the other scale. The same z on a standard deviation of 16 gives 100 + 16 × 1.60 = 125.6. On an old Form L-M report the identical person would have been described with a slightly larger numeral.
- Carry the confidence interval through. The lower bound of 119 is z = 1.267, or the 89.7th percentile. The upper bound of 129 is z = 1.933, or the 97.3rd percentile. So the honest statement is "somewhere between roughly the 90th and the 97th percentile", not "the 94.5th percentile".
That last step is the one most often skipped, and it changes how the result should be described. A confidence interval eight percentile points wide is normal for a well-constructed full-scale score; the point estimate on its own overstates how precisely anyone has been located on the curve.
How to read the percentile and the rarity
Start with the confidence interval, not the point estimate. Every reputable report prints one, because a full-scale score carries a standard error of measurement of two to four points on most batteries. Two scores that differ by five points describe the same performance. If your report does not show an interval, treat the score as accurate to about plus or minus five points and convert both ends.
Read the percentile rather than the numeral when you are explaining the result to anyone. "Above about 95 of every 100 children of the same age" communicates the finding; "124" invites the false impression of a measured quantity on a ratio scale, as if 124 were twice as much of something as 62. It is not. There is no zero point on an IQ scale and no meaningful ratio between two scores.
Rarity is the right frame for selection decisions and the wrong frame for describing a person. Gifted programmes commonly set entry at the 95th, 97th or 98th percentile, which is a score of 125, 128 or 131 on a 15-point scale, and the rarity output tells an administrator how many children per hundred that cutoff will identify. But rarity in the tails is exactly where the normal model is least trustworthy: standardisation samples rarely contain enough people beyond z = 3 to verify that the curve keeps its shape, so a claim of "1 in 30,000" is a property of the mathematical model, not an observed frequency.
Composite scores also hide dispersion. A full-scale score of 100 built from a verbal index of 130 and a processing-speed index of 70 has almost nothing in common with a flat profile of 100 across the board, and most test manuals warn against interpreting the composite at all when index scores are that far apart. Convert each index separately here and compare the percentiles; the spread is usually the clinically interesting part. If you want to see the same standard-score machinery working on an admissions test, the SAT scaled score calculator shows how a fixed reporting scale is built from raw counts.
Score, percentile and rarity on both deviation scales
| z | SD 15 score | SD 16 score | Percentile | 1 in N at or above | Classification (WAIS-IV) |
|---|---|---|---|---|---|
| +3.00 | 145 | 148 | 99.87 | 741 | Very superior |
| +2.67 | 140 | 142.7 | 99.62 | 261 | Very superior |
| +2.00 | 130 | 132 | 97.72 | 44 | Very superior |
| +1.33 | 120 | 121.3 | 90.88 | 11.0 | Superior |
| +1.00 | 115 | 116 | 84.13 | 6.3 | High average |
| +0.67 | 110 | 110.7 | 74.75 | 4.0 | High average |
| 0.00 | 100 | 100 | 50.00 | 2.0 | Average |
| −0.67 | 90 | 89.3 | 25.25 | 1.34 | Average |
| −1.00 | 85 | 84 | 15.87 | 1.19 | Low average |
| −1.33 | 80 | 78.7 | 9.12 | 1.10 | Low average |
| −2.00 | 70 | 68 | 2.28 | 1.02 | Borderline |
| −3.00 | 55 | 52 | 0.13 | 1.001 | Extremely low |
WAIS-5 and WISC-V renamed the bands: very superior became extremely high, superior became very high, and borderline became very low. The cut points are unchanged.
Errors that distort an IQ percentile
- Ignoring the scale's standard deviation. A 132 on a 16-point scale and a 130 on a 15-point scale are the same rank. Comparing the numerals directly, which happens constantly with historic Form L-M scores, overstates the older result by about one point per ten points above the mean.
- Quoting the point estimate without the confidence interval. Full-scale scores carry a standard error of two to four points, so a 95% interval spans about ten points. Convert both ends and describe the range.
- Interpreting a composite over a scattered profile. When index scores differ by more than about 1.5 standard deviations, most manuals advise against interpreting the full-scale score at all. The percentile of a composite that averages a 130 and a 70 describes nobody.
- Treating rarity in the far tails as an observed frequency. Beyond about z = 3 the normal model is extrapolating past the data in the standardisation sample. Ratios like one in a million are model output, not counts.
- Comparing scores from different norm dates. Norms drift — the Flynn effect — so a test standardised decades ago yields higher scores than a freshly normed one for the same performance. Publishers restandardise for exactly this reason.
- Using an online screening score as if it were a battery score. Unstandardised internet tests have no norm sample, no reliability estimate and no age norms, so there is no defensible z-score to convert.
- Reading a percentile as a percentage correct. A percentile rank of 84 does not mean 84% of the items were answered correctly. It means 84% of the norm sample scored no higher.
Where deviation scores show up elsewhere
The same transformation runs through the whole of educational measurement. Subtest scaled scores on the Wechsler batteries use a mean of 10 and a standard deviation of 3, so a scaled score of 13 is z = 1 and sits at the 84th percentile, identical in position to a composite of 115. T-scores, common on behaviour rating scales, use a mean of 50 and a standard deviation of 10. Normal curve equivalents use 50 and 21.06, chosen so the scale is equal-interval and lines up with percentiles at 1, 50 and 99. Every one of these is the same z-score wearing different clothing.
Admissions testing works the same way with different constants. SAT sections are reported on a 200-800 scale, and converting a raw count to that scale is the job of the SAT score calculator; moving between the two admissions tests uses a rank-matching table rather than a formula, which the SAT to ACT conversion calculator applies. Curriculum-based measures used in schools are the exception worth knowing about: oral reading fluency norms, computed by the reading speed calculator, are reported directly as percentiles of words correct per minute because the underlying distribution is not normal enough to justify a standard score.
One caution to carry away. Everything on this page is arithmetic on a normal curve, and it is exact. What it cannot tell you is whether the score itself is a fair estimate of the person's ability, which depends on the test's reliability, its norm sample, the testing conditions, the examinee's language and schooling, and the clinician's judgement. Interpretation of a cognitive assessment belongs to a qualified professional. Use these numbers to understand a report, not to replace one.
Key terms
- Deviation IQ
- A standard score set so that the standardisation sample has a mean of 100 and a fixed standard deviation, usually 15. It expresses relative position, not a measured amount.
- Percentile rank
- The percentage of the norm sample scoring at or below a given score. Percentile ranks are ordinal: the gap between the 50th and 55th is far smaller in score points than the gap between the 94th and 99th.
- Standard error of measurement
- The typical difference between an observed score and the true score it estimates, used to build the confidence interval printed on a report.
- Flynn effect
- The long-run rise in raw test performance across generations, which forces publishers to restandardise. It makes scores from tests normed decades apart non-comparable.
- Normal curve equivalent
- A 1-to-99 scale with a mean of 50 and a standard deviation of 21.06, built so that equal differences represent equal amounts, unlike percentiles.
