What the grade level actually measures
Flesch-Kincaid measures exactly two things: the average number of words in your sentences, and the average number of syllables in your words. Nothing else enters the formula. It does not read your text, does not know what your words mean, and cannot tell whether your argument is clear.
That narrowness is a feature as much as a limitation. Both quantities are proxies for genuine cognitive load. Long sentences hold more clauses in working memory before the reader can resolve them, and long words in English are disproportionately Latin- and Greek-derived, so they are learned later and recognised more slowly. The 1975 Navy study behind the formula fitted these two variables against comprehension test results on real training material, and they explained enough of the variance to be useful.
The output is calibrated to US school grades: 8.0 means an average eighth grader can read it. In practice the scale is more useful as a relative instrument than an absolute one. Rewriting a document from 14.2 to 9.6 is a real and verifiable improvement; insisting that 9.6 is the boundary of acceptability is reading more precision into the number than the regression supports.
Flesch-Kincaid is now embedded in policy. It appears in US federal plain-language guidance, in military specifications for technical documentation, in health-literacy standards for patient information, and it is the readability figure Microsoft Word reports. That institutional weight is why writers need to compute it deliberately rather than discover it after submission.
The two terms, and which one to pull
Write the formula as 0.39·ASL + 11.8·ASW − 15.59 and the structure becomes obvious. Average sentence length carries a coefficient of 0.39, so each extra word per sentence adds 0.39 to the grade. Syllables per word carries 11.8, so each extra tenth of a syllable per word adds 1.18. The constant of −15.59 is a calibration offset with no meaning of its own; it exists to place the output on the school-grade scale.
Those coefficients look lopsided, but the two variables move on very different scales. Average sentence length ranges realistically from about 8 to 35 words, a spread of 27, contributing up to 10.5 grade levels. Syllables per word ranges from about 1.2 to 1.9, a spread of 0.7, contributing about 8.3. In practice the two levers are comparable, with sentence length usually the one you can move furthest and fastest.
This matters when you are editing to a target. Cutting a document's average sentence from 40 words to 20 drops the grade level by 0.39 × 20 = 7.8 — a very large effect. Replacing "utilise" with "use" saves one syllable on one word, which across a 300-word sample changes syllables per word by 0.0033 and the grade by 0.04. Vocabulary edits matter in aggregate, over hundreds of substitutions; sentence surgery matters immediately.
The target inversion works the same algebra backwards. Fix the syllable density you have and solve for the sentence length that produces your target grade: ASL = (target + 15.59 − 11.8·ASW) ÷ 0.39. If that comes out zero or negative, the target is unreachable at your current vocabulary — the words alone put the text above it, and only shorter words will help.
Worked example: a 300-word passage
You take a 300-word sample from a report. Counting carefully, you find 20 sentences and 450 syllables, and your target is grade 8.
- Average sentence length. 300 ÷ 20 = 15.00 words per sentence.
- Syllables per word. 450 ÷ 300 = 1.500 syllables per word.
- Sentence-length term. 0.39 × 15.00 = 5.85.
- Word-length term. 11.8 × 1.500 = 17.70.
- Grade level. 5.85 + 17.70 − 15.59 = 7.96. The passage is at eighth-grade level.
- Reading ease, for comparison. 206.835 − 1.015 × 15.00 − 84.6 × 1.500 = 206.835 − 15.225 − 126.90 = 64.71, which sits in the plain-English band.
- Target check. For grade 8 exactly: (8 + 15.59 − 17.70) ÷ 0.39 = 5.89 ÷ 0.39 = 15.10 words per sentence. You are at 15.00, so you already meet the target with a fraction of a word to spare — 300 ÷ 15.10 = 19.9 sentences.
Now suppose the same 300 words were arranged in 12 sentences instead of 20. Average sentence length becomes 25.0, the first term becomes 9.75, and the grade level rises to 9.75 + 17.70 − 15.59 = 11.86. Identical vocabulary, identical content, nearly four grade levels harder — purely from punctuation.
What grade level to aim for
Match the level to the audience, not to a universal ideal. Public health and consumer material is conventionally targeted at grade 6 to 8, because that range is reachable by most adults reading under stress or in a second language. General newspaper prose sits around grade 9 to 11. Undergraduate textbooks run 11 to 14. Peer-reviewed papers and legal drafting commonly land at 15 or above, and that is appropriate when the readers are specialists reading for precision.
Read the score against a specific document, not against the abstract idea of good writing. The most useful workflow is comparative: measure the draft, edit, measure again. A drop of two or three grade levels between drafts is strong evidence that the revision worked. A single absolute number in isolation tells you far less, particularly since sampling different 300-word passages from the same document routinely varies by a grade either way.
Know what the formula cannot see, because that is where writing goes wrong in ways no score detects. It cannot see whether a short word is a term of art the reader has never met — "the writ shall lie" is monosyllabic and incomprehensible. It cannot see passive voice, buried subjects, missing antecedents, or paragraphs organised backwards. And it can be gamed: chopping every sentence at the nearest comma lowers the grade and can make prose harder to follow, because the connective logic disappears with the conjunctions.
Use the companion Flesch Reading Ease score as a cross-check, since it uses the same two inputs with different weights and is the figure written into several US state plain-language statutes. The Flesch Reading Ease calculator covers that scale and its bands in full. If you are matching a text to a specific reader rather than a grade band, measuring that reader's actual fluency with the reading speed calculator tells you more than any formula.
Grade level for combinations of sentence and word length
| Words per sentence | 1.3 syll/word | 1.5 syll/word | 1.7 syll/word | 1.9 syll/word |
|---|---|---|---|---|
| 10 | 3.65 | 6.01 | 8.37 | 10.73 |
| 12 | 4.43 | 6.79 | 9.15 | 11.51 |
| 15 | 5.60 | 7.96 | 10.32 | 12.68 |
| 18 | 6.77 | 9.13 | 11.49 | 13.85 |
| 20 | 7.55 | 9.91 | 12.27 | 14.63 |
| 25 | 9.50 | 11.86 | 14.22 | 16.58 |
| 30 | 11.45 | 13.81 | 16.17 | 18.53 |
| 35 | 13.40 | 15.76 | 18.12 | 20.48 |
Each step of 0.2 syllables per word adds 2.36 grade levels; each five extra words per sentence adds 1.95. That is why the two levers are worth roughly the same over their realistic ranges.
Counting and interpretation mistakes
- Counting vowel letters instead of vowel sounds. Syllables are spoken units: 'chocolate' is three, 'business' is two, 'fire' is one for most speakers. Software counters use heuristics and disagree by two or three per cent, which shifts the grade by about a tenth.
- Sampling too little text. Under 100 words, one long sentence moves the result by a full grade. Use 200 to 300 words, and for a long document sample three separate passages and average.
- Including headings, tables and references. A heading is not a sentence and a citation string is not prose. Both distort the averages badly. Measure body text only.
- Treating the grade as precise. The formula is a regression fitted on 1970s Navy training material. Reporting 8.37 rather than 'about grade 8' claims accuracy the underlying study does not support.
- Chopping sentences to game the score. Removing conjunctions lowers average sentence length and can make prose harder to follow by deleting the logical connectives that showed how the clauses relate.
- Assuming a short word is an easy word. Terms of art, acronyms and abbreviations score as easy and read as opaque. Readability formulas cannot see vocabulary familiarity at all.
- Comparing scores across formulas. Flesch-Kincaid, SMOG, Gunning Fog and Dale-Chall use different inputs and different scales, and routinely differ by two or three grades on the same text. Pick one and stay with it.
The other readability formulas, and when to use them
Flesch-Kincaid grew directly out of Rudolf Flesch's 1948 Reading Ease score. Kincaid and colleagues refitted the same two variables against comprehension data from Navy enlisted personnel in 1975 and rescaled the output to school grades, which is why the two formulas share their inputs but not their coefficients. Reading Ease returns 0 to 100 with higher meaning easier; Flesch-Kincaid returns a grade with higher meaning harder. Both are computed here so you can see them move together.
Several alternatives use different signals. Gunning Fog counts words of three or more syllables rather than averaging syllables. SMOG counts polysyllabic words in a fixed 30-sentence sample and is widely preferred in health communication because it was calibrated for full comprehension rather than 50 per cent comprehension. Dale-Chall abandons syllables entirely and instead counts words absent from a list of about 3,000 that fourth graders reliably know, which makes it the only common formula that models vocabulary familiarity rather than word length. Where the stakes are high — patient consent forms, medication instructions — running two formulas and reconciling them is better practice than trusting one.
For everyday drafting, the practical pairing is this one and the Flesch Reading Ease calculator, since several US states specify a minimum Reading Ease score for insurance policies while federal plain-language guidance tends to speak in grade levels. If you are estimating how long the finished document takes to consume rather than how hard it is, use the reading time calculator; if you are converting an assignment brief written in pages into a word budget you can measure, the words to pages calculator does that conversion. And when the text is to be spoken rather than read, complexity tolerances are lower still, because a listener cannot go back — plan that with the speech time calculator.
