Comprendo / insights Calibrate

The question

How well can a language model write to a reading level, out of the box?

Reading apps and classroom tools ask models to write “for a 5th grader” or “at a 9th-grade level.” That only works if a model can control the reading difficulty. We measured how well four of them do, with no special prompting or setup.

July 7, 2026

144short passages generated and scored
25%landed in the exact grade band requested
−0.69average miss, in grade bands (− = easier than asked)
0%of 11th-grade requests reached an 11th-grade level

The result

Models track the target for a while, then they flatten out

Asked for increasing reading level, the models climb partway up the scale and then settle into a middle-grade register, no matter how high we aim. The line below is what they produced against the line they were asked to hit.

Reading level requested vs. produced Requested (perfect control) Produced (average of 4 models) Easier than asked Harder than asked
K–1 2–3 4–5 6–8 9–10 11–CCR the gap widens the higher you aim too hard for the youngest 1 3 5 7 9 11 Grade level requested
The bold line is the level the models actually produced, averaged across all four; the dashed line is what a perfect match to the request would look like. For the youngest grades the models overshoot slightly; their writing is a touch harder than asked. From grade 5 up they fall steadily behind, and an 11th-grade request lands around a 4–5 / 6–8 reading level. None of the 24 passages requested at grade 11 reached an 11th-grade level.

The models’ output is a far shallower line than the request; it crosses the target near grade 3 and never recovers. The grade-5, grade-9 and grade-11 passages are nearly interchangeable. The extra difficulty you ask for past the middle grades mostly doesn’t show up in the text.

Miss by requested grade
Grade requested1357911
Average miss (bands)+0.92+0.25−0.58−1.04−1.46−2.25
Hit the exact band8%75%33%21%12%0%

By model

It barely matters which model you choose

Sorted by accuracy, the four models sit within about half a band of one another. The most accurate model surprisingly is the smallest. Gemma is an open source model, so frontier models didn't translate into better grade control.

Models, ranked by average error
ModelAvg missAvg errorExact bandWithin 2 bands
Gemma 4 31Bsmall · open weights−0.390.8931%56%
Gemini 2.5 Flash−0.781.0625%39%
Claude Sonnet 4.6−0.721.2219%44%
GPT-5.1−0.891.2225%42%

With 36 passages per model the order is soft. The takeaway isn’t the ranking, it’s the scale: the spread between grades (over three bands) dwarfs the spread between models (about half a band).

By topic

What you write about moves the result as much as a whole grade

The same models, asked for the same grades, drifted about a full band easier when writing a folktale than when explaining how volcanoes form. Narrative pulls toward simple storybook language; expository science scales with grade more willingly. Content type is a real variable, not noise.

A co-factor worth ruling out: longer passages scored easier (correlation −0.51), but that is mostly grade in disguise; higher grades produce longer text. The difficulty gap holds after accounting for length.

Miss by topic
TopicAvg missAvg errorExact
Folktale (narrative)−1.121.4019%
Volcanoes (expository)−0.260.7931%

How we measured

One prompt, six target grades, an independent judge

Every passage came from the same template, changing only the grade and topic:

“Write an original reading passage of about 200 words for grade N students on the topic of 

We swept 4 models × 6 grades (1, 3, 5, 7, 9, 11) × 2 topics × 3 passages each. The two topics: a folktale about a clever fox, and how volcanoes form. Every model ran with reasoning turned off, so each is judged on its direct, first-pass output. Models that can’t disable reasoning were left out, to keep the four on even ground.

Each finished passage was then read by a separate model, the grade-level judge, which assigns the reading level the text actually lands at, without knowing the grade we asked for. Judging ran through the Learning Commons Evaluator framework, using its GradeLevelAppropriateness evaluator to place each passage on the six-band scale.

The unit

Reading level, in six bands

The judge places each passage on a six-band scale, from kindergarten through college-ready. We score a passage by how many bands its result sits away from the band we requested.

The evaluator is itself validated against a human-scored reading-complexity gold set (the CLEAR Corpus, developed by CommonLit and Georgia State University), where it reports landing the exact band about 80% of the time, our reference point for trusting its calls.

K–1easiest
2–3
4–5
6–8
9–10
11–CCRhardest

Limitations

What this first pass can and can’t say

The direction is clear and consistent but precision is not. Before leaning on any single number, four things to keep in mind:

  • One judge pass per passage. Some of the misses are the judge’s own variability, not the writer’s. Repeated scoring with a majority vote would separate the two.
  • Three passages per cell. Enough to see the gradient, not enough to firm up the model ranking.
  • A six-band scale. Coarse, and it can’t go below K–1 or above 11–CCR, so the ends of the scale can only miss in one direction.
  • Two topics, four models. A probe, not a survey.