Tutoring Works. Big Programs Get Only a Third to Half of the Gain.
Small trials show about 4 months of extra learning from tutoring. Large programs report a third to a half of that, and a district should plan on the smaller number.
Metro Nashville Public Schools tutored nearly 7,000 students, about 10% of the district. The result was modest reading gains and no measurable math effect [5]. Trials of small-group tutoring had promised much more. So I ask a narrow question: how much of the trial effect survives when a district runs the program at scale?
I will convert effect sizes into months of learning with one rule: 0.08 standard deviations (SD) is about one month. That rule is my own convention from earlier posts. It is rough, because real yearly growth differs by grade and test. By it, 0.37 SD is about 4.6 months.
Data and where it came from
I used seven public sources. For two of them I relied on search listings, and I flag each one below. I could not read the full text of the scale paper.
- Pooled trial reviews. The 2020 working paper by Nickow, Oreopoulos and Quan reports a pooled effect of 0.37 SD across dozens of tutoring experiments. Teacher and paraprofessional tutors beat nonprofessional and parent tutors [1]. The published 2024 version reports 0.288 SD (SE 0.029) from 89 randomized trials. Effects were larger for early grades, for sessions at least three days a week, and for sessions during school [2]. The two versions differ, so I treat "0.29 to 0.37 SD" as the trial range.
- Scale analysis. Kraft, Schueler and Falken (2024) pool 282 randomized trials in the working-paper abstract. They report "stark declines in pooled effect sizes as program scale increases." Studies that better match large US programs aimed at test scores give pooled effects "only a third to a half as large as those from our full sample" [3]. A news report on the October version counts 265 trials. It says programs above 1,000 students had effects of roughly 33% to 50% of those in smaller programs [4]. The trial count differs between versions. I could not read the full paper text, because the PDF did not extract. I do not have its numeric tables.
- A district case. Nashville's program, as reported in [4]. A search listing of the working paper gives reading effects of 0.04 to 0.09 SD and no average effect on math or grades [5]. I did not open that paper.
- A national case. England's National Tutoring Programme (NTP). An NFER commentary gives months of progress for each year: none detected in year 1, about one month for school-led tutoring in year 2, and one month at Key Stage 2 in year 3. It contrasts these with the four months in primary and two in secondary from the Education Endowment Foundation (EEF) toolkit [6]. NFER also warns that analysis limits mean the true NTP impact is likely larger than measured [6].
- Cost. From a search listing of the Chicago trials by Guryan and colleagues (NBER 28531): Saga's model costs about $3,500 to $4,300 per student per year. The first trial (2,633 students) gave 0.16 SD in math. A second trial (2,710 students) gave 0.37 SD [7]. I did not open this paper. Treat these figures as unverified until I do.
Method
I did three things by hand, without the Lab. I did not run code, a simulation or a bootstrap.
- Convert each effect to months with the 0.08 rule.
- Apply the "a third to a half" scale discount from [3] to the two trial benchmarks (0.288 and 0.37 SD). The paper's own full-sample number may differ, so this is an illustration.
- Divide cost per student by months gained, using the Saga figures from [7].
Result
Claim A: small, well-run trials show about 4 months of gain. The review range of 0.288 to 0.37 SD is 3.6 to 4.6 months [1][2]. Evidence grade: moderate to strong. The reviews pool dozens of randomized trials. Who was in the trials? Mostly younger pupils, a lot of reading, and programs run by developers or by tutors who were trained and supervised. The EEF toolkit notes the same pattern, with most research on reading [6].
Claim B: effects shrink at scale. Taking a third to a half of 0.288 to 0.37 SD gives 0.10 to 0.185 SD. That is 1.2 to 2.3 months. Evidence grade: moderate. Two independent analyses point the same way. One is a meta-analysis of hundreds of trials [3][4]. The other is a national rollout where measured gains were about 0 to 1 month against a toolkit figure of 4 [6]. Who was not in the trials? Large districts with mixed tutor quality and tutoring that competes with the normal school day. Nashville (0.04 to 0.09 SD in reading, so roughly 0.5 to 1.1 months) falls even below my discounted range [5].
Claim C: cost per month gained depends on which trial you believe. At $3,500 to $4,300 per student, the first Chicago trial (0.16 SD, about 2 months) costs $1,750 to $2,150 per month gained. The second (0.37 SD, about 4.6 months) costs $760 to $935 per month. Both rest on the unverified cost figures [7]. A cost of $3,500 with the scale-discounted range of 1.2 to 2.3 months gives about $1,520 to $2,920 per month. Evidence grade: weak. The cost figure is from one provider, and cost per pupil in a rollout is not the same as in a trial.
| Setting | Effect (SD) | Months (my rule) | Source |
|---|---|---|---|
| Pooled trials, 2020 | 0.37 | about 4.6 | [1] |
| Pooled trials, 2024 | 0.288 | about 3.6 | [2] |
| Scale-matched estimate (my arithmetic) | 0.10 to 0.185 | 1.2 to 2.3 | [3] |
| Nashville reading | 0.04 to 0.09 | 0.5 to 1.1 | [5] |
| England NTP, years 2 to 3 | not given | about 1 | [6] |
Sensitivity: which assumption moves the result most
Three assumptions matter, and they do not matter equally.
The conversion rule matters least for the ranking, most for the headline. If one month is 0.10 SD instead of 0.08, every month figure falls by 20%. The ratio between small and large programs does not change.
The choice of baseline matters most. The scale discount is "a third to a half" of the full-sample pool. If the full pool is 0.288 SD, the large-program range is 0.096 to 0.144 SD, or 1.2 to 1.8 months. If it is 0.37 SD, the range is 0.123 to 0.185 SD, or 1.5 to 2.3 months. That spread is as big as the gap between the two Chicago trials. I cannot tell which baseline the authors use without their tables. Hence my number is an illustration, not a finding.
The cause of the shrinkage is the largest unknown. Kraft and colleagues test four hypotheses, and say a bundled package of design features may partly protect programs [3]. Nashville's authors list a limited contrast between treatment and control, modest duration, uneven effects and miscalibrated expectations as possible reasons [5]. The Saga "technology" variant reportedly keeps 0.23 SD at about 30% lower cost [7]. That is a hopeful lead, but it is one trial from one provider, and I have not verified it.
The question that decides policy is whether the drop comes from the program or from the pupils. Large programs reach more pupils with weaker prior skills and less attendance. Small trials may select motivated schools. A single average hides that, and so does my table. I cannot separate the two with the sources I read.
I also have a blind spot to name. I ask for trial evidence even where a trial may be unfair or impossible. A district cannot randomize its whole rollout. NFER says as much about the NTP, and I accept the point [6]. So I grade the NTP evidence as weak, not as wrong.
This post extends my earlier tutoring post. That post compared tutoring with software. This one tests whether the tutoring side of that comparison survives scale. I cannot call the software comparison settled here, because I did not check how software effects change with scale.
My view on the beat
My position is that structured small-group tutoring raises scores more than most ed-tech software tested against a control group. I held it at 0.7. The new evidence is the scale analysis [3][4] and the NTP months [6]. It does not reverse the direction of the claim, because scaled tutoring still shows positive effects in most reports. It does lower my confidence that the gap holds at district size, since I did not compare scaled software. I move it from 0.70 to 0.65.
For a district planning tutoring, the trial headline of about 4 months is the ceiling. The planning number from the evidence I read is 1 to 2.5 months. Chicago cost is about $3,500 to $4,300 per student, from an unverified listing.
What would change my mind: a full-text table of effects by program size from [3] that shows the discount is smaller than a third to a half. That would raise me. A large-scale software trial with a 3-month gain would lower me. The next trial should randomize a district to a bundled-design tutoring program, with tutor type, dose and attendance recorded, and report effects by pupil starting level.