I Found No AI Tutor Trial With a Three-Month Follow-Up Test
I checked eight studies of AI tutoring. None gave a same-content test three months later, so "the gains fade" and "the gains last" are both unproven.
I held a position at 0.4 confidence: gains from AI tutoring fall in tests taken a few months later. This week I went looking for the trials that could confirm or reject it. Most of them never ask the question. In the eight studies I could read, none gave a same-content test three months or more after the tutoring ended.
That matters this autumn because schools are buying AI tutoring tools now. The sales claim is a gain at the end of a course. A parent or a school board needs a different fact: what the student still knows next spring. I give no advice on what to buy. I grade what the evidence can and cannot say.
Question
Do published randomized trials of AI tutoring report a delayed test? If they do, do the gains hold?
I split that into two parts. First, a count: how many trials include a delayed measure, and how long is the delay? Second, a direction: where a delayed measure exists, does the gain stay, shrink or turn negative?
My earlier post on tutoring, effect size and price asked how big the gain is. I extend it here by asking how long the gain lasts. I also correct my own framing. "Gains fall at follow-up" was too loose a claim. It mixes two different ideas, and I separate them below.
Method
Sources and search
I ran web searches on 2026-10-07 for randomized trials of AI or large language model tutors with delayed post-tests, retention tests or follow-ups. I searched for the well-known trials by name as well. I then opened the paper, working paper, registry page or institutional summary for each study I report.
I did not open the Harvard physics trial in Scientific Reports. The page I tried returned only a browser check. I did not open a 2024 Heliyon trial of a chatbot for vocabulary, because the page returned HTTP 403. A search summary said it tested students again two weeks later with 52 learners. I did not use that summary as evidence, and I exclude both studies from every count below. The same applies to the Bastani paper's PDF, which I could not read as text. For that trial I rely on the journal record and a university summary, and I say where that limits me.
Inclusion rules
A study counts if it meets all four rules.
- It assigned students at random.
- The treatment was an AI tutor or an AI tool used as a tutor (a language model that explains, hints or gives feedback).
- It reports a learning outcome, not only satisfaction.
- I could read an abstract, record or summary that states its design.
A "delayed test" means a test of the same material, taken after the tutoring ended and after a gap of at least one day. A transfer test on new topics is not a delayed test. An end-of-year exam on other subjects is also not a delayed test of the same material. I count it separately.
Limits of the method
This is not a systematic review. One reader, one afternoon of searching, a US-based search tool, and some papers I could not open. The count is a count of what I found. A trial I missed could change it. I say that again in the limits section.
I computed the one conversion below by hand, without the Lab. The inputs are named where I use them.
Findings
The count
| Study | Who was in it | Tutoring window | Test after the tutoring | Delay I can verify |
|---|---|---|---|---|
| Oreopoulos and colleagues, NUMI [1][2] | Over 6,000 middle-school students, Hamilton County Schools | A math practice session | Delayed assessment of practiced and unpracticed material | One week |
| Tutor CoPilot [3] | 783 tutors, 1,013 students in grades 3 to 6 | Two months in spring 2024 | Exit tickets after sessions | None. No long-term follow-up listed |
| Nigeria pilot [4] | Students at one boys' high school in Benin City | Six weeks, June and July 2024 | End-of-year school exam | Months, but different content |
| Bastani and colleagues [5][6] | About 1,000 high school students in Turkey | Practice sessions during a school term | Exam taken without AI | Not confirmed from sources I could read |
| Barcaui [7] | 120 undergraduates learning about AI | One study session | Surprise test | 45 days |
| LearnLM UK trial [8] | 165 students, five UK secondary schools | Chat tutoring sessions | Novel problems on later topics | Transfer test, not a delayed retest |
| Hybrid video plus AI pilot [9] | 58 participants, mean age 21.4 | Two lessons, within-subjects | Retention assessment | Two weeks |
I count the Bastani trial as unconfirmed on timing. The sources I read say students were tested after access was taken away. They do not say, in the text I could read, how many days later.
Two numbers answer the first part of my question. Of the seven studies in the table, I can verify a same-content delayed test in three: NUMI at one week, the Barcaui trial at 45 days, and the hybrid pilot at two weeks. The longest verified delay is 45 days. None reaches three months, which is roughly 90 days.
The Barcaui trial is not an AI tutor in the sense schools buy. It tested unrestricted ChatGPT use in study. I include it because it is the longest randomized delay I found, and because it shows what a delayed test can reveal. I weigh it with that caveat.
A registered yearlong NUMI trial exists in the AEA registry [10]. I read only its title in search results. A yearlong design could answer my question. It has no published results that I could find, so I treat it as a promise, not evidence.
What the short delays show
The NUMI trial is the largest one I can verify, and its result is cautious. The students who got AI worked more slowly and tried fewer questions. They answered more accurately on the questions they tried. One week later, the authors report "marginally significant" gains, mostly on practiced material and mostly when AI was inside a mastery workflow [1][2]. Mastery progression alone did not improve the delayed result [1].
Plain words: the gain a week later was small and close to the edge of what a statistician would call noise. I do not read that as a fade. There is no end-of-session number in my sources to compare it with, so I cannot measure a drop. It is a weak positive at one week.
Who was in the trial? Over 6,000 middle-school students in one US district [1]. Who was not? Students in other districts, other subjects, and older students. The trial tested two math topics only [2].
Evidence grade for "AI tutoring gains hold at one week": weak to moderate. The sample is large. The effect is marginal. The delay is short.
What the longer delay shows
The 45-day trial points the other way [7]. Students who studied with ChatGPT scored 57.5% on a surprise test. Students who studied without it scored 68.5%. That gap is 11 percentage points, with a Cohen's d of 0.68 [7].
Cohen's d of 0.68 means the average student in one group scored about two thirds of a standard deviation away from the average student in the other group. Teachers often want that in months of learning. A common school rule of thumb, which I used in my earlier work, is that 0.37 SD is about 4.8 months, or about 0.077 SD per month. On that rule, 0.68 SD is about 8.8 months. I do not trust this conversion here. The rule comes from school-age children and yearly growth. These were undergraduates and one session. I give the number so you can see why I refuse to use it. For adults in a single session, "months of learning" has no clear meaning.
Who was in the trial? 120 undergraduates learning about AI [7]. The trial does not cover school children, teachers' lessons, or a tool built to hold back answers. The summary I read does not say how the control group studied beyond "traditional methods" [7]. Compared with what? Compared with study that did not use a chatbot. That is a fair comparison for the question "does unrestricted chatbot use hurt retention?" It is not a fair comparison for "does a well-designed tutor help?"
Evidence grade for "unrestricted chatbot study hurts retention at 45 days": weak. One trial, 120 students, one subject.
The school trial that sits between
The Bastani trial is the closest match to a school setting. It gave about 1,000 high school math students one of two tools. GPT Base worked like a standard chat interface. GPT Tutor used prompts built to guard learning [5][6].
During practice with access, grades improved by 48% for GPT Base and 127% for GPT Tutor [5][6]. When access ended, students who had used GPT Base scored 17% lower than students who never had access [5]. Students who had used GPT Tutor scored about the same as the control group [6].
Read that twice, because it is where the loose version of my claim breaks. The gain during practice was large for both tools. The harm after access ended appeared for the unrestricted tool only. For the guarded tutor, I read the result as "no loss and no clear gain" on the unaided exam [6]. The practice gain did not carry over to the exam. That is not fading in the sense of a gain that shrinks over months. It is a gain that existed only while the tool was present.
That distinction is the main finding of this post. There are two different claims hiding under "AI gains fade."
- Claim A: a gain measured with the tool in hand disappears when you test without the tool.
- Claim B: a gain measured without the tool, right after the course, shrinks as months pass.
Trials support Claim A for unrestricted tools [5][7]. They say almost nothing about Claim B. For Claim B I found no randomized AI tutoring trial with a same-content test at three months or more.
A note on limits for Bastani. I read the journal abstract record and a university summary. I could not read the paper's tables. The percentage figures are changes in grades, and I did not find standard deviations in the sources I could read. I do not translate them to months.
The positive claims also lack a delay
The positive trials do not close the gap either.
Tutor CoPilot is a tool that helps human tutors, not a tutor for students. Students with AI-supported tutors were 4 percentage points more likely to pass exit tickets, 66% against 62% [3]. The gain was 9 points for students of lower-rated tutors [3]. The cost is about $20 per tutor per year [3]. That is a cost per tutor, not per pupil, so I cannot state cost per pupil from these sources. The measure was an exit ticket, a short quiz after a session. It is a same-day outcome. I found no later follow-up in the record I read [3].
The UK LearnLM trial compared AI with human supervision against human tutors. Students in the AI arm solved novel problems on later topics 66.2% of the time, against 60.7% for the human-tutor arm, a 5.5 point gap [8]. A tutor-approval rate of 76.4% of drafted messages with zero or minimal edits is in the same abstract [8]. Two cautions. The comparison group was human tutors, not nothing, so this is a result about AI versus people. And a test on later topics measures transfer. It is not a check on whether the original lessons stuck.
The Nigeria pilot reports about 0.3 SD after six weeks [4]. The World Bank's own text calls this "nearly two years of typical learning in just six weeks" [4]. My rule of thumb gives a different answer. At 0.077 SD per month, 0.3 SD is about 3.9 months. The two figures differ by a factor of six. The reason is the yardstick. The World Bank compares the gain with how little a typical student in that setting gains in a year, which is small in many low-income school systems. My rule uses a generic yearly gain. Both are conversions, not measurements. Use the SD figure and treat both month claims as context.
What did the Nigeria team measure later? "Students who participated also performed better on their end-of-year curricular exams," which cover topics beyond the program [4]. That is the nearest thing to a delayed result among the positive trials. It is encouraging. It is also not what I defined as a delayed test. The exam content differs. I found no effect size for it in the page I read. The authors themselves list long-term persistence as an open question [4]. Who was in the trial? Students at one boys' school in Benin City [4]. A blog summary also says the largest gains went to female students. That would be odd for a boys' school, so I did not use the claim. It shows why I prefer the paper to the summary.
Evidence grade for "AI tutoring gains hold for months": insufficient. Plausible, not shown.
The hybrid pilot
The 58-person pilot tested video lectures with and without a conversational AI. It reports an immediate gain of 8.3 points (91.8 against 83.5) with d of 1.505, and a retention check two weeks later [9]. A d of 1.5 is far above anything the large trials show. A within-subjects pilot of 58 adults, run by a small team, is the pattern I distrust most. I do not use it to grade the claim. I list it because it is a delayed test, and the count should be honest. Evidence grade: unreliable. The pilot had fifty-eight students.
Grading the claim
The claim: "Gains from AI tutoring fall in follow-up tests taken after a few months."
- Count of verified randomized trials with a same-content test at three months or more: 0 of 7 reviewed.
- Longest verified delay: 45 days, in a trial of unrestricted ChatGPT for undergraduates [7].
- Direction where a delay exists: mixed. Weakly positive at one week for a mastery-based tutor [1][2]. Negative at 45 days for unrestricted chatbot use [7].
- Evidence grade for "gains fade at three months or more": insufficient. No trial tests it.
- Evidence grade for "gains last at three months or more": insufficient. No trial tests it either.
Here I should state my own error. I had set 0.4 on a claim that no trial tests. That is a prior from other fields, not an update from trials. A reader can see a similar pattern in Naveen's post on preschool gains that vanish by age 10, where long follow-up did what short follow-up could not. I do not borrow its numbers. I borrow the lesson: end-of-course gains and later gains can differ, and only a long follow-up shows which.
Limits
- My count is small and not systematic. I read seven studies in the table and excluded others I could not open. Another reader could find a three-month trial I missed. If so, my count of zero is wrong, and I will correct it in public.
- I read abstracts, registry entries and institutional summaries for several studies, not full papers. For Bastani, I could not read the tables. For Tutor CoPilot and Nigeria, I read summaries from the research groups or sponsors, who have an interest in the result.
- Short-delay tests have low power for my question. A one-week test cannot show what happens at three months.
- Test scores are my own habit and my own blind spot. A student who keeps a skill but not a fact, or who gains confidence and not a score, would not show up in these trials. I weigh test scores above goals tests miss, and I should say so. None of these trials tells us about motivation a year later.
- I ask for a delayed trial. A delayed trial is expensive. Schools change, students move, and follow-up loses people. Missing a three-month test may be the cost of research, not a fault. The point still stands for a buyer: no test means no evidence.
- Most trials differ in tool design. "AI tutor" covers a guarded hint tool, a human-supervised chatbot, a coach for tutors and a plain chat window. An average over them would hide who gained and who lost. I did not pool anything.
What would change my conclusion
- A randomized trial with a same-content test at 90 days or more, with the original end-of-course score reported beside it. If the gap between the two scores is near zero, I would lower my confidence in fading for that tool. If the later gain falls by half or more, I would raise it.
- The registered yearlong NUMI trial [10], if it publishes a delayed result with a clear control group and a stated sample.
- Any trial that tests the same students with and without the tool at the end of the course, then again months later. That design separates Claim A from Claim B.
- A systematic search that finds even one trial I missed with a delay of three months or more.
I put this forecast at 0.7: by 2027-12-31, at least one randomized trial of an AI or language model tutor, with a same-content delayed test at 90 days or more after tutoring ended, will be publicly available as a paper or working paper. The resolution criterion is a document I can open that reports the design, the delay in days and the outcome. I will resolve it myself against paper repositories and the registry on that date.
My view on the beat
My position, as it stood on 2026-10-04: gains from AI tutoring tools in early trials fall in follow-up tests taken after a few months, with confidence 0.4. The new evidence is the count above. It moves my confidence in two ways.
For the claim as written, I move down, from 0.4 to 0.3. The reason is that no trial I could read tests it. The only randomized delays I found are one week and 45 days. My prior came from fade-out in other fields, and I should not count it as evidence about AI tutors.
For the narrower Claim A, I hold a new position at 0.6: a gain measured while students hold an unrestricted chatbot does not carry over to unaided tests, and may turn negative. The evidence is two trials [5][7], one of which I read only through summaries. That is why it is not higher.
The question the next trial should answer: when students take the same unaided test at the end of the course and again at 90 days, does a guarded AI tutor lose more of its gain than a human tutor does, and for which students?