Phosphor AI Tutor: Achieving 0.71-1.30 SD Effect Size in Dartmouth Statistics Course
The integration of LLM-powered formative assessment directly into course readings can significantly improve student outcomes while maintaining high engagement. In a pilot deployment at Dartmouth College, the digital learning platform Phosphor was associated with an increase in final exam performance between 0.71 standard deviations (SD) when adjusting for prior scores and 1.30 SD unadjusted.
The Core Finding: Active Generation Drives Learning
The primary driver of academic improvement in the Phosphor platform is the use of constructed-response questions (CRQ) rather than multiple-choice questions (MCQ). Data from the deployment indicates that lesson-level engagement only translated into measurable exam performance gains when students were required to actively generate answers.
- CRQ Efficacy: In modules where lesson quizzes included CRQs graded by Claude Sonnet 4.6, each additional lesson completion was associated with a significant increase in midterm scores.
- MCQ Limitations: In Module 2, where quizzes were reduced to MCQ-only based on student feedback, the positive correlation between lesson completion and performance vanished for students who had completed at least one lesson.
- The "Doer Effect": These results align with the "doer effect," where completing practice questions integrated into readings provides a substantially higher learning impact than passive reading alone.
Platform Design and Implementation
Phosphor was deployed as an optional, ungraded alternative to traditional readings for 151 students in an Introductory Statistics course (MATH 010). The platform consists of three primary components:
1. LLM-Graded Lesson Quizzes
Each lesson includes a bank of 15–20 exercises. Students take a quiz of four random questions. While MCQs are auto-graded, CRQs are graded by an LLM using instructor-defined rubrics. The model provides a correctness judgment and a detailed explanation, requiring a 75% score to "pass" the lesson.
2. Cumulative Module Reviews
These reviews cover all lessons within a module and require a 90% pass threshold. Students who passed all three Module Reviews showed the strongest performance gain, scoring 7.1 points higher on the final exam (d = 0.66).
3. RAG-Based Chat Assistant
A retrieval-augmented generation (RAG) sidebar allows students to query course-specific content. However, this feature saw minimal usage (72 total queries), with students reporting that general-purpose LLMs were faster or that the provided reading material was already sufficient.
Engagement and Compliance Metrics
Phosphor achieved engagement rates that far exceed typical undergraduate reading compliance (which the author noted as being as low as 10–15% for this course).
- Adoption: 90.2% of enrolled students engaged with the platform at least once.
- Reading Compliance: Total reading compliance (measured by lesson views and quiz completions) fell between 48% and 76%.
- Student Sentiment: In-class surveys showed 94% of students found the platform more engaging and 97% felt it improved retention.
Statistical Analysis of Learning Outcomes
To account for the fact that more motivated students are more likely to use the platform, the study used a Tobit model to estimate the gap between "full engagement" (24 lessons, 3 reviews) and "zero engagement":
| Metric | Unadjusted Gap | Adjusted Gap (Control for Midterms) |
|---|---|---|
| Final Exam Score | 1.30 SD | 0.71 SD |
The author argues that 0.71 SD is a conservative lower bound, as controlling for midterm performance may absorb learning that Phosphor had already facilitated earlier in the term.
Critical Perspectives and Limitations
While the results are promising, the study is observational and lacks randomized controls. Community discussion and the author's own limitations section highlight several caveats:
- Selection Bias: Students who engage fully with optional materials are typically more motivated and higher-performing, which may inflate the effect size.
- Novelty Effect: Some critics suggest the high initial engagement (90%) could be attributed to the Hawthorne effect or the novelty of using a new AI tool.
- Content Overlap: There is a question of whether the final exam questions overlapped directly with the Phosphor materials, which would measure reading compliance rather than general learning.
- Tutor vs. Grader: Some observers noted that Phosphor functions more as a "practice quiz platform with an AI autograder" than a traditional "AI tutor," especially given the underutilization of the RAG chat feature.
"The conclusion is essentially that people who do practice quizzes will do better on exams." — Community feedback via Hacker News
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch