A law professor spent two years testing how an AI ban, unguided AI suggestions, and structured training affect student performance. The group without AI finished last both years. "I was wrong," the researcher writes. He had assumed AI without guidance would do more harm than good.
Thibault Schrepel of Vrije Universiteit Amsterdam randomly split students in his "Law of AI" course into three groups. The task was the same for everyone: working in small teams of four or five, they had 20 minutes to improve a provision of the EU AI Act. Grading covered substance, clarity, proportionality, and innovation.
The first group couldn't use ChatGPT. The second received ChatGPT-generated revision suggestions embedded directly in the text and could keep using the tool but got no guidance on how. The third group got hands-on training in legal prompt engineering and checking AI suggestions for consistency and accuracy.
All students took the same exams: a multiple-choice test and a take-home exam that required them to revise another AI Act provision. The experiment ran in 2024 with 66 students and was repeated in 2025 with 164 participants.
The group without AI ran out of ideas fast
The no-AI group mostly made minor wording tweaks, rephrasing for clarity and cutting redundancies. Only a few subgroups tried substantive legal improvements. After ten to fifteen minutes, many subgroups regularly ran dry, an effect Schrepel calls "idea exhaustion." Going without AI did force deeper discussion among group members, though.
The second group accepted AI suggestions largely without question, often reasoning that they "sounded better," Schrepel writes. Some students replaced the terms "shall" and "individual" with AI-suggested alternatives without showing they understood the legal implications. Every subgroup kept at least one misleading or legally extraneous term from ChatGPT's output.
Only the trained third group engaged in genuine back-and-forth with the AI, testing different phrasings and digging into the substantive questions behind the assignment.
The training advantage nearly vanished by year two
The trained group scored well above the other two in 2024, especially on the more demanding take-home exam. A year later, that gap had almost closed. All three groups performed at roughly the same level in 2025.
Schrepel attributes this to growing chatbot familiarity. Many students already use these tools in daily life, so formal training delivered less of a boost in 2025 than it did a year earlier. Ethical use and legal responsibility still need to be taught, he adds.
One finding held constant across both years: the no-AI group finished last. Banning AI produces worse average outcomes than allowing it, Schrepel concludes. Whether the advantage comes from structured training or just hands-on experience remains an open question.
The author's own assumption turned out wrong
The second group's performance surprised him, Schrepel writes. He expected the uncritical errors from the classroom exercise to carry over into the exam. The opposite happened. The errors didn't repeat, and the group actually scored slightly higher than the AI-free group. The most likely explanation, he says, is that students learn to spot AI weaknesses through their own use, especially when accuracy has real consequences.
That forced Schrepel to abandon his starting assumption. He had been convinced AI only helps with structured training and should otherwise stay out of the classroom. "I was wrong," he writes. Skipping AI instruction is a missed opportunity, but it doesn't cause the educational collapse some had predicted.
Universities should rethink bans
Department leaders should resist blanket AI bans and give instructors room to experiment, Schrepel says. Universities also need to invest in AI skills for faculty, many of whom feel poorly equipped.
The traditional master's thesis needs rethinking, too. Its core value, the long struggle with research and argumentation, can now be handled in minutes with AI tools. Pure literature reviews should no longer count, Schrepel argues, and programs should require practice-oriented or empirical work where thoughtful AI use is part of the grade.
Not all universities are moving in this direction. UC Berkeley Law has banned AI from nearly all graded work, arguing that future lawyers need to build core thinking skills before using AI in any useful way. That position directly conflicts with Schrepel's findings, which suggest keeping AI out of the classroom leaves students worse off regardless of the rationale.
Schrepel acknowledges his study's limits: small sample size, students enrolled in an AI course who were likely more tech-savvy than average, and no way to verify how much AI students actually used on the take-home exam.
That last point matters because other research finds the strongest AI effects on unsupervised work. A UC Berkeley study covering more than 500,000 grades found that the share of A grades in writing- and programming-heavy courses jumped 13 percentage points after ChatGPT launched. Proctored exams showed no comparable shift. A longitudinal study from central China tracking more than 26,000 K-12 students told a similar story from a different angle: AI boosted homework grades while exam performance dropped by up to 24 percent. The damage was greatest where AI replaced independent thinking rather than supporting it.