AI cheating in college has stopped being a story about a few bad actors. At Brown, an economics professor suspected most of his 86-person course: the gap between the take-home exam and the in-person one was wider than anything in his twenty years of teaching it. A day before his story ran in the US press, the university published the first report of its Generative AI committee. This article covers the case and the procedure a university needs for cases like this.

What happened in the Brown economics course

Roberto Serrano has taught Welfare Economics and Social Choice Theory for almost twenty years. The course usually drew up to thirty students. This semester 86 signed up, and the midterm was set as a take-home. The format was chosen out of respect for the students: there had been a shooting on campus in December, and many found it hard to sit in a classroom.

The take-home results came back unusually high, even though Serrano had made the exam harder on purpose. He told the class he suspected large-scale AI use and, with his dean's approval, moved the final to an in-person format. After that, 18 students dropped the course and nine more never came to the final. The gap between the two exam averages was the largest in the history of the course.

Take-home midterm against supervised final in the same course

Serrano cancelled the midterm and gave the final more weight. He also lowered the passing score. Even so, 19 students failed the course.

How the professor checked for AI cheating without a detector

Serrano did a simple thing. He ran the exam questions through ChatGPT and compared the machine's answers with his students' work. They matched in style and in method. One question was easiest to prove with a direct argument, but ChatGPT picked a proof by contradiction, long and clumsy. The same move showed up in many student papers.

The method worked, but it takes too much time to be practical. Manual comparison against a machine answer takes hours, depends on the professor's experience, and stops working on tasks where AI has no recognizable style. In practice, Serrano built by hand what an AI detector does statistically: an estimate of how machine-like a text reads. The difference is that a detector's estimate is repeatable and takes minutes, not a weekend.

Honesty matters here. Serrano himself is skeptical of AI detectors and points to false positives. On the main point he is right: no tool delivers a verdict. A probability score is a reason to talk to the student and to rethink the exam format, which is what happened at Brown. Anyone promising a hundred-percent verdict is the actual problem.

What Brown's Generative AI committee report admits

The Brown committee's report came out on July 7, and it is a rare case of a university discussing the problem in plain words. The committee collected feedback from 105 faculty members and 169 staff. Faculty named two worries above all: damage to students' long-term learning, and a decline in their cognitive capabilities. Cheating came third.

What Brown faculty told the university's AI committee

Brown's numbers match the national picture. The same concern about cheating showed up at the same level in the 2025 nationwide survey by the American Association of Colleges and Universities, which the report cites.

The committee's recommendations are worth any university's time. First: write generative AI into the academic codes and draw "bright lines around misuse." Second, and less expected: "de-emphasize punishment." The report says it outright: "There is no way to check with 100% accuracy whether GenAI has been employed," so betting everything on punishment does not work, and open conversation about the rules has to happen.

We agree with that sentence, and we sell a detector. A check with 100% accuracy does not exist. A vendor who promises one is lying. What a detector gives is a repeatable score with a known error rate. At It's AI, our accuracy data is public. The rest depends on the procedure, not on the tool.

Academic integrity procedure for AI: what to write down

At Brown, four parts of this procedure were missing.

Course AI policy in the syllabus

Most Brown faculty already describe their AI rules in the syllabus and repeat them in the first week of class. This is the baseline. A student has to know in advance what is allowed and what is not. The syllabus also has to say how the work will be checked.

Text checks as routine, not accusation

Serrano started checking after the suspicion appeared. That order creates conflict. A regular AI check of every submission works differently: it applies to all students equally, so no one is accused personally. Plagiarism screening has worked this way for years.

A process for the group case

Brown's committee asked Serrano to file a separate complaint against each student. He refused and called the request ridiculous. Tricia Bertram Gallant, who leads the academic integrity office at UC San Diego, offers two fixes: accept one complaint for a group of students, and let a student admit a violation by email, without a personal hearing. A procedure made for one student does not work for sixty.

Exam format as the final check

The in-person final showed what each student actually knew. No text analysis can do that. A text check and an exam test different things, so a course needs both.

What instructors can do before next semester

If you teach a course with written work, three steps are worth doing before the next semester. Write the AI rules into the syllabus in simple words. Agree with your department on a procedure for group cases before the first case appears. Add a text check to the normal course routine, so the signal comes before the exam and not after a scandal.

Four parts of an academic integrity procedure, and how each one fails

Serrano said it in one sentence: "We cannot choose to become idiots." The midterm grade is a small loss. The bigger one is a generation of students who learn that machine text under their name is normal, at universities that prefer not to notice.

FAQ

How common is AI cheating in college?

Nobody has a reliable prevalence number, and the surveys that exist measure perception rather than counted cases. Faculty surveys report a broad sense that AI use has grown, but they ask instructors what they believe, not what was proven in a hearing. Integrity offices count only the cases that reach them, which undercounts by design: most suspected AI use is settled quietly between the instructor and the student, or never raised at all. Treat any single percentage as a floor rather than a measurement. The number that matters to a department is its own, gathered from its own submissions across a term.

How do professors know if students use AI?

Most instructors work from a mismatch rather than from a single tell. Work that does not match what the student produced earlier in the term is the common trigger. Citations that resolve to nothing are the other. Some instructors run an AI detector across every submission and read the score as one input; others rely on an in-person component that no text analysis can replace. The method that fails is the one applied only after suspicion has already landed on a particular student, because it turns a question about the work into an accusation about the person.

Can a whole class be punished for AI cheating?

No. Academic integrity systems are built on individual findings, and a sanction applied to a class as a group would not survive appeal at most institutions. What can be shared is the procedure rather than the verdict: a single filing covering many students, with one set of evidence behind it. The distinction matters because it is the paperwork, not the principle, that breaks at scale. An instructor facing sixty suspect submissions under a one-student-at-a-time process will usually drop the matter, which is the outcome nobody wants. Departments should agree on the batched route before the first large case arrives.

What should an AI policy in a syllabus include?

A syllabus AI policy needs to answer two questions before the first assignment: what is permitted, and what has to be declared. A third clause, how work will be checked, is the one most policies omit, and its absence is what makes a later AI detector run feel like an ambush. Permission works better per-task than per-course: an instructor may allow AI for brainstorming and forbid it in a graded essay, and the syllabus should say so assignment by assignment. Declaration needs a format rather than a principle, such as a closing line naming the tool and what it did.

Is an AI detector result proof of cheating?

No. An AI detector returns a probability that a text is machine-written, and a probability is evidence rather than a finding. What turns it into a case is the procedure around it: a course policy the student had in advance, and a conversation in which they can account for the work. Detectors also carry a measurable false-positive rate, which is why formal prose from students writing in a second language is the classic misfire. An institution that treats the score as a verdict will eventually sanction someone who did nothing wrong. One that treats it as a prompt for a conversation rarely does.