Sometime in late September, nearly every instructor meets the same moment. A submission arrives that is fluent, organized, grammatically clean—and strangely weightless. The question that rises first is the obvious one: did the student write this? It is also, I want to suggest, the wrong first question. After more than forty years of teaching college writing, nineteen of them entirely online, I have come to believe that the arrival of generative AI has not created a new problem so much as exposed an old one. We built our grading practices on the assumption that a finished text reliably tells us something about the student who submitted it. That assumption was always shakier than we admitted. Now it has collapsed, and no software will rebuild it for us.
The detection tools certainly will not. Independent researchers have repeatedly found that AI detectors are inconsistent across text types, lag behind each new generation of models, and flag the prose of multilingual writers at troubling rates. An accusation you cannot substantiate does lasting damage to the trust a course depends on, and a false one can follow a student for years. The productive move is different: instead of asking whether a text is trustworthy after it arrives, ask what would make it trustworthy before it is assigned. That reframing yields four practical questions, each answerable in the weeks before the term begins.
1. Can I See the Thinking, or Only the Text?
If the final draft is the only thing a student submits, the final draft carries an evidentiary load it can no longer bear. The remedy is to let the assignment generate its own evidence: a short proposal, a rough draft with visible revision, a paragraph of reflection on what changed and why. None of this is surveillance, and none of it requires new technology. A record of thinking is something a student can take pride in and an instructor can actually respond to—which is more than can be said for a verdict from a detector. When the process is visible, the question of authorship largely answers itself.
2. Does My Rubric Still Measure What is Scarce?
Most rubrics in circulation were written for a world in which fluent, well-organized prose was hard to produce and therefore worth rewarding. That world ended. Fluency is now free, and a rubric that awards most of its points to organization, mechanics, and polish is a rubric an algorithm can satisfy. Reweight it toward what remains scarce: a claim the student is willing to defend, a choice explained, a revision that genuinely answers feedback, an interpretation the student can restate in conversation. If your criteria describe qualities a machine produces effortlessly, the machine will earn the grade.
3. What Does My AI Policy Ask Students to Actually Do?
A checkbox—”I did / did not use AI”—teaches compliance and nothing else. Consider asking instead for a brief record of decisions: what the student asked, what came back, what they kept, what they rejected, and the reasoning behind each choice. Framed this way, disclosure stops being a confession and becomes an artifact of judgment—often the most revealing page in the whole submission. Students who must explain why they overruled a machine’s suggestion are practicing exactly the discernment we claim to teach.
4. Who Decides the Hard Cases?
There will be hard cases; there always were. The question is who resolves them. Not a percentage score from a detector, and not a solitary instructor at midnight, second-guessing an intuition. The oldest tool in the profession remains the best one: a colleague. Trade two borderline samples with someone who teaches the same course and talk through them aloud. Judgment that can be explained and defended to a peer is judgment you can stand behind with a student, a chair, or an appeals committee. Judgment that cannot be explained is only a mood.
The exchange need not be elaborate. Here is a version that fits inside thirty minutes: each of you brings one submission you would hesitate to grade alone. Read your colleague’s sample cold, without their comments, and write a one-sentence verdict with one reason. Then compare. Where the two verdicts agree, you have calibration; where they diverge, you have found the exact criterion that needs sharpening—usually a rubric line that means two different things to two experienced readers. Do this twice a term and the hard cases stop feeling like accusations to adjudicate and start feeling like what they are: professional questions with professional answers.
Notice what these four questions do not require: no new platform, no institutional policy, no budget line, no software subscription. They require assignments designed to produce their own evidence, criteria aligned with what still matters, and a firm decision to keep the verdict human. Even at scale—and I have graded at scale for a long time—a brief, genuine human response to each student is worth more than any score an algorithm returns, because it is the only part of the assessment a student remembers.
Syllabi are still open for a few more weeks. The best time to answer these questions is before the first essay arrives.
Roger Ochse, EdD, is Professor and Honors Director Emeritus at Black Hills State University. He has taught college writing for more than forty years, nineteen of them fully online, and is developing the Green Pen Framework, an approach to writing assessment that treats trustworthy texts as something assignments produce rather than assume.