False positive rates in AI detection - what is an acceptable threshold

In my role overseeing people and organizational policies, I have been asked to advise our education division on AI detection tool deployment. The question that keeps coming up from stakeholders is: what false positive rate is acceptable?

In most assessment contexts, we have established tolerances. Drug testing has defined thresholds and confirmation protocols. Background checks have appeals processes. But AI detection seems to be deployed without any agreed-upon standard for acceptable error.

If a tool has a 4% false positive rate and you run 10,000 student submissions through it, that is 400 innocent students flagged. Is that acceptable? What about 1%? Even at 1% that is 100 false accusations.

How should institutions think about this? Is there a threshold that is genuinely “good enough” for academic integrity use cases?

This is the right question and it rarely gets asked with the rigor you are applying. In my experience on academic integrity committees, most institutions adopt detection tools without ever defining their acceptable error tolerance.

My position is that no false positive rate is acceptable if the detection score alone triggers disciplinary action. The tool should be a screening mechanism that identifies submissions for human review, not a decision-making tool.

The analogy I use with colleagues: a metal detector at an airport flags items for inspection. No one is arrested because the detector beeped. The same framework should apply to AI detection. Flag, review, decide with human judgment and additional evidence.

Practically speaking, what I have seen work best is a tiered response system:

  • Low confidence flag (below 50%): No action taken, logged for pattern monitoring
  • Medium confidence flag (50-80%): Informal conversation with the student, request for process evidence
  • High confidence flag (above 80%): Formal review involving the student, their advisor, and committee

At no tier does the detection score alone result in a penalty. It always triggers a process, not a verdict.

The institutions getting this wrong are the ones using detection scores as binary pass/fail and skipping the human review step entirely. That approach will eventually produce a lawsuit that reshapes the entire field.

And relatedly, I keep hearing students ask why don’t schools like Grammarly. The answer is nuanced - most schools are fine with grammar tools. The issue is when editing tools normalize writing patterns to the point that detectors mistake polished writing for AI. It is a tool distinction problem, not a cheating problem.

the missing piece in this conversation is the asymmetric cost of errors. a false positive (accusing an innocent student) and a false negative (missing actual AI use) do not carry equal weight.

a false positive can damage a students academic record, mental health, and career prospects. a false negative means someone got away with using AI, which in many contexts is not even a clear-cut ethical violation depending on the assignment parameters.

when the cost of false positives is dramatically higher than the cost of false negatives, you need to set a very aggressive threshold that strongly favors avoiding false positives. most tools are not calibrated this way. they are calibrated to catch as much AI use as possible.

I am coming at this from outside academia but the parallels to my industry are interesting. In real estate we use fraud detection systems that flag suspicious transactions. The false positive rate in our systems is managed through mandatory human review of every flag.

The key insight: the technology is the first filter, not the final decision. Any organization deploying AI detection without a robust human review process is accepting liability for every false positive. And eventually someone will hold them accountable for it.