Less than the interface implies, and something quite specific. The number is a classifier's confidence about writing style. It has no source document, no chain of custody, and a threshold somebody chose. Knowing exactly what it measures makes it a useful prompt to look closer, and a poor basis for a decision on its own.
This page has the two tools worth having open: a reading of an actual submission, and the arithmetic that tells you what a flag means in a class of your size. Both run in your browser. Nothing you paste is uploaded, stored, or sent to a detector.
Detector interfaces present a number the way a thermometer presents a temperature. It is not that kind of number. Here is what is actually being reported, including by the tool on this page.
A 0–100 reading of how machine-made the prose is, composed from four measurements: how densely the 44 known tells occur, how uniform the sentence rhythm is, how much stock lexicon is in play, and how heavily the text hedges. It is a description of style. It is not a probability that a person cheated, and it is not evidence.
A classifier's confidence, produced from writing style alone. The vendor picks the threshold at which confidence becomes a red flag and generally does not publish it. Two detectors run on the same paragraph on the same afternoon can disagree, because they are different models trained on different corpora.
A match against a source. A plagiarism checker can name the document it matched and quote the line; that claim is checkable. A detector has no source document, because there is no source document. The number is an inference about style, all the way down.
One consequence is worth stating plainly, because it decides how much weight a flag can carry. A detector reads the finished text and nothing else. Two identical paragraphs, one typed by a student over four evenings and one generated in a second, produce the same score, because the score is computed from the words alone. How this site composes its score →
This is the part that survives every argument about which detector is best, because it is arithmetic rather than a claim about any product. Most of your class did the work. The false positives are drawn from that large group and the true positives from a small one, so even a modest error rate can put more honest students in the flagged pile than dishonest ones.
Set the three rates yourself and watch the result move. The share who used AI is your estimate of your own class. The other two belong to the detector, and if your vendor publishes figures, put theirs in. Nothing here is a measurement of any product, and the starting values are placeholders rather than findings.
The one measured figure on this page is in the next section, and it is the reason the false-positive slider is not hypothetical.
If detector errors were random, a flag would be bad luck. They are not random. They concentrate on the students who are already the least able to absorb an accusation.
A detector estimates how predictable each word is given the words before it and reads low surprise as machine authorship. Second-language instruction is built to make writing predictable: a fixed set of linking words, a fixed paragraph shape, the safe construction over the idiomatic one. That is competence, and the classifier reads it as a machine.
Grammar assistance, dictation and translation passes all smooth sentence rhythm, and rhythm uniformity is one of the properties detectors measure. A student using accommodations they are entitled to can move up a band without changing a single idea in the essay.
The structures you rewarded all term, the announced conclusion, the signposted transition, the balanced tricolon, are on every tell list including ours. A student who followed the rubric closely has written the most predictable text in the pile.
The mechanism in full is on why detectors flag human writing, and the same ground from the student's side is on for second-language writers.
Each one is cheap to answer and each one changes what the flag is worth. Together they are the difference between a suspicion and something you can act on.
A percentage is meaningless without the cut-off that turned it red. If the vendor does not publish the threshold, or the institution has not chosen one, the flag is a colour rather than a finding.
A score in isolation has no baseline. Earlier work from the same student, ideally written under conditions nobody disputes, is the most useful comparison available and costs one folder search.
Careful, formulaic, heavily taught writing scores high. So does writing produced with grammar assistance or speech-to-text, both of which flatten sentence rhythm. Second-language writers are flagged far more often than first-language writers by every detector that has been measured.
If the same tool flags a tenth of every cohort, the flag is telling you about the tool. Run the base-rate arithmetic below with your own class numbers before treating any single flag as a signal.
Drafts, version history, notes, sources, a conversation about the argument. A score is the weakest item on that list and the only one that cannot be examined. If it is the only item, there is no case yet.
You know how to have a difficult conversation with a student. The only thing this one adds is a piece of evidence that cannot be examined, which tends to push the exchange towards the tool and away from the work.
What follows is what tends to keep it on the writing. It is also, usefully, the approach that costs you nothing if the flag turns out to be wrong.
Open with the writing rather than the tool. "Walk me through how you built this argument" gets you further in two minutes than a percentage does in twenty, and it works whether or not anything was generated. A student who did the work can usually tell you which source changed their mind and why they cut the paragraph that used to be third.
Being unable to find out what you are accused of is the part students describe as unbearable. Naming the specific thing, a flag from a named tool, a section that reads unlike the rest of the term, keeps the conversation about something concrete rather than about their character.
Whether the prose reads as machine-made and whether a machine wrote it are different questions with different evidence. Keeping them apart lets you give useful feedback on the first without having settled the second, which is often where the conversation should end.
A student may have used a spellchecker, a grammar tool, a translation pass, or a friend who edits heavily. Some of that is permitted under most policies and some is not, and few students know which is which. Asking what they used, without the framing of a confession, usually gets you an accurate answer.
A student cannot prove a negative, and asking them to tends to produce panic rather than information. What they can produce is the trail their writing left, and their ability to talk about the work.
If a student drafts in our editor, it keeps a local history of how the document was built and exports that as a report signed in their browser with a key their browser generated. Anyone can check the signature on our verification page against the public key the report carries, without an account and without contacting us. It shows the record has not been altered since it was signed. It does not prove authorship, and no tool that reads finished text can.
Your institution's policy is the authority on that, and it is worth reading before the meeting rather than during it. What is true of the number itself: it is a classifier's confidence produced from writing style, it has no source document to examine, it carries no chain of custody, and the threshold that turned it red is usually unpublished. It can reasonably prompt you to look more closely. It cannot stand in for what you find when you do.
It is the classifier saying it is very confident, which is a claim about the model rather than about the student. Confidence saturates: the features that produce a maximum reading are the same ones careful formulaic writing produces, so a diligent student following the structure you taught can land at the ceiling. The number does not tell you how it got there. The highlighted sentences do, which is why this site shows those instead.
Most institutional policies treat assistive writing tools differently from generative ones, and most detectors cannot tell them apart. Grammar assistance and speech-to-text both smooth sentence rhythm, which is one of the properties detectors measure. That is a large part of why assistive-technology users appear so often in false-positive research.
That is your call and your institution's. One practical thing worth knowing before you do: pasting a student's unpublished work into a commercial detector sends that work to a vendor, under whatever data terms the vendor sets. The tools on this site do not, because they run in your browser and have nowhere to send it.
No. A low reading means the text does not carry the statistical fingerprints the tool looks for, which is a fact about the text rather than about its author. Prose can be generated and then edited into a low reading in a few minutes. The tools are useful for finding out what a text reads like, not for certifying who produced it.
No. Nothing on this page is uploaded. It uses 41 pattern rules, 3 whole-document statistics and some arithmetic. There is no analysis endpoint behind them, nothing is stored, and no student work is sent to a detector. Load the page once and switch the network off if you want to watch that hold.
Whatever record their writing process left, plus their ability to discuss the work. If they draft in our editor, it keeps a local history of how the document was built and exports it as a report signed in their browser, which anyone can check on our verification page against the key inside the report. That shows the record has not been altered since it was signed. It does not prove authorship, and no tool that reads finished text can.
Free, no account, and scoring runs in your browser.