How to Check if an Assignment Is AI-Generated
There is no test that proves who wrote a piece of text. What you can do is gather evidence: check whether the citations exist, look for specifics, compare the work against the course, open the file's version history, and run a detector beside a known sample. Several signals pointing one way settle it. A single number does not.
Two people read a page like this. One has to decide what to do about a document that reads a little too well. The other wrote every word and has just been told a machine disagrees. The checks below serve both, because they are the same checks.
How to check if an assignment is AI-generated
There is no single test, so run five checks and see whether they agree:
- Check the citations. Pick three references and try to find them.
- Look for specifics. Dates, figures, named researchers, anything only this writer would know.
- Compare the work against the course, not against good writing.
- Open the file properties and the version history. Ask where it was written.
- Run a detector on the whole document, with a known sample of the same person's writing beside it.
Four of those five cost nothing and need no tool. None of them is proof on its own, and a detector score is the weakest of the five, not the strongest. The rest of this article is how to run each check and how to read what comes back.
What an AI detector actually measures
A detector does not recognise ChatGPT's handwriting. It holds no list of sentences the model has produced. It measures how predictable the text is.
Two properties do most of the work. The first is how surprised a language model is by each word, given the words before it. The second is how much the writing varies: sentence length, sentence structure, the rhythm of long against short. Human prose wanders. A twenty-eight word sentence sits next to a five-word one. There is a clumsy clause the writer never went back to fix, and a word chosen because it was that writer's word rather than the obvious one.
Generated text sits closer to the average at every step, because choosing a likely next word is the thing a language model does.
So the score answers one question: how closely does this resemble the most predictable way to write this? That question is related to authorship. It is not the same question, and the gap between the two is where every wrongful accusation lives. A methodical writer with a plain style produces predictable prose. So does an academic writing to a template. So does anyone working in a second language, from a smaller vocabulary and safer constructions.
What to look for in the assignment itself
The strongest signals are in the work itself, and they cost nothing but attention. No single one of them is conclusive. Together they usually settle it.
- Check the citations. This is the best test there is. Pick three references and try to find them. Generated citations are plausible in every respect except existence: a real-sounding author, a real journal, a sensible year, and a DOI that resolves to nothing. There is a harder variant, where the paper exists but says something other than what it is cited for.
- Look for anything specific. Generated writing is fluent about the general and thin about the particular. Watch for what is missing: dates, figures, place names, named researchers. "Many scholars have argued" costs nothing to produce and names nobody.
- Compare it against the course, not against good writing. The work often answers the question competently and answers the module not at all. It skips the framework taught in week six, cites nothing from the reading list, and uses terminology the class never used.
- Look at the shape of it. Paragraphs of near-identical length, each opening with its topic sentence. Lists that keep arriving in threes. Headings imposed on a nine-hundred-word essay. A conclusion that restates the introduction and adds no claim of its own.
- Open the file properties. People forget this one. A .docx file records total editing time and a revision count, under File then Info in Word, and a three-thousand-word essay with four minutes of editing time was written somewhere else. Be careful about what that does and does not show: Word counts only the time the file was open in Word, so a student who drafts in Google Docs and exports at the end produces the same four minutes honestly. Ask where it was written before you conclude anything. Google Docs version history is the stronger record, because it shows whether a document was built over days or arrived in one paste.
How to run a detector without fooling yourself
If you are going to use a tool, use it properly. Most bad conclusions come from four avoidable mistakes.
Run the whole document. Detectors are unreliable on short passages. Turnitin will not return an AI score at all below 300 words of prose, and it raised that floor from 150 words because short samples produced too many false positives. Feeding in the one paragraph that felt wrong is the fastest route to a false answer.
Get a baseline first. Run something the same person certainly wrote, long enough to clear the same word count: an in-class piece, last term's assignment, a draft chapter. A 70% score on the suspect document means very little by itself. A 70% sitting next to a 9% on their known work means a great deal.
Check the close calls twice. Run the document again, and where it matters, run it through a second tool. The ones you will meet are Turnitin, which most universities use and most students never see the settings of, and the public tools around it: GPTZero, Copyleaks, Originality.ai. What each one costs and what it reads compares them on the things you can check before paying. They disagree with each other more than any of them advertises. In the study below, seven detectors read the same 91 human-written essays: 89 were flagged by at least one of them, and only 18 by all seven. Two tools that disagree wildly are telling you the text sits near the decision boundary, which is exactly where you should not be drawing conclusions.
Mind where you paste it. Another person's unpublished work going into a free web box is a confidentiality problem before it is anything else. Use something that states what happens to the file, in writing, before you upload anything. Ours is set out in our FAQ.
What a high AI score proves, and what it does not
It does not prove authorship, and it fails in both directions.
False positives are ordinary, not rare. A 2023 study in the journal Patterns ran 91 TOEFL essays, all written by non-native English speakers, through seven detectors. The detectors called 61% of them AI-generated on average, and 89 of the 91 essays were flagged by at least one detector. Every one had been written by a person. The mechanism is the one described above: simpler constructions and a narrower vocabulary read as predictable, and predictable is what the tool scores. Heavy grammar-tool rewriting does the same thing to a score, and so does text that has been translated and translated back.
A low score clears nobody. Generated text that somebody has edited by hand for twenty minutes drops a long way down the scale. Ask the model for a rougher register and it drops further. Turnitin treats its own low scores with the same suspicion: a result above 0% but below 20% is shown as an asterisk rather than a number, because that band produced too many false positives to report plainly.
That asymmetry is the useful part. Treat a high score as a reason to look harder, and a low score as no information at all.
If you are the teacher or the marker
Somebody hands you an assignment. Two thousand words, and they read well. They read a little too well, in a way you cannot quite point at, and now you have to decide what to do about a feeling.
Never open with the number. A teacher who leads with a percentage has started an argument about a tool. A score is a prompt to investigate, and the investigation is a conversation.
Ask the writer to talk about the work. Have them explain the argument in their own words, find one of their own sources, say why they chose that method over the other one. Somebody who wrote it handles this in two minutes and is usually glad to. Somebody who did not cannot, and you will both know it within a few questions.
Then follow your institution's procedure. Turnitin's own guidance says its AI indicator "should not be used as the sole basis for adverse actions against a student", and institutional policy increasingly says the same. Keep the record as you go: the score, the citation checks, the file metadata, and what was said.
If you are the writer being checked
Your position is the reverse one. You wrote every word of the assignment yourself, a detector has just called it AI-generated, and now you have to prove a negative.
Build the evidence while you write, because you cannot build it afterwards. Draft in something that keeps version history, such as Google Docs, or Word with history switched on. Keep your notes, your outline, and the messy second draft. That record is the only thing that reconstructs how a document actually came together, and it settles the question faster than any argument about percentages.
Then check early rather than late. A surprising score a week before the deadline is a conversation with your supervisor. The same score the night before is a panel meeting.
One more thing, and it is what people get wrong most often. If you used AI for any part of the work, including outlining or language editing, declare it wherever your institution's policy asks. Declared assistance is a policy question with a short answer. Undeclared assistance is a misconduct question, and those do not end quickly.
One honest caveat
An AI score is one input among the signals above, and anyone who tells you their tool is the last word on authorship is selling a certainty that does not exist. What a check gives you is a documented starting point: a score for the whole document, the text it was calculated on, and a plain statement of what the number can and cannot support.
That is what our AI and plagiarism check is built to be. Upload the file, run both checks in the same pass, and download the report as a PDF, since the two questions almost always turn up together. How it works is the four-step version, and pricing is per document, with no account and no subscription. Each check is its own purchase, so if you plan to re-check after edits, budget for two. If you are unsure whether your file will work, ask us before you pay.
Frequently asked questions
Can a teacher prove an assignment was written by AI?
No tool proves authorship. A detector estimates how predictable the text is, which is related to authorship without being the same question. What settles a case is several signals agreeing: citations that do not exist, no specifics anywhere, work that ignores the course, thin version history, and a score far from that writer's known baseline.
Are AI detectors accurate enough to act on?
They are accurate enough to prompt a look and not enough to decide anything. In a 2023 study in Patterns, seven detectors read 91 TOEFL essays written by non-native English speakers and flagged 89 of them as AI at least once, at an average false positive rate of 61%. Every essay had a human author. That population is the point: predictable prose is what the tools score.
What should I do if I am wrongly accused of using AI?
Ask which tool produced the score, and what your institution's policy says about acting on one. Then produce your record: version history, notes, outlines, drafts. Offer to talk through the argument in person, because that conversation is your strongest evidence. Turnitin's own guidance says the indicator should not be the sole basis for action.
Does Turnitin show students their AI writing score?
No. The indicator is visible to instructors and administrators, and its highlights do not appear in the Similarity Report at all. An instructor can choose to download the AI report and share it, and the report itself covers why nobody can sell you that figure.
Is an AI score the same as a similarity score?
No. They are two measurements from two different systems, and a document can score low on one and high on the other. Similarity counts text that matches something already in a database, and what percentage is safe is the number your university applies to that one.
Can both checks run on the same upload?
Yes. The two questions usually arrive together, so both run in one pass on the same file and come back in one PDF. The AI result is a document-level score shown with the text it was calculated on. The similarity result adds the matched passages and the source each one came from. How it works walks through the four steps, and pricing is charged per document.