ai-code-detector

Methodology & limits

How accurate is an AI code detector?

Useful enough to help you choose what to review. Not reliable enough to decide a case by itself. That is the standard this product is built around.

Useful for screening, not proof

An AI code detector can help you sort a large set of files and find the ones that deserve a closer read. It should not be treated as proof of misconduct. A high score means "review this first." It does not mean "this person cheated."

Why code is hard to classify

Short programs, starter templates, common libraries, and strict style rules can make human code look generic. Rewrites and mixed authorship can make AI-assisted code look more human. The same submission can look different depending on language, assignment, and prompt design.

A 2024 code-specific study found that existing AIGC detectors performed poorly when separating human-written and AI-generated code.
arXiv:2401.03676 · ICSE-SEET 2024 · research context, not a competitor attack

What our report shows

The report is designed to make the score inspectable. Instead of one number and a shrug, it should show the file-level result, the snippets that contributed most, and enough context for a human to decide what to ask next.

  1. Input handling: file type, language, and assignment context.
  2. Signals: code patterns associated with AI-generated output.
  3. Output: per-file score, snippet highlights, exportable report.
  4. Human step: review, compare, ask, document.

Our current evaluation status

We still do not publish a single accuracy percentage. What we can publish is where our evaluation stands today: a named test set, a date, a method summary, and the failure modes we already know about.

Test set

Our most recent evaluation ran on CoDET-M4, a public code-detection benchmark, plus our own held-out test set — 1,922 code samples in total, human-written and AI-generated, across six languages: Python, Java, C++, JavaScript, HTML, and Lua.

Evaluation date

October 8, 2026. We re-run the evaluation when the model, the scoring threshold, or the input pipeline changes, and this section is updated when a new round completes.

Method

Every sample went through the same input normalization our production scanner uses, and every sample was scored at the same production threshold. That threshold is calibrated so that human-written code in the evaluation set is not flagged. The trade-off is deliberate: some AI-generated code will pass, because a tool that flags honest work causes more harm than a tool that asks for a second look.

Known failure modes

  • Very short files are not scored. Submissions with fewer than 10 non-empty lines of code are rejected rather than guessed at.
  • HTML, JavaScript, and Lua are less validated on the human side. This evaluation round contained AI-generated samples in these languages but no human-written ones, so false-positive behavior there is extrapolated, not measured. Treat flags in these languages with extra caution.
  • Heavily rewritten AI code is a known weak spot. Deliberate rewrites, paraphrasing, and mixed human-AI authorship all reduce the signal.
  • Boilerplate can look generic. Starter templates, auto-formatters, and strict style rules can make honest human code resemble model output.

This is a snapshot, not a certificate. The next evaluation round will update this section — including human-side validation for HTML, JavaScript, and Lua.

The human review workflow

  1. Read the flagged snippets before you read the score.
  2. Check whether starter code or required libraries explain the pattern.
  3. Compare with earlier work from the same author when available.
  4. Ask the author to explain, modify, or extend the code.
  5. Record the review outcome separately from the detector output.

Get the report template

What we will not claim

We will not claim a fixed accuracy percentage without evidence. We will not say a report proves cheating. We will not tell students how to make AI-generated code pass screening. Those claims make the category less trustworthy and make real review harder.

Questions

Why not publish one accuracy number?

Because one number hides the part that matters: which kinds of code fail, on which languages, and under which edits. Until we can publish that honestly, we will not publish a number.

Does the detector work equally well in every language?

No. In our most recent evaluation (October 8, 2026), Python, Java, and C++ were tested with both human-written and AI-generated samples. HTML, JavaScript, and Lua had AI-generated samples but no human-written ones, so results in those languages are extrapolated and carry more uncertainty.

What is a false positive here?

A false positive is human-written code that gets flagged. Short assignments, starter templates, and strict formatting can all raise that risk.

What is a false negative?

A false negative is AI-assisted code that does not get flagged. Rewrites, mixed authorship, and heavy editing can all reduce the signal.

Should we stop using detectors completely?

Not necessarily. If you treat the output as a screening signal and keep a human review step, it can still save time and focus attention.

Can students check their code before submitting?

Students can use a free check to understand what might get flagged and to prepare an explanation of their work. We do not provide advice for hiding AI use.

Use the detector where it helps. Keep the decision with a human.