Skip to content

AI-Assistance Forensics

Integrity Suite · Access: University partnersIncluded free · CorporatePaid · Government / workforce boardsAdd-on · K-12 / microschoolsPaid

Software-only proctoring — no camera, no microphone. Per-paragraph authorship classification, paste-provenance tracking, a cross-site domain ledger, and phone-bound biometric continuity, all running inside the tools students already write in.

Synergy co-owns two awarded U.S. utility patents in educational technology and holds an unlimited license to that technology.

Session Inputs

Total assessments this student has finished in the course to date.

keystrokes: 0
paste bursts: 0
ops: 0
Live keystroke timeline is being recorded for Draftback-style replay.

Type directly here to capture a live keystroke timeline for Draftback-style replay, or upload a .txt, .md, .docx, or .pdf — we extract text in your browser; the file itself is not uploaded to the server. Uploaded/pasted text has no keystroke history to replay.

45 min
0

Cross-site spider · 20 sites tracked

·

AI assistants

Study / answer-bank sites

Proctoring mode

Standard coursework

Software-only: stylometry, paste ledger, cross-site spider, and phone biometric assertions. Recommended for 90% of coursework.

Phone biometric continuity

Not paired

No passkeys paired on this browser yet. Tap Pair phone above to register one.

Revoking removes the passkey server-side immediately. Add device registers another phone or platform authenticator under the same pairing id.

Idle

Pings this session

0

Total signed

0

Ledger root

Pairing id: Signed assertions: 0

Live session prompts your phone's Face ID / fingerprint at the chosen interval. Each verified ping is SHA-256 chained to the previous one and written to the forensic Merkle ledger — no video, no microphone.

Pairing is optional. You can run an analysis without a paired phone — the report will simply carry no biometric-continuity signal. If a phone is paired, a freshly verified assertion (within the last 2 minutes) is required before the run starts.
Run the analysis to see the authorship heatmap, provenance ledger, and risk score.
How this works — what each analysis does, and what to look for

1. Per-paragraph authorship heatmap

What it does: Each paragraph is classified as Original, AI-referenced, Lightly edited AI, or Verbatim AI paste using keystroke vs. paste ratio, edit churn, and stylometric AI-likeness.

Why it matters: Catches the most common cheating pattern: students pasting an LLM answer and lightly rewording it. Sentence-level signals are harder to game than a single document-level score.

Expected result: Green = student-written. Yellow = AI-referenced. Orange/red = pasted from an AI tool with little editing. Mixed colors across paragraphs is normal; a wall of red is the smoking gun.

2. Paste-provenance ledger

What it does: Every paste event is recorded with character count, source domain (when available), a one-way hash of the payload, and a short preview.

Why it matters: Proves where text actually came from rather than guessing. Hashes let you re-verify later without storing the student's content.

Expected result: A short list of paste events tied to specific paragraphs and source domains (e.g. chatgpt.com). Many large pastes from AI domains = high risk.

3. Cross-site domain spider

What it does: Tracks which AI assistants and answer-bank sites (Chegg, CourseHero, Quizlet, Brainly, etc.) were open during the session and for how long.

Why it matters: AI-likeness alone is circumstantial; a 3-minute ChatGPT dwell mid-session turns suspicion into evidence.

Expected result: Timeline rows for each visited domain with dwell time. Visits to AI/answer-bank sites are flagged; the doc editor itself should dominate.

4. Phone-bound biometric continuity

What it does: A paired phone sends periodic WebAuthn assertions (Face ID / fingerprint) during the session. Each ping proves the same human is still present.

Why it matters: Closes the impersonation gap without a webcam. If the phone stops responding, authorship can no longer be bound to a specific person.

Expected result: “Paired · continuous” with coverage ≥ 60% of expected pings. Unpaired or 0 pings = continuity gap, even if the writing looks clean.

5. Evidence timeline

What it does: Merges paste events, domain visits, and biometric pings onto a single session timeline.

Why it matters: Lets a reviewer see causality at a glance — e.g. ChatGPT visit → large paste → no biometric ping in that window.

Expected result: Clusters of red (pastes) lining up with AI-domain visits are the strongest signal. Evenly spaced green pings across the session = healthy.

6. AI-assistance risk score & verdict

What it does: A 0–100 score combining heatmap distribution, paste provenance, domain ledger, and pairing coverage, plus a Likely Authentic / Needs Review / High Risk verdict.

Why it matters: Gives faculty a single triage number while keeping all underlying evidence inspectable.

Expected result: 0–30 Likely Authentic, 31–69 Needs Review (open the heatmap + timeline), 70–100 High Risk (escalate with the ledger as evidence).

Bias safeguards

How this tool avoids unfairly flagging any group — including non-native English writers. Detection weighs process signals (how the text was produced), not language-proficiency signals (how polished the English reads). We flag for human review; we never accuse.

Process signals outweigh language signals

The strongest evidence — paste bursts, typing pace, keystroke timeline, source domains, biometric continuity — describes HOW the document was produced, not how polished the English is. These signals do not correlate with a writer's native language, so they do not disadvantage non-native or dialect-varied writers.

Stylometry is an indicator, never a verdict

Our stylometric AI-likeness score is explicitly labeled an indicator only and can never, on its own, produce a misconduct determination. Fluent-but-non-native prose can raise stylometry; because it cannot stand alone, that alone will not flag a student.

We flag, we don't accuse

Every result routes to human review. The tool surfaces evidence for an instructor to weigh; it never issues an academic-integrity finding automatically. A high score is a prompt to look closer, with all underlying evidence inspectable.

No proficiency, grammar, or dialect scoring

We deliberately do NOT score vocabulary richness, grammatical 'correctness', spelling, or dialect. Those metrics are exactly where language bias creeps in, so they are excluded from every score on this report.

Honest 'insufficient evidence' handling

When authorship telemetry wasn't captured (e.g. a pasted or uploaded file), we say 'insufficient evidence' rather than inferring guilt from writing style. This protects writers whose style might otherwise be misread.

Which signals were captured is disclosed

Each result states which signals backed it (keystroke op-log, paste events, domain visits, biometric continuity). When strong signals like a captured writing-process op-log are absent, the tool says so plainly instead of overstating certainty. We publish no accuracy percentage, because no detector can certify an exact cheating rate.

How to read these results

Every number on this page — the AI-assistance risk score, per-paragraph classifications, and per-writer indicators — is a statistical indicator, not a determination of misconduct. An exact “cheating percentage” cannot be scientifically guaranteed by any detector on the market, and we make no such claim. Where authorship telemetry was not captured, the tool says so plainly (“insufficient evidence”) rather than fabricate a verdict. Treat these signals as a way to prioritize human review — never as standalone proof.