AI-Assistance Forensics
Integrity Suite · Access: University partners — Included free · Corporate — Paid · Government / workforce boards — Add-on · K-12 / microschools — Paid
Software-only proctoring — no camera, no microphone. Per-paragraph authorship classification, paste-provenance tracking, a cross-site domain ledger, and phone-bound biometric continuity, all running inside the tools students already write in.
Synergy co-owns two awarded U.S. utility patents in educational technology and holds an unlimited license to that technology.
Total assessments this student has finished in the course to date.
Type directly here to capture a live keystroke timeline for Draftback-style replay, or upload a .txt, .md, .docx, or .pdf — we extract text in your browser; the file itself is not uploaded to the server. Uploaded/pasted text has no keystroke history to replay.
Cross-site spider · 20 sites tracked
AI assistants
Study / answer-bank sites
Proctoring mode
Software-only: stylometry, paste ledger, cross-site spider, and phone biometric assertions. Recommended for 90% of coursework.
Phone biometric continuity
No passkeys paired on this browser yet. Tap Pair phone above to register one.
Revoking removes the passkey server-side immediately. Add device registers another phone or platform authenticator under the same pairing id.
Pings this session
0
Total signed
0
Ledger root
—
Live session prompts your phone's Face ID / fingerprint at the chosen interval. Each verified ping is SHA-256 chained to the previous one and written to the forensic Merkle ledger — no video, no microphone.
1. Per-paragraph authorship heatmap
What it does: Each paragraph is classified as Original, AI-referenced, Lightly edited AI, or Verbatim AI paste using keystroke vs. paste ratio, edit churn, and stylometric AI-likeness.
Why it matters: Catches the most common cheating pattern: students pasting an LLM answer and lightly rewording it. Sentence-level signals are harder to game than a single document-level score.
Expected result: Green = student-written. Yellow = AI-referenced. Orange/red = pasted from an AI tool with little editing. Mixed colors across paragraphs is normal; a wall of red is the smoking gun.
2. Paste-provenance ledger
What it does: Every paste event is recorded with character count, source domain (when available), a one-way hash of the payload, and a short preview.
Why it matters: Proves where text actually came from rather than guessing. Hashes let you re-verify later without storing the student's content.
Expected result: A short list of paste events tied to specific paragraphs and source domains (e.g. chatgpt.com). Many large pastes from AI domains = high risk.
3. Cross-site domain spider
What it does: Tracks which AI assistants and answer-bank sites (Chegg, CourseHero, Quizlet, Brainly, etc.) were open during the session and for how long.
Why it matters: AI-likeness alone is circumstantial; a 3-minute ChatGPT dwell mid-session turns suspicion into evidence.
Expected result: Timeline rows for each visited domain with dwell time. Visits to AI/answer-bank sites are flagged; the doc editor itself should dominate.
4. Phone-bound biometric continuity
What it does: A paired phone sends periodic WebAuthn assertions (Face ID / fingerprint) during the session. Each ping proves the same human is still present.
Why it matters: Closes the impersonation gap without a webcam. If the phone stops responding, authorship can no longer be bound to a specific person.
Expected result: “Paired · continuous” with coverage ≥ 60% of expected pings. Unpaired or 0 pings = continuity gap, even if the writing looks clean.
5. Evidence timeline
What it does: Merges paste events, domain visits, and biometric pings onto a single session timeline.
Why it matters: Lets a reviewer see causality at a glance — e.g. ChatGPT visit → large paste → no biometric ping in that window.
Expected result: Clusters of red (pastes) lining up with AI-domain visits are the strongest signal. Evenly spaced green pings across the session = healthy.
6. AI-assistance risk score & verdict
What it does: A 0–100 score combining heatmap distribution, paste provenance, domain ledger, and pairing coverage, plus a Likely Authentic / Needs Review / High Risk verdict.
Why it matters: Gives faculty a single triage number while keeping all underlying evidence inspectable.
Expected result: 0–30 Likely Authentic, 31–69 Needs Review (open the heatmap + timeline), 70–100 High Risk (escalate with the ledger as evidence).
How this tool avoids unfairly flagging any group — including non-native English writers. Detection weighs process signals (how the text was produced), not language-proficiency signals (how polished the English reads). We flag for human review; we never accuse.
Process signals outweigh language signals
The strongest evidence — paste bursts, typing pace, keystroke timeline, source domains, biometric continuity — describes HOW the document was produced, not how polished the English is. These signals do not correlate with a writer's native language, so they do not disadvantage non-native or dialect-varied writers.
Stylometry is an indicator, never a verdict
Our stylometric AI-likeness score is explicitly labeled an indicator only and can never, on its own, produce a misconduct determination. Fluent-but-non-native prose can raise stylometry; because it cannot stand alone, that alone will not flag a student.
We flag, we don't accuse
Every result routes to human review. The tool surfaces evidence for an instructor to weigh; it never issues an academic-integrity finding automatically. A high score is a prompt to look closer, with all underlying evidence inspectable.
No proficiency, grammar, or dialect scoring
We deliberately do NOT score vocabulary richness, grammatical 'correctness', spelling, or dialect. Those metrics are exactly where language bias creeps in, so they are excluded from every score on this report.
Honest 'insufficient evidence' handling
When authorship telemetry wasn't captured (e.g. a pasted or uploaded file), we say 'insufficient evidence' rather than inferring guilt from writing style. This protects writers whose style might otherwise be misread.
Which signals were captured is disclosed
Each result states which signals backed it (keystroke op-log, paste events, domain visits, biometric continuity). When strong signals like a captured writing-process op-log are absent, the tool says so plainly instead of overstating certainty. We publish no accuracy percentage, because no detector can certify an exact cheating rate.
How to read these results
Every number on this page — the AI-assistance risk score, per-paragraph classifications, and per-writer indicators — is a statistical indicator, not a determination of misconduct. An exact “cheating percentage” cannot be scientifically guaranteed by any detector on the market, and we make no such claim. Where authorship telemetry was not captured, the tool says so plainly (“insufficient evidence”) rather than fabricate a verdict. Treat these signals as a way to prioritize human review — never as standalone proof.
