Scam AI raised $2.6M to combat AI-powered scams→
scam.ai
← Blog

Insights ·

Deepfake detection accuracy: how to interpret the evidence

Learn how to read detector scores, verdicts, localized evidence and abstentions, then route media to approval, step-up checks or human review.

Deepfake detectors turn complex forensic signals into a score, and that simplicity invites overreach. A business may read a score as certainty, treat a low-risk verdict as proof, or send every inconclusive result to the same queue. Each move strips context from the evidence. A safer decision accounts for the result, the media and the consequences of being wrong. At Scam AI, we return the layers of evidence needed for that process.

At a glance

  • A score is a likelihood signal for one input. It is not the detector's accuracy or proof of fraud.
  • A verdict selects a policy path. Your organization owns the action at the end of that path.
  • Localized evidence tells a reviewer where to inspect. It does not make the highlighted area ground truth.
  • A null value or abstention is a gap in the evidence, never a favorable result.
  • High-impact decisions need representative testing, corroboration and accountable review.

Deepfake detection accuracy starts with the test conditions

Accuracy is a property of an evaluation. It is not a property of the score attached to one file.

A reported accuracy figure only becomes useful when its methodology names:

  • the image, video or audio dataset and how ground truth was established
  • the mix of genuine and manipulated media
  • the manipulation methods and generation tools represented
  • benign changes such as compression, cropping, resizing, background noise and screen recording
  • the model version and decision rule used for the test
  • how unscored and abstained cases were counted

Without that context, two accuracy figures may describe very different tasks. A detector tested on clean, balanced research data faces a different problem from a detector receiving recompressed uploads, short voice notes and unfamiliar generation methods.

The AI Risk Management Framework from the National Institute of Standards and Technology (NIST) calls for accuracy measurements to use realistic test sets that represent expected conditions, with the methodology documented. NIST's OpenMFC evaluations also separate file-level detection from manipulation localization and define the operating point used to assess false alarms. Those are useful design principles for a buyer evaluation, even when the production workflow uses different measures.

This is why a single headline number cannot answer the operational question: what should we do with this file?

Read a detection result as a stack of evidence

A useful result has several layers. Each layer answers a different question.

Evidence layerWhat it tells youWhat it cannot establish
ScoreHow strongly the model's learned signals align with AI generation or manipulation patterns for this inputOverall detector accuracy, the sender's intent or proof of fraud
VerdictThe product's condensed interpretation of the resultYour organization's final business decision
Localized evidenceWhich frame, audio window or image region deserves inspection, when supportedThat every other part is clean or that the highlighted area is ground truth
Missing evidence or abstentionWhere the system could not produce a valid result and whyA clean result
Provenance and capture evidenceWhat is known about the file's source, history or capture pathWhether the depicted event is true
Case contextWhether the media is consistent with known channels, records and behaviorA forensic assessment on its own

The layers remain distinct, even though they feed the same policy route and human decision.

Five layers of detection evidence feeding a policy route and human decision

Our result structure exposes these layers directly. The result documentation defines verdict, score and a human-readable summary. Video responses add per-frame evidence and a cause beside an unscored frame. Audio responses add per-window values, an abstention status and reason codes. The response also records the model, detection ID and completion timestamp for traceability.

Score: a likelihood signal for this input

The score is model evidence. It can help order a review queue, compare cases inside the same tested workflow and study where a routing policy should place more scrutiny.

It does not tell you the percentage of all deepfakes the detector catches. It should not be translated into a statement such as “this customer is that percent likely to be committing fraud.” Interpreting a score as a real-world probability requires a documented calibration study on representative deployment data. Keep it as a likelihood signal when that evidence is unavailable.

Our video results can also include threshold_used. This is the line used to interpret per-frame scores for that run. It is separate from the business rules that determine approval, review or decline.

Verdict: a compact model interpretation

We use three verdict labels across Scam AI media results:

  • LIKELY_REAL means the detector found no strong signs of AI generation or manipulation. It does not authenticate the person, source or event.
  • ALERT means AI-related signals appeared and the result is inconclusive. This belongs in a review or step-up path.
  • LIKELY_AI means the detector found strong signs of AI generation or manipulation. Consequential action still follows the customer's policy and supporting evidence.

Use a verdict to select a policy path. The customer remains responsible for the decision at the end of that path.

Localized evidence: where to look

A whole-file score compresses time and space. Local evidence restores some of that detail.

For video, a reviewer can inspect the frames and timestamps that produced evidence, then watch the surrounding sequence. For audio, the reviewer can listen to a flagged window with the speech immediately before and after it. An image detector may return a region or heatmap when that capability is supported.

Localization makes a result easier to investigate, but it does not turn the highlighted area into proof. A model may respond to compression, compositing or capture artifacts that overlap with synthetic-media patterns. Reviewers need the original file, the local evidence and the case context together.

The documented example below shows Scam AI's result schema rather than a production detection. It illustrates how the response explains a missing frame score.

Scam AI result documentation showing an unscored frame cause

Missing evidence and abstention: an explicit gap

An abstention is a valid system state. It means the detector did not have enough usable evidence to score all or part of the input.

In our result structure, a video frame can carry a null score with a cause such as a timeout. An audio window can be marked abstained with reason codes such as no voice content. Null does not mean zero, genuine or low risk.

Route that gap according to its cause. A corrupted or unsupported input may need a new capture in a supported format. A missing speech segment may require another recording. A high-impact case with incomplete evidence belongs in review. Do not silently replace an abstention with a favorable score.

Map the evidence to an action

The detector supplies evidence. Your policy assigns the action. A practical routing matrix looks like this:

Evidence stateCandidate actionRequired safeguard
Complete result, no strong synthetic-media signals, corroborating checks are consistentApprove when policy permitsRecord that the detector found no strong signals; do not label the media “proven authentic”
Inconclusive result, conflicting checks or unusual case contextStep upRequest a fresh capture, use a known-channel callback, or obtain another independent record
Strong or localized synthetic-media signals in a consequential caseHuman reviewPreserve the original, inspect the localized evidence and document the reviewer rationale
Strong signals plus corroboration, within a workflow validated for this input typeDecline when the organization's policy permitsApply appeal, exception and audit procedures appropriate to the decision
Missing result, abstention or unsupported inputStep up or reviewResolve the evidence gap; never treat missing analysis as a favorable result

Result evidence is one input into the route. Corroboration and the consequence of a wrong outcome also belong in the customer-owned policy.

Result evidence, corroboration and consequence feeding a customer policy with four review paths

This matrix should vary by consequence. A content-moderation label, a request for recapture and the denial of a financial service do not carry the same harm if the detector is wrong.

A false alert sends genuine media into friction or review. A missed detection allows manipulated media to proceed. Lowering one type of error can raise the other, so there is no universal routing line. Choose a policy through representative testing, review capacity and the cost of each error in the actual workflow.

Corroborate high-impact results

Independent checks answer questions that a detector score cannot.

Provenance. Content Credentials can record signed information about a file's origin and edit history. The C2PA explainer from the Coalition for Content Provenance and Authenticity (C2PA) is explicit that valid provenance does not establish whether the depicted content is true or factual. Missing credentials are also inconclusive because many legitimate capture and editing tools do not add them.

Capture integrity. Device attestation, virtual-camera checks, liveness controls and a fresh in-app capture can show whether media traveled through the expected path. These controls address delivery and presentation attacks. They answer a different question from content analysis.

Known-channel verification. A callback to an established number, confirmation inside an existing account session or direct contact with the purported sender can test the identity and request behind the media.

Case records. Transaction history, timestamps, prior submissions and source documents can reveal contradictions or support a benign explanation.

Run these checks as complementary evidence, rather than as a vote among similar scores. Do not assume multiple scores are independent when detectors use overlapping training data or signals. NIST's synthetic-content report likewise treats provenance tracking and content-based detection as distinct approaches within a broader transparency system.

Validate the workflow with a representative shadow pilot

Vendor benchmark results cannot replace testing on the media your organization receives. Use a representative shadow-mode pilot, where detection runs beside the current process without changing live outcomes.

  1. Define the decision and the harms. Name the action the signal may influence, who can be affected and what happens after an incorrect route.
  2. Build a labeled sample. Include genuine and known manipulated media across the real mix of modalities, sources, capture devices, quality levels and benign edits. Keep a documented labeling and adjudication process.
  3. Hold out evaluation cases. Keep part of the labeled set separate from policy design so the final check measures the chosen workflow on unseen cases.
  4. Run blind and log every state. Record the model version, score, verdict, local evidence, abstentions, processing failures and current human outcome without exposing the label to the detector or reviewer.
  5. Study errors by slice. Review false alerts and missed detections separately by modality, source, quality and attack type. Count abstentions as their own outcome. A single aggregate accuracy figure can hide the route that creates operational harm.
  6. Set customer-owned actions. Assign approve, step-up, review and decline paths based on the pilot evidence, decision severity and available review capacity. Validate the complete policy on the holdout set.
  7. Monitor after launch. Track reviewer overrides, appeals, evidence gaps and shifts in input quality or attack methods. Re-run the evaluation when the model or workflow changes.

For remote identity proofing, NIST SP 800-63A provides a useful example of this method. It calls for testing automated image analysis against genuine and manipulated media, documenting the attack artifacts tested and augmenting automated decisioning with manual review to address errors.

Keep a result record that another reviewer can reconstruct

A decision becomes defensible when another qualified person can understand what happened without guessing. Retain the fields your policy and privacy requirements allow:

  • source or capture method and a stable identifier for the original media
  • detection ID, model version and completion timestamp
  • score, verdict and any result-level interpretation line
  • localized evidence, missing values, causes and abstention reasons
  • provenance, capture-integrity and known-channel checks
  • action taken, reviewer identity and rationale
  • recapture, exception and appeal outcomes

Retention should match the sensitivity of the media and the purpose of the review. Our Scam AI Trust Center explains our current security, privacy and media-handling approach. Your audit record should also capture the customer-owned policy version that produced the action.

Deepfake detection accuracy matters, but operational reliability comes from the whole evidence path: representative evaluation, transparent result fields, explicit gaps, corroboration and accountable review.

Book a Scam AI demo to see how scores, verdicts and abstentions can support your review policy.