Scam AI raised $2.6M to combat AI-powered scams→
scam.ai

Independent applied research · 2025—2026

Research for a world where seeing is no longer believing.

We study how synthetic media and automated scams behave outside the lab—and build the evidence needed to detect them in the real world.

Publications
21
Datasets
13
Programs
7

Program 01 / Documents

Document Forgery

7 publications · newest first

Program 02 / Generated media

AI-Generated Detection

3 publications · newest first

Program 03 / Deepfakes

Deepfake Detection

3 publications · newest first

Program 04 / Age assurance

Age Estimation

2 publications · newest first

Program 05 / Calls

Voice & Phone Scams

4 publications · newest first

Program 06 / Behavioral signals

Interview Tech

1 publication · newest first

Program 07 / Language models

LLM Studies

1 publication · newest first

Research resources

Datasets

Free to download. Needs a scam.ai account with an initial deposit.

Real-World Faceswap Dataset (RWFS)

A real-world faceswap collection used to evaluate deepfake detectors against in-the-wild manipulations rather than lab-only synthetics.

Log in to access

AI-edit document forgery dataset (AIForge-Doc)

Financial and form documents tampered by a suite of AI editing tools, paired with originals — used to benchmark document forgery detectors.

Log in to access

AIForge-Doc v2.0 (GPT-Image-2 document forgeries)

A v2 expansion of AIForge-Doc covering GPT-Image-2 generated tampering — used in the 'When the Forger Is the Judge' benchmark.

Log in to access

GPT-Image-2 Twitter Dataset

Self-reported AI-generated images collected from Twitter during the first week of GPT-Image-2 deployment — captures real-world distribution shift.

Log in to access

Adversarial age estimation attack dataset

Faces with low-cost cosmetic adversarial perturbations designed to defeat age estimation systems — used in 'Can a Teenager Fool an AI?'.

Log in to access

Fully-synthetic AI-generated receipt (GPT-4o-receipt)

Fully synthetic receipts generated by GPT-4o, paired with a human-study evaluation of detectability — used for AI-generated document forensics research.

Log in to access

Real scam & spam call dataset (English)

English-language recordings of real scam and spam phone calls — the corpus behind 'Anatomy of a Scam Call'.

Log in to access

Simulated gaze estimation for reading dataset

Synthetic eye-movement trajectories rendered through a 3D eye simulator, replaying real reading paths — used for script reading detection research.

Log in to access

AgentForge-Bench

162 PDFs: 81 publicly filed financial documents, each paired with one verified agent-made forgery, plus task cards, the verifier, and per-run verdicts.

Log in to access

ChatGPT Images 2.5 in the Wild

3,478 images and 12 videos drawn from 2,440 launch-period posts across eight platforms, with source metadata for detector evaluation.

Log in to access

AIForge-Doc v3

A controlled benchmark of ChatGPT Images 2.5 document-forgery tasks, including prompts, masks, scorers, and score files.

Log in to access

The Machines Are Calling dataset

Derived features and human and detector labels for 6,192 unwanted inbound calls captured by a voice honeypot.

Log in to access

Open-Jev CallScreenBench judgments

82 generated test calls covering 41 CallScreenBench scenarios and two scripted secretaries, with 577 per-turn labels.

Log in to access

Interested in collaborating?

We partner with academic institutions and industry labs on deepfake detection research.