AI vs Real Image Detector β 3-Branch Ensemble
Full write-up, methodology, and deployment code: https://github.com/ammarsapru/ai-detector
Binary classifier that predicts whether an image is a real photograph or AI-generated. It fuses three complementary views of an image:
- CLIP ViT-L-14 (frozen) + trainable adapter β semantic plausibility (768-d)
- EfficientNet-B3 (fine-tuned) β local texture / spatial artifacts (1536-d)
- FFT-CNN (custom, on the Fourier spectrum) β frequency-domain fingerprints left by generator upsampling (512-d)
The three feature vectors (2816-d) are concatenated and passed to a fusion MLP
(2816 β 512 β 128 β 2).
- Input: RGB image, resized to 224Γ224
- Output: 2 logits β softmax. Label map:
0 = real,1 = ai - Params: ~317.9M total, ~13.9M trainable
Full architecture write-up (per-branch design rationale): MODEL_ARCHITECTURE.md.
How to use
This is a custom nn.Module, so AutoModel / pipeline() will not load it.
Download model.py + best_model.pt from this repo and:
import torch, torch.nn.functional as F
import torchvision.transforms as T
from PIL import Image
from huggingface_hub import hf_hub_download
from model import load_detector # model.py from this repo
repo = "YOUR-USERNAME/ai-detector"
ckpt = hf_hub_download(repo, "best_model.pt")
device = "cuda" if torch.cuda.is_available() else "cpu"
model, cfg = load_detector(ckpt, device=device)
tf = T.Compose([T.Resize((224, 224)), T.ToTensor()])
img = Image.open("example.jpg").convert("RGB")
with torch.no_grad():
probs = F.softmax(model(tf(img).unsqueeze(0).to(device)), dim=1)[0]
print({"real": float(probs[0]), "ai": float(probs[1])})
Training data
- AI images: Tiny-GenImage β 8 generators (Midjourney, SD v1.4/v1.5, SDXL, DALLΒ·E 2/3, VQDM, Wukong)
- Real images: Tiny-ImageNet
- Balanced 14,000 real / 14,000 AI; 85/15 train/val split; images resized to 224Γ224.
Results
In-distribution: 98.5% accuracy / 0.985 macro-F1 (98% real recall, 99% AI recall), n=4,200, 3 epochs.
Out-of-distribution testing revealed an asymmetric generalization profile β both classes have a genuine strength, with one narrow gap:
| Probe | Class | Condition | Result | Provenance |
|---|---|---|---|---|
| AIGC-Detection-Benchmark, standard (n=900, 17 gen.) | AI | full/variable native res, no resize trick | 94.0% recall | recovered chat transcript |
| AIGC-Detection-Benchmark, standard (n=900, 17 gen.) | real | full/variable native res, no resize trick | 8.0% recall | recovered chat transcript |
Tiny-ImageNet valid, held-out (n=200) |
real | native 64Γ64 | 97.0% recall | saved classification report |
| AIGC-Detection-Benchmark, forced to 64Γ64 (n=900) | AI | forced downscale-then-upscale | 68.0% recall | saved classification report |
| AIGC-Detection-Benchmark, forced to 64Γ64 (n=900) | real | forced downscale-then-upscale | 36.0% recall | saved classification report |
AI-image detection already generalizes well to full-resolution external
images (94.0% recall, no resize trick needed). Real-image detection
performs extremely well within its native 64Γ64 resolution domain (97.0%
recall on a held-out split never seen in training) β essentially matching
in-distribution performance. The gap is narrow: that native-resolution
strength doesn't yet extend to full-resolution photos (8.0%). Forcing inputs
to 64Γ64 before inference is not a fix: it partially helps real
(8.0%β36.0%) but hurts AI (94.0%β68.0%), a net loss. Full methodology,
mechanism, and per-generator breakdown: see report.pdf in this repo, or the
technical report on GitHub.
Limitations & intended use
- Real-image detection does not generalize to full-resolution photos
(8.0% recall) β the central limitation. Works well on natively low-res real
photos (97.0%), not on full-resolution ones. Traces to the real-image
training source (Tiny-ImageNet) being natively 64Γ64 β not an architectural
flaw, since the AI branch does not show the same failure. See
report.pdfΒ§6. - Do not force inputs to 64Γ64 before inference as a workaround β it trades a partial real-recall gain (8.0%β36.0%) for a larger AI-recall loss (94.0%β68.0%), net negative.
- One data point (8.0% real / 94.0% AI recall) is recovered from a
preserved chat transcript, not the notebook's saved output β the
evaluation cell was edited in place to produce the forced-resize version,
overwriting the original run's output. See
report.pdfΒ§5.1/Β§7. - Distribution note: real and AI images came from different source datasets, so in-distribution numbers may be optimistic.
- Cross-generator generalization has not been properly validated β an
intended train/test generator exclusion in the OOD benchmark code was not
actually enforced (see
report.pdfΒ§5.3). Treat any apparent generalization in this report as provisional. - Intended for research / educational use in synthetic-media detection, not as a sole authority for high-stakes decisions.
Training config
AdamW (per-branch LRs: CLIP adapter 1e-5, EfficientNet 3e-5, FFT 1e-4, head 1e-4), CosineAnnealingLR, grad-clip 1.0, weight decay 1e-4, CrossEntropy loss, augmentation (flip, color jitter, grayscale, blur).

