Model Card
Transparency report for the core detection engine powering DeepShield AI.
Model Overview
Model NameMAPF-Lite (Method-Aware Prompting and Fusion - Lite)
Primary Objective
Multi-modal classification of digital video content to detect synthetic manipulations (deepfakes) across visual and auditory domains.
Architecture Details
MAPF-Lite leverages a highly optimized, multi-stream architecture designed for CPU-efficient inference:
- Visual Backbone: Frozen CLIP ViT-B/32. Extracts semantic visual features.
- Audio Backbone: Frozen Whisper Base. Extracts semantic auditory features.
- Forgery Signature Gate: Trainable module utilizing Spatial Rich Model (SRM) filters and 2D Discrete Cosine Transform (DCT) to capture low-level artifact traces (noise inconsistencies, blending boundaries).
- Lite Prompt Generator: Dynamically generates manipulation-method-aware prompts to guide the model's attention.
- Fusion Mechanism: Frame-Level Fusion with Cross-Modal Feature Matching (CMFM) to align audio and visual representations and detect temporal desynchronization.
Known Limitations
While highly accurate, MAPF-Lite has known bounds of operational reliability:
- Demographic Variance: Performance may exhibit minor variance across different ethnic groups, skin tones, or ages, largely inherited from the pre-training distributions of the backbone models.
- Environmental Factors: Extreme lighting conditions, significant physical occlusions, or heavy makeup can degrade visual feature extraction.
- Compression: Severe video compression artifacts (e.g., highly compressed social media re-uploads) can destroy the low-level noise patterns relied upon by the Forgery Signature Gate.
- Legal Standing: Output scores are probabilistic. The model is intended as a triage and investigative tool, not as sole, definitive evidence in legal proceedings.
Detection Categories
- Authentic (Real)
- GAN Generated
- Diffusion Model
- Face Swap
- Audio Lip-Sync
Training Data
Trained and validated primarily on:
FakeAVCeleb v1.2
A comprehensive dataset featuring diverse, high-quality multi-modal deepfakes generated via multiple methods.
Optimization Profile
Quantization: Dynamic INT8
Target Hardware: General-purpose CPUs
Efficiency: Reduced memory footprint suitable for edge and standard cloud deployment without dedicated GPU acceleration.