ms eff gcvit deepfake b0 ff plus plus
Overview
Highlights
- Optimized GCViT architecture for efficient video classification
- Specialized in detecting synthetic facial manipulations
- Low-latency inference suitable for real-time moderation
- Permissive MIT license for flexible commercial integration
- High precision in identifying deepfake temporal artifacts
Usage
# Install Hugging Face transformers
pip install transformers torch
# Load model with transformers
from transformers import AutoModel, AutoTokenizer
model = AutoModel.from_pretrained("KoreaPeter/ms-eff-gcvit-deepfake-b0-ff-plus-plus")
tokenizer = AutoTokenizer.from_pretrained("KoreaPeter/ms-eff-gcvit-deepfake-b0-ff-plus-plus")
Hugging Face Download
We recommend downloading the model via the Hugging Face CLI or Hub SDK.
Guidance:Before downloading, install huggingface_hub with:
pip install -U huggingface_hub
CLI Download
Download the full repository
huggingface-cli download KoreaPeter/ms-eff-gcvit-deepfake-b0-ff-plus-plus
Download a single file to a local folder (e.g. config.json into ./dir)
huggingface-cli download KoreaPeter/ms-eff-gcvit-deepfake-b0-ff-plus-plus config.json --local-dir ./dir
See the official docs for more CLI options
SDK Download
# 模型下载
from huggingface_hub import snapshot_download
model_dir = snapshot_download('KoreaPeter/ms-eff-gcvit-deepfake-b0-ff-plus-plus')
Git Download
Make sure git-lfs is installed first
git lfs install
git clone https://huggingface.co/KoreaPeter/ms-eff-gcvit-deepfake-b0-ff-plus-plus
To skip LFS large-file downloads, use:
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/KoreaPeter/ms-eff-gcvit-deepfake-b0-ff-plus-plus
Model files are hosted on the Hugging Face Hub — download directly via HF CLI / SDK / Git, not through this site.
PyTorch / Transformers Usage
Install Transformers
pip install -U transformers torch
Load the model and run inference
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained('KoreaPeter/ms-eff-gcvit-deepfake-b0-ff-plus-plus')
tokenizer = AutoTokenizer.from_pretrained('KoreaPeter/ms-eff-gcvit-deepfake-b0-ff-plus-plus')
Full Documentation
---
license: mit
datasets:
- ILSVRC/imagenet-1k
metrics:
- accuracy
- roc_auc
base_model:
- timm/tf_efficientnet_b0.ns_jft_in1k
pipeline_tag: video-classification
library_name: transformers
tags:
- PyTorch
- vision
- DeepFake-Detection
- DeepGuard
- tf-efficientnet
- global-context-vision-transformer
---
🚀 Multi Scale Efficient Global Context Vision Transformer
!Task
!Image Classification
!Video Classification
!FaceForensics++
!Celeb-DF(v2)-00C853?style=flat-square)
!KODF
<img src="./ms_eff_gcvit.JPG" width="900">
> 🔗 GitHub Repository: HanMoonSub/DeepGuard
> 🤗 Live demo: DeepFake Video Detection
> 🤗 Live demo: DeepFake Image Detection
> 🤗 Live demo: DeepFake Detection XAI
Multi-Scale Efficient Global Context Vision Transformer (MS-EffGCViT) is a hybrid CNN-ViT
architecture for deepfake detection. It fuses CNN-driven spatial inductive bias with
hierarchical global-context attention to catch both *local* artifacts (textures, blending seams)
and *global* artifacts (lighting, structural inconsistency).
A single architecture ships in two sizes and three domain-tuned checkpoints, working on both
static images and video at the frame level.
✨ Core Features
- 🎞️ Frame-level — one model handles both images and videos (frame-level inference + aggregation).
- 🌍 Cross-domain — robust on both East-Asian (KoDF) and Western (Celeb-DF-v2, FaceForensics++) faces.
- ⚡🔥 Two variants — Fast (b0) for real-time/edge, Pro (b5) for enterprise accuracy.
- 🧩 timm-compatible — load via the
timminterface or thedeepguardpackage.
⚙️ Model Specifications
| Spec | Detail | |---|---| | Task | Binary deepfake detection (real / fake) | | Domain | Frame-level, spatial-domain | | Input | Image or video (face-cropped) | | Output | Sigmoid probability in[0, 1] — higher = more likely fake |
| Backbone | EfficientNet (ImageNet-1K pretrained) |
| Framework | PyTorch / timm |
🧬 Model Zoo
⚡ ms_eff_gcvit_b0 is Optimized for real-time inference and mobile deployment.
🔥 ms_eff_gcvit_b5 is Engineered for high-fidelity analysis and enterprise-grade accuracy.
| Config | ⚡ Fast (b0) | 🔥 Pro (b5) |
|---|---|---|
| Model name | ms_eff_gcvit_b0 | ms_eff_gcvit_b5 |
| Backbone | tf_efficientnet_b0.ns_jft_in1k | tf_efficientnet_b5.ns_jft_in1k |
| Resolution | 224×224 | 384×384 |
| Params (M) | 8.7 | 50.3 |
| FLOPs (G) | 0.87 | 13.64 |
📚 Dataset: FaceForensics++
Learning to detect manipulated facial images [[Paper]](https://arxiv.org/abs/1901.08971) [[Download]](https://github.com/ondyari/FaceForensics),
featuring 1,000 original YouTube videos manipulated by 5 face forgery methods.
- [x] Source: 1,000 original videos (from 977 YouTube videos)
- [x] Manipulation Methods: 5 types (Deepfakes, Face2Face, FaceSwap, FaceShifter, NeuralTextures)
- [x] Faces: Trackable, mostly frontal, no occlusion
| Source | Real/Fake | Videos | Description |
| ------ | --------- | ------ | ----------- |
| Deepfakes | !Fake | 1,000 | Autoencoder-based face replacement |
| Face2Face | !Fake | 1,000 | Expression transfer (reenactment) |
| FaceSwap | !Fake | 1,000 | Graphics-based face replacement |
| FaceShifter | !Fake | 1,000 | High-fidelity swap with occlusion handling |
| NeuralTextures | !Fake | 1,000 | Neural-texture-based reenactment |
| Original | !Real | 1,000 | Unaltered authentic YouTube videos |
> 📎 Available at GitHub or Kaggle
📈 Test Evaluation
Trained and tested on the same dataset.
| Dataset | Variant | Accuracy | AUC | Log Loss |
|---|---|---|---|---|
| FaceForensics++ | ⚡ Fast | 0.9808 | 0.9969 | 0.0637 |
| FaceForensics++ | 🔥 Pro | 0.9850 | 0.9974 | 0.0492 |
📈 Cross-Dataset Evaluation (Trained on FaceForensics++)
Generalization to unseen domains — trained on FaceForensics++
| Tested on | Variant | Accuracy | AUC | Log Loss |
|---|---|---|---|---|
| Celeb-DF-v2 | ⚡ Fast | 0.7259 | 0.6999 | 0.6794 |
| Celeb-DF-v2 | 🔥 Pro | 0.7722 | 0.7309 | 0.6657 |
| KoDF | ⚡ Fast | 0.7544 | 0.8620 | 0.8829 |
| KoDF | 🔥 Pro | 0.7695 | 0.8821 | 0.7635 |
🚀 Model Usage
pip install deepguard
from transformers import pipeline🖼️ Image Classification
clf = pipeline(
"image-classification",
model="KoreaPeter/ms-eff-gcvit-deepfake-b0-ff-plus-plus",
trust_remote_code=True,
)
── Basic Inference ───────────────────────────────────────────────
result = clf("face.jpg")
[{'label': 'fake', 'score': 0.9712}, {'label': 'real', 'score': 0.0288}]
── Custom Parameters ─────────────────────────────────────────────
result = clf(
"face.jpg",
margin_ratio=0.2, # Margin ratio around the detected face bbox (default: 0.2)
conf_thres=0.5, # Confidence threshold for YOLO face detection (default: 0.5)
min_face_ratio=0.01, # Minimum face-to-frame area ratio to process (default: 0.01)
tta_hflip=0.0, # Probability of horizontal flip for TTA (default: 0.0)
top_k=1, # Number of top labels to return (default: all)
)
[{'label': 'fake', 'score': 0.9712}]
🎬 Video Classification
clf = pipeline(
"video-classification",
model="KoreaPeter/ms-eff-gcvit-deepfake-b0-ff-plus-plus",
trust_remote_code=True,
)
── Basic Inference ───────────────────────────────────────────────
result = clf("video.mp4")
[{'label': 'fake', 'score': 0.9634}, {'label': 'real', 'score': 0.0366}]
── Custom Parameters ─────────────────────────────────────────────
result = clf(
"video.mp4",
num_frames=20, # Number of frames to sample (default: 20)
margin_ratio=0.2, # Margin ratio around the detected face bbox (default: 0.2)
conf_thres=0.5, # Confidence threshold for YOLO face detection (default: 0.5)
min_face_ratio=0.01, # Minimum face-to-frame area ratio to process (default: 0.01)
tta_hflip=0.0, # Probability of horizontal flip for TTA (default: 0.0)
agg_mode="conf", # Aggregation mode: 'conf' | 'mean' | 'vote' (default: 'conf')
return_frame_scores=True, # Return per-frame scores (default: False)
)
[{'label': 'fake', 'score': 0.9634},
{'label': 'real', 'score': 0.0366},
{'frame_scores': [0.97, 0.95, 0.98, ...], 'agg_mode': 'conf'}]
Deep Dive into Model
Part 1: CNN-based Patch Embedding for Spatial Inductive Bias
While traditional Vision Transformers (ViTs) utilize a Linear Projection for patch embedding, our proposed model adopts a CNN-based Patch Embedding module incorporating MBConvBlocks.
- Injecting Inductive Bias : Standard ViTs often suffer from a lack of inherent spatial inductive bias, typically