ms eff gcvit deepfake b5 celeb df v2

提供商KoreaPeter
分类video-classification
许可证mit
下载量13.4K
星标0

简介

这是一个基于微软 GCViT 架构的深度伪造(Deepfake)检测模型,专门针对 Celeb-DF v2 数据集进行了优化。它能够通过分析视频帧中的细微异常,识别出由 AI 生成或篡改的人脸视频。对于需要构建自动化内容审核系统、验证视频真实性的开发者来说,这是一个高效的分类工具。由于其基于视觉 Transformer 架构,对细节捕捉能力较强,上手难度较低,可直接集成到视频预处理流水线中,作为反欺诈或真伪鉴定的核心模块。

核心亮点

  • 基于 GCViT 架构,精准识别人脸篡改痕迹
  • 针对 Celeb-DF v2 优化,检测鲁棒性强
  • 适用于自动化视频审核与真伪鉴定场景
  • MIT 协议开源,方便开发者快速集成部署

使用方法

安装依赖
# 安装 Hugging Face transformers
pip install transformers torch
SDK 使用
# 使用 transformers 加载模型
from transformers import AutoModel, AutoTokenizer

model = AutoModel.from_pretrained("KoreaPeter/ms-eff-gcvit-deepfake-b5-celeb-df-v2")
tokenizer = AutoTokenizer.from_pretrained("KoreaPeter/ms-eff-gcvit-deepfake-b5-celeb-df-v2")

Hugging Face 下载

我们推荐使用命令行或者 Hugging Face Hub SDK 来进行模型的下载。

操作指引:在下载前,请先通过如下命令安装 huggingface_hub:

操作指引
pip install -U huggingface_hub

命令行下载

下载完整模型库

下载完整模型库
huggingface-cli download KoreaPeter/ms-eff-gcvit-deepfake-b5-celeb-df-v2

下载单个文件到指定本地文件夹(以下载 config.json 到当前路径下 ./dir 目录为例)

下载单个文件到指定本地文件夹(以下载 config.json 到当前路径下 ./dir 目录为例)
huggingface-cli download KoreaPeter/ms-eff-gcvit-deepfake-b5-celeb-df-v2 config.json --local-dir ./dir

更多命令行下载选项,可参见官方文档

SDK 下载

SDK 下载
# 模型下载
from huggingface_hub import snapshot_download
model_dir = snapshot_download('KoreaPeter/ms-eff-gcvit-deepfake-b5-celeb-df-v2')

Git 下载

请确保 lfs 已经被正确安装

Git 下载
git lfs install
git clone https://huggingface.co/KoreaPeter/ms-eff-gcvit-deepfake-b5-celeb-df-v2

如果您希望跳过 lfs 大文件下载,可以使用如下命令

跳过 LFS
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/KoreaPeter/ms-eff-gcvit-deepfake-b5-celeb-df-v2

模型文件托管在 Hugging Face Hub,使用 HF CLI / SDK / Git 直接下载,不经过本站。

PyTorch / Transformers 使用

安装 Transformers

安装 Transformers
pip install -U transformers torch

模型加载和推理

模型加载和推理
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained('KoreaPeter/ms-eff-gcvit-deepfake-b5-celeb-df-v2')
tokenizer = AutoTokenizer.from_pretrained('KoreaPeter/ms-eff-gcvit-deepfake-b5-celeb-df-v2')

完整文档

来源: HuggingFace

---
license: mit
datasets:

  • ILSVRC/imagenet-1k

metrics:
  • accuracy

  • roc_auc

base_model:
  • timm/tf_efficientnet_b5.ns_jft_in1k

pipeline_tag: video-classification
library_name: transformers
tags:
  • PyTorch

  • vision

  • DeepFake-Detection

  • DeepGuard

  • tf-efficientnet

  • global-context-vision-transformer

---

🚀 Multi Scale Efficient Global Context Vision Transformer

!Task
!Image Classification
!Video Classification

!FaceForensics++
!Celeb-DF(v2)-00C853?style=flat-square)
!KODF

<img src="./ms_eff_gcvit.JPG" width="900">

> 🔗 GitHub Repository: HanMoonSub/DeepGuard

> 🤗 Live demo: DeepFake Video Detection

> 🤗 Live demo: DeepFake Image Detection

> 🤗 Live demo: DeepFake Detection XAI

Multi-Scale Efficient Global Context Vision Transformer (MS-EffGCViT) is a hybrid CNN-ViT
architecture for deepfake detection. It fuses CNN-driven spatial inductive bias with
hierarchical global-context attention to catch both *local* artifacts (textures, blending seams)
and *global* artifacts (lighting, structural inconsistency).

A single architecture ships in two sizes and three domain-tuned checkpoints, working on both
static images and video at the frame level.

✨ Core Features

  • 🎞️ Frame-level — one model handles both images and videos (frame-level inference + aggregation).
  • 🌍 Cross-domain — robust on both East-Asian (KoDF) and Western (Celeb-DF-v2, FaceForensics++) faces.
  • ⚡🔥 Two variantsFast (b0) for real-time/edge, Pro (b5) for enterprise accuracy.
  • 🧩 timm-compatible — load via the timm interface or the deepguard package.

⚙️ Model Specifications

| Spec | Detail | |---|---| | Task | Binary deepfake detection (real / fake) | | Domain | Frame-level, spatial-domain | | Input | Image or video (face-cropped) | | Output | Sigmoid probability in [0, 1] — higher = more likely fake | | Backbone | EfficientNet (ImageNet-1K pretrained) | | Framework | PyTorch / timm |

🧬 Model Zoo

ms_eff_gcvit_b0 is Optimized for real-time inference and mobile deployment.

🔥 ms_eff_gcvit_b5 is Engineered for high-fidelity analysis and enterprise-grade accuracy.

| Config | ⚡ Fast (b0) | 🔥 Pro (b5) |
|---|---|---|
| Model name | ms_eff_gcvit_b0 | ms_eff_gcvit_b5 |
| Backbone | tf_efficientnet_b0.ns_jft_in1k | tf_efficientnet_b5.ns_jft_in1k |
| Resolution | 224×224 | 384×384 |
| Params (M) | 8.7 | 50.3 |
| FLOPs (G) | 0.87 | 13.64 |

📚 Dataset: Celeb-DF-v2

A large-scale challenging dataset for deepfake forensics [[Paper]](https://openaccess.thecvf.com/content_CVPR_2020/papers/Li_Celeb-DF_A_Large-Scale_Challenging_Dataset_for_DeepFake_Forensics_CVPR_2020_paper.pdf) [[Download]](https://github.com/yuezunli/celeb-deepfakeforensics/tree/master),
featuring 590 YouTube celebrity videos with diverse ages, ethnic groups, and genders.

  • [x] Source: 590 original YouTube videos (celebrities)
  • [x] Synthesis: 5,639 deepfake videos generated from real videos
  • [x] Subjects: Diverse ages, ethnicities, and genders

| Source | Real/Fake | Videos | Description |
| ------ | --------- | ------ | ----------- |
| celeb-real | !Real | 590 | Celebrity videos from YouTube |
| youtube-real | !Real | 300 | Additional YouTube videos |
| celeb-synthesis | !Fake | 5,639 | Synthesized from celeb-real |

> 📎 Available at GitHub or Kaggle

📈 Test Evaluation

Trained and tested on the same dataset.

<img src="./celeb_df_v2_gcvit.png" width="900">

| Dataset | Variant | Accuracy | AUC | Log Loss |
|---|---|---|---|---|
| Celeb-DF-v2 | ⚡ Fast | 0.9842 | 0.9965 | 0.0283 |
| Celeb-DF-v2 | 🔥 Pro | 0.9981 | 0.9984 | 0.0089 |

📈 Cross-Dataset Evaluation (Trained on Celeb DF(v2))

Generalization to unseen domains — trained on Celeb DF(v2)

| Tested on | Variant | Accuracy | AUC | Log Loss |
|---|---|---|---|---|
| KoDF | ⚡ Fast | 0.4935 | 0.7258 | 1.3459 |
| KoDF | 🔥 Pro | 0.4832 | 0.7160 | 1.4897 |
| FaceForensics++ | ⚡ Fast | 0.5492 | 0.7301 | 1.0556 |
| FaceForensics++ | 🔥 Pro | 0.5825 | 0.7307 | 0.8897 |


🚀 Model Usage

python
pip install deepguard
from transformers import pipeline

🖼️ Image Classification

python
clf = pipeline(
    "image-classification",
    model="KoreaPeter/ms-eff-gcvit-deepfake-b5-celeb-df-v2", 
    trust_remote_code=True,
)

── Basic Inference ───────────────────────────────────────────────

result = clf("face.jpg")

[{'label': 'fake', 'score': 0.9712}, {'label': 'real', 'score': 0.0288}]

── Custom Parameters ─────────────────────────────────────────────

result = clf( "face.jpg", margin_ratio=0.2, # Margin ratio around the detected face bbox (default: 0.2) conf_thres=0.5, # Confidence threshold for YOLO face detection (default: 0.5) min_face_ratio=0.01, # Minimum face-to-frame area ratio to process (default: 0.01) tta_hflip=0.0, # Probability of horizontal flip for TTA (default: 0.0) top_k=1, # Number of top labels to return (default: all) )

[{'label': 'fake', 'score': 0.9712}]

🎬 Video Classification

python
clf = pipeline(
    "video-classification",
    model="KoreaPeter/ms-eff-gcvit-deepfake-b5-celeb-df-v2",  
    trust_remote_code=True,
)

── Basic Inference ───────────────────────────────────────────────

result = clf("video.mp4")

[{'label': 'fake', 'score': 0.9634}, {'label': 'real', 'score': 0.0366}]

── Custom Parameters ─────────────────────────────────────────────

result = clf( "video.mp4", num_frames=20, # Number of frames to sample (default: 20) margin_ratio=0.2, # Margin ratio around the detected face bbox (default: 0.2) conf_thres=0.5, # Confidence threshold for YOLO face detection (default: 0.5) min_face_ratio=0.01, # Minimum face-to-frame area ratio to process (default: 0.01) tta_hflip=0.0, # Probability of horizontal flip for TTA (default: 0.0) agg_mode="conf", # Aggregation mode: 'conf' | 'mean' | 'vote' (default: 'conf') return_frame_scores=True, # Return per-frame scores (default: False) )

[{'label': 'fake', 'score': 0.9634},

{'label': 'real', 'score': 0.0366},

{'frame_scores': [0.97, 0.95, 0.98, ...], 'agg_mode': 'conf'}]

Deep Dive into Model

Part 1: CNN-based Patch Embedding for Spatial Inductive Bias

While traditional Vision Transformers (ViTs) utilize a Linear Projection for patch embedding, our proposed model adopts a CNN-based Patch Embedding module incorporating MBConvBlocks.

  • Injecting Inductive Bias : Standard ViTs often suffer from a lack of inherent spatial inductive bias, typically necessitating massive datasets to learn fundamental visual structures from scratch. In contrast, our CNN-based module leverages overlapping receptive fields to facilitate information sharing between neighboring patches.