ms eff gcvit deepfake b0 ff plus plus

提供商KoreaPeter
分类video-classification
许可证mit
下载量13.7K
星标0

简介

这是一个基于 GCViT 架构的轻量化视频分类模型,专门用于 Deepfake(深度伪造)检测。它通过高效的视觉 Transformer 机制捕捉视频帧间的细微异常,旨在识别 AI 换脸或篡改的痕迹。对于开发者而言,该模型在保持低计算开销的同时提供了较高的检测精度,非常适合集成到需要实时性或端侧部署的真伪验证场景中。相比于通用视觉模型,它更像是一个精准的“数字鉴伪插件”,上手难度较低,可直接用于视频质量分析或安全审核流水线。

核心亮点

  • 专注于 Deepfake 检测,精准识别视频伪造
  • 基于 GCViT 架构,兼顾推理速度与检测精度
  • 轻量化设计,适合端侧部署或实时审核场景
  • 采用 MIT 协议,对开发者极其友好且易于集成

使用方法

安装依赖
# 安装 Hugging Face transformers
pip install transformers torch
SDK 使用
# 使用 transformers 加载模型
from transformers import AutoModel, AutoTokenizer

model = AutoModel.from_pretrained("KoreaPeter/ms-eff-gcvit-deepfake-b0-ff-plus-plus")
tokenizer = AutoTokenizer.from_pretrained("KoreaPeter/ms-eff-gcvit-deepfake-b0-ff-plus-plus")

Hugging Face 下载

我们推荐使用命令行或者 Hugging Face Hub SDK 来进行模型的下载。

操作指引:在下载前,请先通过如下命令安装 huggingface_hub:

操作指引
pip install -U huggingface_hub

命令行下载

下载完整模型库

下载完整模型库
huggingface-cli download KoreaPeter/ms-eff-gcvit-deepfake-b0-ff-plus-plus

下载单个文件到指定本地文件夹(以下载 config.json 到当前路径下 ./dir 目录为例)

下载单个文件到指定本地文件夹(以下载 config.json 到当前路径下 ./dir 目录为例)
huggingface-cli download KoreaPeter/ms-eff-gcvit-deepfake-b0-ff-plus-plus config.json --local-dir ./dir

更多命令行下载选项,可参见官方文档

SDK 下载

SDK 下载
# 模型下载
from huggingface_hub import snapshot_download
model_dir = snapshot_download('KoreaPeter/ms-eff-gcvit-deepfake-b0-ff-plus-plus')

Git 下载

请确保 lfs 已经被正确安装

Git 下载
git lfs install
git clone https://huggingface.co/KoreaPeter/ms-eff-gcvit-deepfake-b0-ff-plus-plus

如果您希望跳过 lfs 大文件下载,可以使用如下命令

跳过 LFS
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/KoreaPeter/ms-eff-gcvit-deepfake-b0-ff-plus-plus

模型文件托管在 Hugging Face Hub,使用 HF CLI / SDK / Git 直接下载,不经过本站。

PyTorch / Transformers 使用

安装 Transformers

安装 Transformers
pip install -U transformers torch

模型加载和推理

模型加载和推理
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained('KoreaPeter/ms-eff-gcvit-deepfake-b0-ff-plus-plus')
tokenizer = AutoTokenizer.from_pretrained('KoreaPeter/ms-eff-gcvit-deepfake-b0-ff-plus-plus')

完整文档

来源: HuggingFace

---
license: mit
datasets:

  • ILSVRC/imagenet-1k

metrics:
  • accuracy

  • roc_auc

base_model:
  • timm/tf_efficientnet_b0.ns_jft_in1k

pipeline_tag: video-classification
library_name: transformers
tags:
  • PyTorch

  • vision

  • DeepFake-Detection

  • DeepGuard

  • tf-efficientnet

  • global-context-vision-transformer

---

🚀 Multi Scale Efficient Global Context Vision Transformer

!Task
!Image Classification
!Video Classification

!FaceForensics++
!Celeb-DF(v2)-00C853?style=flat-square)
!KODF

<img src="./ms_eff_gcvit.JPG" width="900">

> 🔗 GitHub Repository: HanMoonSub/DeepGuard

> 🤗 Live demo: DeepFake Video Detection

> 🤗 Live demo: DeepFake Image Detection

> 🤗 Live demo: DeepFake Detection XAI

Multi-Scale Efficient Global Context Vision Transformer (MS-EffGCViT) is a hybrid CNN-ViT
architecture for deepfake detection. It fuses CNN-driven spatial inductive bias with
hierarchical global-context attention to catch both *local* artifacts (textures, blending seams)
and *global* artifacts (lighting, structural inconsistency).

A single architecture ships in two sizes and three domain-tuned checkpoints, working on both
static images and video at the frame level.

✨ Core Features

  • 🎞️ Frame-level — one model handles both images and videos (frame-level inference + aggregation).
  • 🌍 Cross-domain — robust on both East-Asian (KoDF) and Western (Celeb-DF-v2, FaceForensics++) faces.
  • ⚡🔥 Two variantsFast (b0) for real-time/edge, Pro (b5) for enterprise accuracy.
  • 🧩 timm-compatible — load via the timm interface or the deepguard package.

⚙️ Model Specifications

| Spec | Detail | |---|---| | Task | Binary deepfake detection (real / fake) | | Domain | Frame-level, spatial-domain | | Input | Image or video (face-cropped) | | Output | Sigmoid probability in [0, 1] — higher = more likely fake | | Backbone | EfficientNet (ImageNet-1K pretrained) | | Framework | PyTorch / timm |

🧬 Model Zoo

ms_eff_gcvit_b0 is Optimized for real-time inference and mobile deployment.

🔥 ms_eff_gcvit_b5 is Engineered for high-fidelity analysis and enterprise-grade accuracy.

| Config | ⚡ Fast (b0) | 🔥 Pro (b5) |
|---|---|---|
| Model name | ms_eff_gcvit_b0 | ms_eff_gcvit_b5 |
| Backbone | tf_efficientnet_b0.ns_jft_in1k | tf_efficientnet_b5.ns_jft_in1k |
| Resolution | 224×224 | 384×384 |
| Params (M) | 8.7 | 50.3 |
| FLOPs (G) | 0.87 | 13.64 |

📚 Dataset: FaceForensics++

Learning to detect manipulated facial images [[Paper]](https://arxiv.org/abs/1901.08971) [[Download]](https://github.com/ondyari/FaceForensics),
featuring 1,000 original YouTube videos manipulated by 5 face forgery methods.

  • [x] Source: 1,000 original videos (from 977 YouTube videos)
  • [x] Manipulation Methods: 5 types (Deepfakes, Face2Face, FaceSwap, FaceShifter, NeuralTextures)
  • [x] Faces: Trackable, mostly frontal, no occlusion

| Source | Real/Fake | Videos | Description |
| ------ | --------- | ------ | ----------- |
| Deepfakes | !Fake | 1,000 | Autoencoder-based face replacement |
| Face2Face | !Fake | 1,000 | Expression transfer (reenactment) |
| FaceSwap | !Fake | 1,000 | Graphics-based face replacement |
| FaceShifter | !Fake | 1,000 | High-fidelity swap with occlusion handling |
| NeuralTextures | !Fake | 1,000 | Neural-texture-based reenactment |
| Original | !Real | 1,000 | Unaltered authentic YouTube videos |

> 📎 Available at GitHub or Kaggle

📈 Test Evaluation

Trained and tested on the same dataset.

| Dataset | Variant | Accuracy | AUC | Log Loss |
|---|---|---|---|---|
| FaceForensics++ | ⚡ Fast | 0.9808 | 0.9969 | 0.0637 |
| FaceForensics++ | 🔥 Pro | 0.9850 | 0.9974 | 0.0492 |

📈 Cross-Dataset Evaluation (Trained on FaceForensics++)

Generalization to unseen domains — trained on FaceForensics++

| Tested on | Variant | Accuracy | AUC | Log Loss |
|---|---|---|---|---|
| Celeb-DF-v2 | ⚡ Fast | 0.7259 | 0.6999 | 0.6794 |
| Celeb-DF-v2 | 🔥 Pro | 0.7722 | 0.7309 | 0.6657 |
| KoDF | ⚡ Fast | 0.7544 | 0.8620 | 0.8829 |
| KoDF | 🔥 Pro | 0.7695 | 0.8821 | 0.7635 |

🚀 Model Usage

python
pip install deepguard
from transformers import pipeline

🖼️ Image Classification

python
clf = pipeline(
    "image-classification",
    model="KoreaPeter/ms-eff-gcvit-deepfake-b0-ff-plus-plus", 
    trust_remote_code=True,
)

── Basic Inference ───────────────────────────────────────────────

result = clf("face.jpg")

[{'label': 'fake', 'score': 0.9712}, {'label': 'real', 'score': 0.0288}]

── Custom Parameters ─────────────────────────────────────────────

result = clf( "face.jpg", margin_ratio=0.2, # Margin ratio around the detected face bbox (default: 0.2) conf_thres=0.5, # Confidence threshold for YOLO face detection (default: 0.5) min_face_ratio=0.01, # Minimum face-to-frame area ratio to process (default: 0.01) tta_hflip=0.0, # Probability of horizontal flip for TTA (default: 0.0) top_k=1, # Number of top labels to return (default: all) )

[{'label': 'fake', 'score': 0.9712}]

🎬 Video Classification

python
clf = pipeline(
    "video-classification",
    model="KoreaPeter/ms-eff-gcvit-deepfake-b0-ff-plus-plus", 
    trust_remote_code=True,
)

── Basic Inference ───────────────────────────────────────────────

result = clf("video.mp4")

[{'label': 'fake', 'score': 0.9634}, {'label': 'real', 'score': 0.0366}]

── Custom Parameters ─────────────────────────────────────────────

result = clf( "video.mp4", num_frames=20, # Number of frames to sample (default: 20) margin_ratio=0.2, # Margin ratio around the detected face bbox (default: 0.2) conf_thres=0.5, # Confidence threshold for YOLO face detection (default: 0.5) min_face_ratio=0.01, # Minimum face-to-frame area ratio to process (default: 0.01) tta_hflip=0.0, # Probability of horizontal flip for TTA (default: 0.0) agg_mode="conf", # Aggregation mode: 'conf' | 'mean' | 'vote' (default: 'conf') return_frame_scores=True, # Return per-frame scores (default: False) )

[{'label': 'fake', 'score': 0.9634},

{'label': 'real', 'score': 0.0366},

{'frame_scores': [0.97, 0.95, 0.98, ...], 'agg_mode': 'conf'}]

Deep Dive into Model

Part 1: CNN-based Patch Embedding for Spatial Inductive Bias

While traditional Vision Transformers (ViTs) utilize a Linear Projection for patch embedding, our proposed model adopts a CNN-based Patch Embedding module incorporating MBConvBlocks.

  • Injecting Inductive Bias : Standard ViTs often suffer from a lack of inherent spatial inductive bias, typically