ms eff gcvit deepfake b5 kodf

提供商KoreaPeter
分类video-classification
许可证mit
下载量16.3K
星标0

简介

这是一个由 KoreaPeter 提供的视频分类模型,基于微软的 GCViT 架构并针对 Deepfake(深度伪造)检测进行了深度优化。该模型专注于识别视频中的面部篡改痕迹,能够有效区分真实拍摄与 AI 生成的虚假视频。对于需要构建内容审核系统或反欺诈应用的开发者来说,它提供了一个高性能的预训练基准,上手难度中等,可直接集成到 PyTorch 等深度学习流水线中,作为视频真伪校验的核心插件使用。

核心亮点

  • 基于 GCViT 架构,高效捕捉视频空间-时间特征
  • 专注于 Deepfake 检测,识别 AI 篡改视频
  • 采用 MIT 开源协议,商业集成灵活且无压力
  • 适用于内容审核、身份验证等反欺诈场景

使用方法

安装依赖
# 安装 Hugging Face transformers
pip install transformers torch
SDK 使用
# 使用 transformers 加载模型
from transformers import AutoModel, AutoTokenizer

model = AutoModel.from_pretrained("KoreaPeter/ms-eff-gcvit-deepfake-b5-kodf")
tokenizer = AutoTokenizer.from_pretrained("KoreaPeter/ms-eff-gcvit-deepfake-b5-kodf")

Hugging Face 下载

我们推荐使用命令行或者 Hugging Face Hub SDK 来进行模型的下载。

操作指引:在下载前,请先通过如下命令安装 huggingface_hub:

操作指引
pip install -U huggingface_hub

命令行下载

下载完整模型库

下载完整模型库
huggingface-cli download KoreaPeter/ms-eff-gcvit-deepfake-b5-kodf

下载单个文件到指定本地文件夹(以下载 config.json 到当前路径下 ./dir 目录为例)

下载单个文件到指定本地文件夹(以下载 config.json 到当前路径下 ./dir 目录为例)
huggingface-cli download KoreaPeter/ms-eff-gcvit-deepfake-b5-kodf config.json --local-dir ./dir

更多命令行下载选项,可参见官方文档

SDK 下载

SDK 下载
# 模型下载
from huggingface_hub import snapshot_download
model_dir = snapshot_download('KoreaPeter/ms-eff-gcvit-deepfake-b5-kodf')

Git 下载

请确保 lfs 已经被正确安装

Git 下载
git lfs install
git clone https://huggingface.co/KoreaPeter/ms-eff-gcvit-deepfake-b5-kodf

如果您希望跳过 lfs 大文件下载,可以使用如下命令

跳过 LFS
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/KoreaPeter/ms-eff-gcvit-deepfake-b5-kodf

模型文件托管在 Hugging Face Hub,使用 HF CLI / SDK / Git 直接下载,不经过本站。

PyTorch / Transformers 使用

安装 Transformers

安装 Transformers
pip install -U transformers torch

模型加载和推理

模型加载和推理
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained('KoreaPeter/ms-eff-gcvit-deepfake-b5-kodf')
tokenizer = AutoTokenizer.from_pretrained('KoreaPeter/ms-eff-gcvit-deepfake-b5-kodf')

完整文档

来源: HuggingFace

---
license: mit
datasets:

  • ILSVRC/imagenet-1k

metrics:
  • accuracy

  • roc_auc

base_model:
  • timm/tf_efficientnet_b5.ns_jft_in1k

pipeline_tag: video-classification
library_name: transformers
tags:
  • PyTorch

  • vision

  • DeepFake-Detection

  • DeepGuard

  • tf-efficientnet

  • global-context-vision-transformer

---

🚀 Multi Scale Efficient Global Context Vision Transformer

!Task
!Image Classification
!Video Classification

!FaceForensics++
!Celeb-DF(v2)-00C853?style=flat-square)
!KODF

<img src="./ms_eff_gcvit.JPG" width="900">

> 🔗 GitHub Repository: HanMoonSub/DeepGuard

> 🤗 Live demo: DeepFake Video Detection

> 🤗 Live demo: DeepFake Image Detection

> 🤗 Live demo: DeepFake Detection XAI

Multi-Scale Efficient Global Context Vision Transformer (MS-EffGCViT) is a hybrid CNN-ViT
architecture for deepfake detection. It fuses CNN-driven spatial inductive bias with
hierarchical global-context attention to catch both *local* artifacts (textures, blending seams)
and *global* artifacts (lighting, structural inconsistency).

A single architecture ships in two sizes and three domain-tuned checkpoints, working on both
static images and video at the frame level.

✨ Core Features

  • 🎞️ Frame-level — one model handles both images and videos (frame-level inference + aggregation).
  • 🌍 Cross-domain — robust on both East-Asian (KoDF) and Western (Celeb-DF-v2, FaceForensics++) faces.
  • ⚡🔥 Two variantsFast (b0) for real-time/edge, Pro (b5) for enterprise accuracy.
  • 🧩 timm-compatible — load via the timm interface or the deepguard package.

⚙️ Model Specifications

| Spec | Detail | |---|---| | Task | Binary deepfake detection (real / fake) | | Domain | Frame-level, spatial-domain | | Input | Image or video (face-cropped) | | Output | Sigmoid probability in [0, 1] — higher = more likely fake | | Backbone | EfficientNet (ImageNet-1K pretrained) | | Framework | PyTorch / timm |

🧬 Model Zoo

ms_eff_gcvit_b0 is Optimized for real-time inference and mobile deployment.

🔥 ms_eff_gcvit_b5 is Engineered for high-fidelity analysis and enterprise-grade accuracy.

| Config | ⚡ Fast (b0) | 🔥 Pro (b5) |
|---|---|---|
| Model name | ms_eff_gcvit_b0 | ms_eff_gcvit_b5 |
| Backbone | tf_efficientnet_b0.ns_jft_in1k | tf_efficientnet_b5.ns_jft_in1k |
| Resolution | 224×224 | 384×384 |
| Params (M) | 8.7 | 50.3 |
| FLOPs (G) | 0.87 | 13.64 |

📚 Dataset: KoDF (Korean Deepfake Dataset)

Large-scale Korean deepfake dataset provided by AI-Hub in 2020

featuring 400 Korean participants across 6 deepfake synthesis methods

  • [x] Number of Subjects: 400 participants
  • [x] Videos per Subject: 150+ videos
  • [x] Total Data Duration: 88.5 days
  • [x] Deepfake Model Variants: 6 types(DeepFaceLab, FaceSwap, FSGAN, FOMM, 3DMM, Wav2Lip)

| Metric | Original Data | Fake Data |
| ------ | ------------- | --------- |
| Total Videos | 62,166 | 175,776 |
| Average Video Length | 90+ second | 15+ second |
| Total Duration | 1,500+ hours | 625+ hours|
| Resolution | 1920 X 1080 | 1920 X 1080 |
| FPS | 30 FPS | 30 FPS |
| Total Frames | 162,000,000+ | -- |

📈 Test Evaluation

Trained and tested on the same dataset.

| Dataset | Variant | Accuracy | AUC | Log Loss |
|---|---|---|---|---|
| KoDF | ⚡ Fast | 0.9655 | 0.9792 | 0.1237 |
| KoDF | 🔥 Pro | 0.9792 | 0.9831 | 0.0692 |

📈 Cross-Dataset Evaluation (Trained on KoDF)

Generalization to unseen domains — trained on KoDF, evaluated on western-face datasets.

| Tested on | Variant | Accuracy | AUC | Log Loss |
|---|---|---|---|---|
| Celeb-DF-v2 | ⚡ Fast | 0.5579 | 0.4719 | 1.2605 |
| Celeb-DF-v2 | 🔥 Pro | 0.5946 | 0.5400 | 1.0078 |
| FaceForensics++ | ⚡ Fast | 0.4875 | 0.5341 | 1.6178 |
| FaceForensics++ | 🔥 Pro | 0.4525 | 0.5902 | 1.5321 |

🚀 Model Usage

python
pip install deepguard
from transformers import pipeline

🖼️ Image Classification

python
clf = pipeline(
    "image-classification",
    model="KoreaPeter/ms-eff-gcvit-deepfake-b5-kodf", 
    trust_remote_code=True,
)

── Basic Inference ───────────────────────────────────────────────

result = clf("face.jpg")

[{'label': 'fake', 'score': 0.9712}, {'label': 'real', 'score': 0.0288}]

── Custom Parameters ─────────────────────────────────────────────

result = clf( "face.jpg", margin_ratio=0.2, # Margin ratio around the detected face bbox (default: 0.2) conf_thres=0.5, # Confidence threshold for YOLO face detection (default: 0.5) min_face_ratio=0.01, # Minimum face-to-frame area ratio to process (default: 0.01) tta_hflip=0.0, # Probability of horizontal flip for TTA (default: 0.0) top_k=1, # Number of top labels to return (default: all) )

[{'label': 'fake', 'score': 0.9712}]

🎬 Video Classification

python
clf = pipeline(
    "video-classification",
    model="KoreaPeter/ms-eff-gcvit-deepfake-b5-kodf",
    trust_remote_code=True,
)

── Basic Inference ───────────────────────────────────────────────

result = clf("video.mp4")

[{'label': 'fake', 'score': 0.9634}, {'label': 'real', 'score': 0.0366}]

── Custom Parameters ─────────────────────────────────────────────

result = clf( "video.mp4", num_frames=20, # Number of frames to sample (default: 20) margin_ratio=0.2, # Margin ratio around the detected face bbox (default: 0.2) conf_thres=0.5, # Confidence threshold for YOLO face detection (default: 0.5) min_face_ratio=0.01, # Minimum face-to-frame area ratio to process (default: 0.01) tta_hflip=0.0, # Probability of horizontal flip for TTA (default: 0.0) agg_mode="conf", # Aggregation mode: 'conf' | 'mean' | 'vote' (default: 'conf') return_frame_scores=True, # Return per-frame scores (default: False) )

[{'label': 'fake', 'score': 0.9634},

{'label': 'real', 'score': 0.0366},

{'frame_scores': [0.97, 0.95, 0.98, ...], 'agg_mode': 'conf'}]

Deep Dive into Model

Part 1: CNN-based Patch Embedding for Spatial Inductive Bias

While traditional Vision Transformers (ViTs) utilize a Linear Projection for patch embedding, our proposed model adopts a CNN-based Patch Embedding module incorporating MBConvBlocks.

  • Injecting Inductive Bias : Standard ViTs often suffer from a lack of inherent spatial inductive bias, typically necessitating massive datasets to learn fundamental visual structures from scratch. In contrast, our CNN-based module leverages overlapping receptive fields to facilitate information sharing between neighboring patches. By explicitly injecting this spatial bias into the architecture, the model achieves more stable and accelerated convergence during the training process.

Part 2: Long-Short Range Spatial Interaction

We utilizes two distinct types of self-attention to capture both long-range and short-range information across feature maps.

<img src="./window_attention.JPG" width="900">

  • Local Window Attention: this model efficiently captures local textures and precise spatial details while maintaining linear c