nsfw image detection

提供商Falconsai
分类image-classification
许可证apache-2.0
下载量4.6M
星标0

简介

这是一个由 Falconsai 提供的轻量级图像分类模型,专门用于识别图片中的非安全内容(NSFW)。对于需要构建社区产品、社交应用或内容管理系统的开发者来说,它提供了一个低门槛的自动化审核方案。该模型上手极其简单,无需复杂的预处理即可快速判断图片是否合规。相比于调用昂贵的商业 API,这款基于 Apache-2.0 协议的开源模型允许开发者在本地部署,在保证内容过滤效率的同时,能更好地兼顾用户隐私和推理成本。

核心亮点

  • 快速识别图片合规性,适用于自动化内容审核
  • 开源协议友好,支持本地部署以降低调用成本
  • 轻量化设计,推理速度快,集成难度极低
  • 有效过滤非安全内容,保障用户端视觉体验

使用方法

安装依赖
# 安装 Hugging Face transformers
pip install transformers torch
SDK 使用
# 使用 transformers 加载模型
from transformers import AutoModel, AutoTokenizer

model = AutoModel.from_pretrained("Falconsai/nsfw_image_detection")
tokenizer = AutoTokenizer.from_pretrained("Falconsai/nsfw_image_detection")

Hugging Face 下载

我们推荐使用命令行或者 Hugging Face Hub SDK 来进行模型的下载。

操作指引:在下载前,请先通过如下命令安装 huggingface_hub:

操作指引
pip install -U huggingface_hub

命令行下载

下载完整模型库

下载完整模型库
huggingface-cli download Falconsai/nsfw_image_detection

下载单个文件到指定本地文件夹(以下载 config.json 到当前路径下 ./dir 目录为例)

下载单个文件到指定本地文件夹(以下载 config.json 到当前路径下 ./dir 目录为例)
huggingface-cli download Falconsai/nsfw_image_detection config.json --local-dir ./dir

更多命令行下载选项,可参见官方文档

SDK 下载

SDK 下载
# 模型下载
from huggingface_hub import snapshot_download
model_dir = snapshot_download('Falconsai/nsfw_image_detection')

Git 下载

请确保 lfs 已经被正确安装

Git 下载
git lfs install
git clone https://huggingface.co/Falconsai/nsfw_image_detection

如果您希望跳过 lfs 大文件下载,可以使用如下命令

跳过 LFS
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/Falconsai/nsfw_image_detection

模型文件托管在 Hugging Face Hub,使用 HF CLI / SDK / Git 直接下载,不经过本站。

PyTorch / Transformers 使用

安装 Transformers

安装 Transformers
pip install -U transformers torch

模型加载和推理

模型加载和推理
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained('Falconsai/nsfw_image_detection')
tokenizer = AutoTokenizer.from_pretrained('Falconsai/nsfw_image_detection')

完整文档

来源: HuggingFace

---
license: apache-2.0
pipeline_tag: image-classification
---

Model Card: Fine-Tuned Vision Transformer (ViT) for NSFW Image Classification

Model Description

The Fine-Tuned Vision Transformer (ViT) is a variant of the transformer encoder architecture, similar to BERT, that has been adapted for image classification tasks. This specific model, named "google/vit-base-patch16-224-in21k," is pre-trained on a substantial collection of images in a supervised manner, leveraging the ImageNet-21k dataset. The images in the pre-training dataset are resized to a resolution of 224x224 pixels, making it suitable for a wide range of image recognition tasks.

During the training phase, meticulous attention was given to hyperparameter settings to ensure optimal model performance. The model was fine-tuned with a judiciously chosen batch size of 16. This choice not only balanced computational efficiency but also allowed for the model to effectively process and learn from a diverse array of images.

To facilitate this fine-tuning process, a learning rate of 5e-5 was employed. The learning rate serves as a critical tuning parameter that dictates the magnitude of adjustments made to the model's parameters during training. In this case, a learning rate of 5e-5 was selected to strike a harmonious balance between rapid convergence and steady optimization, resulting in a model that not only learns swiftly but also steadily refines its capabilities throughout the training process.

This training phase was executed using a proprietary dataset containing an extensive collection of 80,000 images, each characterized by a substantial degree of variability. The dataset was thoughtfully curated to include two distinct classes, namely "normal" and "nsfw." This diversity allowed the model to grasp nuanced visual patterns, equipping it with the competence to accurately differentiate between safe and explicit content.

The overarching objective of this meticulous training process was to impart the model with a deep understanding of visual cues, ensuring its robustness and competence in tackling the specific task of NSFW image classification. The result is a model that stands ready to contribute significantly to content safety and moderation, all while maintaining the highest standards of accuracy and reliability.

Intended Uses & Limitations

Intended Uses

  • NSFW Image Classification: The primary intended use of this model is for the classification of NSFW (Not Safe for Work) images. It has been fine-tuned for this purpose, making it suitable for filtering explicit or inappropriate content in various applications.

How to use

Here is how to use this model to classifiy an image based on 1 of 2 classes (normal,nsfw):
markdown
# Use a pipeline as a high-level helper
from PIL import Image
from transformers import pipeline

img = Image.open("<path_to_image_file>")
classifier = pipeline("image-classification", model="Falconsai/nsfw_image_detection")
classifier(img)

<hr>

`` markdown

Load model directly

import torch from PIL import Image from transformers import AutoModelForImageClassification, ViTImageProcessor

img = Image.open("<path_to_image_file>")
model = AutoModelForImageClassification.from_pretrained("Falconsai/nsfw_image_detection")
processor = ViTImageProcessor.from_pretrained('Falconsai/nsfw_image_detection')
with torch.no_grad():
inputs = processor(images=img, return_tensors="pt")
outputs = model(inputs)
logits = outputs.logits

predicted_label = logits.argmax(-1).item()
model.config.id2label[predicted_label]

code
<hr>
Run Yolo Version
markdown

import os
import matplotlib.pyplot as plt
from PIL import Image
import numpy as np
import onnxruntime as ort
import json # Added import for json

Predict using YOLOv9 model

def predict_with_yolov9(image_path, model_path, labels_path, input_size): """ Run inference using the converted YOLOv9 model on a single image.

Args:
image_path (str): Path to the input image file.
model_path (str): Path to the ONNX model file.
labels_path (str): Path to the JSON file containing class labels.
input_size (tuple): The expected input size (height, width) for the model.

Returns:
str: The predicted class label.
PIL.Image.Image: The original loaded image.
"""
def load_json(file_path):
with open(file_path, "r") as f:
return json.load(f)

# Load labels
labels = load_json(labels_path)

# Preprocess image
original_image = Image.open(image_path).convert("RGB")
image_resized = original_image.resize(input_size, Image.Resampling.BILINEAR)
image_np = np.array(image_resized, dtype=np.float32) / 255.0
image_np = np.transpose(image_np, (2, 0, 1)) # [C, H, W]
input_tensor = np.expand_dims(image_np, axis=0).astype(np.float32)

# Load YOLOv9 model
session = ort.InferenceSession(model_path)
input_name = session.get_inputs()[0].name
output_name = session.get_outputs()[0].name # Assuming classification output

# Run inference
outputs = session.run([output_name], {input_name: input_tensor})
predictions = outputs[0]

# Postprocess predictions (assuming classification output)
# Adapt this section if your model output is different (e.g., detection boxes)
predicted_index = np.argmax(predictions)
predicted_label = labels[str(predicted_index)] # Assumes labels are indexed by string numbers

return predicted_label, original_image

Display prediction for a single image

def display_single_prediction(image_path, model_path, labels_path, input_size): """ Predicts the class for a single image and displays the image with its prediction.

Args:
image_path (str): Path to the input image file.
model_path (str): Path to the ONNX model file.
labels_path (str): Path to the JSON file containing class labels.
input_size (tuple): The expected input size (height, width) for the model.
"""
try:
# Run prediction
prediction, img = predict_with_yolov9(image_path, model_path, labels_path, input_size)

# Display image and prediction
fig, ax = plt.subplots(1, 1, figsize=(8, 8)) # Create a single plot
ax.imshow(img)
ax.set_title(f"Prediction: {prediction}", fontsize=14)
ax.axis("off") # Hide axes ticks and labels

plt.tight_layout()
plt.show()

except FileNotFoundError:
print(f"Error: Image file not found at {image_path}")
except Exception as e:
print(f"An error occurred: {e}")

--- Main Execution ---

Paths and parameters - MODIFY THESE

single_image_path = "path/to/your/single_image.jpg" # <--- Replace with the actual path to your image file model_path = "path/to/your/yolov9_model.onnx" # <--- Replace with the actual path to your ONNX model labels_path = "path/to/your/labels.json" # <--- Replace with the actual path to your labels JSON file input_size = (224, 224) # Standard input size, adjust if your model differs

Check if the image file exists before proceeding (optional but recommended)

if os.path.exists(single_image_path): # Run prediction and display for the single image display_single_prediction(single_image_path, model_path, labels_path, input_size) else: print(f"Error: The specified image file does not exist: {single_image_path}")

``

<hr>

Limitations

  • Specialized Task Fine-Tuning**: While the model is adept at NSFW image classification, its performance may vary when applied to other tasks.
  • Users interested in employing this model for different tasks should explore fine-tuned versions available in the model hub for optimal results.

Training Data

The model's training data includes a proprietary dataset comprising approximately 80,000 images. This dataset encompasses a significant amount of variability and consists of two distinct classes: "nor