en PP OCRv5 mobile rec

提供商PaddlePaddle
分类image-to-text
许可证apache-2.0
下载量34.2K
星标1

简介

PP-OCRv5 mobile rec 是百度飞桨(PaddlePaddle)推出的轻量化文本识别模型,专门针对移动端和边缘设备优化。它在保持高识别精度的同时,极大地降低了推理延迟和内存占用,非常适合集成到手机 App 或嵌入式设备中。对于开发者来说,该模型上手门槛低,能够与 PP-OCR 系列的检测模型无缝衔接,构建完整的端到端 OCR 流水线。相比于大型通用模型,它在处理工业场景下的单行文本识别时,速度优势明显,是追求实时性识别任务的理想选择。

核心亮点

  • 轻量化设计,极速适配移动端与边缘设备
  • 百度飞桨生态支持,部署链路成熟且便捷
  • 专注单行文本识别,推理延迟极低
  • Apache-2.0 协议,商业化集成无压力

使用方法

安装依赖
# 安装 Hugging Face transformers
pip install transformers torch
SDK 使用
# 使用 transformers 加载模型
from transformers import AutoModel, AutoTokenizer

model = AutoModel.from_pretrained("PaddlePaddle/en_PP-OCRv5_mobile_rec")
tokenizer = AutoTokenizer.from_pretrained("PaddlePaddle/en_PP-OCRv5_mobile_rec")

Hugging Face 下载

我们推荐使用命令行或者 Hugging Face Hub SDK 来进行模型的下载。

操作指引:在下载前,请先通过如下命令安装 huggingface_hub:

操作指引
pip install -U huggingface_hub

命令行下载

下载完整模型库

下载完整模型库
huggingface-cli download PaddlePaddle/en_PP-OCRv5_mobile_rec

下载单个文件到指定本地文件夹(以下载 config.json 到当前路径下 ./dir 目录为例)

下载单个文件到指定本地文件夹(以下载 config.json 到当前路径下 ./dir 目录为例)
huggingface-cli download PaddlePaddle/en_PP-OCRv5_mobile_rec config.json --local-dir ./dir

更多命令行下载选项,可参见官方文档

SDK 下载

SDK 下载
# 模型下载
from huggingface_hub import snapshot_download
model_dir = snapshot_download('PaddlePaddle/en_PP-OCRv5_mobile_rec')

Git 下载

请确保 lfs 已经被正确安装

Git 下载
git lfs install
git clone https://huggingface.co/PaddlePaddle/en_PP-OCRv5_mobile_rec

如果您希望跳过 lfs 大文件下载,可以使用如下命令

跳过 LFS
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/PaddlePaddle/en_PP-OCRv5_mobile_rec

模型文件托管在 Hugging Face Hub,使用 HF CLI / SDK / Git 直接下载,不经过本站。

PyTorch / Transformers 使用

安装 Transformers

安装 Transformers
pip install -U transformers torch

模型加载和推理

模型加载和推理
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained('PaddlePaddle/en_PP-OCRv5_mobile_rec')
tokenizer = AutoTokenizer.from_pretrained('PaddlePaddle/en_PP-OCRv5_mobile_rec')

模型下载

我们推荐使用命令行或者 ModelScope SDK 来进行模型的下载。

操作指引:在下载前,请先通过如下命令安装 ModelScope:

操作指引
pip install modelscope

命令行下载

下载完整模型库

下载完整模型库
modelscope download --model PaddlePaddle/en_PP-OCRv5_mobile_rec

下载单个文件到指定本地文件夹(以下载 README.md 到当前路径下 dir 目录为例)

下载单个文件到指定本地文件夹(以下载 README.md 到当前路径下 dir 目录为例)
modelscope download --model PaddlePaddle/en_PP-OCRv5_mobile_rec README.md --local_dir ./dir

更多更丰富的命令行下载选项,可参见具体文档

SDK 下载

SDK 下载
# 模型下载
from modelscope import snapshot_download
model_dir = snapshot_download('PaddlePaddle/en_PP-OCRv5_mobile_rec')

Git 下载

请确保 lfs 已经被正确安装

Git 下载
git lfs install
git clone https://www.modelscope.cn/PaddlePaddle/en_PP-OCRv5_mobile_rec.git

如果您希望跳过 lfs 大文件下载,可以使用如下命令

跳过 LFS
GIT_LFS_SKIP_SMUDGE=1 git clone https://www.modelscope.cn/PaddlePaddle/en_PP-OCRv5_mobile_rec.git

ModelScope 模型页直接下载模型文件;无需将模型文件放在本站服务器。

Notebook 快速开发

下载并安装 ModelScope library

下载并安装 ModelScope library
pip install "modelscope[audio,cv,nlp,multi-modal,science]" -f https://modelscope.oss-cn-beijing.aliyuncs.com/releases/repo.html

模型加载和推理

模型加载和推理
from modelscope.pipelines import pipeline
from modelscope.utils.constant import Tasks

p = pipeline('text-generation', 'PaddlePaddle/en_PP-OCRv5_mobile_rec')

完整文档

来源: HuggingFace

---
license: apache-2.0
library_name: PaddleOCR
language:

  • en

pipeline_tag: image-to-text
tags:
  • OCR

  • PaddlePaddle

  • PaddleOCR

  • textline_recognition

---

en_PP-OCRv5_mobile_rec

Introduction

en_PP-OCRv5_mobile_rec is one of the PP-OCRv5_rec that are the latest generation text line recognition models developed by PaddleOCR team. It aims to efficiently and accurately support the recognition of English. The key accuracy metrics are as follow:

| Model | Accuracy (%) |
|-|-|
| en_PP-OCRv5_mobile_rec | 85.3|

Note: If any character (including punctuation) in a line was incorrect, the entire line was marked as wrong. This ensures higher accuracy in practical applications.

Quick Start

Installation

1. PaddlePaddle

Please refer to the following commands to install PaddlePaddle using pip:

bash
# for CUDA11.8
python -m pip install paddlepaddle-gpu==3.0.0 -i https://www.paddlepaddle.org.cn/packages/stable/cu118/

for CUDA12.6

python -m pip install paddlepaddle-gpu==3.0.0 -i https://www.paddlepaddle.org.cn/packages/stable/cu126/

for CPU

python -m pip install paddlepaddle==3.0.0 -i https://www.paddlepaddle.org.cn/packages/stable/cpu/

For details about PaddlePaddle installation, please refer to the PaddlePaddle official website.

2. PaddleOCR

Install the latest version of the PaddleOCR inference package from PyPI:

bash
python -m pip install paddleocr

Model Usage

You can quickly experience the functionality with a single command:

bash
paddleocr text_recognition \
    --model_name en_PP-OCRv5_mobile_rec \
    -i https://cdn-uploads.huggingface.co/production/uploads/681c1ecd9539bdde5ae1733c/QmaPtftqwOgCtx0AIvU2z.png

You can also integrate the model inference of the text recognition module into your project. Before running the following code, please download the sample image to your local machine.

python
from paddleocr import TextRecognition
model = TextRecognition(model_name="en_PP-OCRv5_mobile_rec")
output = model.predict(input="QmaPtftqwOgCtx0AIvU2z.png", batch_size=1)
for res in output:
    res.print()
    res.save_to_img(save_path="./output/")
    res.save_to_json(save_path="./output/res.json")

After running, the obtained result is as follows:

json
{'res': {'input_path': '/root/.paddlex/predict_input/QmaPtftqwOgCtx0AIvU2z.png', 'page_index': None, 'rec_text': 'the number of model parameters and FLOPs get larger, it', 'rec_score': 0.993655264377594}}

The visualized image is as follows:

!image/jpeg

For details about usage command and descriptions of parameters, please refer to the Document.

Pipeline Usage

The ability of a single model is limited. But the pipeline consists of several models can provide more capacity to resolve difficult problems in real-world scenarios.

#### PP-OCRv5

The general OCR pipeline is used to solve text recognition tasks by extracting text information from images and outputting it in string format. And there are 5 modules in the pipeline:

  • Document Image Orientation Classification Module (Optional)

  • Text Image Unwarping Module (Optional)

  • Text Line Orientation Classification Module (Optional)

  • Text Detection Module

  • Text Recognition Module

Run a single command to quickly experience the OCR pipeline:

bash
paddleocr ocr -i https://cdn-uploads.huggingface.co/production/uploads/681c1ecd9539bdde5ae1733c/c3hSldnYVQXp48T5V0Ze4.png \
    --text_recognition_model_name en_PP-OCRv5_mobile_rec \
    --use_doc_orientation_classify False \
    --use_doc_unwarping False \
    --use_textline_orientation True \
    --save_path ./output \
    --device gpu:0

Results are printed to the terminal:

json
{'res': {'input_path': '/root/.paddlex/predict_input/c3hSldnYVQXp48T5V0Ze4.png', 'page_index': None, 'model_settings': {'use_doc_preprocessor': True, 'use_textline_orientation': False}, 'doc_preprocessor_res': {'input_path': None, 'page_index': None, 'model_settings': {'use_doc_orientation_classify': False, 'use_doc_unwarping': False}, 'angle': -1}, 'dt_polys': array([[[252, 172],
        ...,
        [254, 241]],

...,

[[665, 566],
...,
[663, 601]]], dtype=int16), 'text_det_params': {'limit_side_len': 64, 'limit_type': 'min', 'thresh': 0.3, 'max_side_limit': 4000, 'box_thresh': 0.6, 'unclip_ratio': 1.5}, 'text_type': 'general', 'textline_orientation_angles': array([-1, ..., -1]), 'text_rec_score_thresh': 0.0, 'return_word_box': False, 'rec_texts': ['The moon tells the sky', 'The sky tells the sea', 'The sea tells the tide', 'And the tide tells me', 'Lemn Sissay'], 'rec_scores': array([0.98405874, ..., 0.9837752 ]), 'rec_polys': array([[[252, 172],
...,
[254, 241]],

...,

[[665, 566],
...,
[663, 601]]], dtype=int16), 'rec_boxes': array([[252, ..., 241],
...,
[663, ..., 612]], dtype=int16)}}

If save_path is specified, the visualization results will be saved under save_path. The visualization output is shown below:

!image/jpeg

The command-line method is for quick experience. For project integration, also only a few codes are needed as well:

python
from paddleocr import PaddleOCR

ocr = PaddleOCR(
text_recognition_model_name="en_PP-OCRv5_mobile_rec",
use_doc_orientation_classify=False, # Use use_doc_orientation_classify to enable/disable document orientation classification model
use_doc_unwarping=False, # Use use_doc_unwarping to enable/disable document unwarping module
use_textline_orientation=True, # Use use_textline_orientation to enable/disable textline orientation classification model
device="gpu:0", # Use device to specify GPU for model inference
)
result = ocr.predict("https://cdn-uploads.huggingface.co/production/uploads/681c1ecd9539bdde5ae1733c/6KQKOS42DKVEUnrticvhd.png")
for res in result:
res.print()
res.save_to_img("output")
res.save_to_json("output")

The default model used in pipeline is PP-OCRv5_server_rec, so it is needed that specifing to en_PP-OCRv5_mobile_rec by argument text_recognition_model_name. And you can also use the local model file by argument text_recognition_model_dir. For details about usage command and descriptions of parameters, please refer to the Document.

Links

PaddleOCR Repo

PaddleOCR Documentation