table transformer structure recognition v1.1 all

提供商microsoft
分类object-detection
许可证mit
下载量422
星标0

简介

Table Transformer (TATR) 是微软推出的一款专门用于文档表格结构识别的深度学习模型。它将表格识别视为目标检测任务,能够精准地从复杂的 PDF 或图像文档中定位表格区域,并解析出其中的行、列及单元格结构。对于需要处理大量非结构化报表、财报或学术论文的开发者来说,它是将视觉图像转化为可编辑结构化数据的关键环节。该模型上手难度中等,通常作为 OCR 流水线中的中间件,与 PaddleOCR 或 Tesseract 等文本识别工具配合使用,共同完成“定位-识别-解析”的完整链路。

核心亮点

  • 精准识别文档中表格的边界与内部行列结构
  • 由微软开源且采用 MIT 协议,商业集成无压力
  • 适配复杂布局文档,提升数据提取的自动化率
  • 可与主流 OCR 引擎无缝衔接构建文档解析管线

使用方法

安装依赖
# 安装 Hugging Face transformers
pip install transformers torch
SDK 使用
# 使用 transformers 加载模型
from transformers import AutoModel, AutoTokenizer

model = AutoModel.from_pretrained("microsoft/table-transformer-structure-recognition-v1.1-all")
tokenizer = AutoTokenizer.from_pretrained("microsoft/table-transformer-structure-recognition-v1.1-all")

Hugging Face 下载

我们推荐使用命令行或者 Hugging Face Hub SDK 来进行模型的下载。

操作指引:在下载前,请先通过如下命令安装 huggingface_hub:

操作指引
pip install -U huggingface_hub

命令行下载

下载完整模型库

下载完整模型库
huggingface-cli download microsoft/table-transformer-structure-recognition-v1.1-all

下载单个文件到指定本地文件夹(以下载 config.json 到当前路径下 ./dir 目录为例)

下载单个文件到指定本地文件夹(以下载 config.json 到当前路径下 ./dir 目录为例)
huggingface-cli download microsoft/table-transformer-structure-recognition-v1.1-all config.json --local-dir ./dir

更多命令行下载选项,可参见官方文档

SDK 下载

SDK 下载
# 模型下载
from huggingface_hub import snapshot_download
model_dir = snapshot_download('microsoft/table-transformer-structure-recognition-v1.1-all')

Git 下载

请确保 lfs 已经被正确安装

Git 下载
git lfs install
git clone https://huggingface.co/microsoft/table-transformer-structure-recognition-v1.1-all

如果您希望跳过 lfs 大文件下载,可以使用如下命令

跳过 LFS
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/microsoft/table-transformer-structure-recognition-v1.1-all

模型文件托管在 Hugging Face Hub,使用 HF CLI / SDK / Git 直接下载,不经过本站。

PyTorch / Transformers 使用

安装 Transformers

安装 Transformers
pip install -U transformers torch

模型加载和推理

模型加载和推理
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained('microsoft/table-transformer-structure-recognition-v1.1-all')
tokenizer = AutoTokenizer.from_pretrained('microsoft/table-transformer-structure-recognition-v1.1-all')

模型下载

我们推荐使用命令行或者 ModelScope SDK 来进行模型的下载。

操作指引:在下载前,请先通过如下命令安装 ModelScope:

操作指引
pip install modelscope

命令行下载

下载完整模型库

下载完整模型库
modelscope download --model microsoft/table-transformer-structure-recognition-v1.1-all

下载单个文件到指定本地文件夹(以下载 README.md 到当前路径下 dir 目录为例)

下载单个文件到指定本地文件夹(以下载 README.md 到当前路径下 dir 目录为例)
modelscope download --model microsoft/table-transformer-structure-recognition-v1.1-all README.md --local_dir ./dir

更多更丰富的命令行下载选项,可参见具体文档

SDK 下载

SDK 下载
# 模型下载
from modelscope import snapshot_download
model_dir = snapshot_download('microsoft/table-transformer-structure-recognition-v1.1-all')

Git 下载

请确保 lfs 已经被正确安装

Git 下载
git lfs install
git clone https://www.modelscope.cn/microsoft/table-transformer-structure-recognition-v1.1-all.git

如果您希望跳过 lfs 大文件下载,可以使用如下命令

跳过 LFS
GIT_LFS_SKIP_SMUDGE=1 git clone https://www.modelscope.cn/microsoft/table-transformer-structure-recognition-v1.1-all.git

ModelScope 模型页直接下载模型文件;无需将模型文件放在本站服务器。

Notebook 快速开发

下载并安装 ModelScope library

下载并安装 ModelScope library
pip install "modelscope[audio,cv,nlp,multi-modal,science]" -f https://modelscope.oss-cn-beijing.aliyuncs.com/releases/repo.html

模型加载和推理

模型加载和推理
from modelscope.pipelines import pipeline
from modelscope.utils.constant import Tasks

p = pipeline('text-generation', 'microsoft/table-transformer-structure-recognition-v1.1-all')

完整文档

来源: HuggingFace

---
license: mit
---

Table Transformer (pre-trained for Table Structure Recognition)

Table Transformer (TATR) model trained on PubTables1M and FinTabNet.c. It was introduced in the paper Aligning benchmark datasets for table structure recognition by Smock et al. and first released in this repository.

Disclaimer: The team releasing Table Transformer did not write a model card for this model so this model card has been written by the Hugging Face team.

Model description

The Table Transformer is equivalent to DETR, a Transformer-based object detection model. Note that the authors decided to use the "normalize before" setting of DETR, which means that layernorm is applied before self- and cross-attention.

Usage

You can use the raw model for detecting tables in documents. See the documentation for more info.