Model card
Nougat-base is a specialized vision-to-text transformer architecture designed to bridge the gap between complex visual document layouts and machine-readable text. Unlike general-purpose OCR engines that struggle with mathematical notation or multi-column academic structures, Nougat is optimized for parsing scientific papers and structured PDFs into clean Markdown. For developers, this means moving away from fragile heuristic-based parsing and toward a streamlined end-to-end pipeline. It integrates natively with the Hugging Face Transformers ecosystem, making it straightforward to deploy within existing PyTorch workflows. While it excels at converting dense academic content into structured data, developers should note its non-commercial license and evaluate its performance on specific document densities before scaling. It is an ideal choice for building RAG (Retrieval-Augmented Generation) pipelines where high-fidelity document ingestion is critical for downstream LLM accuracy.
Model files and versions
Download this model
We recommend using the ModelScope CLI or SDK. Install ModelScope first, then choose a full snapshot, single file, SDK or Git LFS workflow.
facebook/nougat-baseInstall the CLI and SDK dependency before downloading.
pip install modelscopeDownload the complete weights, configuration and model card.
modelscope download --model facebook/nougat-baseREADME.md is used as an example; replace it with another repository file when needed.
modelscope download --model facebook/nougat-base README.md --local_dir ./dirUseful in Python projects and automation scripts.
from modelscope import snapshot_download
model_dir = snapshot_download('facebook/nougat-base')Make sure Git LFS is installed correctly.
git lfs install
git clone https://www.modelscope.cn/facebook/nougat-base.gitFetch the repository structure first, then pull large files when needed.
GIT_LFS_SKIP_SMUDGE=1 git clone https://www.modelscope.cn/facebook/nougat-base.gitHow to use
- 01Step 1
Read the model card and source information.
- 02Step 2
Start with a small, non-sensitive evaluation.
- 03Step 3
Review quality, licensing and usage limits.
- 04Step 4
Adopt it only after validation.
Discussions
Use this space to keep checking source information, usage experience and maintenance status.
Open source page