Categories

AI Models Marketplace

Explore the latest and most popular AI models — LLM, vision, audio and more — to pick and integrate the right one.

PP LCNet x1 0 doc ori

apache-2.0
PP-LCNet x1.0 is a lightweight convolutional neural network optimized for high-efficiency image processing. Designed spe...
PaddlePaddle image-to-text
650.0K 0

PP OCRv5 server det

apache-2.0
PP OCRv5 server det is a high-performance text detection model designed for industrial-scale OCR pipelines. Unlike light...
PaddlePaddle image-to-text
543.4K 5

manga ocr base

apache-2.0
Manga OCR Base is a specialized image-to-text model engineered specifically for the complexities of Japanese manga types...
kha-white image-to-text
520.4K 0

pix2text mfr

mit
pix2text mfr is a specialized image-to-text model designed for high-accuracy mathematical formula recognition. Unlike ge...
breezedeus image-to-text
417.4K 0

UVDoc

apache-2.0
UVDoc is a specialized image-to-text model developed by PaddlePaddle, designed to handle complex document parsing and vi...
PaddlePaddle image-to-text
408.2K 1

PP OCRv5 server rec

apache-2.0
PP OCRv5 is a high-performance optical character recognition framework designed for scalable production environments. Un...
PaddlePaddle image-to-text
365.3K 3

Qwen3 VL 8B Instruct NVFP4

apache-2.0
Qwen3 VL 8B Instruct NVFP4 is a vision-language model optimized for high-throughput deployment via 4-bit Normal Float (N...
JEILDLWLRMA image-to-text
289.4K 0

en PP OCRv5 mobile rec

apache-2.0
The PP-OCRv5 mobile recognition model is a lightweight, high-efficiency text recognition engine optimized for edge deplo...
PaddlePaddle image-to-text
34.2K 1

NuExtract3

apache-2.0
NuExtract3 is a specialized image-to-text model designed for structured information extraction. Unlike general-purpose V...
numind image-to-text
14.5K 2

blip image captioning base

bsd-3-clause
BLIP (Bootstrapping Language-Image Pre-training) is a versatile vision-language model designed to bridge the gap between...
Salesforce image-to-text
3.8K 2

blip image captioning large

bsd-3-clause
BLIP (Bootstrapping Language-Image Pre-training) Large is a versatile vision-language model designed for high-fidelity i...
Salesforce image-to-text
876 1

trocr base printed

Apache-2.0
TrOCR-base-printed is a transformer-based optical character recognition (OCR) model designed specifically for printed te...
microsoft image-to-text
412 0

granite vision 3.3 2b

apache-2.0
Granite Vision 3.3 2B is a lightweight vision-language model designed for efficient image-to-text processing. At 2 billi...
ibm-granite image-to-text
330 0

NuMarkdown 8B Thinking

mit
NuMarkdown 8B Thinking is a specialized vision-language model optimized for high-fidelity image-to-markdown conversion. ...
numind image-to-text
287 0

trocr base handwritten

mit
TrOCR-base-handwritten is a transformer-based optical character recognition (OCR) model specifically optimized for handw...
microsoft image-to-text
276 1

trocr small handwritten

Apache-2.0
TrOCR-small is a lightweight transformer-based model designed specifically for optical character recognition of handwrit...
microsoft image-to-text
193 0

kosmos 2 patch14 224

mit
Kosmos-2 (Patch14 224) is a multimodal model designed to bridge the gap between visual perception and natural language p...
microsoft image-to-text
190 0

trocr large handwritten

Apache-2.0
TrOCR-Large is a transformer-based optical character recognition model designed specifically for handwritten text recogn...
microsoft image-to-text
186 0
Join our Telegram