Global AI chat room · 17 online now Join now
DIRECTORY / 02

AI Models | Open-Source LLM Directory

Discover and compare open-source LLMs, language models and multimodal models by capability, scale, license, downloads and provenance.

Compare modelsFind the right building block for your next workflow
Directory overview
18
curated entries
24 topic groupsLive
02 / MODEL INDEX

Find the right model for the job

Context first, better decisions. Every entry keeps the signal that matters.

CURATED DIRECTORY18 results

GLM-OCR

zai-org
Not specified

GLM OCR is a specialized vision-language model designed to bridge the gap between raw image data and structured text. Unlike general-purpose OCR engines that often struggle with complex layouts or handwritten notes, this model leverages the GLM architecture to maintain spatial awareness and semantic context. For developers, this means higher accuracy in digitizing multi-column documents, tables, and mixed-media assets without requiring extensive pre-processing pipelines. It integrates easily into RAG workflows where document parsing is a bottleneck, offering a more robust alternative to traditional Tesseract-based solutions. Whether you are building automated invoice processing or digitizing archival records, GLM OCR provides the precision needed for downstream LLM consumption.

image to textmit
2.1K starsView details

blip-image-captioning-large

Salesforce
Not specified

BLIP (Bootstrapping Language-Image Pre-training) Large is a versatile vision-language model designed for high-fidelity image captioning and visual question answering. Unlike basic image-to-text models, BLIP is trained to bridge the gap between noisy web data and clean synthetic captions, resulting in descriptions that are more contextually accurate and descriptive. For developers, it serves as a robust backbone for automating alt-text generation, indexing visual libraries, or building accessibility tools. It integrates well into Python-based ML pipelines via Hugging Face Transformers, offering a strong balance between inference speed and descriptive quality compared to smaller CLIP-based encoders.

image to textbsd-3-clause
1.5K starsView details

vit-gpt2-image-captioning

nlpconnect
Not specified

The vit-gpt2-image-captioning model offers a streamlined pipeline for converting visual data into descriptive natural language. By leveraging a Vision Transformer (ViT) encoder paired with a GPT-2 language model decoder, it bridges the gap between computer vision and sequence generation. For developers, this means a robust architecture for tasks like automated image tagging, accessibility enhancements for the visually impaired, and generating metadata for large-scale visual datasets. Unlike massive multi-modal models that require significant compute, this architecture is relatively lightweight, making it easier to integrate into existing transformer-based workflows via the Hugging Face ecosystem. While it excels at generating coherent, contextually relevant captions for standard imagery, developers should benchmark its performance against domain-specific datasets—such as medical or satellite imagery—to ensure accuracy before moving to production. Its Apache-2.0 license provides the flexibility needed for both commercial and open-source deployments.

image to textapache-2.0
935 starsView details

blip-image-captioning-base

Salesforce
Not specified

BLIP (Bootstrapping Language-Image Pre-training) is a versatile vision-language model designed to bridge the gap between image understanding and natural language generation. For developers, its primary value lies in its ability to generate descriptive, contextually accurate captions from raw image data. Unlike basic tagging models, BLIP leverages a unified framework for both understanding and generation, making it highly effective for building automated alt-text generators, image search indexing tools, and visual question-answering (VQA) systems. It integrates easily into PyTorch-based pipelines and offers a strong balance between inference speed and descriptive quality, serving as a reliable baseline for image-to-text tasks before moving to heavier, multi-modal LLMs.

image to textbsd-3-clause
889 starsView details

trocr-base-handwritten

microsoft
Not specified

TrOCR-base-handwritten is a transformer-based optical character recognition (OCR) model specifically optimized for handwritten text. Unlike traditional OCR engines that rely on separate CNN and RNN components, TrOCR employs a unified encoder-decoder architecture, using a Vision Transformer (ViT) to process images and a RoBERTa-like decoder to generate text. This end-to-end approach eliminates the need for complex language modeling post-processing. For developers, this model is ideal for digitizing archives, automating form processing, or building accessibility tools. It integrates seamlessly via the Hugging Face Transformers library, allowing for rapid deployment in Python environments. While it offers high accuracy on clear handwriting, performance varies based on script legibility compared to printed text models.

image to textmit
519 starsView details

NuMarkdown-8B-Thinking

numind
Not specified

NuMarkdown 8B Thinking is a specialized vision-language model optimized for high-fidelity image-to-markdown conversion. Unlike general-purpose OCR, this model focuses on structural integrity, accurately translating complex visual layouts—such as nested tables, mathematical formulas, and hierarchical headers—into clean, semantic Markdown. For developers, this means significantly less post-processing when digitizing documentation or converting legacy PDFs into LLM-ready datasets. It bridges the gap between raw visual data and structured text, offering a streamlined pipeline for RAG systems that rely on precise document parsing. Its 8B parameter scale provides a balanced trade-off between inference latency and reasoning capabilities, making it suitable for integration into automated data ingestion workflows.

image to textmit
496 starsView details

NuExtract3

numind
Not specified

NuExtract3 is a specialized image-to-text model designed for structured information extraction. Unlike general-purpose VLMs that often struggle with precision or hallucinate during data parsing, NuExtract3 focuses on transforming unstructured visual data into machine-readable formats. It is particularly effective for developers building automated pipelines for invoice processing, form digitization, and document analysis where schema adherence is critical. With an Apache-2.0 license, it offers the flexibility for commercial integration without restrictive overhead. Developers can integrate it into existing OCR workflows to replace brittle rule-based parsing with a more robust, neural extraction layer that maintains high fidelity to the source document.

image to textapache-2.0
349 starsView details

LightOnOCR-1B-1025

lightonai
Not specified

LightOnOCR-1B-1025 is an open image-to-text model from LightOn AI, built for optical character recognition and document understanding. It takes images as input and returns extracted text, making it useful for digitizing scanned documents, forms, receipts, or any visual content containing text. The model is published under the permissive Apache-2.0 license, so it can be freely used in commercial and research applications without licensing concerns. It integrates smoothly with Hugging Face transformers, allowing developers to load and run it with standard pipelines for token classification or sequence-to-sequence tasks. While the parameter count isn't specified, its focus on OCR suggests it's optimized for accuracy in reading text rather than general image understanding. Compared to larger multimodal models, LightOnOCR-1B-1025 is lightweight and specialized, offering faster inference and lower resource usage for text extraction workflows. Developers should evaluate the model card for supported languages, input resolution limits, and recommended preprocessing steps before deployment.

image to textapache-2.0
256 starsView details

trocr-base-printed

microsoft
Not specified

TrOCR-base-printed is a transformer-based optical character recognition (OCR) model designed specifically for printed text. Unlike traditional OCR pipelines that rely on separate text detection and recognition stages, TrOCR leverages a vision transformer (ViT) encoder and a language model decoder to map image patches directly to text sequences. This end-to-end architecture makes it particularly effective for high-accuracy transcription of documents, labels, and digitized archives. For developers, it offers a streamlined integration path via the Hugging Face ecosystem, providing a robust alternative to Tesseract or cloud-based OCR APIs when local deployment and Apache-2.0 licensing are priorities.

image to textSee model card
220 starsView details

nougat-base

facebook
Not specified

Nougat-base is a specialized vision-to-text transformer architecture designed to bridge the gap between complex visual document layouts and machine-readable text. Unlike general-purpose OCR engines that struggle with mathematical notation or multi-column academic structures, Nougat is optimized for parsing scientific papers and structured PDFs into clean Markdown. For developers, this means moving away from fragile heuristic-based parsing and toward a streamlined end-to-end pipeline. It integrates natively with the Hugging Face Transformers ecosystem, making it straightforward to deploy within existing PyTorch workflows. While it excels at converting dense academic content into structured data, developers should note its non-commercial license and evaluate its performance on specific document densities before scaling. It is an ideal choice for building RAG (Retrieval-Augmented Generation) pipelines where high-fidelity document ingestion is critical for downstream LLM accuracy.

image to textcc-by-nc-4.0
189 starsView details

kosmos-2-patch14-224

microsoft
Not specified

Kosmos-2 (Patch14 224) is a multimodal model designed to bridge the gap between visual perception and natural language processing. Unlike traditional image-to-text models that rely on separate encoders and decoders, Kosmos-2 treats visual patches as discrete tokens, allowing it to process images and text within a unified transformer architecture. For developers, this means stronger capabilities in visual grounding and spatial reasoning, making it particularly effective for tasks like image captioning, visual question answering (VQA), and identifying specific object coordinates within a frame. Integration is streamlined for those already utilizing PyTorch or Hugging Face ecosystems. Compared to larger proprietary models, it offers a more lightweight footprint while maintaining high precision in multimodal alignment, providing a flexible baseline for building specialized vision-language agents.

image to textmit
183 starsView details

manga-ocr-base

kha-white
Not specified

Manga OCR Base is a specialized image-to-text model engineered specifically for the complexities of Japanese manga typesetting. Unlike general-purpose OCR, this model is optimized to handle vertical text flow, stylized fonts, and the overlapping visual noise common in comic panels. For developers, it serves as a robust backend for translation pipelines, archival tools, or accessibility plugins. It integrates easily into Python-based workflows via standard OCR wrappers, offering a focused alternative to monolithic vision models by prioritizing high accuracy in niche typographic layouts over general scene recognition.

image to textapache-2.0
180 starsView details

trocr-large-handwritten

microsoft
Not specified

TrOCR-Large is a transformer-based optical character recognition model designed specifically for handwritten text recognition. Unlike traditional OCR pipelines that rely on separate detection and recognition stages, TrOCR utilizes a vision transformer (ViT) encoder and a language model decoder to map image pixels directly to text sequences. For developers, this means superior performance on curved or irregular handwriting where standard OCR often fails. It is an ideal choice for digitizing archives, automating form processing, or building accessibility tools. The model is released under the Apache-2.0 license, making it highly flexible for commercial integration via Hugging Face or custom PyTorch pipelines.

image to textSee model card
166 starsView details

granite-vision-3.3-2b

ibm-granite
Not specified

Granite Vision 3.3 2B is a lightweight vision-language model designed for efficient image-to-text processing. At 2 billion parameters, it targets developers who need a low-latency solution for visual understanding without the overhead of massive frontier models. It is particularly effective for OCR tasks, image captioning, and visual question answering (VQA) where local deployment or edge computing is a priority. Licensed under Apache-2.0, it offers significant flexibility for commercial integration. Compared to larger VLM alternatives, it prioritizes a smaller memory footprint and faster inference speeds, making it an ideal candidate for embedding into real-time applications or pipelines requiring rapid visual analysis.

image to textapache-2.0
88 starsView details

PP-OCRv5_server_det

PaddlePaddle
Not specified

PP OCRv5 server det is a high-performance text detection model designed for industrial-scale OCR pipelines. Unlike lightweight mobile versions, the server-side architecture prioritizes precision and robustness across complex backgrounds and varied font styles. It serves as the critical first stage in an OCR workflow, isolating text regions with high spatial accuracy before passing them to a recognition engine. For developers, this model is ideal for automating document digitizing, invoice processing, and license plate recognition where reliability outweighs latency constraints. It integrates seamlessly into PaddlePaddle-based environments and is released under the Apache-2.0 license, offering significant flexibility for commercial deployment and custom fine-tuning.

image to textapache-2.0
75 starsView details

trocr-small-handwritten

microsoft
Not specified

TrOCR-small is a lightweight transformer-based model designed specifically for optical character recognition of handwritten text. Unlike traditional OCR engines that rely on separate text detection and recognition stages, TrOCR uses an end-to-end encoder-decoder architecture, leveraging a Vision Transformer (ViT) to process images and a language model to generate text. This makes it particularly effective for curved or irregular handwriting where standard OCR often fails. For developers, the 'small' variant offers a critical balance between inference speed and accuracy, making it suitable for edge deployment or real-time applications. It integrates easily into Python pipelines via the Hugging Face Transformers library, allowing for rapid implementation of digitizing workflows, form processing, and archival automation without requiring massive compute resources.

image to textSee model card
67 starsView details

mgp-str-base

alibaba-damo
Not specified

mgp-str-base is a specialized vision-language model developed by Alibaba DAMO Academy, designed specifically for high-accuracy image-to-text transduction tasks. Unlike general-purpose multimodal LLMs that prioritize conversational reasoning, this model is architected for structured visual perception and precise textual extraction. For developers working on OCR-heavy pipelines, document intelligence, or automated metadata generation, mgp-str-base offers a streamlined alternative to massive, resource-intensive models. It integrates natively with the Hugging Face Transformers ecosystem, making it easy to drop into existing PyTorch or TensorFlow workflows. While the exact parameter count is undisclosed, its 'base' designation suggests a balance between inference latency and descriptive accuracy, making it suitable for edge deployments or high-throughput microservices where rapid visual parsing is critical. When integrating, focus on evaluating its performance against your specific domain datasets to ensure the alignment meets your production requirements.

image to textSee model card
65 starsView details

pix2text-mfr

breezedeus
Not specified

pix2text mfr is a specialized image-to-text model designed for high-accuracy mathematical formula recognition. Unlike general OCR, this model focuses on the structural complexity of LaTeX-style notation, making it ideal for digitizing academic papers, technical documentation, and educational content. For developers, it serves as a robust bridge between raw image data and editable markup, streamlining the pipeline for converting screenshots or PDFs into structured mathematical text. It integrates efficiently into document processing workflows where precision in symbolic representation is critical, offering a lightweight alternative to heavy multimodal LLMs for specific formula extraction tasks.

image to textmit
58 starsView details
Email