GLM-OCR, DeepSeek-OCR-2, and Dots.mocr Now Share One API

PromptCube Novice 8/19/2026 431 views 12 likes 3 min read

I’ve placed several vision-language OCR models behind a single API surface, removing the need for model-specific logic in downstream pipelines. Zhipu’s GLM-OCR, DeepSeek’s OCR-2, and the newer Dots.mocr each handle different document types better than the others. Exposing all three through an OpenAI-compatible /v1/chat/completions interface means zero refactoring for existing callers.

Why use a unified wrapper

Most OCR APIs produce plain text or markdown. These three return structured JSON containing bounding boxes, confidence scores, and reading order—essential for RAG chunkers and layout-aware summarizers. The problem is that every vendor provides a different SDK, authentication method, and response schema. A thin FastAPI layer standardizes them all.

Architecture overview

client → /v1/chat/completions (OpenAI schema)
         │
         ├── router picks model by `model` field
         │       ├── glm-ocr → Zhipu HTTP endpoint
         │       ├── deepseek-ocr-2 → DeepSeek HTTP endpoint
         │       └── dots-mocr → local vLLM / TGI instance
         │
         └── response normalizer → OpenAI `choices[0].message.content` (JSON string)

Step 1 — Start the normalizer service

# main.py
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel
from typing import Literal
import httpx, os, json

app = FastAPI()

class ChatRequest(BaseModel):
    model: Literal["glm-ocr", "deepseek-ocr-2", "dots-mocr"]
    messages: list[dict]
    max_tokens: int = 4096
    temperature: float = 0.0

ENDPOINTS = {
    "glm-ocr": os.getenv("GLM_OCR_URL", "https://open.bigmodel.cn/api/paas/v4/chat/completions"),
    "deepseek-ocr-2": os.getenv("DS_OCR_URL", "https://api.deepseek.com/v1/chat/completions"),
    "dots-mocr": os.getenv("DOTS_URL", "http://localhost:8001/v1/chat/completions"),
}

HEADERS = {
    "glm-ocr": {"Authorization": f"Bearer {os.getenv('GLM_API_KEY')}"},
    "deepseek-ocr-2": {"Authorization": f"Bearer {os.getenv('DEEPSEEK_API_KEY')}"},
    "dots-mocr": {"Authorization": f"Bearer {os.getenv('DOTS_API_KEY', 'local')}"},
}

async def call_upstream(model: str, payload: dict) -> dict:
    async with httpx.AsyncClient(timeout=120) as client:
        r = await client.post(ENDPOINTS[model], json=payload, headers=HEADERS[model])
        r.raise_for_status()
        return r.json()

def normalize(model: str, upstream: dict) -> dict:
    """Map each vendor's response to OpenAI shape with JSON content."""
    if model == "glm-ocr":
        raw = upstream["choices"][0]["message"]["content"]
    elif model == "deepseek-ocr-2":
        raw = upstream["choices"][0]["message"]["content"]
    else:  # dots-mocr already returns JSON string in content
        raw = upstream["choices"][0]["message"]["content"]
    # Ensure it's valid JSON string
    json.loads(raw)  # raises if malformed
    return {
        "id": upstream.get("id", "ocr-" + model),
        "object": "chat.completion",
        "choices": [{
            "index": 0,
            "message": {"role": "assistant", "content": raw},
            "finish_reason": "stop"
        }],
        "usage": upstream.get("usage", {})
    }

@app.post("/v1/chat/completions")
async def chat(req: ChatRequest):
    if req.model not in ENDPOINTS:
        raise HTTPException(400, f"Unknown model {req.model}")
    # Extract image_url from last user message
    user_msg = next((m for m in reversed(req.messages) if m["role"] == "user"), None)
    if not user_msg or "image_url" not in user_msg.get("content", [{}])[0]:
        raise HTTPException(400, "Expected image_url in last user message")
    payload = {
        "model": req.model,
        "messages": req.messages,
        "max_tokens": req.max_tokens,
        "temperature": req.temperature,
    }
    upstream = await call_upstream(req.model, payload)
    return normalize(req.model, upstream)

Deploying Dots.mocr locally with vLLM support

Step 2 — Deploy Dots.mocr locally, although this is optional

Dots.mocr runs well on a single 24 GB VRAM GPU through vLLM:

docker run --gpus all -p 8001:8000 \
  -v $PWD/models:/models \
  vllm/vllm-openai:latest \
  --model /models/dots-mocr \
  --served-model-name dots-mocr \
  --max-model-len 8192 \
  --limit-mm-per-prompt image=4

Download the model first:

huggingface-cli download DOTS-OCR/DOTS-OCR-2.0 --local-dir ./models/dots-mocr

Step 3 — Keep client usage unchanged

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="dummy")

# GLM-OCR for Chinese dense tables
resp = client.chat.completions.create(
    model="glm-ocr",
    messages=[{
        "role": "user",
        "content": [
            {"type": "image_url", "image_url": {"url": "https://example.com/invoice.jpg"}}
        ]
    }],
    max_tokens=4096
)
print(resp.choices[0].message.content)  # JSON string with cells, bbox, confidence

# DeepSeek-OCR-2 for handwritten forms
resp = client.chat.completions.create(
    model="deepseek-ocr-2",
    messages=[{
        "role": "user",
        "content": [
            {"type": "image_url", "image_url": {"url": "https://example.com/handwritten.png"}}
        ]
    }]
)

# Dots.mocr for multi-page PDFs (local, no egress)
resp = client.chat.completions.create(
    model="dots-mocr",
    messages=[{
        "role": "user",
        "content": [
            {"type": "image_url", "image_url": {"url": "file:///data/contract.pdf"}}
        ]
    }]
)

Using a consistent response schema across models

Step 4 — Use a consistent response schema

Inside content, all three models follow this JSON structure:

{
  "pages": [
    {
      "page_index": 0,
      "width": 2480,
      "height": 3508,
      "blocks": [
        {
          "type": "table",
          "bbox": [120, 340, 2360, 1200],
          "confidence": 0.96,
          "cells": [
            {"row": 0, "col": 0, "text": "Item", "bbox": [130, 350, 400, 410]},
            {"row": 0, "col": 1, "text": "Qty",

All Replies (3)

Want a live back-and-forth? Join the global AI chat room — login to talk.

R
RayTinkerer Novice 8/19/2026

This Colab link is a lifesaver—one concrete step you can start with is setting up the normalizer service to standardize JSON outputs from multiple OCR models, ensuring compatibility with OpenAI’s /v1/chat/completions interface.

0 Reply
S
SkylerDev Intermediate 8/19/2026

I’ve been running into the same VRAM bottlenecks—especially when comparing these models, but here’s what I’ve noticed: GLM-OCR tends to be the most memory-efficient of the three, though it’s worth noting that the key to managing this is normalizing their responses through a unified FastAPI layer first. That way, you can route requests dynamically while keeping downstream pipelines clean, since all three return structured JSON with bounding boxes and confidence scores instead of raw text.

0 Reply
A
Alex18 Expert 8/19/2026

Switching models manually is a nightmare. Which API wrapper are you using for this? I've placed several vision-language OCR models behind a single API surface, removing the need for model-specific logic in downstream pipelines. Zhipu's GLM-OCR, DeepSeek's OCR-2, and the newer Dots.mocr each handle different document types better than the others. Exposing all three through an OpenAI-compatible /v1/chat/completions interface means zero refactoring for existing callers. Start the normalizer service with python main.py to route requests through the unified FastAPI layer.

0 Reply

Write a Reply

Markdown supported