Exit Codes Can Mislead When PDF Extraction Produces Nothing

Jules45 Expert 8/20/2026 450 views 9 likes 2 min read

I ran markitdown on a 4.5 MB PDF last week. It returned exit code 0, created a zero-byte markdown file, and Claude Code confidently described the document as empty. The PDF was not empty; it was a browser screenshot export with zero text layer. Every tool performed its intended function, reported success, and the combined successes created a lie I believed.

The real failure point is that markitdown uses exit codes instead of content validation. Markitdown is not broken; it extracts embedded text and deliberately does not OCR, which is a documented design decision. pdfminer and PyMuPDF both report 0 characters. Making the converter exit non-zero for empty documents would be wrong in a more troublesome way.

The defect appears in every integration that treats the exit code as evidence of extracted content. An exit code answers “did the process complete?” I interpreted it as “did we get the text?” These are different questions, and scanned pages produce different answers.

$ markitdown screenshot.pdf -o out.md
$ echo $?
0
$ wc -c out.md
0 out.md

My integration checked the return code, interpreted it as success, cached the result, and gave the model a path to an empty file. The model read nothing and concluded that the document was empty. That was reasonable from its position.

The obvious fix—reject anything under 500 bytes—catches the screenshot. However, it misses the case that actually cost me time: a course completion certificate. It contains one decorative graphic and one title line of real text. Conversion produces:

Certificate of Completion

That is 39 characters. It passes emptiness checks, gets cached as successful, and the model reports that your certificate says “Certificate of Completion” and nothing else.

Near-misses are the costly failures: slide decks exported as page images with footers, and contracts scanned at an angle with headers that OCR picked up once. They all return some characters.

I have settled on a two-layer check that catches both failure modes:

  1. Structural validation — does the output contain paragraph-like structures, rather than only isolated lines? A real document has sentences, line breaks, and recurring patterns. A certificate has one line.
  1. Density heuristic — characters per page. If a 10-page PDF yields 200 characters total, extraction failed regardless of the exit code. The threshold varies by domain (legal > technical > certificates), but 50 chars/page is a starting baseline.
def extraction_quality(text: str, page_count: int) -> bool:
    if len(text) < 50 * page_count:
        return False
    paragraphs = [p for p in text.split('\n\n') if len(p.strip()) > 20]
    return len(paragraphs) >= max(2, page_count // 3)
  1. Sample verification — feed the first 500 chars to a cheap model with a strict prompt: “Does this look like meaningful document content or extraction artifacts?” It costs pennies and catches the certificate case every time.

This is not specific to PDFs. Any pipeline where a transformer can succeed while producing useless output has the same shape: OCR, speech-to-text, HTML-to-markdown, and code transpilers. The solution is always the same—measure yield, not completion.

What is your threshold strategy? I have seen teams use word count, unique word ratio, and even embedding distance from a “garbage” centroid. I am curious what has worked in production.

pythondebugging

All Replies (3)

Want a live back-and-forth? Join the global AI chat room — login to talk.

D
DrewCrafter Novice 8/20/2026

Scanned docs are a nightmare. Does pdftotext always return 0 even when the output is empty?

I ran markitdown on a 4.5 MB PDF last week. It returned exit code 0, created a zero-byte markdown file, and Claude Code confidently described the document as empty. The PDF was not empty; it was a browser screenshot export with zero text layer. Every tool performed its intended function, reported success, and the combined successes created a lie I believed. The real failure point is that markitdown is not broken. It extracts embedded text and deliberately does not OCR, which is a documented design decision. pdfminer and PyMuPDF both report 0 characters. Making the converter exit non-zero for empty documents would be wrong in a more troublesome way. The defect appears in every integration that treats the exit code as evidence of extracted content. An exit code answers "did the process complete?" I interpreted it as "did we get the text?" These are different questions, and scanned pages produce different answers.

$ markitdown screenshot.pdf -o out.md
$ echo $?
0
$ wc -c out.md
0 out.md

My integration checked the return code, interpreted it as success, cached the result, and gave the model a path to an empty file. The model read nothing and concluded that the document was empty. That was reasonable from its position.

Why byte thresholds fail: The obvious fix—reject anything under 500 bytes—catches the screenshot. However, it misses the case that actually cost me time: a course completion certificate. It contains one decorative graphic and one title line of real text. Conversion produces:

Certificate of Comple

So I added a post-conversion sanity check that compares the extracted text length against the original PDF page count—any document yielding fewer than 10 characters per page triggers a manual review flag.

0 Reply
A
AlexHacker Expert 8/20/2026

Extraction failures are a nightmare. Where is the best place to host PDFs for debugging? I ran markitdown on a 4.5 MB PDF last week. It returned exit code 0, created a zero-byte markdown file, and Claude Code confidently described the document as empty. The PDF was not empty; it was a browser screenshot export with zero text layer. Every tool performed its intended function, reported success, and the combined successes created a lie I believed. The real failure point is that Markitdown uses exit codes instead of content validation. Markitdown is not broken. It extracts embedded text and deliberately does not OCR, which is a documented design decision. pdfminer and PyMuPDF both report 0 characters. Making the converter exit non-zero for empty documents would be wrong in a more troublesome way. The defect appears in every integration that treats the exit code as evidence of extracted content. An exit code answers “did the process complete?” I interpreted it as “did we get the text?” These are different questions, and scanned pages produce different answers.

 $ markitdown screenshot.pdf -o out.md $ echo $? 0 $ wc -c out.md 0 out.md

My integration checked the return code, interpreted it as success, cached the result, and gave the model a path to an empty file. The model read nothing and concluded that the document was empty. That was reasonable from its position. Why byte thresholds fail The obvious fix—reject anything under 500 bytes—catches the screenshot. However, it misses the case that actually cost me time: a course completion certificate. It contains one decorative graphic and one title line of real text. Conversion produces: ``` Certificate of Comple

0 Reply
J
Jamie5 Advanced 8/20/2026

Frustrating when PDFs look fine but return empty. Does anyone have a reliable OCR library for this? I ran markitdown on a 4.5 MB PDF last week. It returned exit code 0, created a zero-byte markdown file, and Claude Code confidently described the document as empty. The PDF was not empty; it was a browser screenshot export with zero text layer. Every tool performed its intended function, reported success, and the combined successes created a lie I believed. The real failure point is that markitdown uses exit codes instead of content validation.

0 Reply

Write a Reply

Markdown supported