Strict Pydantic contracts and local models reduce ad reviews to five seconds

小Max爱学习 Advanced 8/25/2026 495 views 10 likes 2 min read

Before CreativeAudit, verifying brand rules for ad creatives took 20–40 minutes per submission. This guide replaces that process with a five-second automated review.

Plain LLM chats fail in this context because their output formats drift, constraints are ignored, and results are unstructured text that may expose client data to cloud services. The system enforces five key requirements to address these issues:

  1. Strict output contracts – Pydantic v2 schemas act as an unbreakable validation layer, rejecting any LLM response that doesn’t conform to the expected JSON structure.
  2. Local inference – All processing occurs on-premise via Ollama, ensuring no data leaves the machine.
  3. Prompts as code – Jinja2 templates are stored in a versioned folder, allowing edits without modifying application logic.
  4. Binary compliance – Brand-book violations trigger immediate 0 or 10 scores, eliminating subjective judgment for high-risk rules.
  5. Human-in-the-loop – An initial LLM step extracts structured data from messy input, which users can correct before final evaluation.

The architecture follows this flow:

Brief + Creatives → Jinja2 Prompt → Local LLM (Ollama)
↓
Pydantic Validation
↓
Score + Verdict + Feedback

Core components include:

  • app/schemas.py – Defines exact JSON shapes using strict Pydantic contracts
  • prompts/*.j2 – Versioned prompt templates treated as source code
  • app/main.py – Handles prompt assembly, LLM calls, validation, and scoring
  • demo/streamlit_app.py – A lightweight UI for testing
Strict Pydantic contracts and local models reduce ad reviews to five seconds

Engineering decisions ensure reliability through:

  • Pydantic validation – Treating the schema as an absolute requirement prevents malformed responses from being treated as valid scores
  • Binary scoring (0/10) – Removes ambiguity for critical rules by enforcing absolute compliance
  • Jinja2 prompts – Separate files enable A/B testing and clean diff reviews without touching application code
  • Local Ollama deployment – Switching between models like qwen2.5:7b or llama3.2 requires only a one-line config change
  • Smart input normalization – Handles messy pasted fragments by first converting them to structured JSON before final audit

Each creative receives three scores:

  • Brand alignment (1–10)
  • Constraint compliance (0–10)
  • Message clarity (1–10)

The total combines them using these weights:
total = brand × 0.4 + compliance × 0.3 + clarity × 0.3

Verdicts derive from the total score and any critical failures detected by the binary check.

The system includes 38 deterministic unit tests that mock the LLM for sub-2-second execution, covering schema edge cases, malformed responses, and scoring logic. The implementation follows a practical tutorial structure with clear deployment instructions, making it accessible while demonstrating real-world LLM agent workflows.

Example prompt:

You are an ad compliance analyst. Given a creative description and a brand brief, output ONLY a JSON object that matches the schema below. Do not add any extra text.
{
  "verdict": "PASS|NEEDS_REVISION|FAIL",
  "brand_alignment": 1-10,
  "constraint_compliance": 0-10,
  "message_clarity": 1-10,
  "feedback": "short sentence explaining why"
}

CreativeAudit replaces hours of manual review with a deterministic process that maintains data privacy and enforces strict output contracts. The step-by-step setup provides a reproducible foundation for building reliable LLM pipelines.

pythonPromptPydantic

All Replies (3)

Want a live back-and-forth? Join the global AI chat room — login to talk.

S
SoloSmith Expert 8/25/2026

Curious if the Pydantic schema breaks when ads have overlapping brand elements. How is that handled? When brand elements overlap—like a logo partially covering text or conflicting color schemes—the schema enforces strict validation by requiring each element to be classified independently before scoring. If the LLM cannot cleanly separate the elements, the validation fails and returns a 0 score rather than a compromised result, ensuring binary compliance even in ambiguous cases.

0 Reply
S
Sam46 Advanced 8/25/2026

Hallucinations are driving me crazy. The key step is defining the exact JSON shapes via strict Pydantic v2 contracts, then validating the LLM output so a contract violation fails the call instead of returning fabricated scores. Which Pydantic version are you using to stop the fabrications?

0 Reply
J
JulesCrafter Novice 8/25/2026

Curious about motion content. Do these contracts actually work when elements shift mid-frame? To ensure strict validation, the schema acts as a hard wall; if the LLM disobeys, the call fails and no fake scores appear.

0 Reply

Write a Reply

Markdown supported