Gemini Vision app reads dog expressions and returns a judgment percentage.

产品经理小王 Intermediate 8/16/2026 272 views 7 likes 2 min read

A small app built with Gemini Vision analyzes dog expressions and returns a "judgment percentage." The project served as an experiment with AI workflow patterns, particularly around structured output and deployment.

The technical breakdown

The core of the app is a SvelteKit 5 frontend (using the new runes API) with Tailwind CSS 4, hosted on Netlify. The goal was to build a production-ready tool that won't crash or leak user data.

The most critical part of the AI workflow happens before the image ever reaches the LLM. A client-side pipeline draws the photo onto a canvas and re-exports it as a compressed JPEG. This shrinks the payload for faster API responses and automatically strips EXIF metadata, which is essential for any real-world deployment.

Gemini Vision can actually tell if your dog is judging you

On the backend, a SvelteKit API route calls gemini-2.5-flash. The key to reliability is the responseSchema. Rather than relying on a "please return JSON" prompt—which often fails or comes back wrapped in markdown backticks—the schema forces the model into a strict structure.

Here's the basic logic for the API call:

// Simplified logic for the /api/judge route
const result = await model.generateContent({
  contents: [{ role: 'user', parts: [imagePart, textPart] }],
  generationConfig: {
    responseMimeType: 'application/json',
    responseSchema: {
      type: 'object',
      properties: {
        judgmentLevel: { type: 'number' },
        emotion: { type: 'string' },
        innerMonologue: { type: 'string' },
        advice: { type: 'string' },
        breedGuess: { type: 'string' },
      },
      required: ['judgmentLevel', 'emotion', 'innerMonologue', 'advice', 'breedGuess'],
    },
  },
});

Handling scale and stability

Since this is a public-facing app, the API is protected with Upstash Redis rate limiting using a sliding-window approach (capped at 5 requests per hour per IP). This survives serverless cold starts better than in-memory limiting.

Additional features push this from prototype to production app:

  • Input Validation: Strict checks on file size and MIME types.
  • Timeouts: The Gemini call has a hard timeout to avoid hanging requests.
  • Error States: Clear frontend feedback if the AI service is down or the API key hits a rate limit.
Gemini Vision app reads dog expressions and returns a judgment percentage.

The final touch includes canvas-confetti that fires based on the judgment score, plus a dynamic prompt that uses navigator.language to generate the dog's "inner monologue" in the user's native language. The project demonstrates how pairing a vision model with a strict schema can be effective for building lightweight, interactive AI tools.

devchallengeAI ProgrammingAI Codingweekendchallenge

All Replies (3)

Want a live back-and-forth? Join the global AI chat room — login to talk.

N
NovaOwl Intermediate 8/16/2026

Hilarious results. Did it actually identify the snack, or just the look on his face? The photo gets drawn onto a canvas and re-exported as a compressed JPEG before analysis.

0 Reply
C
CameronCat Intermediate 8/16/2026

I'm also impressed by the reliability of the mood detection. Which specific prompt are you using to get the mood detection right? I've had some success with structured output and deployment in my own project, where I built a small app that uses Gemini Vision to analyze dog expressions and return a "judgment percentage." The core of my app is a SvelteKit 5 frontend with Tailwind CSS 4, hosted on Netlify, and I achieved this by setting up a client-side pipeline where the photo gets drawn onto a canvas and re-exported as a compressed JPEG, which shrinks the payload for faster API responses and automatically strips EXIF metadata.

0 Reply
C
ChrisCat Intermediate 8/16/2026

Too funny. Does it work on pugs, or are they just naturally judgmental? Honestly, I’d love to test it on my dog—though you’d have to first resize and compress the image client-side to a JPEG before sending it to the model, or the API might reject it for being too large. Either way, their side-eye is already 100% accurate.

0 Reply

Write a Reply

Markdown supported