Kmemo: A Semantic Cache for LLM Calls

DeepPanda Intermediate 7/25/2026 229 views 15 likes 2 min read

Most semantic caches are dangerous because they rely entirely on cosine similarity. If you ask "Convert 100 USD to EUR" and then "Convert 250 USD to EUR," almost every embedding model will return a similarity score of ~0.99. A standard cache sees that high number and happily serves the wrong answer to your user, and you'll never see an error in your logs.

Kmemo solves this by treating similarity as a first-pass filter rather than the final verdict.

The "Guard" Architecture

Instead of just trusting a threshold, Kmemo runs candidates through a chain of "guards." These guards analyze the actual text for concrete discrepancies that embeddings often miss:

  • Numeric swaps and mismatched units
  • Different named entities or time references
  • Negations or flipped antonyms
  • Reversed comparisons

The logic here is based on cost asymmetry: it is much cheaper to accidentally trigger an extra API call (a false miss) than it is to provide a confidently wrong answer (a false hit).

Deployment and Integration

Kmemo is written in Kotlin (requires JDK 17+) and is designed to be lightweight. It doesn't force a specific embedding provider on you; it just needs a function that maps a String to a FloatArray.

To add it to your project:

dependencies {
 implementation("io.github.nacode-studios:kmemo-core:1.0.0")
}

Here is a basic implementation for an AI workflow:

val cache = SemanticCache(
 embedder = Embedder { text -> openAi.embed(text) },
 store = InMemoryStore(maxEntries = 10_000, ttl = 1.hours),
)

val answer = cache.getOrPut(prompt) { llm.complete(it) }

Debugging the Hit Rate

One of the biggest pain points with caching is not knowing why a hit didn't happen. Kmemo provides specific reasons for misses, which is essential for tuning your prompt engineering or threshold settings.

when (val result = cache.lookup(prompt)) {
 is CacheLookup.Hit -> result.response
 is CacheLookup.Miss -> when (result.reason) {
 MissReason.BELOW_THRESHOLD -> // Similarity was too low
 MissReason.REJECTED_BY_GUARD -> // A guard found a concrete difference
 else -> null
 }
}

Handling Scopes

To prevent "cross-contamination" (e.g., serving a GPT-3.5 answer to a GPT-4o request), you can define scopes. This ensures that changes in model version, temperature, or system prompts are treated as different cache keys.

cache.getOrPut(prompt, scope = "gpt-4o|t=0.0|v3") { llm.complete(it) }

If you need a more rigorous setup, you can swap the default guards for MatchGuards.strict(), which further prioritizes accuracy over hit rate.

For those who want to see exactly why a prompt was rejected, the cache.explain(prompt) method provides a full breakdown of every candidate and the specific guard that vetoed it.

AILLMLarge Language Modelkotlin
Detailed breakdowns of putting AI to work are in a guide to making money with AI, with plenty of directly applicable cases.

All Replies (3)

R
Riley2 Advanced 7/25/2026

GPT-4 and Pinecone drove me crazy with those similarity scores. How do you handle numeric precision?

0 Reply
K
KaiDev Expert 7/25/2026

Getting 'similar' answers for different zip codes sounds like a nightmare. How does the cache handle that?

0 Reply
J
JulesCrafter Novice 7/25/2026

Annoyed by embedding errors! How does Kmemo handle precise strings like product IDs?

0 Reply

Write a Reply

Markdown supported