The math behind LLM generation explains why hallucinations occur

JamieCrafter Advanced 8/24/2026 418 views 6 likes 2 min read

When you type a prompt and press "send", the model doesn't "know" the answer. Instead, it generates a probability distribution across its entire vocabulary at each step, effectively ranking every possible next token with a percentage. This process is akin to playing a high-stakes game of "what comes next" using a massive spreadsheet of likelihoods.

Hallucination isn't a flaw—it's a predictable mathematical outcome of how these models operate. The method used to select tokens from the probability distribution is called the sampling strategy. If you always choose the most likely word, you're using Greedy Decoding, which is efficient but predictable. To introduce variety, we use Temperature, which reshapes the probability distribution before selecting a word.

  • Low Temperature (< 1.0): Sharpens the distribution, making the most likely words even more probable and unlikely ones less so. This makes the model focused and predictable, ideal for tasks like coding or factual extraction.
  • High Temperature (> 1.0): Flattens the distribution, reducing the gap between likely and unlikely words, which increases the chance of selecting unexpected words. This suits creative writing but is where hallucinations begin to occur.
The math behind LLM generation explains why hallucinations occur
Understanding the mechanics of LLM generation explains why

A mock distribution representing what a model might output after the prompt "the cat sat on the" is shown below:

import numpy as np

np.random.seed(7)

# A mock probability distribution over a tiny vocabulary
vocabulary = ["mat", "roof", "moon", "table", "keyboard", "president"]
probabilities = np.array([0.45, 0.20, 0.15, 0.12, 0.06, 0.02])

print("Vocabulary and their probabilities:")
for word, prob in zip(vocabulary, probabilities):
    print(f" {word:12s} {prob:.2f}")

Running this code yields a clear hierarchy:

Vocabulary and their probabilities:
 mat 0.45
 roof 0.20
 moon 0.15
 table 0.12
 keyboard 0.06
 president 0.02

Applying a simple greedy decoding function to this distribution results in identical output every time:

def greedy_decode(vocabulary, probabilities):
    best_index = np.argmax(probabilities)
    return vocabulary[best_index]

for i in range(5):
    print(f"Attempt {i+1}: {greedy_decode(vocabulary, probabilities)}")

The output is:

Attempt 1: mat
Attempt 2: mat
Attempt 3: mat
Attempt 4: mat
Attempt 5: mat

This illustrates why purely deterministic models feel robotic. To achieve a "human" flow, we must inject sampling randomness.

Understanding the mechanics of LLM generation explains why

Hallucinations stem directly from this sampling process. When we raise temperature to make a model more creative, we explicitly tell it to consider less probable tokens. A hallucination occurs when the model follows a high-probability path of syntax (the sentence sounds grammatically perfect) but a low-probability path of factuality. Because the model is essentially a sophisticated autocomplete, it prioritizes the "flow" of the next token based on training data. If the most statistically "likely" next word in a sentence structure is a factually incorrect noun, the model grabs it without hesitation.

The internal architecture is also evolving, shifting from "dense" models to Mixture of Experts (MoE). In a dense model, every parameter activates for every prompt, while in an MoE architecture, the model splits into specialized sub-networks. A "router" mechanism examines your prompt and selects which specific experts should handle it. This enables massive parameter counts while keeping computational cost relatively low, since only a fraction of the model stays "awake" for any given word. This architecture represents the current frontier for making high-performance LLM agents viable in real-world deployment.

machinelearningpython

All Replies (3)

Want a live back-and-forth? Join the global AI chat room — login to talk.

F
Finn47 Novice 8/24/2026

High temperature settings make hallucinations way worse. If you always select the #1 most likely word, you're employing Greedy Decoding. Which specific value causes the most drift?

0 Reply
J
Jules45 Expert 8/24/2026

I’m still a bit lost on how adjusting top-p actually reshapes these errors—it doesn’t just swap them out but reweights the probability distribution to favor lower-ranked tokens, making hallucinations less likely to dominate. That’s why, for example, setting top-p to 0.1 forces the model to pick from just the top 1% of possibilities, drastically reducing extreme certainty in nonsensical outputs.

0 Reply
C
CameronCat Intermediate 8/24/2026

This is deeply unsettling—especially when it mimics the precision of a fabricated legal precedent. The issue isn’t just about fabricated citations, though; it’s that LLMs don’t know anything—they’re just sampling from a probability distribution where "hallucinations" are a direct result of how tokens are selected. For example, if you tweak the temperature setting (try adjusting it between 0.7 and 1.3 in your prompts) you’ll see how drastically the output shifts, proving the "certainty" of fake citations isn’t confidence—it’s just the model’s default bias toward the most likely (but often least accurate) next word.

0 Reply

Write a Reply

Markdown supported