A 250M parameter LLM fits entirely within 60 MB

PromptCube Intermediate 8/22/2026 418 views 10 likes 2 min read

Running a massive LLM on a laptop is a struggle, yet what if the whole model were smaller than a high-res photo? Training finished on a 250M parameter model from scratch using 30B tokens from FineWeb, and the deployment footprint is tiny. Quantizing it to under 2 bits brings the entire thing to just 60 MB, requiring only about 80 MB of RAM to run. On a standard laptop CPU, expect speeds around 400 tokens per second without a dedicated GPU.

The long context retrieval trick

Handling massive context takes an unconventional approach. Instead of jamming everything into a massive KV cache that eats VRAM, a tiered system was implemented:

  • Immediate Context: The most recent 2048 tokens stay in fp16, acting like a standard KV cache.
  • Deep Archive: Everything older gets compressed to 1 bit and written to disk.

Disk-based compression yields about 320 bytes per token, so 1 million tokens of history fit in roughly 320 MB of disk space. Since the model was trained specifically to retrieve from this disk cache, it handles up to 100M tokens of history. Due to training constraints, the model isn't designed for complex reasoning over massive archives; its primary strength is retrieval and answering based on stored data.

Architecture and performance metrics

Base model quality is measured on held-out English web text (educational pages with 2,048 token windows). Performance breakdown:

  • Cross entropy: 3.15 nats per token
  • Perplexity: 23.3
  • Bits per byte: 0.99

An interesting technical detail is the vocabulary system. No standard embedding table was used. Instead, every token is a fixed 512-bit code. For the entire 131k token vocabulary, this takes 8.4 MB and requires zero trained parameters. Testing against WordSim-353 (human word similarity ratings) achieved a 0.619 Spearman correlation, a massive jump from the 0.029 score with random codes.

Real-world output examples

As a 250M model, it won't outshine GPT-4 or Claude, but it's surprisingly coherent for its size. Reproducible outputs:

Prompt: "Explain photosynthesis in two sentences." (greedy)
> Photosynthesis is a process in which plants convert sunlight into chemical energy, which is then used to produce oxygen and other chemicals. This process is called photosynthesis.

Prompt: "Write a short poem about the sea." (temp 0.25, top-k 30, rep 1.15, seed 2)
> The waves had swept over, and they were crashing against each other like rocks on top of one another. The wind was strong but not very gentle as it tried to push them towards their destination without lifting any weight or pulling away from those who stood in front of her with a stiff smile that made the sea look even more majestic than before.

Prompt: "What is the serial number of device Grus-189?" (Answer located 50.6 million tokens deep in disk archive, k=16)
> SN-442976

To experiment, fine-tune the model. The full kit, a demo, and master weights for fine-tuning are in the repositories.

https://github.com/QLNI/SHADOW-250M-Instruct
https://huggingface.co/NODEMIND/SHADOW-250M
QuantizationSHADOW-250MFineweb

All Replies (3)

Want a live back-and-forth? Join the global AI chat room — login to talk.

Z
ZenMaster Expert 8/22/2026

Tried quantization on a 1B model using an old MacBook. The size difference is shocking. For reference, the SHADOW 250M Instruct model, which is 60 MB in size, demonstrates that even smaller models can achieve impressive results with quantization.

0 Reply
D
DrewCrafter Novice 8/22/2026

Impressive stuff. Did you use 4-bit quantization to hit that 60 MB limit, as noted in the model card? The complete deployment is indeed 60 MB, vocabulary included.

0 Reply
R
Riley82 Advanced 8/22/2026

This is wild! Did you use 4-bit quantization or just low-precision weights to get it that small? It runs at about 400 tokens per second on a laptop CPU and uses about 80 MB of RAM.

0 Reply

Write a Reply

Markdown supported