BDH-CQ achieves 29.5% on ARC-AGI-1 using only 150M parameters
The cost-accuracy trade-off in AGI benchmarks is shifting as BDH-CQ tackles ARC-AGI-1 tasks for a fraction of a cent per run. Its appeal lies not only in efficiency but in its use of recurrent latent reasoning. Rather than relying on standard Chain of Thought, which consumes tokens and introduces linguistic noise by writing out reasoning in plain English, BDH-CQ conducts its thinking within a high-dimensional latent workspace.
How does BDH-CQ achieve in-context learning?
It manages in-context learning by updating a recurrent memory upon encountering a new task. To solve a query, the model iterates internally. The vital takeaway is that intermediate reasoning states are never decoded into language, resulting in essentially silent reasoning.
How the latent workspace actually functions
Can BDH-CQ outperform prompt engineering in LLM agents?
While typical LLM agents use prompt engineering to force verbal step-by-step breakdowns, BDH-CQ integrates memory and inference into a single computational fabric. The process follows these steps:
- Memory Update: After receiving demonstrations of an unseen task, the model updates its recurrent memory instead of simply storing them in a KV cache.
- Latent Iteration: The query undergoes iterative computation, cycling through the latent space to refine the solution.
- Direct Output: The model moves straight to the solution without verbalizing any scratchpad steps.
Does BDH-CQ's architecture eliminate the need for task identifiers?
This architecture removes the need for task identifiers or specific demonstration pairs during training. Remarkably, no parameters are updated during inference; the entire process occurs through the recurrent state.
Performance and Efficiency
The performance metrics are striking given the model's scale. A 150M-parameter setup, which is tiny compared to industry behemoths, reached a 29.5% pass@2 on ARC-AGI-1.
How does BDH-CQ's cost compare to larger models?
- Model Size: 150M parameters
- ARC-AGI-1 Pass@2: 29.5%
- Cost per task: $0.00070
Compared to larger models that attempt to brute-force AGI benchmarks using massive prompt windows, BDH-CQ demonstrates that recurrent latent states can outperform token-heavy reasoning in efficiency. This serves as a real-world example of how moving away from thinking out loud may improve generalization on abstract reasoning tasks.
The full paper is available here:
https://arxiv.org/abs/2608.09888All Replies (3)
Want a live back-and-forth? Join the global AI chat room — login to talk.
Stunned by that accuracy. Does it actually hold up on the ARC-AGI-2 set or is it just overfitted?