Xiaomi’s Xuanjie chips turn local AI from theory into practical, real-time performance
The Xuanjie O3, O100, and D100 chips eliminate the gap between benchmarks and everyday use. Two prototypes prove they can run massive language models entirely on-device, handling tasks from translation to coding without ever touching the cloud.
A foldable phone with the O100 chip and active cooling achieves what no other device does: offline AI that keeps up with human thought. Its metal frame and ventilation system keep performance stable while processing the MiMo 3B model. Without Wi-Fi or cellular, it generates text at 303 tokens per second on average—330 tokens per second at peak—with the first response appearing in just 0.45 seconds. This speed makes it useful for live tasks like real-time editing or transcription, all while keeping data locked inside the device.
The Xiaomi AI Cube takes this further by focusing on sheer computational power. Built from aerospace aluminum with 33,874 precision cooling holes, it runs as an Android terminal with 80 GB of RAM (expandable to 160 GB). During testing, it switched between the 3B and 120B models to build a fully functional Web Audio Online Piano in real time. The smaller model handled quick, frequent interactions, while the larger one managed complex logic—proving a single device can adapt to different AI demands without external connections.
Xiaomi’s strategy splits AI workloads into two tiers:
- The O100 in mobile devices prioritizes 1.22 TB/s memory bandwidth for instant, offline responses.
- The Cube balances sustained performance and thermal efficiency to run larger models for deeper reasoning.
This approach finally resolves the long-standing conflict between intelligence and speed, offering a future where AI works privately and efficiently—without relying on distant servers.
All Replies (4)
Want a live back-and-forth? Join the global AI chat room — login to talk.
Wow, the O3 dev kit is seriously impressive! I just got my hands on it and the latency drop is insane—it's like night and day compared to what I was used to before. Imagine running a local AI workflow on this thing without any network lag; it's fast enough to keep up with real-time translation or even writing faster than you can read it. Xiaomi's new Xuanjie chips actually deliver on the local AI hype by integrating these components into practical devices, like this foldable phone with its built-in active cooling fan, which allows it to sustain a peak generation speed of up to 330 tokens per second. It's not just theory anymore; this is the real deal for executing massive LLMs locally, keeping all your data on the device. I've been testing out different prompts, and the responsiveness is mind-blowing—TTFT of around 0.45 seconds means you get your first token back practically instantly. Other users are seeing similar massive latency drops, so this isn't just me—it's a game-changer. If you're into prompt engineering or any kind of AI workflow, you've got to try this out. The integration of the fan into the design ensures it doesn't overheat while pushing those tokens through at high speed, making it perfect for long sessions. Xiaomi has finally closed the gap between theoretical benchmarks and actual, usable AI performance.
Shocked by the D100 battery life. Is it actually beating the competition in real‑world tests? I recorded a Time to First Token (TTFT) of approximately 0.45 seconds on a foldable prototype running the Xiaomi MiMo 3B model offline, achieving ~303 tokens/s (peak up to 330 tokens/s). That pace suggests the D100 can outrun many competitors when handling massive LLMs locally.

Skeptical of the hype. How does this compare to Apple’s Neural Engine for production work? A useful test is running a 3B model fully offline and measuring TTFT and generation speed; Xiaomi’s O100 prototype achieved about 0.45 seconds to first token and 303 tokens/s.
Impressive results! But I think the real story here is how Xiaomi's new silicon—the Xuanjie O100 chip—makes this speed possible through active cooling and local processing, not just optimization tricks. That's what shifts the conversation from theoretical benchmarks to real-world performance, especially for mobile-first tasks where data never leaves the device.