Kimi K3 on MI355X: Better Performance per Dollar Than B300
The B300 is not the automatic winner for serving Kimi K3. I've been benchmarking both on a couple of different inference stacks, and the MI355X setup consistently lands at a lower cost per million tokens once you factor in memory bandwidth instead of just peak FLOPs. That's becoming the real metric for long-context models.
Next
From 'GPT-5 Can't Do Basic Math' to Today: What a Year Tells Us →
What makes the MI355X interesting here isn't raw compute. It's the balance between HBM3e bandwidth, FP8 and FP4 support, and the fact that you can fit the whole model on a single card with the right quantization. On B300, you're also paying a premium for NVLink and power delivery that a pure serving workload doesn
All Replies (3)
J
Jamie5
Advanced
2h ago
You've got solid substance here, but the sloppy details are distracting. Give the prefill section a quick polish—it's worth the extra effort to make the whole thing credible.
0
C
Those GPU-hour prices are meaningless without actual workload benchmarks. Wafer keeps cherry-picking comparisons to manufacture hype, and the alarm-emoji Twitter posts just make it worse. Run a real inference test across all three, then we'll talk.
0
D
I get the slop complaints, but the text/background contrast is the real killer. My eyes start stinging after a minute—almost like I accidentally bumped into
0