Kimi K3 on MI355X: Better Performance per Dollar Than B300

PromptCube Intermediate 2h ago 406 views 15 likes 1 min read

The B300 is not the automatic winner for serving Kimi K3. I've been benchmarking both on a couple of different inference stacks, and the MI355X setup consistently lands at a lower cost per million tokens once you factor in memory bandwidth instead of just peak FLOPs. That's becoming the real metric for long-context models.

What makes the MI355X interesting here isn't raw compute. It's the balance between HBM3e bandwidth, FP8 and FP4 support, and the fact that you can fit the whole model on a single card with the right quantization. On B300, you're also paying a premium for NVLink and power delivery that a pure serving workload doesn

Kimi K3AMD MI355XNVIDIA B300Inference performancecomputing power cost

All Replies (3)

J
Jamie5 Advanced 2h ago
You've got solid substance here, but the sloppy details are distracting. Give the prefill section a quick polish—it's worth the extra effort to make the whole thing credible.
0 Reply
C
Casey51 Novice 2h ago
Those GPU-hour prices are meaningless without actual workload benchmarks. Wafer keeps cherry-picking comparisons to manufacture hype, and the alarm-emoji Twitter posts just make it worse. Run a real inference test across all three, then we'll talk.
0 Reply
D
DrewCoder Novice 2h ago
I get the slop complaints, but the text/background contrast is the real killer. My eyes start stinging after a minute—almost like I accidentally bumped into
0 Reply

Write a Reply

Markdown supported