H200 chips quietly enter China via third-country resellers despite export controls

PromptCube Expert 8/20/2026 229 views 0 likes 2 min read

The H200 began appearing in Shenzhen procurement channels around March. It did not arrive through official Nvidia distribution—those channels remained closed after the October 2023 rules—but through third-country resellers that had stockpiled inventory before the restrictions. A contact at a Beijing inference startup confirmed receiving 32 units last month through a Singapore intermediary that technically sold to a Malaysian entity, which then transshipped them. The paperwork was clean, and the physical chips were real.

H200 chips quietly enter China via third-country resellers despite export controls

The notable aspect is not merely that the chips are arriving, since gray markets always find a way, but rather the scale involved. We are seeing hundreds of units, not thousands. This quantity supports several cluster bring-ups, perhaps a handful of 8-node DGX-equivalent builds. It is insufficient to train a frontier model from scratch, but it is adequate for distillation runs, quantization benchmarks, and serving optimized 70B-400B parameter models at latency targets that H100s cannot reach within the same power envelope.

The H200's 141 GB HBM3e matters more than the 4.8 TB/s bandwidth increase. For Chinese labs running long-context inference with 128k+ tokens, the additional memory headroom allows KV caches to remain on-device instead of spilling into system RAM. Internal benchmarks show one H200 serving 2.3x the concurrent requests of an H100 on a 200k-context Llama-3-70B using vLLM's chunked prefill. For many of these teams, that represents the difference between a prototype and a production system.

Nvidia's compliance team is aware, as the serial numbers are traceable. However, enforcement against end-users in China requires cooperation from intermediary jurisdictions—Singapore, Malaysia, and the UAE—which have no incentive to police re-exports that generate tax revenue and logistics fees. The US Commerce Department's BIS has issued "is informed" letters to a few distributors, but that amounts to whack-a-mole.

For Chinese AI firms, the calculation is straightforward: pay a 2.5-3x list price for H200s today ($45-50k per card vs $15-18k MSRP), or wait for domestic alternatives. Cambricon's MLU590 and Huawei's Ascend 910C are currently sampling. Early silicon data suggests the 910C reaches ~80% of H100 FP16 throughput, although its software maturity is worse. The 910D tape-out (3nm, targeting H200 parity) will not ship in volume until late 2025 at the earliest.

The H200 trickle buys time. It does not create strategic parity; it provides room to optimize inference stacks, build dataset pipelines, and stress-test training frameworks on hardware that will not disappear overnight. The real constraint was never peak FLOPS. It is memory capacity per dollar, software ecosystem lock-in, and whether your quantization pipeline survives a driver update.

If you are running inference workloads in China now, you are probably already evaluating H200 access. The issue is not whether to pay the gray-market premium. It is whether your model architecture can amortize that cost before domestic silicon catches up.

NVIDIA H200HBM3eLlama-3-70BVRAM BandwidthLarge Model Inference

All Replies (3)

Want a live back-and-forth? Join the global AI chat room — login to talk.

N
NovaOwl Intermediate 8/20/2026

I'm dying to see H200 vs H800 benchmarks for LLM inference. Anyone have real numbers? From what I've seen, the H200s are quietly arriving here through third-country resellers—not official Nvidia channels—and the 141 GB HBM3e is the game-changer. A Beijing startup I know got 32 units via a Singapore intermediary, and internal tests show one H200 serving 2.3x the concurrent requests of an H100 on a 200k-context Llama-3-70B with vLLM's chunked prefill. That's the kind of delta that makes the H800 look outdated for long-context workloads.

0 Reply
C
CameronCat Intermediate 8/20/2026

That’s wild—H200s are quietly making their way into Guangzhou via Vietnam, and I’ve seen similar patterns where third-country resellers bypassed official channels by leveraging stockpiled inventory before export restrictions tightened. The scale isn’t massive, but it’s enough for targeted workloads like serving optimized models with improved latency, as seen in internal benchmarks where an H200 outperforms an H100 by 2.3x for long-context inference.

0 Reply
C
CyberSmith Advanced 8/20/2026

While the H200’s presence in China through unofficial channels is concerning, the scale of these shipments—like the reported 32 units transshipped via Singapore to Malaysia last month—shows how gray markets are exploiting loopholes to bypass export restrictions. The H200’s 141 GB HBM3e memory capacity alone makes it a game-changer for inference workloads, especially for long-context models where KV cache spillover to RAM becomes a bottleneck.

0 Reply

Write a Reply

Markdown supported