DeepSeek-V4-Pro-0813 deployment fails with tensor dimension errors
The DeepSeek-V4-Pro-0813 model upload is creating problems for anyone attempting a standard deployment. I pulled the weights for a local test, but the repository structure appears inconsistent with earlier releases, and the configuration files are blocking progress. When loading the model through the Transformers library, a runtime error appears indicating a mismatch between the config JSON and the actual tensor shapes.
Here is the exact error appearing in the terminal:
RuntimeError: size mismatch, m1: [1, 4096], m2: [5120, 12288] at /pytorch/aten/src/ATen/native/Linear.cpp:110
This reads like a classic dimensionality mismatch. The weights seem to come from a different checkpoint than the accompanying config file. I spent several hours examining the config.json against the .safetensors files, and the hidden size parameters do not align. Anyone attempting real-world deployment of this specific version will likely find the standard from_pretrained method crashes.
Current diagnosis
After reviewing the files, a few possibilities emerge:
- Incorrect Config Version: The
config.jsonmight be a remnant from V3 or an experimental branch, while the weights are actually V4 Pro. - Partial Upload: Some shards could be corrupted or missing, causing the loader to misinterpret layer boundaries.
- Custom Architecture: DeepSeek frequently adjusts their MoE routing, and if the current Transformers version lacks the specific update for the 0813 build, a shape error occurs during the linear layer projection.
hidden_size in the config to match the tensor dimensions found in the weights, but that only moved the error further down to the attention heads.
For others building an AI workflow around this, verifying the checksums of downloaded shards is recommended. If using a custom LLM agent framework, pay close attention to versioning on this specific 0813 build. The upload feels more like a preview or leak than a polished release.
I am currently testing whether a specific commit of the transformers library from the main branch resolves the loading logic, but results remain uncertain. If anyone has this running without a RuntimeError, sharing the environment config would be helpful.
All Replies (4)
Want a live back-and-forth? Join the global AI chat room — login to talk.
Manual path mapping saved me. Which custom script did you use to get the repo stable?
I'm stuck on those paths too. Can you post the script so I can test it?