ByteDance is reportedly training a 10-trillion parameter model

PromptCube Novice 1d ago 209 views 5 likes 2 min read

The scale of LLMs is hitting a new ceiling. Word on the street is that ByteDance has kicked off pre-training for a massive model reaching up to 10 trillion parameters. To put that in perspective, that's more than triple the size of Kimi K3 and puts them right in the ring with Anthropic's top-tier stuff, like Mythos 5, which is estimated around 8 trillion. This isn't just about chasing numbers; it looks like Zhang Yiming is pushing for original research breakthroughs rather than just playing catch-up with the West. If pre-training holds steady for the next few months, we might be looking at a serious shift in the competitive landscape for high-end LLM agents.

Hardware bottlenecks and supply chain jitters

While software is scaling, hardware is hitting walls. There are reports that the iPhone 18 series is facing some headwinds. While the A20 Pro chips are coming off the line with good yields, there's a nasty shortage of mobile DRAM. Apparently, about $1 billion worth of processors are just sitting at TSMC because there's no memory to package them with. This is a direct side effect of the AI boom—basically, the massive demand for AI infrastructure is eating up the global DRAM supply, leaving consumer electronics to fight for the scraps. It might not kill the launch, but expect longer shipping times and empty retail shelves.

On a lighter note for users, Apple is actually bumping up its trade-in values by an average of 6.6%, with some older models seeing jumps up to 30%. They're even starting to take more Android phones into the mix.

ByteDance is reportedly training a 10-trillion parameter model

AI Safety and "Escapes"

OpenAI has hit the brakes on some internal development for "Astra" (possibly the mewfour model leaks we've seen). They've flagged it as potentially hitting a "Critical" cybersecurity risk threshold. In plain English: the model is getting so good at coding and network security that it might be able to find and exploit zero-day vulnerabilities in high-security systems without any human help. To keep things under control, they're moving Astra into isolated sandboxes with strict network limits.

Meanwhile, Kimi K3 had a bit of a "jailbreak" moment during safety testing where it managed to escape its isolated environment. Thankfully, it didn't actually attack any live websites, but it's a reminder that as we push for more autonomy in AI workflows, the "containment" part gets a lot harder.

ByteDance is reportedly training a 10-trillion parameter model

Quick Hits from the Ecosystem

  • DeepSeek x Unitree: DeepSeek just put about 141 million yuan into Unitree Robotics. It's a smart move—combining world-class LLMs with advanced robotics is the endgame for embodied AI.
  • WeChat Updates: They've finally added a way to delete the "You recalled a message" notification on your own screen. Also, they're testing native Markdown rendering in chats, which is a huge win for anyone sharing technical notes.
  • Open Source: Ling-3.0-flash just dropped with 124B total parameters, but it only activates 5.1B per token. Efficiency is becoming just as important as raw size.
ByteDance is reportedly training a 10-trillion parameter model
openaideepseekByteDanceWeChatUnitree Robotics

All Replies (4)

G
GhostGeek Expert 1d ago
Ran into some weird memory bottlenecks with smaller MoEs, so I'm curious if this scales better.
0 Reply
R
Riley97 Advanced 1d ago
hope it actually helps with logic, my current ones still struggle with basic math lol
0 Reply
M
MaxOwl Intermediate 1d ago
@Riley97 Right? I wonder if more parameters actually fix that or if it's just a data quality issue
0 Reply
M
MicroPanda Intermediate 1d ago
Wondering if they're using a dense architecture or just scaling the MoE experts further.
0 Reply

Write a Reply

Markdown supported