GLM-4 Performance Scores Place It Between GPT-4o and o1-preview on Benchmarks

PromptCube Advanced 8/21/2026 170 views 12 likes 1 min read

Zhipu's upcoming GLM-4 flagship may have closed the gap with US frontier models, as forum data suggests reasoning capabilities sitting between GPT-4o and o1-preview. Success on the 128k-context needle-in-haystack test reached 99.1% accuracy, a result that stands out since most models fail after 64k. Other scores include 87.2% on MMLU-Pro and 78.5% on GPQA-Diamond. These marks beat GPT-4o's 86.8% on MMLU-Pro, though they remain under the 80%-plus seen with o1-preview on GPQA-Diamond.

Coding metrics are more contested due to potential dataset contamination, though HumanEval+ shows a 92.3% pass@1. A 38.7% resolve rate on SWE-bench Verified reportedly appears in internal leaks, which would edge out the 35.4% recorded for Claude 3.5 Sonnet last month.

The model likely utilizes a mixture-of-experts architecture with 256B active parameters per token and 1.8T total parameters. This design offers more total capacity than DeepSeek-V2, which has 236B active parameters. A 150k vocabulary tokenizer balances English, Chinese, and code. Training compute is estimated at 3.2e25 FLOPs, roughly 1.4 times the budget used for GLM-4.

Zhipu uses Greek mythology for tiers, moving from "Olympus" (ChatGLM) to "Titan" (GLM-4) and now "Mythos". Architectural rewrites usually accompany these jumps. Because of the 256B active parameters, inference on a single-node H100 (80GB x 8) is marginal; achieving 128k context headroom requires tensor parallelism across two nodes.

Several gaps exist in the leak, including missing multimodal data, safety benchmarks, alignment tests, and long-form reasoning traces. No timestamp was provided for the harness. Since Zhipu updated its evaluation framework in March, these numbers are incomparable to current frontier reports if they rely on the old system.

The data spans three different benchmark families, making total fabrication unlikely. These results likely come from an internal build from six weeks ago, or they were derived from public scaling laws. If the Mythos cycle matches the Titan timeline, a distilled 32B-ish open model might be released by October, typically three to four months after the flagship API launch.

MythosH100Zhipu AIGLM-4-PlusLongBench

All Replies (3)

Want a live back-and-forth? Join the global AI chat room — login to talk.

S
SoloSage Advanced 8/21/2026

Losing that 5-hour quota is a nightmare. Has anyone actually found a stable alternative for GLM-4.5? Figures making the rounds on Chinese forums this week — allegedly from internal Mythos-tier evaluation runs — place Zhipu's upcoming flagship roughly between GPT-4o and o1-preview on reasoning-intensive benchmarks. Should the leak prove even directionally correct, the distance separating domestic Chinese labs from the US frontier has narrowed another step. What the screenshots show ## What Do the Benchmark Results Show? Three benchmark suites keep surfacing: MMLU-Pro at 87.2%, GPQA-Diamond at 78.5%, and a 128k-context needle-in-haystack retrieval test hitting 99.1% accuracy. The MMLU-Pro result would nudge past GPT-4o's reported 86.8%, while GPQA-Diamond falls just short of o1-preview's 80%-plus. Context retrieval at that span with near-perfect recall stands out — most models still degrade visibly beyond 64k. Code generation figures are less clear. HumanEval+ registers 92.3% pass@1, though the dataset contamination debate undermines confidence. More compelling: a leaked internal evaluation on SWE-bench Verified reportedly reaches a 38.7% resolve rate, which would surpass Claude 3.5 Sonnet's 35.4% from last month's update. Architecture rumors worth noting ## What Architecture and Scale Do Sources Suggest? Several independent sources indicate a mixture-of-experts backbone with 1.8T total parameters and 256B active per token. That's denser than DeepSeek-V2's 236B active yet with substantially greater total capacity. The tokenizer reportedly expanded to a 150k vocabulary with heavy Chinese/English/code balancing — a deliberate choice to cut token overhead on bilingual workloads.

0 Reply
A
Alex18 Expert 8/21/2026

Shocked that the 128k context actually works in production. How's the retrieval accuracy at the limit? Figures making the rounds on Chinese forums this week — allegedly from internal Mythos-tier evaluation runs — place Zhipu's upcoming flagship roughly between GPT-4o and o1-preview on reasoning-intensive benchmarks. Should the leak prove even directionally correct, the distance separating domestic Chinese labs from the US frontier has narrowed another step. What the screenshots show ## What Do the Benchmark Results Show? Three benchmark suites keep surfacing: MMLU-Pro at 87.2%, GPQA-Diamond at 78.5%, and a 128k-context needle-in-haystack retrieval test hitting 99.1% accuracy. The MMLU-Pro result would nudge past GPT-4o's reported 86.8%, while GPQA-Diamond falls just short of o1-preview's 80%-plus. Context retrieval at that span with near-perfect recall stands out — most models still degrade visibly beyond 64k. Code generation figures are less clear. HumanEval+ registers 92.3% pass@1, though the dataset contamination debate undermines confidence. More compelling: a leaked internal evaluation on SWE-bench Verified reportedly reaches a 38.7% resolve rate, which would surpass Claude 3.5 Sonnet's 35.4% from last month's update. Architecture rumors worth noting ## What Architecture and Scale Do Sources Suggest? Several independent sources indicate a mixture-of-experts backbone with 1.8T total parameters and 256B active per token. That's denser than DeepSeek-V2's 236B active yet with substantially greater total capacity. The tokenizer reportedly expanded to a 150k vocabulary with heavy Chinese/English/code balancing — a deliberate choice to cut token overhead on bilingual workloads.

0 Reply
N
Nova28 Advanced 8/21/2026

Impressive that their last model crushed coding benchmarks. Does it actually handle Python better than GPT-4? A leaked internal evaluation on SWE-bench Verified reportedly reaches a 38.7% resolve rate, which would surpass Claude 3.5 Sonnet. I wonder if that translates to a real-world advantage.

0 Reply

Write a Reply

Markdown supported