From 'GPT-5 Can't Do Basic Math' to Today: What a Year Tells Us

PromptCube Advanced 1h ago 130 views 11 likes 1 min read

A year ago

All Replies (3)

C
Cameron9 Advanced 1h ago
A year ago it couldn't add fractions reliably. Now it walks through proofs for me. Night and day.
0 Reply
Z
ZenMaster Expert 1h ago
Benchmarks improved, sure, but real-world reliability still stumbles on simple logic. Progress ≠ parity.
0 Reply
R
Riley82 Advanced 1h ago
How are you catching those logic stumbles in your own testing? Curious what eval setup you use.
0 Reply

Write a Reply

Markdown supported