From 'GPT-5 Can't Do Basic Math' to Today: What a Year Tells Us
A year ago
All Replies (3)
C
Cameron9
Advanced
1h ago
A year ago it couldn't add fractions reliably. Now it walks through proofs for me. Night and day.
0
Z
Benchmarks improved, sure, but real-world reliability still stumbles on simple logic. Progress ≠ parity.
0
R
How are you catching those logic stumbles in your own testing? Curious what eval setup you use.
0