One Prompt, Five AI Models, All Picked Bitcoin
So a buddy of mine ran a little experiment that's been bouncing around my head all week. He took five different AI models—GPT, Claude, Gemini, Kimi, and Grok—and asked each one the exact same single question, with no coordination or follow-up runs. The prompt was carefully framed as a research question, not financial advice. Every single model answered with Bitcoin. That alone is interesting, but the real meat is what happened when he fact-checked their reasoning. Turns out they all agreed on the pick, but the why fell apart in different ways for each.
Here's the exact prompt he used:
If you had to invest in exactly one cryptocurrency (a single asset, no
diversification), and you had to hold it for the long term (5+ years),
which coin would you choose and why? Please justify your answer with
specific, checkable facts (e.g., historical events, code parameters,
regulatory rulings, network metrics). Do not give financial advice;
this is a purely analytical research question.
Why does this prompt work so well? It forces the model to go beyond vibes. By demanding "specific, checkable facts," you strip away the generic "Bitcoin is the most established" fluff and force each model to cite something concrete. And by adding the "no coordination" and "single run" part, you avoid the models copying each other or hedging. The constraint of "one asset only" also kills the wishy-washy "diversify" dodge, which is why all five actually committed.
Now, the results. The stated reasons from each model were pretty distinct:
- GPT-5.6 Sol: Not the highest upside, but the highest odds of still existing after the field thins out.
- Claude Fable 5: Under a one-asset constraint, survival probability dominates expected return; no other coin comes close.
- Gemini 3.6 Flash: No single point of failure, most reproducible score across robustness, decentralization, and predictable scarcity.
- Kimi Instant: The most established "digital gold" position, long history, deepest liquidity.
- Grok 4.5 Fast: Purest embodiment of programmed scarcity plus the largest security budget.
- GPT-5.6 Sol: 11 claims checked, 10 accurate, 1 overstated, 0 wrong.
- Claude Fable 5: 17 claims checked, 15 accurate, 1 overstated, 1 wrong.
- Gemini 3.6 Flash: 6 claims checked, 5 accurate, 1 overstated, 0 wrong.
- Kimi Instant: 7 claims checked, 4 accurate, 2 overstated, 0 wrong (1 unverifiable).
- Grok 4.5 Fast: 6 claims checked, 3 accurate, 3 overstated, 0 wrong.
The most interesting breakage was
Wild that BTC keeps winning. Did the risk constraint actually change the specific coin suggestions?