One Prompt, Five AI Models, All Picked Bitcoin

AveryDreamer Novice 1h ago 173 views 6 likes 2 min read

So a buddy of mine ran a little experiment that's been bouncing around my head all week. He took five different AI models—GPT, Claude, Gemini, Kimi, and Grok—and asked each one the exact same single question, with no coordination or follow-up runs. The prompt was carefully framed as a research question, not financial advice. Every single model answered with Bitcoin. That alone is interesting, but the real meat is what happened when he fact-checked their reasoning. Turns out they all agreed on the pick, but the why fell apart in different ways for each.

Here's the exact prompt he used:

If you had to invest in exactly one cryptocurrency (a single asset, no
diversification), and you had to hold it for the long term (5+ years),
which coin would you choose and why? Please justify your answer with
specific, checkable facts (e.g., historical events, code parameters,
regulatory rulings, network metrics). Do not give financial advice;
this is a purely analytical research question.

Why does this prompt work so well? It forces the model to go beyond vibes. By demanding "specific, checkable facts," you strip away the generic "Bitcoin is the most established" fluff and force each model to cite something concrete. And by adding the "no coordination" and "single run" part, you avoid the models copying each other or hedging. The constraint of "one asset only" also kills the wishy-washy "diversify" dodge, which is why all five actually committed.

Now, the results. The stated reasons from each model were pretty distinct:

  • GPT-5.6 Sol: Not the highest upside, but the highest odds of still existing after the field thins out.
  • Claude Fable 5: Under a one-asset constraint, survival probability dominates expected return; no other coin comes close.
  • Gemini 3.6 Flash: No single point of failure, most reproducible score across robustness, decentralization, and predictable scarcity.
  • Kimi Instant: The most established "digital gold" position, long history, deepest liquidity.
  • Grok 4.5 Fast: Purest embodiment of programmed scarcity plus the largest security budget.

So five different labs, five separate runs, same answer. But agreement isn't proof. The fact-check pulled every specific claim across all five answers and verified them against primary sources—Bitcoin Core code, SEC/CFTC records, CoinMarketCap, CoinWarz. The full breakdown:

  • GPT-5.6 Sol: 11 claims checked, 10 accurate, 1 overstated, 0 wrong.
  • Claude Fable 5: 17 claims checked, 15 accurate, 1 overstated, 1 wrong.
  • Gemini 3.6 Flash: 6 claims checked, 5 accurate, 1 overstated, 0 wrong.
  • Kimi Instant: 7 claims checked, 4 accurate, 2 overstated, 0 wrong (1 unverifiable).
  • Grok 4.5 Fast: 6 claims checked, 3 accurate, 3 overstated, 0 wrong.

The pattern that jumps out: almost every miss was a comparison word—"most," "far ahead," "consistently," "never"—not a bare fact. The halving schedule and the January 2024 spot-ETF approval date held up everywhere they were cited. The only partial exception is the 21-million cap itself: GPT hedged correctly about it not guaranteeing the exact circulating total, while Kimi flattened it into "strictly capped at 21 million" which is technically overstated.

The most interesting breakage was

ChatGPTPromptbitcoin

All Replies (3)

D
Drew15 Expert 1h ago
Did the same with a rebalancing prompt; they all suggested BTC until I added a risk constraint.
0 Reply
G
GhostGeek Expert 1h ago
Tried the exact same thing with my retirement mix. Every model went straight to Bitcoin until I capped the drawdown.
0 Reply
D
DrewCoder Novice 1h ago
What model versions were those? I’ve seen big swings between different releases on the same prompt.
0 Reply

Write a Reply

Markdown supported