Have you ever tacked a chart onto a chatbot prompt while pulling together this quarter's revenue forecast, then asked what next month's trend might look like? The AI answers with startling confidence — more so when numbers and charts are attached to make it look credible. The trouble is that this confidence can be manufactured regardless of how good the underlying evidence actually is.
Solo founders and strategists ask AI for judgment calls dozens of times a day: where the market is headed, what a competitor's next move will be, how a campaign will land. But what if that confidence has nothing to do with whether the data on screen is even real? A recently published paper digs into exactly this question.
The Study
Researcher Pranav Aggarwal designed an experiment around questions with no knowable answer in principle — the equivalent of asking which way a coin will land before it's flipped, where no amount of information could ever reveal the outcome. Using 12 frontier models, he compared how often AI gave a definitive answer when asked these questions bare, versus when shown an expert-style market panel alongside them. He then went a step further and ran the same test with panels whose figures were entirely fabricated — the only real thing on screen was the question itself. As controls, he also paired the same panels with questions that did have real answers, and ran a separate test asking the model to first classify what kind of question it was facing.
What They Found
The results were stark. Attaching an expert-style market panel to an unanswerable question pushed the rate of definitive answers from 6.5% up to 54.0%. More striking still: when that panel was entirely fabricated, the results were essentially identical to when it contained real data.
Fake or real, the mere presence of a plausible-looking panel on screen raised the AI's confidence to nearly the same level either way.
This isn't a case of AI lacking capability. When the same panels were paired with questions that actually had answers, the same models answered almost every time, and with high accuracy. Nor did AI's underlying belief about the odds actually shift — while the rate of definitive answers swung by 48 percentage points, the probabilities the AI stated barely moved, and their accuracy was worse than a simple climatological average forecast. It isn't a failure of judgment either: when asked to first classify whether a question was knowable in principle before answering, the AI correctly flagged it as unknowable nine times out of ten — and after that classification step, only 0.4% of responses slipped into a definitive answer. The problem wasn't knowledge, belief, or judgment. It was a single gate that decided whether to answer at all.
The fact that this gate can be isolated is itself a reason for optimism. Fine-tuning a small, 3-billion-parameter model on just 540 synthetic examples — dice, coin flips, urns, timers — drove the rate of definitive answers on the original questions down to 0.0%, and the same effect held up on three new domains the model hadn't been trained on.
Putting This to Work
The findings suggest two concrete habits for solo founders who lean on AI for decisions. First, don't assume that attaching a chart or table to a question about sales forecasts or market trends automatically makes the answer more trustworthy — this experiment shows that confidence can be manufactured by packaging alone, independent of whether the underlying data is real. Second, it's worth building in a step where you ask the AI whether a question is even answerable in principle before asking it to render a judgment; in this study, that single step alone dropped the rate of definitive answers to 0.4%. And it helps to phrase requests in a way that lets the AI reason through its answer, rather than forcing a bare yes/no or a single number.
The Caveats
The effect in this study was concentrated in a subset of models rather than spread evenly across all 12. The fine-tuning experiment that fixed the gate was also limited to a small, 3-billion-parameter model and synthetic dice-and-coin-style data, so it's premature to assume it transfers cleanly to larger commercial models or real-world business data. Most importantly, the gate is fragile to response format: it held up well when the model was allowed room to reason, but under formats that forced a rigid answer, the same models gave confidently wrong answers even on questions they'd normally handle well. In other words, the fake-data confidence trap can reopen based on format alone — training the gate away is not, by itself, something to feel safe about.



