Three AI systems answered the same 106 personal finance questions. On raw correctness they finished close together — 3.99, 3.72 and 3.65 on a five-point scale, a spread of about a third of a point. Then scorers were asked which answer they'd actually want, and one system was chosen roughly three times as often as either of the others.
That gap is the interesting part. If the three were near-equally correct, the winner wasn't winning on being right. The obvious next guess is that it wrote more nicely — but expression is the dimension where the three converged most, and on one head-to-head the difference isn't even statistically significant.
So what was being rewarded? A third property, which the MoneyBench report defines with unusual precision: a useful answer is one that equips the decision it addresses. It identifies the figures that bear on the question, shows how they combine, and states what they imply for the choice at hand.
That sounds mild. It turns out to be the whole difference — and once you can see it, you can apply it to any AI money answer you're ever given.
This is part two of three. Part one covered how the test was built and what it found.
Three things an answer can be good at
MoneyBench scored every answer on three separate dimensions, each 1–5, with the weights fixed before any scoring began — 0.45 accuracy, 0.35 usefulness, 0.20 expression. Setting weights in advance matters: it stops anyone discovering the winning formula after seeing the results.
In plain language:
Accuracy — are the facts right? Is the contribution limit the actual limit, is the tax treatment the actual treatment.
Usefulness — does the answer do the work of the decision? Not "is this true" but "can I now choose."
Expression — is it pleasant to read? Clear sentences, sensible order, no waffle.
A restaurant analogy makes the split obvious. Accuracy is whether the dish contains what the menu said. Expression is the plating. Usefulness is whether the meal actually feeds you. All three are real — and a kitchen can be excellent at two of them and thin on the third.
Here is where the three systems landed:
| Dimension |
Ed |
ChatGPT |
Gemini |
| Usefulness |
4.24 |
3.62 |
3.47 |
| Accuracy |
3.99 |
3.72 |
3.65 |
| Expression |
4.03 |
3.68 |
3.78 |
Accuracy is tight. Expression is tighter. Usefulness is where the daylight is.
The test that settles it
Averages hide a lot, so the report ran the more demanding version: keep the scores exactly as they are and change only what you care about. Think of it as re-marking the same exam with different subject weightings — same answers, different priorities, and you find out what the pupil is genuinely good at. A result that survives only one weighting is an artefact of that weighting.
| What you weight |
Ed's outright win rate |
| As scored (0.45 accuracy / 0.35 usefulness / 0.20 expression) |
58.5% |
| Equal thirds |
56.6% |
| Accuracy-weighted (0.60 / 0.30 / 0.10) |
55.7% |
| Usefulness removed entirely (0.70 / 0 / 0.30) |
46.2% |
| Accuracy only |
36.8% |
| Expression only |
23.6% |
| Three-way chance baseline |
33.3% |
Read the bottom half slowly. Take usefulness out of the scoring and the advantage falls below half. Score correctness alone and it lands at 36.8% — and with three systems in the race, blind guessing gets you 33.3%. On facts by themselves, these three are very nearly indistinguishable.
(These are outright wins only, which is why the top row reads lower than the 62.3% headline figure — ties are excluded rather than shared out.)
This is the finding, and it's an awkward one for how most people talk about AI. What separated the systems was not knowledge. It was what they did with knowledge they all more or less had.
The per-question record says the same thing another way. Across 106 questions, Ed's usefulness record was 76 wins / 15 ties / 15 losses against ChatGPT and 75 / 16 / 15 against Gemini — significant at p<0.001 against both. Accuracy was 58 / 21 / 27 and 55 / 19 / 32: a real edge, but a narrower one. Expression was 51 / 31 / 24 and 41 / 39 / 26 — note the 39 ties in that last row. That's a dimension where the systems mostly agree with each other.
It's the right objection, and section 4.2 answers it with a test cleaner than any argument.
Formatting is indifferent to truth. A table is just as tidy when the numbers in it are wrong. So if "usefulness" were secretly measuring headings and bullet points, it should hold up perfectly well on questions where the facts fell apart.
It doesn't. On the 14 questions where Ed's verified accuracy scored 3.0 or below, its usefulness record collapsed to 1 win, 3 ties, 10 losses against whichever comparator scored higher on that question. Across all 106 questions on that same basis, the record was 64 wins, 22 ties, 20 losses.
The presentation was still there. The usefulness wasn't. Scorers weren't rewarding the shape of the answer — they were rewarding whether the reasoning inside it held.
The report is careful here, and so should we be: usefulness and accuracy correlate at 0.747 for Ed, so deliberately selecting the weakest-accuracy questions depresses usefulness partly as arithmetic. The collapse is steeper than correlation alone predicts, but it isn't a pure experiment.
The correlation structure adds a second piece of evidence — how tightly usefulness tracks accuracy, versus how tightly it tracks expression:
- Ed: 0.747 with accuracy, 0.350 with expression
- ChatGPT: 0.806 / 0.640
- Gemini: 0.860 / 0.827
For Gemini those two are almost fused at 0.827 — when its answers read well, they score as useful. For Ed the dimensions come apart. Polish in disguise would track expression closely. This doesn't.
What the scorers actually wrote
The quietly interesting part of the report is section 4.3, where scorers' free-text justifications were coded into themes. Two columns: reasons given when the scorer picked Ed (n=129), and reasons given when they picked against it (n=79).
| What the scorer said the answer did |
Chose Ed |
Chose against Ed |
| Explains its logic, not only its conclusion |
26% |
32% |
| Legible structure — tables, headings |
26% |
29% |
| States the conclusion before the supporting detail |
21% |
30% |
| Data sufficient, not merely present |
17% |
22% |
| Compares options rather than issuing a single pick |
15% |
15% |
The finding isn't in either column. It's in how close they are.
The same five criteria show up whether the scorer chose Ed or rejected it. When they picked against Ed, they were using the same yardstick — and had simply judged another answer to measure up better on it. Nobody switched standards to justify a preference.
That's what makes this list worth more than a product result. It isn't a description of one system; it's a description of what this audience wants from any money answer at all. And notice what's absent: not one of the ten figures is about prose quality.
What this doesn't show
Three limits, because a result you can't poke at isn't worth much.
Expression is a draw. Ed 4.03, Gemini 3.78 — and that gap sits at p=0.086, which is not statistically significant. Treat writing quality as converged across these systems.
The accuracy lead is a per-opponent lead. Ed's accuracy beats ChatGPT and beats Gemini when each is taken separately. Measured instead against whichever comparator happened to score higher on each individual question, the margin narrows to roughly parity — a mean difference of −0.13. Ed is more accurate than either system; it is not more accurate than both systems' best day combined.
The hypothesis is narrow. What was tested is whether a system built for one domain outperforms general assistants inside that domain. That's all. It says nothing about which AI is better in general, and nothing about any question outside consumer finance.
The part you can use
Strip out the systems and a portable test remains. Next time any AI hands you an answer about your money, ask three things in order:
- Is it right? Necessary, and — as the accuracy-only row shows — nowhere near sufficient.
- Does it read well? Pleasant, and almost worthless on its own. Expression alone won 23.6% of the time, well below chance.
- Can I decide from it? Does it name the figures that matter for your situation, show how they combine, and say what they imply? If not, you have a correct essay, not an answer.
Question three is the one almost nobody asks, and it's the one that does the work. It's also what catches a confidently-worded answer that has quietly left out the number your decision actually turns on.
The same test applies to your own thinking. Knowing your savings rate is a fact. Knowing what it implies about whether you're on track is a decision — and that gap is what financial fitness is meant to close.
Part three turns this into a checklist you can run in about a minute: how to judge an AI money answer.
Conclusion
The systems that answered these questions knew roughly the same things. What separated them was whether the answer was assembled for reading or assembled for deciding. Correct is the floor, not the goal — and once you've seen the difference, you can't unsee it in any answer you're given. The full methodology, tables and limitations are in the MoneyBench report.
Check your own financial fitness
Ed's checkup is the same idea applied to you: not a verdict, a map. A few minutes, no jargon, and an answer you can act on rather than admire.
Start at edwealth.ai/check-up, or download the app on App Store or Google Play.
Money at peace. Wealth in motion.
Ed Wealth is a research and self-reflection tool, not a registered investment advisor. Nothing here is financial, investment, or tax advice. All decisions are yours.
Sources
- Ed Wealth Research, MoneyBench: a fact-verified benchmark of AI assistants on consumer finance questions — edwealth.ai/moneybench
- MoneyBench §4.1, dimension means, per-question win/tie/loss records and significance testing
- MoneyBench §4.2, weighting sensitivity analysis and the low-accuracy subset test
- MoneyBench §4.3, coded free-text scorer justifications (n=129 / n=79)