A correct answer and a useful one are not the same thing

EdWealth
· Sep 04 2026
A correct answer and a useful one are not the same thing

Three AI systems answered the same 106 personal finance questions. On raw correctness they finished close together — 3.99, 3.72 and 3.65 on a five-point scale, a spread of about a third of a point. Then scorers were asked which answer they'd actually want, and one system was chosen roughly three times as often as either of the others.

That gap is the interesting part. If the three were near-equally correct, the winner wasn't winning on being right. The obvious next guess is that it wrote more nicely — but expression is the dimension where the three converged most, and on one head-to-head the difference isn't even statistically significant.

So what was being rewarded? A third property, which the MoneyBench report defines with unusual precision: a useful answer is one that equips the decision it addresses. It identifies the figures that bear on the question, shows how they combine, and states what they imply for the choice at hand.

That sounds mild. It turns out to be the whole difference — and once you can see it, you can apply it to any AI money answer you're ever given.

This is part two of three. Part one covered how the test was built and what it found.

Three things an answer can be good at

MoneyBench scored every answer on three separate dimensions, each 1–5, with the weights fixed before any scoring began — 0.45 accuracy, 0.35 usefulness, 0.20 expression. Setting weights in advance matters: it stops anyone discovering the winning formula after seeing the results.

In plain language:

Accuracy — are the facts right? Is the contribution limit the actual limit, is the tax treatment the actual treatment.

Usefulness — does the answer do the work of the decision? Not "is this true" but "can I now choose."

Expression — is it pleasant to read? Clear sentences, sensible order, no waffle.

A restaurant analogy makes the split obvious. Accuracy is whether the dish contains what the menu said. Expression is the plating. Usefulness is whether the meal actually feeds you. All three are real — and a kitchen can be excellent at two of them and thin on the third.

Here is where the three systems landed:

Dimension Ed ChatGPT Gemini
Usefulness 4.24 3.62 3.47
Accuracy 3.99 3.72 3.65
Expression 4.03 3.68 3.78

Accuracy is tight. Expression is tighter. Usefulness is where the daylight is.

The test that settles it

Averages hide a lot, so the report ran the more demanding version: keep the scores exactly as they are and change only what you care about. Think of it as re-marking the same exam with different subject weightings — same answers, different priorities, and you find out what the pupil is genuinely good at. A result that survives only one weighting is an artefact of that weighting.

What you weight Ed's outright win rate
As scored (0.45 accuracy / 0.35 usefulness / 0.20 expression) 58.5%
Equal thirds 56.6%
Accuracy-weighted (0.60 / 0.30 / 0.10) 55.7%
Usefulness removed entirely (0.70 / 0 / 0.30) 46.2%
Accuracy only 36.8%
Expression only 23.6%
Three-way chance baseline 33.3%

Read the bottom half slowly. Take usefulness out of the scoring and the advantage falls below half. Score correctness alone and it lands at 36.8% — and with three systems in the race, blind guessing gets you 33.3%. On facts by themselves, these three are very nearly indistinguishable.

(These are outright wins only, which is why the top row reads lower than the 62.3% headline figure — ties are excluded rather than shared out.)

This is the finding, and it's an awkward one for how most people talk about AI. What separated the systems was not knowledge. It was what they did with knowledge they all more or less had.

The per-question record says the same thing another way. Across 106 questions, Ed's usefulness record was 76 wins / 15 ties / 15 losses against ChatGPT and 75 / 16 / 15 against Gemini — significant at p<0.001 against both. Accuracy was 58 / 21 / 27 and 55 / 19 / 32: a real edge, but a narrower one. Expression was 51 / 31 / 24 and 41 / 39 / 26 — note the 39 ties in that last row. That's a dimension where the systems mostly agree with each other.

"Isn't this just nicer formatting?"

It's the right objection, and section 4.2 answers it with a test cleaner than any argument.

Formatting is indifferent to truth. A table is just as tidy when the numbers in it are wrong. So if "usefulness" were secretly measuring headings and bullet points, it should hold up perfectly well on questions where the facts fell apart.

It doesn't. On the 14 questions where Ed's verified accuracy scored 3.0 or below, its usefulness record collapsed to 1 win, 3 ties, 10 losses against whichever comparator scored higher on that question. Across all 106 questions on that same basis, the record was 64 wins, 22 ties, 20 losses.

The presentation was still there. The usefulness wasn't. Scorers weren't rewarding the shape of the answer — they were rewarding whether the reasoning inside it held.

The report is careful here, and so should we be: usefulness and accuracy correlate at 0.747 for Ed, so deliberately selecting the weakest-accuracy questions depresses usefulness partly as arithmetic. The collapse is steeper than correlation alone predicts, but it isn't a pure experiment.

The correlation structure adds a second piece of evidence — how tightly usefulness tracks accuracy, versus how tightly it tracks expression:

  • Ed: 0.747 with accuracy, 0.350 with expression
  • ChatGPT: 0.806 / 0.640
  • Gemini: 0.860 / 0.827

For Gemini those two are almost fused at 0.827 — when its answers read well, they score as useful. For Ed the dimensions come apart. Polish in disguise would track expression closely. This doesn't.

What the scorers actually wrote

The quietly interesting part of the report is section 4.3, where scorers' free-text justifications were coded into themes. Two columns: reasons given when the scorer picked Ed (n=129), and reasons given when they picked against it (n=79).

What the scorer said the answer did Chose Ed Chose against Ed
Explains its logic, not only its conclusion 26% 32%
Legible structure — tables, headings 26% 29%
States the conclusion before the supporting detail 21% 30%
Data sufficient, not merely present 17% 22%
Compares options rather than issuing a single pick 15% 15%

The finding isn't in either column. It's in how close they are.

The same five criteria show up whether the scorer chose Ed or rejected it. When they picked against Ed, they were using the same yardstick — and had simply judged another answer to measure up better on it. Nobody switched standards to justify a preference.

That's what makes this list worth more than a product result. It isn't a description of one system; it's a description of what this audience wants from any money answer at all. And notice what's absent: not one of the ten figures is about prose quality.

What this doesn't show

Three limits, because a result you can't poke at isn't worth much.

Expression is a draw. Ed 4.03, Gemini 3.78 — and that gap sits at p=0.086, which is not statistically significant. Treat writing quality as converged across these systems.

The accuracy lead is a per-opponent lead. Ed's accuracy beats ChatGPT and beats Gemini when each is taken separately. Measured instead against whichever comparator happened to score higher on each individual question, the margin narrows to roughly parity — a mean difference of −0.13. Ed is more accurate than either system; it is not more accurate than both systems' best day combined.

The hypothesis is narrow. What was tested is whether a system built for one domain outperforms general assistants inside that domain. That's all. It says nothing about which AI is better in general, and nothing about any question outside consumer finance.

The part you can use

Strip out the systems and a portable test remains. Next time any AI hands you an answer about your money, ask three things in order:

  1. Is it right? Necessary, and — as the accuracy-only row shows — nowhere near sufficient.
  2. Does it read well? Pleasant, and almost worthless on its own. Expression alone won 23.6% of the time, well below chance.
  3. Can I decide from it? Does it name the figures that matter for your situation, show how they combine, and say what they imply? If not, you have a correct essay, not an answer.

Question three is the one almost nobody asks, and it's the one that does the work. It's also what catches a confidently-worded answer that has quietly left out the number your decision actually turns on.

The same test applies to your own thinking. Knowing your savings rate is a fact. Knowing what it implies about whether you're on track is a decision — and that gap is what financial fitness is meant to close.

Part three turns this into a checklist you can run in about a minute: how to judge an AI money answer.

Conclusion

The systems that answered these questions knew roughly the same things. What separated them was whether the answer was assembled for reading or assembled for deciding. Correct is the floor, not the goal — and once you've seen the difference, you can't unsee it in any answer you're given. The full methodology, tables and limitations are in the MoneyBench report.

Check your own financial fitness

Ed's checkup is the same idea applied to you: not a verdict, a map. A few minutes, no jargon, and an answer you can act on rather than admire.

Start at edwealth.ai/check-up, or download the app on App Store or Google Play.

Money at peace. Wealth in motion.

Ed Wealth is a research and self-reflection tool, not a registered investment advisor. Nothing here is financial, investment, or tax advice. All decisions are yours.

Sources

  • Ed Wealth Research, MoneyBench: a fact-verified benchmark of AI assistants on consumer finance questions — edwealth.ai/moneybench
  • MoneyBench §4.1, dimension means, per-question win/tie/loss records and significance testing
  • MoneyBench §4.2, weighting sensitivity analysis and the low-accuracy subset test
  • MoneyBench §4.3, coded free-text scorer justifications (n=129 / n=79)
Recommend
Six checks you can run on any AI — ChatGPT, Gemini, Ed, whatever you already use — to judge whether its money answers are worth acting on. Drawn from what a 106-question benchmark actually found.

How to tell if an AI is giving you good money answers

An AI money answer arrives fast, reads well, and gives you almost no way to tell whether it's right. That's the whole problem. A stale contribution limit and a current one look identical on screen — same tone, same formatting, same confidence. You can't audit the model. But you can audit the answer, and six checks do most of the work. Ask for one current number you can verify yourself. Check whether it separates rules from opinions. See whether it shows its reasoning, not just its verdict. Notice whether the conclusion comes first. Ask whether it handed you options or a single instruction. And look at whether whoever built it publishes results when their own tool loses. These checks came out of MoneyBench, a benchmark Ed Wealth Research ran on real money questions across Ed, ChatGPT and Gemini. They aren't about Ed. They work on whatever you already have open — and running them once will tell you more in ten minutes than any leaderboard tells you in a year. Adoption surveys disagree sh
EdWealth
·
Sep 07 2026
We ran 106 real money questions through Ed, ChatGPT and Gemini, had them judged blind by an outside model, and fact-checked every answer against live sources. Ed won 62.3% of them. Here's how the test worked.

We tested three AIs on real money questions. Here's what won.

In July 2026 we put 106 real money questions to three AI systems — Ed, ChatGPT and Gemini — and had every answer scored blind, by a judge outside all three systems' model families, with every key fact checked against live sources. Ed won 62.3% of the questions. Gemini took 20.8%. ChatGPT took 17.0%. The full study is published as MoneyBench, including the methodology, the scoring rubric, the limitations and the round we lost. Two things about that number are worth knowing before you read anything else. First, the questions were not written for the test. They were drawn from real queries people had already sent to a personal finance AI — the messy, specific kind, not textbook prompts. Second, the fact-checking pass moved Ed's score down. Before verification Ed sat at 67.6%. After every key fact in every answer was checked against live sources, Ed sat at 62.3% — a 5.3-point cut — while both competitors moved up. Ed still finished first. This article explains what MoneyBench measures and
EdWealth
·
Sep 03 2026
Owning five ETFs doesn't mean you're diversified — 73% of holdings in popular growth ETFs overlap. Here's how to check if your portfolio is secretly one concentrated bet.

Your 5 ETFs might all be making the same bet

Here's a number that'll make you look at your portfolio differently: 94.8% of QQQ's holdings — by weight — are stocks that already live inside VOO. Not a small overlap. Almost complete overlap. If you own both, you're not doubling your diversification. You're mostly just doubling your exposure to the same names — and paying two sets of fund fees to do it. Toss VGT into the mix and it gets stranger. Your top three positions — NVIDIA, Microsoft, and Apple — are now each appearing in three separate funds simultaneously. Three ETFs. Three expense ratios. One concentrated bet on the same handful of companies. This is the ETF overlap problem. It's quiet, it looks like diversification on paper, and it catches a lot of careful people off guard. The simplest way to see what's happening is to pull the top holdings of the three most popular growth and broad-market ETFs side by side. Here's what's sitting inside them as of mid-2026:
EdWealth
·
Sep 02 2026
Feel behind on money? The data says you're probably not. Here are 7 signs of real financial health — each backed by an actual benchmark, from the Fed's $400 test to what most people's debt really looks like.

7 signs you're doing better with money than you think

Short answer: if you have any cash buffer at all, roughly know what you spend, put anything toward retirement, and have never missed a rent or mortgage payment — you're ahead of a large share of American adults on every one of those counts. Feeling behind and being behind are different things, and the data measures only one of them. Here's the strange part about money anxiety: the people who feel it most are often the people doing the work. You compare yourself to a coworker's new car, a cousin's kitchen renovation, a stranger's vacation photos — a highlight reel with no balance sheet attached. Nobody posts their credit card statement. So instead of comparing you to an imaginary person who has it all figured out, this piece compares you to the actual data: what the Federal Reserve, FINRA, and the New York Fed can verify about how Americans really handle money. Not to make anyone feel superior — but because reassurance is only worth something when it's built on evidence. Think of it as
EdWealth
·
Sep 01 2026
Avoiding your bank balance isn't laziness — it's an anxiety response with a name: the ostrich effect. Here's what not looking quietly costs, and the 90-second habit that makes checking feel safe again.

Why you avoid looking at your bank account (and what it's costing you)

The short answer: you avoid your bank account because looking feels like a verdict, and your brain protects you from verdicts. Behavioral economists have a name for this — the ostrich effect — and it's so normal that researchers can measure it at population scale. But avoidance has a quiet price: overdraft fees that only hit people who don't know their balance, subscriptions that bill unnoticed for months, small problems compounding into big ones. The fix is not a full budget audit. It's a 90-second weekly glance at three numbers — enough to shrink the fear without triggering it. You know the move. The banking app sits on your home screen and you scroll past it. A balance alert comes in and you swipe it away without reading the number. Someone asks "can you afford it?" and you say "probably" — because probably doesn't require opening the app. If that's you, here's the first thing to know: nothing is wrong with you. You're not lazy, you're not irresponsible, and you're not uniquely bad
EdWealth
·
Aug 31 2026
43% of Gen Z say their view of their own money doesn't match reality — many feel broke with five figures in savings. The gap has a name: money dysmorphia. Here's why it happens, and how to separate the feeling from the math.

Money dysmorphia: why you feel broke when the numbers say you're not

Short answer: the anxious gap between how your finances feel and what they actually are has a name — money dysmorphia. It's common (roughly 4 in 10 Gen Z and millennials report it), it's not a character flaw, and it doesn't reliably shrink when your balance grows. In one survey, over a third of people who felt this way had more than $10,000 saved. The fix usually isn't more saving — it's calibration: measuring your finances on axes that separate the feeling from the math. You check your accounts and the numbers are... fine. Savings exist. Bills get paid. Nothing is on fire. And yet the background hum doesn't stop: I'm behind. Everyone else is further along. One bad month and it all goes. If your financial anxiety refuses to match your financial data, you're not imagining it — and you're very much not alone. There's a name for the gap, and a financial fitness lens that makes it visible. Money dysmorphia is a distorted perception of your own finances — most often, feeling significantly w
EdWealth
·
Aug 29 2026
US median net worth is $39,000 under 35 and $364,500 at 55–64 (Federal Reserve SCF 2022). But the median can't tell you if you're okay. Here's the better read by decade.

Financial fitness by age: what okay looks like at 25, 35, 45, and 55

Here are the numbers you came for. In the Federal Reserve's latest Survey of Consumer Finances — 2022 data, the most recent vintage of a survey that runs every three years — median US household net worth is $39,000 for households under 35, $135,600 for ages 35–44, $247,200 for 45–54, and $364,500 for 55–64. Now the uncomfortable part: age-based net worth benchmarks are probably the most-searched and least-useful numbers in personal finance. Not because the data is wrong — the Fed's survey is as good as this data gets. But the median only tells you where the middle of the country is. It doesn't tell you whether you are okay. A single net worth number can't see whether you're trending up or down, whether one bad month would break you, or whether your money runs on a system or on stress. That's what financial fitness measures — and at each age, different parts of it do the heavy lifting. So this article does both jobs. The honest benchmarks first, since that's what you searched for. Then
EdWealth
·
Aug 28 2026
A 12-point financial fitness checklist you can run in one sitting: safety, control, progress, upside, and mental load. Each check is a yes-or-no question with a clear next step.

The financial fitness checklist: 12 checks, one honest hour

You don't need a financial advisor to check your financial fitness. You need one honest hour. This is a financial fitness checklist — 12 yes-or-no questions, grouped into five areas: safety, control, progress, upside, and mental load. No scoring, no formulas, no spreadsheet. Just questions with honest answers, and for every "no," a concrete next step. The rules are simple. Sit down once a season — quarterly is the right cadence — with your phone, your bank app, and nothing else on the calendar. Answer each question honestly. "Sort of" counts as no. You're not trying to pass; you're trying to find out where the gaps are while they're still cheap to fix. If you're new to the idea of financial fitness as a concept — health you maintain, not a number you hit — the full picture is in what is financial fitness. And if you'd rather get an actual score than work a list, take the financial fitness test instead. This page is for people who want to do something today. Here's the whole list up fro
EdWealth
·
Aug 27 2026
A free 10-question financial fitness test you can score in five minutes. Rate yourself across five areas — Safety, Control, Progress, Upside, Mental Load — and find out which one needs you first.

Financial fitness test: 10 questions to score yourself honestly

You searched for a financial fitness test, so let's skip the throat-clearing: the test is right below. Ten questions, five minutes, scored out of 20. No email required, no "results are waiting for you" gate. One thing before you start. The best-known version of this idea is the CFPB's Financial Well-Being Scale — a 10-question survey the U.S. government's consumer finance agency built to measure how people are actually doing with money. When the CFPB ran it nationally, the average American adult scored 54 out of 100, and roughly a third of adults scored 50 or below — the range where struggling to cover basic needs becomes more likely than not. So if this test stings a little, you're in large company. The test below covers five areas — Safety, Control, Progress, Upside, and Mental Load — two questions each. (If you want the full argument for why financial fitness is measured in dimensions rather than one number, that's here. This article is just the test.) Answer honestly. Nobody's watc
EdWealth
·
Aug 26 2026
Financial fitness isn't being rich, and it isn't a good credit score. It's whether your money can cover your life, take a hit, keep building, and stay off your mind. Here are the five signs.

What is financial fitness? The 5 signs you're actually okay

Financial fitness is your money's ability to do three things at once: cover your life without strain, absorb a shock without falling apart, and stay out of your head the rest of the time. It's a measure of resilience and direction — not a measure of how much you have. That last part is what trips people up. Most of the signals we use to judge our finances measure something else. A bank balance tells you what you have today. A salary tells you what flows in. A credit score tells you how lenders feel about you. None of them tell you whether your finances would survive a bad month, or whether this year's income is building anything. The fitness comparison isn't just a cute name. Body weight is a number; physical fitness is a capability — can you climb the stairs, recover from a cold, carry the groceries. Money works the same way: two people with the same income, even the same net worth, can be in completely different shape. One takes a surprise car repair in stride. The other spirals for
EdWealth
·
Aug 25 2026

Money at peace.Wealth in motion.

Your money, finally handled. Your life, finally unhurried.