We tested three AIs on real money questions. Here's what won.

EdWealth
· Sep 03 2026
We tested three AIs on real money questions. Here's what won.

In July 2026 we put 106 real money questions to three AI systems — Ed, ChatGPT and Gemini — and had every answer scored blind, by a judge outside all three systems' model families, with every key fact checked against live sources. Ed won 62.3% of the questions. Gemini took 20.8%. ChatGPT took 17.0%.

The full study is published as MoneyBench, including the methodology, the scoring rubric, the limitations and the round we lost.

Two things about that number are worth knowing before you read anything else. First, the questions were not written for the test. They were drawn from real queries people had already sent to a personal finance AI — the messy, specific kind, not textbook prompts. Second, the fact-checking pass moved Ed's score down. Before verification Ed sat at 67.6%. After every key fact in every answer was checked against live sources, Ed sat at 62.3% — a 5.3-point cut — while both competitors moved up. Ed still finished first.

This article explains what MoneyBench measures and how the test was built. Two companion pieces go deeper: what actually makes an AI money answer useful, and how to judge an AI money answer yourself.

Why this benchmark had to exist

Coding and maths have public benchmarks. You can look up how models rank, see the test set, and argue about the scoring. Consumer finance answering had no equivalent — which is strange, because it is one of the domains where being wrong costs the most.

A wrong answer about money is actionable. Someone reads a contribution limit and moves cash. Someone reads a stale share price and adjusts a position. Money answers also decay: a figure that was right in March can be wrong in July, and nothing about the way an AI writes tells you which one you're getting. The answer arrives instantly, in the same confident register, either way.

That gap shows up in behaviour. In an Intuit Credit Karma survey of 1,019 US adults, 52% of those who had acted on AI financial advice reported it led to a poor decision.

So we built the test we wanted to be measured against: real questions, blind scoring, an outside judge, and every key fact checked. Then we published it, protocol included, so anyone can run it — including on us.

What we actually measured

Most AI evaluations score whether an answer is correct. Correct is necessary. It is not the job.

People don't consult an AI about money to fact-check a number they already know. They consult it to decide something: whether to pay down the loan or top up the retirement account, whether a bonus should go to cash or the market, what a tax rule means for their specific situation. An answer can be flawless and still leave you exactly where you started.

So MoneyBench scores three dimensions on a 1–5 scale, weighted: accuracy 0.45, usefulness 0.35, expression 0.20. Accuracy is whether the facts hold. Usefulness is whether the answer moves a person closer to a decision. Expression is whether it's readable.

There is one hard rule sitting above the weights: an accuracy veto. An answer containing a fact that verification proved false cannot win its question, no matter how useful or well-written it is. Usefulness never buys its way past a wrong number.

The results, across three rounds

Round Date Ed Gemini ChatGPT How it was scored
Panel I May 2026 17.4% 46.1% 36.5% 37 blind human scorers, 1,308 scored responses
Panel II June 2026 54.4% 28.3% 17.3% 25 blind human scorers, 90 questions, 375 votes
Verified Evaluation July 2026 62.3% 20.8% 17.0% Outside model judge, per-question fact verification, 106 questions

Read the first row again. In May, Ed came last — third of three, by a wide margin.

What the loss-cause coding showed was specific, and not what we expected. The capability existed; it wasn't reaching users. On questions that routed to Ed's extended-reasoning path, Ed placed first — 39.8% against Gemini's 30.5% and ChatGPT's 29.7%. But that path carried only around 10% of traffic. Most questions never got there.

What we rebuilt was retrieval, task orchestration and answer construction — how a question is understood, what data gets pulled, which path it takes, and how the answer is assembled. Not a model upgrade. The engine was fine. The plumbing wasn't.

Ed has led every round since.

Where the July gap came from

Dimension (1–5 mean) Ed ChatGPT Gemini
Usefulness 4.24 3.62 3.47
Accuracy 3.99 3.72 3.65
Expression 4.03 3.68 3.78

Expression is close to a three-way tie — all three write well. Accuracy separates, but modestly. The gap is in usefulness, and it is wide: Ed scored higher on usefulness than ChatGPT on 76 of 106 questions, and higher than Gemini on 75 of 106.

That pattern is the finding. The general assistants are not bad at money. They are good at answering money questions and less good at resolving them — the numbers arrive, the decision doesn't.

Why this result is hard to wave away

Any company can publish a test it wins. What makes a result worth anything is the conditions it was won under.

The judge was outside all three systems. Scoring was done by Claude, built by Anthropic — outside ChatGPT's model family, outside Gemini's, and outside the undisclosed base model Ed itself runs on. Nobody was marking their own homework, us included.

Presentation was blind and randomised. Answers were stripped of anything identifying which system produced them, and shown in randomised positions so ordering couldn't tilt the score.

Every key fact was verified against live sources. Not sampled — every key fact in every answer, checked against live sources as of 23 July 2026. Verification succeeded on 106 of 106 questions.

The competitors were not handicapped. ChatGPT ran GPT-5.6 Sol at Pro effort — above its own default. Gemini ran 3.6 Flash at its default. Ed ran in ordinary production configuration, the same one users get.

Then there are the two facts we'd point to first if we were trying to poke holes in this ourselves.

Verification cost us. Ed's share fell from 67.6% to 62.3% while Gemini rose from 17.1% to 20.8% and ChatGPT from 15.2% to 17.0%. A test designed to flatter us would not include a step that takes 5.3 points off our own score and hands them to the competition. The win survived the step that hurt us most.

And May is in the report. Ed placed third at 17.4%, and that round is published with the same detail as the rounds we won — the scores, the loss-cause coding, the diagnosis. A benchmark you only publish when you win isn't a benchmark; it's a press release.

For what it's worth statistically: the 95% confidence interval on Ed's July share is [52.8%, 71.7%], against a three-way chance baseline of 33.3%. The bottom of that range still clears chance comfortably.

What this doesn't prove

The three studies are related tests, not one clean trajectory — the report itself states that only Panel I and the Verified Evaluation are a supported comparison, and we've left that caveat in rather than drawing a tidy line through three dots. In July, 47% of questions were close calls. And the hypothesis on trial was narrow: that a system built for one domain outperforms general assistants inside that domain. Not that Ed reasons better than a frontier model. It doesn't, and that was never the claim.

What the test says is smaller and more useful than "Ed is smarter." It says that for money questions specifically, built-for-purpose beat general-purpose — under an outside judge, blind, with the facts checked.

Conclusion

Ed won 62.3% of 106 fact-verified real money questions, against 20.8% for Gemini and 17.0% for ChatGPT, judged blind by a model outside all three families. The margin came from usefulness — whether an answer gets a person to a decision — not from writing quality, where all three are close.

We think that's the right thing to measure, and the right way to measure it: publish the protocol, use an outside judge, check the facts even when checking costs you points, and publish the round you lost. The full report is open. So is the protocol.

Check your own financial fitness

A benchmark measures the answer. Your own money is a question about you — your cash flow, your buffer, your goals. Ed's checkup walks you through the five dimensions of financial fitness in a few minutes, no jargon, and shows where you're strong and where you're exposed. It's the difference between a general assistant and a money person of your own.

Start at edwealth.ai/check-up, or download the app on App Store or Google Play.

Money at peace. Wealth in motion.

Ed Wealth is a research and self-reflection tool, not a registered investment advisor. Nothing here is financial, investment, or tax advice. All decisions are yours.

Sources

  • Ed Wealth Research, MoneyBench: measuring usefulness and factual accuracy on real money questions (August 2026) — edwealth.ai/moneybench
  • Intuit Credit Karma, survey of 1,019 US adults, fielded 7–14 August 2025
Recommend
Employer stock is the one holding that arrives as a reward, grows on autopilot, and doubles a risk you already carry. Here's how to size it against your net worth, and why deciding your rhythm in advance beats deciding in the moment.

Your job and your savings are betting on the same company

The short answer: if a meaningful slice of your net worth sits in your employer's stock, you're exposed to that company twice — once through your paycheck, once through your portfolio. That's the part most RSU writing skips, because almost all of it is about taxes. The structural problem isn't the tax bill; it's the correlation. A bad year at your company can shrink your bonus, freeze your raise, thin out your team and mark down your holdings in the same quarter — two risks moving together in exactly the way diversification exists to prevent. And unlike most concentration, this one builds itself: every vest adds more of the same stock, so the position grows unless someone actively decides otherwise. Doing nothing is a decision to concentrate further. A workable frame: size the position against your whole net worth rather than your brokerage account alone, treat under 10% as unremarkable and above 20% as worth a hard look, adjust that line for how you honestly feel about being all-in on
EdWealth
·
Sep 12 2026
A largest holding under about 10% of your portfolio reads as healthy. Over about 20% gets flagged. But the bands move with your horizon — and the better test is a question you can actually answer.

If your biggest holding dropped 30% tomorrow, would you be okay?

If your biggest holding dropped 30% tomorrow, would you be okay? Not "would you be annoyed." Would the things you've actually committed to survive it — the deposit, the retirement date, the ability to sleep through the week without selling at the bottom. That question is a better test of concentration risk than any ratio, because you can answer it. Most people cannot tell you what percentage of their money sits in their single largest position. Almost everyone can tell you whether a 30% haircut on it would hurt. Here's the rough shape of an answer anyway. As a general read, a largest holding under about 10% of your portfolio is unremarkable. Over about 20% is worth a serious look. But those lines move with the person: someone investing aggressively with a ten-year-plus horizon can sit nearer 30% without being reckless, while someone conservative who needs the money in three years should probably be tighter than 15%. The rest of this is why that number sneaks up on people, what actually
EdWealth
·
Sep 11 2026
How much cash is too much? The honest answer in months of expenses — plus the arithmetic to work out what your idle cash quietly costs you every year.

How much cash is too much to keep in savings?

The short answer: cash stops being prudent and starts being expensive somewhere past three months of your expenses sitting in an account that pays close to nothing. Under about one month of idle cash is normal operating float. Between one and three months is a grey zone worth a look. Past three months, the balance isn't cautious anymore — it's costing you a number you could write down. Notice what that sentence is measuring. Not how much cash you hold. How much cash you hold that isn't earning anything. Those are different questions, and mixing them up is why this topic stays confusing. Six months of expenses in an account paying 4% is a well-built emergency fund. Six months of expenses in a checking account paying 0.07% is the same money doing a much worse job. Same balance, same safety, wildly different outcome. So the real work here is two steps: figure out how many months of expenses you're actually holding idle, then run the arithmetic on what the idle part costs per year. Both ta
EdWealth
·
Sep 10 2026
A dated, sourced comparison of financial advisor apps in 2026 — Betterment, Ed, Empower, Facet, Fidelity Go, Origin, Range, Schwab, Vanguard, Wealthfront, Zoe. Every price checked 1 September 2026, plus the arithmetic on when a flat fee actually beats a percentage.

The best financial advisor apps in 2026, and what each one actually costs

"Best financial advisor app" is three different questions wearing one coat, which is why most roundups of them are useless. Some want software that moves the money — invests it, rebalances it, harvests the losses. That's a robo-advisor: Betterment, Wealthfront, Fidelity Go, Schwab, Vanguard Digital Advisor. Some want to talk to a human who is licensed and accountable: Facet, Empower, Vanguard Personal Advisor Select, Range, Zoe. And some want guidance without handing over a percentage of their assets every year for as long as they hold them — a smaller lane, where Ed and Origin sit. Those three needs don't share a winner. A list that ranks them together has to pretend they do. So this page separates them, then prices everything. Every fee below was checked on 1 September 2026 against the company's own pricing page or its SEC filing. Where a figure couldn't be confirmed, it says so instead of guessing — and one product's price is missing entirely, for a reason we explain. Transparency:
EdWealth
·
Sep 08 2026
Six checks you can run on any AI — ChatGPT, Gemini, Ed, whatever you already use — to judge whether its money answers are worth acting on. Drawn from what a 106-question benchmark actually found.

How to tell if an AI is giving you good money answers

An AI money answer arrives fast, reads well, and gives you almost no way to tell whether it's right. That's the whole problem. A stale contribution limit and a current one look identical on screen — same tone, same formatting, same confidence. You can't audit the model. But you can audit the answer, and six checks do most of the work. Ask for one current number you can verify yourself. Check whether it separates rules from opinions. See whether it shows its reasoning, not just its verdict. Notice whether the conclusion comes first. Ask whether it handed you options or a single instruction. And look at whether whoever built it publishes results when their own tool loses. These checks came out of MoneyBench, a benchmark Ed Wealth Research ran on real money questions across Ed, ChatGPT and Gemini. They aren't about Ed. They work on whatever you already have open — and running them once will tell you more in ten minutes than any leaderboard tells you in a year. Adoption surveys disagree sh
EdWealth
·
Sep 07 2026
Three AI systems scored within a third of a point of each other on accuracy — and one was still chosen roughly three times as often. Here's what the MoneyBench data says actually separates a money answer that helps from one that's merely right.

A correct answer and a useful one are not the same thing

Three AI systems answered the same 106 personal finance questions. On raw correctness they finished close together — 3.99, 3.72 and 3.65 on a five-point scale, a spread of about a third of a point. Then scorers were asked which answer they'd actually want, and one system was chosen roughly three times as often as either of the others. That gap is the interesting part. If the three were near-equally correct, the winner wasn't winning on being right. The obvious next guess is that it wrote more nicely — but expression is the dimension where the three converged most, and on one head-to-head the difference isn't even statistically significant. So what was being rewarded? A third property, which the MoneyBench report defines with unusual precision: a useful answer is one that equips the decision it addresses. It identifies the figures that bear on the question, shows how they combine, and states what they imply for the choice at hand. That sounds mild. It turns out to be the whole differenc
EdWealth
·
Sep 04 2026
Owning five ETFs doesn't mean you're diversified — 73% of holdings in popular growth ETFs overlap. Here's how to check if your portfolio is secretly one concentrated bet.

Your 5 ETFs might all be making the same bet

Here's a number that'll make you look at your portfolio differently: 94.8% of QQQ's holdings — by weight — are stocks that already live inside VOO. Not a small overlap. Almost complete overlap. If you own both, you're not doubling your diversification. You're mostly just doubling your exposure to the same names — and paying two sets of fund fees to do it. Toss VGT into the mix and it gets stranger. Your top three positions — NVIDIA, Microsoft, and Apple — are now each appearing in three separate funds simultaneously. Three ETFs. Three expense ratios. One concentrated bet on the same handful of companies. This is the ETF overlap problem. It's quiet, it looks like diversification on paper, and it catches a lot of careful people off guard. The simplest way to see what's happening is to pull the top holdings of the three most popular growth and broad-market ETFs side by side. Here's what's sitting inside them as of mid-2026:
EdWealth
·
Sep 02 2026
Feel behind on money? The data says you're probably not. Here are 7 signs of real financial health — each backed by an actual benchmark, from the Fed's $400 test to what most people's debt really looks like.

7 signs you're doing better with money than you think

Short answer: if you have any cash buffer at all, roughly know what you spend, put anything toward retirement, and have never missed a rent or mortgage payment — you're ahead of a large share of American adults on every one of those counts. Feeling behind and being behind are different things, and the data measures only one of them. Here's the strange part about money anxiety: the people who feel it most are often the people doing the work. You compare yourself to a coworker's new car, a cousin's kitchen renovation, a stranger's vacation photos — a highlight reel with no balance sheet attached. Nobody posts their credit card statement. So instead of comparing you to an imaginary person who has it all figured out, this piece compares you to the actual data: what the Federal Reserve, FINRA, and the New York Fed can verify about how Americans really handle money. Not to make anyone feel superior — but because reassurance is only worth something when it's built on evidence. Think of it as
EdWealth
·
Sep 01 2026
Avoiding your bank balance isn't laziness — it's an anxiety response with a name: the ostrich effect. Here's what not looking quietly costs, and the 90-second habit that makes checking feel safe again.

Why you avoid looking at your bank account (and what it's costing you)

The short answer: you avoid your bank account because looking feels like a verdict, and your brain protects you from verdicts. Behavioral economists have a name for this — the ostrich effect — and it's so normal that researchers can measure it at population scale. But avoidance has a quiet price: overdraft fees that only hit people who don't know their balance, subscriptions that bill unnoticed for months, small problems compounding into big ones. The fix is not a full budget audit. It's a 90-second weekly glance at three numbers — enough to shrink the fear without triggering it. You know the move. The banking app sits on your home screen and you scroll past it. A balance alert comes in and you swipe it away without reading the number. Someone asks "can you afford it?" and you say "probably" — because probably doesn't require opening the app. If that's you, here's the first thing to know: nothing is wrong with you. You're not lazy, you're not irresponsible, and you're not uniquely bad
EdWealth
·
Aug 31 2026
43% of Gen Z say their view of their own money doesn't match reality — many feel broke with five figures in savings. The gap has a name: money dysmorphia. Here's why it happens, and how to separate the feeling from the math.

Money dysmorphia: why you feel broke when the numbers say you're not

Short answer: the anxious gap between how your finances feel and what they actually are has a name — money dysmorphia. It's common (roughly 4 in 10 Gen Z and millennials report it), it's not a character flaw, and it doesn't reliably shrink when your balance grows. In one survey, over a third of people who felt this way had more than $10,000 saved. The fix usually isn't more saving — it's calibration: measuring your finances on axes that separate the feeling from the math. You check your accounts and the numbers are... fine. Savings exist. Bills get paid. Nothing is on fire. And yet the background hum doesn't stop: I'm behind. Everyone else is further along. One bad month and it all goes. If your financial anxiety refuses to match your financial data, you're not imagining it — and you're very much not alone. There's a name for the gap, and a financial fitness lens that makes it visible. Money dysmorphia is a distorted perception of your own finances — most often, feeling significantly w
EdWealth
·
Aug 29 2026

Money at peace.Wealth in motion.

Your money, finally handled. Your life, finally unhurried.