In July 2026 we put 106 real money questions to three AI systems — Ed, ChatGPT and Gemini — and had every answer scored blind, by a judge outside all three systems' model families, with every key fact checked against live sources. Ed won 62.3% of the questions. Gemini took 20.8%. ChatGPT took 17.0%.
The full study is published as MoneyBench, including the methodology, the scoring rubric, the limitations and the round we lost.
Two things about that number are worth knowing before you read anything else. First, the questions were not written for the test. They were drawn from real queries people had already sent to a personal finance AI — the messy, specific kind, not textbook prompts. Second, the fact-checking pass moved Ed's score down. Before verification Ed sat at 67.6%. After every key fact in every answer was checked against live sources, Ed sat at 62.3% — a 5.3-point cut — while both competitors moved up. Ed still finished first.
This article explains what MoneyBench measures and how the test was built. Two companion pieces go deeper: what actually makes an AI money answer useful, and how to judge an AI money answer yourself.
Why this benchmark had to exist
Coding and maths have public benchmarks. You can look up how models rank, see the test set, and argue about the scoring. Consumer finance answering had no equivalent — which is strange, because it is one of the domains where being wrong costs the most.
A wrong answer about money is actionable. Someone reads a contribution limit and moves cash. Someone reads a stale share price and adjusts a position. Money answers also decay: a figure that was right in March can be wrong in July, and nothing about the way an AI writes tells you which one you're getting. The answer arrives instantly, in the same confident register, either way.
That gap shows up in behaviour. In an Intuit Credit Karma survey of 1,019 US adults, 52% of those who had acted on AI financial advice reported it led to a poor decision.
So we built the test we wanted to be measured against: real questions, blind scoring, an outside judge, and every key fact checked. Then we published it, protocol included, so anyone can run it — including on us.
What we actually measured
Most AI evaluations score whether an answer is correct. Correct is necessary. It is not the job.
People don't consult an AI about money to fact-check a number they already know. They consult it to decide something: whether to pay down the loan or top up the retirement account, whether a bonus should go to cash or the market, what a tax rule means for their specific situation. An answer can be flawless and still leave you exactly where you started.
So MoneyBench scores three dimensions on a 1–5 scale, weighted: accuracy 0.45, usefulness 0.35, expression 0.20. Accuracy is whether the facts hold. Usefulness is whether the answer moves a person closer to a decision. Expression is whether it's readable.
There is one hard rule sitting above the weights: an accuracy veto. An answer containing a fact that verification proved false cannot win its question, no matter how useful or well-written it is. Usefulness never buys its way past a wrong number.
The results, across three rounds
| Round |
Date |
Ed |
Gemini |
ChatGPT |
How it was scored |
| Panel I |
May 2026 |
17.4% |
46.1% |
36.5% |
37 blind human scorers, 1,308 scored responses |
| Panel II |
June 2026 |
54.4% |
28.3% |
17.3% |
25 blind human scorers, 90 questions, 375 votes |
| Verified Evaluation |
July 2026 |
62.3% |
20.8% |
17.0% |
Outside model judge, per-question fact verification, 106 questions |
Read the first row again. In May, Ed came last — third of three, by a wide margin.
What the loss-cause coding showed was specific, and not what we expected. The capability existed; it wasn't reaching users. On questions that routed to Ed's extended-reasoning path, Ed placed first — 39.8% against Gemini's 30.5% and ChatGPT's 29.7%. But that path carried only around 10% of traffic. Most questions never got there.
What we rebuilt was retrieval, task orchestration and answer construction — how a question is understood, what data gets pulled, which path it takes, and how the answer is assembled. Not a model upgrade. The engine was fine. The plumbing wasn't.
Ed has led every round since.
Where the July gap came from
| Dimension (1–5 mean) |
Ed |
ChatGPT |
Gemini |
| Usefulness |
4.24 |
3.62 |
3.47 |
| Accuracy |
3.99 |
3.72 |
3.65 |
| Expression |
4.03 |
3.68 |
3.78 |
Expression is close to a three-way tie — all three write well. Accuracy separates, but modestly. The gap is in usefulness, and it is wide: Ed scored higher on usefulness than ChatGPT on 76 of 106 questions, and higher than Gemini on 75 of 106.
That pattern is the finding. The general assistants are not bad at money. They are good at answering money questions and less good at resolving them — the numbers arrive, the decision doesn't.
Why this result is hard to wave away
Any company can publish a test it wins. What makes a result worth anything is the conditions it was won under.
The judge was outside all three systems. Scoring was done by Claude, built by Anthropic — outside ChatGPT's model family, outside Gemini's, and outside the undisclosed base model Ed itself runs on. Nobody was marking their own homework, us included.
Presentation was blind and randomised. Answers were stripped of anything identifying which system produced them, and shown in randomised positions so ordering couldn't tilt the score.
Every key fact was verified against live sources. Not sampled — every key fact in every answer, checked against live sources as of 23 July 2026. Verification succeeded on 106 of 106 questions.
The competitors were not handicapped. ChatGPT ran GPT-5.6 Sol at Pro effort — above its own default. Gemini ran 3.6 Flash at its default. Ed ran in ordinary production configuration, the same one users get.
Then there are the two facts we'd point to first if we were trying to poke holes in this ourselves.
Verification cost us. Ed's share fell from 67.6% to 62.3% while Gemini rose from 17.1% to 20.8% and ChatGPT from 15.2% to 17.0%. A test designed to flatter us would not include a step that takes 5.3 points off our own score and hands them to the competition. The win survived the step that hurt us most.
And May is in the report. Ed placed third at 17.4%, and that round is published with the same detail as the rounds we won — the scores, the loss-cause coding, the diagnosis. A benchmark you only publish when you win isn't a benchmark; it's a press release.
For what it's worth statistically: the 95% confidence interval on Ed's July share is [52.8%, 71.7%], against a three-way chance baseline of 33.3%. The bottom of that range still clears chance comfortably.
What this doesn't prove
The three studies are related tests, not one clean trajectory — the report itself states that only Panel I and the Verified Evaluation are a supported comparison, and we've left that caveat in rather than drawing a tidy line through three dots. In July, 47% of questions were close calls. And the hypothesis on trial was narrow: that a system built for one domain outperforms general assistants inside that domain. Not that Ed reasons better than a frontier model. It doesn't, and that was never the claim.
What the test says is smaller and more useful than "Ed is smarter." It says that for money questions specifically, built-for-purpose beat general-purpose — under an outside judge, blind, with the facts checked.
Conclusion
Ed won 62.3% of 106 fact-verified real money questions, against 20.8% for Gemini and 17.0% for ChatGPT, judged blind by a model outside all three families. The margin came from usefulness — whether an answer gets a person to a decision — not from writing quality, where all three are close.
We think that's the right thing to measure, and the right way to measure it: publish the protocol, use an outside judge, check the facts even when checking costs you points, and publish the round you lost. The full report is open. So is the protocol.
Check your own financial fitness
A benchmark measures the answer. Your own money is a question about you — your cash flow, your buffer, your goals. Ed's checkup walks you through the five dimensions of financial fitness in a few minutes, no jargon, and shows where you're strong and where you're exposed. It's the difference between a general assistant and a money person of your own.
Start at edwealth.ai/check-up, or download the app on App Store or Google Play.
Money at peace. Wealth in motion.
Ed Wealth is a research and self-reflection tool, not a registered investment advisor. Nothing here is financial, investment, or tax advice. All decisions are yours.
Sources
- Ed Wealth Research, MoneyBench: measuring usefulness and factual accuracy on real money questions (August 2026) — edwealth.ai/moneybench
- Intuit Credit Karma, survey of 1,019 US adults, fielded 7–14 August 2025