The Receipts Index
Why We Grade the Machines: The Case for a Dated Honesty Record
Something quiet has happened to the way ordinary people get answers about money. A generation ago you asked a person. A banker, a broker, a parent, a friend who seemed to know things. Then you asked a search engine, which at least showed you where the answer came from. Now, more and more, you ask a machine, and the machine simply answers. No byline. No date. No sense of where the answer was born or how old it is.
This is not a complaint about AI. These tools are genuinely useful, and they are getting better. The point is narrower and more uncomfortable. The machine is becoming the middle layer between people and their money. It sits between the question and the decision. And almost nobody is checking its work in a way that lasts.
The gap nobody owns
When a machine gives money advice, three things are true at once, and none of them belongs to anyone.
First, it can be confidently wrong. A model does not hesitate the way an unsure person does. A wrong answer arrives in the same calm, fluent voice as a right one, and that fluency is exactly what makes the error dangerous. A human advisor who is guessing usually sounds like it. A machine never does.
Second, it can be out of date. Rules change. Rates change. Products appear and vanish. A model trained or tuned at one moment can carry that moment forward long after the world has moved on, and it will rarely volunteer that its information has an expiration date.
Third, it can be manipulated. Anything that learns from public text can be nudged by whoever writes the most public text. Marketers, promoters, and outright scammers all have an interest in shaping what the middle layer says about money.
Now notice what is missing. Reviewers rate these tools on features. Speed, price, interface, cleverness. Benchmarks measure puzzle-solving. But almost nobody keeps a dated, public, honest record of how well these tools actually answer the money questions ordinary people ask. Nobody grades honesty over time. That is the gap, and it is unowned because it is unglamorous. It requires patience instead of hot takes.
Why a dated record, specifically
A fair question: why a record? Why not just test the tools once and publish the results?
Because a single answer proves nothing. Ask a model one question and you get one draw from a machine that can answer the same question differently tomorrow. Anyone can screenshot a brilliant response and call the tool a genius. Anyone can screenshot a blunder and call it a menace. Both screenshots are real. Neither one is evidence of anything except that day, that phrasing, that moment. The internet is full of both, and they cancel out.
What cannot be faked is a time series. The same questions, asked the same way, graded by the same standard, cycle after cycle, with every entry dated when it happened. A record like that shows drift. It shows improvement. It shows the gap between what a tool claims and what it delivers, not once, but as a pattern. The value is never in any single entry. The value is the record itself.
And a record has one property that no clever writing can replicate. It cannot be back-dated. You cannot decide in three years that you wish you had been measuring all along. The only way to have a five year record in five years is to start now, keep it honestly, and let it grow. That is why starting is the whole point. The first entries will look thin. Every honest record looks thin at the beginning. The dishonest ones are the only ones that arrive fully formed.
What honest looks like in practice
Honest is not a mood. It is a set of habits, and they are checkable.
Every claim carries a source and a date. Not "studies show" but which source, retrieved when. If a claim cannot carry its receipt, it does not go in the record.
Forecasts are labeled as forecasts. A guess about the future dressed up as a fact is the oldest trick in financial writing, human or machine. The record separates what is known from what is expected, and says so in plain words.
And the keeper corrects its own mistakes out loud. Any record kept by people will contain errors, because people make errors. The test of honesty is not a spotless page. It is what happens when a mistake is found. A correction, dated and visible, is worth more to a reader's trust than a hundred entries that were simply right. A record that never admits fault is not a clean record. It is an unaudited one.
Notice that these are the same standards the record applies to the machines. That symmetry is deliberate. It would be absurd to grade AI tools on sourcing, freshness, and candor while exempting the grader. The discipline runs both directions or it is theater.
The counterweight
There is a harder reason this matters, beyond curiosity about which tool answers best.
The same shift that puts a machine between people and their money also puts machines in the hands of people who want to take that money. The technology that writes a helpful explanation of index funds also writes a persuasive pitch for a fund that does not exist. Fluency is cheap now, for everyone, including the people who used to be undone by their own bad grammar. The old tells are dying.
When fluency is no longer evidence of legitimacy, the question "does this sound right" stops protecting anyone. What replaces it is the question "can this be verified, by whom, dated when." That is a receipts question. A public habit of grading machine answers against sources and dates is not an academic exercise in that world. It is defense. It keeps alive the reflex that answers must be checked, precisely at the moment when answers have never been easier to fake.
The honest limits
Here is what this record is not, said plainly so nobody can read this essay as a bigger claim than it makes.
It is one practice, not the whole truth. A dated record of graded answers will not tell you everything about a tool, and it will not tell you what to do with your money. It is a lens, and every lens crops.
It grades answers as given, on the questions it asks, on the dates it runs. A tool might answer other questions better or worse. It might answer the same question differently an hour later. The record can only witness what it actually tested, and it should never pretend to have seen more.
And it publishes conclusions only when enough dated cycles exist to mean something. Two data points make a line, not a trend. Until the record is deep enough to support a claim, the honest move is to keep collecting and say so. A record keeper who rushes to verdicts is just a pundit with a spreadsheet.
These limits are not fine print. They are the argument. A record that admits exactly what it can and cannot show is the only kind worth keeping, because it is the only kind a reader can weigh for themselves.
The standing idea
Strip everything else away and the idea is old. Write it down. Date it. Show your sources. Fix your errors where people can see. That standard predates every machine and it will outlast the current crop of them. What is new is only the subject: the machines themselves, graded by the same rules they should be held to, in a record that grows one honest cycle at a time.
Real numbers. No hype. Receipts. If that standard sounds like the bare minimum, good. It is. It is also rarer than it should be, and the only way a record like this ever gets long enough to matter is if it starts small and public and keeps going. You are welcome to watch it from the beginning. That is the best seat there is.
Real numbers. No hype. Receipts.