Corva › Blog › Checking AI financial numbers
Using AI on filings
Most wrong numbers from a language model are not invented. They are real figures lifted from the wrong line of a real filing, which is why "just check EDGAR" does not help: the number is in EDGAR. Here are the five fields to check instead.
10 min read· Corva Research
The short version
Ask a chatbot for Apple's earnings per share and you will get a number. Usually the right one. When it is wrong, it tends to be wrong in a small number of repeatable ways, and that is more useful than it first sounds.
Every guide to this question ends in the same place. It hallucinates, so open the filing and check. That advice is true and close to useless, because if you have to re-verify everything then you have done the work you used the model to avoid. It also misdescribes the failure. In our own use the wrong figure is rarely invented. It is a real number, printed in the filing, taken from a line that answers a slightly different question.
So this is the narrower version of how to check if ChatGPT got the financial numbers right: five fields, five checks, and nothing else. Each one is a place where the filing itself contains more than one defensible answer.
Three different things, and only one of them is hard.
Fabrication, where the figure appears nowhere in the document, is the failure everyone talks about and the easiest to catch: one search of the filing kills it. Arithmetic errors are next, and you catch those by asking for the two inputs rather than the ratio.
The third survives a spot check. The model returns a number that is genuinely in the filing but answers a different question than the one you asked. Ask for "shares outstanding" and a 10-K offers four candidates. Ask for "2025 revenue" at a company with a June year end and there are two honest answers. The retrieval was correct. The interpretation was not.
That is where a taxonomy earns its keep. If the failure is unpredictable, you check everything. If it clusters, you check five things.
A share count is not one number. Apple's Form 10-K for the fiscal year ended 27 September 2025 contains at least four, all correct, all serving different purposes.
| Figure | Shares | Where it lives |
|---|---|---|
| Issued and outstanding, 17 Oct 2025 | 14,776,353 | Cover page |
| Outstanding, ending balance FY2025 | 14,773,260 | Shareholders' equity note |
| Weighted average, basic | 14,948,500 | Income statement |
| Weighted average, diluted | 15,004,697 | Income statement |
The spread from lowest to highest is 231,437 thousand shares, about 1.6%. Source: Apple Inc. Form 10-K, FY2025, cover page, Consolidated Statements of Operations and Note 10. Filed 31 October 2025.
Apple reported net income of $112,010m and diluted EPS of $7.46 for the year. Divide that same net income by the cover-page count and you get $7.58. Twelve cents, entirely a question of which real number you used, and it matters most where you would least expect it: the cover-page count falls with buybacks while the weighted average lags them by up to a year, so the two tell different stories about dilution.
The check: ask which of the four the model used. Per-share earnings need weighted average diluted. Market capitalisation needs the cover-page count. Dilution needs the ending balance across several years. If the answer names none of them, the number is unverified.
Which raises the question of what "the year" even means.
Microsoft's fiscal 2026 ended on 30 June 2026. Half of it happened in 2025.
So "Microsoft's 2025 revenue" has two honest answers and they are far apart. Fiscal 2025, the twelve months to 30 June 2025, was $281,724m. Fiscal 2026, the twelve months to 30 June 2026, was $331,839m. Both figures appear in the FY2026 Form 10-K, side by side, and calendar 2025 is neither of them.
The pattern repeats across the market, and it is worst where the label is furthest from the calendar:
Snowflake is the sharp case. Asked for "Snowflake's 2025 results", a model can reasonably return fiscal 2025, which ended in January 2025 and is mostly calendar 2024, or fiscal 2026, which is mostly calendar 2025. Two answers, a full year apart, both defensible.
The check: make the model state the period end date, not the year label. "Fiscal 2026" is ambiguous. "The twelve months ended 30 June 2026" is not. One follow-up question removes the whole class of error.
Dates fixed, the next problem is that the parts of a company do not always add up to the whole.
Segment disclosure is where a reader most often builds a number the company never reported. Add up the segments, call the result the company total, and you land wrong in either direction depending on what was left out.
Alphabet's 10-K for the year ended 31 December 2025 shows this cleanly, and in both directions at once.
| Segment | Revenue | Operating income |
|---|---|---|
| Google Services | 342,721 | 139,404 |
| Google Cloud | 58,705 | 13,910 |
| Other Bets | 1,537 | (7,515) |
| Hedging gains (losses) | (127) | n/a |
| Alphabet-level activities | n/a | (16,760) |
| Consolidated total | 402,836 | 129,039 |
Adding the three named segments gives revenue of $402,963m against a reported $402,836m, and operating income of $145,799m against a reported $129,039m, an overstatement of 13.0%. Source: Alphabet Inc. Form 10-K, FY2025, Note 15 and Segment Profitability in Item 7. Filed 5 February 2026.
The revenue gap is small and comes from hedging. The operating income gap is $16.8bn of shared artificial intelligence research, philanthropy and corporate cost that Alphabet deliberately does not allocate to any segment, and the filing says so in plain language. A reader who sums the three segments has not uncovered a hidden number. They have built one Alphabet took care not to report.
Other structures do the same thing for other reasons. Intersegment sales are eliminated on consolidation, so segment revenue can exceed the total. A "Corporate and Other" bucket can hold costs, a real operating business, or both.
The check: ask for the reconciling line by name and value. A model that cannot name what sits between the segments and the total has read the segment table, not the segment note.
An earnings press release carries two sets of books. One is prepared under GAAP. The other is management's preferred view, with items it considers unrepresentative removed. Both are legitimate. Only one is the accounting number.
Snowflake's fiscal 2026, the year ended 31 January 2026, shows how far apart they get.
Source: Snowflake Inc., fourth quarter and full year fiscal 2026 results, furnished as Exhibit 99.1 to a Form 8-K on 25 February 2026. GAAP net loss cross-checked against the company's XBRL filing data.
The sign flips. A company that lost $1.44bn at the operating line reports $490m of operating income on the adjusted basis, and the difference is very largely the cost of paying employees in shares. That is a real cost to existing holders even though no cash left the building, which is why both presentations exist and why the argument between them is genuine rather than cosmetic.
The risk is not that either number is wrong. It is that the release leads with the adjusted figure, the summary of the release keeps it, and the word "adjusted" falls off somewhere in between.
The check: ask whether the figure is GAAP or non-GAAP, and for the largest reconciling item by name and value. A company reporting a non-GAAP measure must reconcile it to the nearest GAAP measure, so the answer is always in the same document.
Even with the basis settled, the number may have changed since it was first reported.
The figure for a completed year is not fixed. It can be restated for error, recast for a new accounting standard, or reclassified for a presentation change, and the filing from three years ago still shows the old number. Both are right on their own basis.
Kraft Heinz gives the clearest illustration available, because its Form 10-K for fiscal 2018 prints all three versions of the prior year in one table.
| Line | As previously reported | As restated | As restated and recast |
|---|---|---|---|
| Net sales | 26,232 | 26,076 | 26,076 |
| Gross profit | 9,703 | 9,591 | 9,033 |
| Operating income | 6,773 | 6,693 | 6,057 |
Operating income for the same twelve months reads 10.6% lower in the later filing than in the earlier one. Source: The Kraft Heinz Company Form 10-K, fiscal 2018, Note 2, Restatement of Previously Issued Consolidated Financial Statements. Filed 7 June 2019.
Two different mechanisms moved those two columns. The restatement corrected misstatements in procurement accounting. The recast applied a new pension presentation standard retrospectively, shifting $558m out of gross profit without changing net income at all. One is a correction. The other is not.
A model reading the 2017 annual report will tell you Kraft Heinz earned $6,773m of operating income that year. A model reading the 2018 report will say $6,057m. Neither is hallucinating.
The check: ask which filing the prior-year figure came from, by document and date. Take a historical number from the most recent filing that presents it, and expect the older one to disagree.
Here are the five, on one screen.
| Field | Why it goes wrong | Ask for |
|---|---|---|
| Share count | Four valid counts in one filing | Which count, and its purpose |
| Fiscal year | Label does not match the calendar | The period end date |
| Segments | Unallocated costs and eliminations | The reconciling line, by name |
| Adjusted vs GAAP | The qualifier gets dropped | Basis, and the largest add-back |
| Restated priors | Two filings, two right answers | Source filing and its date |
None of the five requires you to read the filing yourself. Each requires the model to say where it looked, which is a much smaller ask, and a wrong answer usually collapses on the first follow-up.
A taxonomy of failures is a biased sample, so it is worth saying what works. Corva uses language models, so we have an obvious interest in this section reading well and you should discount it accordingly. Here is the honest version anyway.
Models are strong at finding and explaining, weak at selecting and computing. Everything above is a selection failure. These are not:
The measured record on the numerical half is poor, and it deserves stating plainly. FinanceBench, a 2023 benchmark built on questions answerable from public filings, reported that GPT-4-Turbo paired with a retrieval system answered incorrectly or refused on 81% of questions. Models have improved considerably since then and that figure should not be read as current. The shape of the weakness has not changed, which is why the five checks are worth having.
The useful division of labour is old and boring. Let the model find the page. Do the selecting yourself, or use something that shows you which line it selected.
Five fields is a filter, not an audit, and it has real gaps.
It depends entirely on whether it retrieved the filing or is answering from memory. With a live document in context it is reasonably good at finding figures and poor at choosing between similar ones, which is the failure this post is about. Without a document it should not be trusted for any specific figure, because a plausible number is exactly what a language model produces when it has nothing to read.
Weighted average diluted for anything per-share, the cover-page count for market capitalisation, and the ending balance for measuring dilution over time. Apple's FY2025 filing carries all three plus the basic weighted average, and they range across 231 million shares.
Usually intersegment eliminations, unallocated corporate items, or a reconciling line such as hedging. Alphabet's three named segments produced $402,963m of 2025 revenue against a consolidated $402,836m, with the difference being hedging losses, and its segment operating income overshoots the consolidated figure by $16.8bn of unallocated cost.
Ask, and ask for the largest reconciling item alongside it. Companies reporting non-GAAP measures must present the most directly comparable GAAP measure and reconcile the two, so both numbers sit in the same document. Where they diverge by a lot, as at Snowflake, the gap is normally stock-based compensation.
Yes, and it happens more often through recasting and reclassification than through restatement for error. A new accounting standard applied retrospectively, or a change in segment structure, will move prior-year figures in the current filing while the original filing stays on EDGAR unchanged. Read more on how presentation changes move a line without moving the business in our post on what causes gross margin to decline.
No. It means using it for retrieval and explanation, and treating every specific figure as a claim that carries a source. If you want the underlying document skills as well, start with reading a 10-K and the rest of the Corva blog.
Corva uses language models, and says so. What it does not do is let one choose a number unsupervised. Figures are computed from the filed statements rather than generated, cross-checked against SEC EDGAR, and carry a graded citation you can open. Where three sources disagree, the disagreement is flagged rather than averaged away, and where a figure cannot be found the report says it is unsure instead of inventing one.
Research a company free →Live financials on any listed company are free. No card.
Corva is a research tool, not a broker or investment adviser, and nothing here is a recommendation to buy, sell or hold any security. Apple, Microsoft, Alphabet, Snowflake and Kraft Heinz are cited only as documented examples of disclosure structures that are easy to misread, and for no other reason. All figures are taken from the filings linked in the text and were checked against those documents on 7 September 2026; per-share and percentage figures not printed in a filing are computed from the figures shown. Verify anything you intend to act on against the primary filing. See terms and disclaimer.