7. How to audit an AI-assisted financial report
An audit should begin with the report’s purpose. What decision or understanding is the report intended to support? Was the analytical question framed neutrally, or was the system directed toward a preferred result?
The reviewer should then examine the report across several dimensions:
| Audit dimension | Controlling question |
|---|---|
| Source integrity | Are the material quantitative and qualitative sources used by the models identified? |
| Platform neutrality | Did the model selection, programming, or feature set favor a predetermined conclusion? |
| Analytical separation | Are observations distinguished from transformations, classifications, and inferences? |
| Reproducibility | Can the calculations and classifications be reconstructed? |
| Contrary evidence | Does the report identify material facts supporting another interpretation? |
| Confidence | Are uncertainty and methodological limits expressed proportionately? |
| Comprehension | Can the intended reader identify the controlling inference? |
| Decision boundary | Does the report distinguish intelligence from a recommendation? |
| Accountability | Has a qualified human reviewed the evidence and approved publication? |
This review should compare the final prose with the analysis that produced it. Editing may inadvertently remove a qualification, generalize a limited observation, or convert a conditional inference into an apparent prediction. A polished sentence should fail the audit when it overstates the underlying evidence.
Stylistic review remains part of the process. Formulaic language should be revised when it substitutes rhetoric for analysis. It should not be removed mechanically when it expresses the analysis accurately.
Human review remains essential because the market meaning of a fact cannot always be reduced to syntax or calculation. An experienced stock trader may recognize that an apparent change in available supply resulted from a change in reporting coverage, timing, or classification rather than from a real change in lendable inventory. The trader may also detect that a fee comparison uses incompatible populations or that an event date changes the appropriate observation window. Those judgments require domain knowledge and responsibility for the consequences.
8. When the human auditor becomes the choke point
The audit framework described above still depends upon a human reviewer. That person must inspect the sources, reconstruct the calculations, identify embedded assumptions, evaluate contrary evidence, and determine whether the conclusion exceeds its support. As AI-generated reports become longer and their analytical relationships become more complex, the human auditor can become the production choke point. The AI report may be generated in minutes, while a responsible human review still requires hours or even days.
That human constraint is easy to underestimate. In his MIT lecture How to Speak, Patrick Henry Winston, a longtime MIT professor and former director of its Artificial Intelligence Laboratory, observed that people absorb ideas at roughly the speed of writing on a blackboard. He added that an audience cannot keep pace with a rapid sequence of slides.9
Inferential layering helps the executive consume the analysis, but it does not eliminate the auditor’s burden. Each layer may contain transformations, classifications, thresholds, probability estimates, or dependencies that must be checked against the layer below it. A human reviewer can confirm a limited number of such relationships carefully. At scale, however, the reviewer may begin sampling the chain rather than verifying it.
A possible solution is to program the inferential proposition down to its mathematical and logical fundamentals. Each material conclusion would be expressed as a formal relationship among defined premises, transformations, constraints, and decision rules. Here, formal means expressed in a precisely defined logical or mathematical language so that each step can be checked mechanically. The prose would remain available to the executive, but the underlying inference would also exist in a machine-verifiable form.
Stock-loan traders must decide when and to whom their institutional loans are made. Then they must rerate and recall those loans when conditions are right. Consider this simplified securities-finance inference:
- New-loan activity exceeded returns during the observation window.
- Reported available supply declined.
- The marginal new-loan fee rose above the seasoned-book fee.
- The approved classification rule associates that combination with prospective tightening.
- Therefore, the evidence supports a prospective-tightening classification, subject to the stated confidence limits.
A formal machine auditor could confirm that the stated conclusion follows from the encoded premises and approved classification rule. It could recalculate the relationships, test whether the thresholds were satisfied, locate inconsistent definitions, and reject a conclusion that did not follow logically. The system would not merely produce another opinion about the report. It would generate or verify a proof of the inferential proposition. But the executive or trader still has to make the final decision.
Tudor Achim describes the broader ambition as “mathematical superintelligence.” His approach draws upon formal verification, through which reasoning is translated into a mathematical language whose steps can be checked rather than merely accepted.10 Harmonic, the company he co-founded, describes its objective as producing transparent and automatically verifiable reasoning traces. Its Aristotle system uses the Lean proof environment to construct and verify formal proofs.11
Achim’s work suggests a possible direction for financial-report auditing. A sufficiently capable AI-based agent auditor could test thousands of inferential relationships without relying upon fluency, intuition, or statistical resemblance to earlier reports. It could establish whether a conclusion follows from the premises, whether every required condition was met, and whether the report’s prose accurately represents the formal result.
The boundary is important. Mathematical verification cannot prove that a market observation was recorded correctly, that a selected feature captures the intended economic behavior, or that an empirical relationship will persist. It proves deductive validity within a specified formal system. Human responsibility therefore moves toward validating the data, definitions, objectives, and assumptions presented to that system.
Under a defined governance policy, the AI agent auditor should have authority to block publication when a material inference fails verification. It should also be able to initiate withdrawal or correction of a distributed report when later review identifies a material analytical defect. That authority would require documented thresholds, escalation procedures, human accountability, and a preserved record of the failed proof obligation.
The resulting division of responsibility is more credible than either human-only or AI-only review. For board-level reporting, human experts should apply courtroom-level rigor when determining whether the premises correspond to the market and whether the analytical question is properly framed. A formally grounded AI auditor can determine whether the inference follows from those premises. Harari’s alignment point adds a separate requirement: the system must remain aligned with the authorized human objective, not merely execute its assigned logic correctly.12 Executives then decide how much practical weight the verified inference deserves within their organization’s obligations, policies, and risk limits. Only then should a board member accept an AI-enabled report in its context.
Footnotes
- Patrick Henry Winston, “How to Speak,” MIT OpenCourseWare, January IAP 2018, “The Tools: Boards, Props, and Slides,” beginning at 13:24; see also the official transcript, 4-5. ↩
- Tudor Achim, “The Path to Mathematical Superintelligence,” TEDAI San Francisco 2025, YouTube video. ↩
- Harmonic, “Introducing Harmonic: Our Mission and First Results,” June 10, 2024. ↩
- Yuval Noah Harari, Nexus: A Brief History of Information Networks from the Stone Age to AI (New York: Random House, 2024), chap. 6. ↩
