
A financial model that balances can still destroy millions in decision quality.That is the core problem with financial AI today.
A financial model that balances can still destroy millions in decision quality.
That is the core problem with financial AI today.
Large language models can generate spreadsheets that look indistinguishable from analyst work. They can follow structure, label rows, populate numbers, and produce outputs that pass a visual scan. But appearance is not analytical integrity. A DCF can calculate correctly while using the wrong cash flow definition. A debt schedule can hold in the base case and collapse under stress. A sensitivity table can be entirely hardcoded while appearing dynamic. A working capital schedule can carry the wrong sign convention through five projection years before anyone notices.
The failure is not cosmetic. It is structural. And it is invisible to any evaluation system that is not built by people who know what to look for.
This is what the field calls surface-valid modeling: outputs that satisfy every visual and structural check while being analytically wrong at the level of formula logic, assumption sourcing, or model integration. Surface-valid modeling is the defining challenge in financial AI right now. It cannot be solved by scale alone. It cannot be solved by better base models alone. It requires a different kind of infrastructure entirely.
For the past two years, the dominant conversation around AI in finance has been about automation: summarizing 10-Ks, extracting earnings data, generating first-draft investment memos, screening public companies. That wave is still underway and still valuable.
The next wave is more consequential.
AI systems are now being trained and evaluated on actual analyst workflows: three-statement models, DCF builds, LBO structures, credit analysis, SEC filing reviews, and structured research outputs. OpenAI has described internal testing where ChatGPT Agent was evaluated on first to third-year investment banking analyst tasks, graded across hundreds of criteria covering correctness, formula logic, and workflow quality.
This signals a structural shift in what financial AI is being asked to do. The question is no longer whether AI can accelerate information retrieval. The question is whether AI can replicate judgment-heavy analytical work reliably enough to trust at the point of decision.
That is a fundamentally different problem. And the answer, as the benchmark data makes clear, is that the field is not close.
Despite billions in AI investment, frontier models still fail the majority of real-world finance workflows.
The Finance Agent Benchmark covers 537 expert-authored tasks spanning retrieval, market research, financial projections, and complex modeling drawn from recent SEC filings. The best-performing model achieved 46.8% accuracy. That is not partial analyst replacement. That is evidence that financial reasoning remains fundamentally unreliable without the right evaluation infrastructure underneath it.
Finch, a benchmark covering real-world finance and accounting workflows including modeling, validation, cross-file retrieval, and reporting, found that GPT-4.1 Pro passed only 38.4% of composite workflows. These are not adversarial edge cases or exotic stress tests. They are the tasks analysts handle routinely, on ordinary deals, for ordinary clients.
The ResearchRubrics benchmark, built with more than 2,800 hours of human labor and over 2,500 expert-written criteria, found that leading deep research AI agents achieved under 68% average rubric compliance. The primary failure modes were missing implicit context and inadequate reasoning over retrieved information. Finance is dense with exactly this kind of implicit context. The assumptions that senior analysts carry in their heads, the ones that never appear in any prompt, are precisely what separates a model that looks right from one that is right.
A 46.8% benchmark score does not mean AI handles roughly half of analyst work correctly. It means that in a structured evaluation built around real tasks, the model fails more often than it succeeds. For an institution relying on that system to support valuation, credit, or acquisition analysis, the error rate is not a statistic. It is a liability.
The dominant assumption in AI development remains that enough scale eventually solves quality. Finance breaks that assumption.
Financial modeling is not primarily a language problem. It is a systems integrity problem.
A three-statement model, a DCF, an LBO: the scaffolding is well-documented. Models trained on enough financial content understand the structure. What they do not reliably replicate is the judgment embedded inside that structure, because that judgment does not appear in any document. It lives in the mental checklist of a trained analyst who has built enough models to know what breaks and where.
Consider a leveraged buyout model where interest expense accrues on beginning debt rather than average debt during a rapid paydown period. The model balances. The returns appear reasonable. The output passes every structural check. But the error compounds across projection years, systematically overstating sponsor returns and distorting acquisition pricing. A generic AI evaluator does not detect it. A trained analyst notices it on first review.
That is the analytical integrity gap: the distance between what a model produces and what a trained analyst would trust. Closing that gap is the central challenge in financial AI development. And it cannot be approached without evaluation infrastructure built to the same standard as the work itself.
The questions that expose this gap are not exotic:
Does revenue growth link properly across all periods? Does working capital flow through the cash flow statement with the correct sign convention? Is the debt schedule circularity handled or broken? Is the DCF built on unlevered free cash flow rather than the wrong metric? Are sensitivity tables dynamic or hardcoded? Does the model produce the right output for the wrong reason, which is arguably more dangerous than producing the wrong output transparently?
Checking these things requires analyst-level evaluation criteria. Those criteria do not emerge automatically from task descriptions. They have to be built deliberately by people who know what to look for.
That process runs on three assets: the prompt, the gold model, and the rubric.
A prompt in AI training is not a user instruction. It is the task definition. Everything downstream the gold model, the rubric, the evaluation score — is a function of how precisely the prompt describes what the analyst is supposed to build.
A weak prompt says: "Build a DCF model for this company."
A strong prompt says: "Using the attached 10-K, build a five-year unlevered DCF with a UFCF build showing EBITDA, SBC, unlevered taxes, capex, and change in NWC. Apply mid-year convention with a stub factor for the partial first year. Calculate terminal value using an exit multiple method with an NTM toggle. Bridge from enterprise value to equity value using the balance sheet inputs provided. Keep all hardcoded assumptions in a separate input section, use formulas for all projected periods, and include a balance check and sensitivity table linked dynamically to WACC and terminal multiple."
The difference is not length. The difference is specificity. The strong prompt tells the AI system exactly what a real analyst deliverable looks like, what it must contain, how it must be structured, and what quality standards it must meet. Without that precision, evaluation becomes subjective. Two reviewers looking at the same output will disagree because the task definition left room for interpretation, and that ambiguity compounds through every downstream scoring step.
Good financial AI prompts share several properties. They specify the source material and what must be drawn from it. They define the deliverable type and scope. They set explicit modeling conventions: whether to use beginning or average balances for interest, whether to apply a stub period, whether working capital should be calculated as a change in operating assets net of operating liabilities. They include formatting standards. And they are written at the level of precision a senior analyst would use when delegating to a junior analyst who has never seen the model before.
The gaps that cause the most downstream damage are modeling convention gaps. Whether a stub period applies. Whether interest accrues on beginning or average debt. Whether working capital is defined net of cash and revolver. These are never specified in a weak prompt. AI systems fill them with a plausible default, the output looks structurally complete, and the error only becomes visible when a reviewer checks the formula logic rather than the final number.
A prompt written without modeling experience produces tasks that are technically complete and analytically ambiguous. That ambiguity does not stay in the prompt. It propagates into every gold model and every rubric built against it, embedding the same analytical uncertainty into the evaluation infrastructure itself.
The Gold Model: The Standard Everything Is Measured Against
The gold model is the expert-built answer. It is the benchmark against which AI-generated outputs are compared, and it is the single most consequential asset in the evaluation pipeline.
Most people underestimate what makes a gold model actually gold. A clean-looking Excel file is not sufficient. A model that produces correct final outputs is not sufficient. A gold model used in AI training must be correct at every layer: historically accurate inputs, properly sourced assumptions, formula-driven schedules with no hardcoded projected values, correct cash flow logic, dynamic valuation outputs, and formatting that makes every cell auditable by a reviewer who did not build it.
Gold model failures are often subtle. An assumption hardcoded in a projected period rather than linked to an input cell. A working capital change calculated correctly in isolation but carrying the wrong sign into the cash flow statement. A debt schedule that rolls forward accurately in a base case but breaks when a paydown assumption is changed. A DCF that applies the right discount rate to the wrong cash flow definition. Each of these passes visual review. Each of them corrupts the evaluation.
The most common gold model failure is not a formula error. It is a hardcoded assumption in a projected period that passes visual inspection and only surfaces when a downstream input is changed and the output does not move. When that flaw exists in the gold model and an AI system replicates the surface behavior, the rubric cannot catch it. The evaluation scores the wrong behavior as correct, and that incorrect signal feeds back into training.
This is why gold model construction is a distinct discipline from financial modeling for client delivery. A client model needs to be correct and presentable. A gold model needs to be correct, auditable, internally consistent under perturbation, and built with the explicit awareness that it will be used to define what right looks like for an AI system. Every cell is a decision about what behavior the evaluation will reward.
The gold model also determines the boundary of the evaluation. If a task asks an analyst to build a dashboard pulling metrics from an existing three-statement model, the gold model defines exactly which cells source from which tabs, with what formula logic, formatted to what standard. The rubric is then written against that specific output. Any deviation from the gold model's structure, sourcing, or logic becomes a scoreable failure mode. And any rubric item that tests something the gold model does not define is testing a standard that was never established, producing scores that measure nothing.
A rubric is not a checklist. It is an evaluation system designed to convert analyst judgment into criteria precise enough for consistent scoring, whether by a human reviewer or an LLM judge.
A well-designed financial modeling rubric has several properties that distinguish it from a generic evaluation framework.
It is scoped to the actual task. Before writing a single criterion, the reviewer compares the input and output workbooks. The difference between them defines the task surface area. If the only new element is a dashboard tab, the rubric tests the dashboard tab. It does not test the income statement, the balance sheet, or any tab that was not modified. Rubric items that check unchanged areas of a model award points regardless of whether the actual deliverable was built correctly. These are free points, and they corrupt the scoring signal.
Every criterion is atomic. A rubric item tests one thing. A criterion that asks whether a current ratio is linked to the balance sheet and calculated correctly is two questions compressed into one. Split it. One item checks the linkage. A separate item checks the formula logic. Binary scoring only works if the question being scored is genuinely binary.
Every criterion is self-contained. A reviewer should be able to grade any item using only the criterion and the workbook, without referencing the prompt or applying a general sense of whether the model looks reasonable. Criteria that use language like "consistent with best practices" or "formatted appropriately" cannot be graded consistently. A criterion that specifies the exact source cell, the expected formula structure, and the acceptable output range can.
It distinguishes between formula correctness and model integration. Formula correctness checks whether a calculation is mathematically right. Model integration checks whether the formula draws inputs from the right source: does the operating income figure on the dashboard link back to the income statement, or is it hardcoded or recalculated from scratch? These are different failure modes. A model can have correct formulas applied to wrong inputs. Treating both under the same rubric category produces ambiguous scoring and misses the integration failure entirely.
It includes perturbation testing. Perturbation is the most revealing test in a financial model rubric. Change an input assumption and verify that every downstream output that depends on it updates correctly. If the model is genuinely dynamic, the change propagates. If outputs were hardcoded to match expected values rather than built through linked formulas, they stay fixed.
A model can pass output validation entirely and fail perturbation entirely. That tells you the outputs were reverse engineered to match rather than constructed through correct logic. That distinction matters enormously for AI training because it separates a model that understands the underlying structure from one that has only learned to reproduce the answer. Training on the former builds capability. Training on the latter builds the appearance of capability.
Weights are proportional to analytical consequence. A structural logic error in a core financial schedule is a fundamentally different failure from a minor formatting inconsistency. A hardcoded output where a formula should exist is more dangerous than a mislabeled row header. A well-designed rubric reflects this hierarchy explicitly in its point allocation, assigning high weights to formula logic, model integration, and output accuracy, and lower weights to presentation and structural completeness checks. The result is a scoring system that rewards analytical correctness over surface appearance. That is the right signal for AI training.
Building a prompt is difficult. Building a gold model is time-intensive. Designing the rubric is the hardest part, because it requires the reviewer to hold two things in mind simultaneously: what the task asks for, and what a trained analyst would notice if it were done wrong.
The single most frequent failure is scope bleed: rubric items written against parts of the workbook the task never touched, awarding points that have no relationship to whether the actual deliverable was built correctly. A rubric item that checks for error values across an entire workbook is wrong if the task only creates a new tab. The correct item restricts the check to the task area. Checking unchanged tabs adds no information and dilutes the scoring signal.
The next most common failure is criteria stacking. A rubric item that tests whether the current ratio is correctly calculated and properly linked to the balance sheet is two items, not one. Stacked criteria introduce ambiguity into every downstream scoring step, including LLM-as-a-judge pipelines that depend on clean binary inputs to produce reliable scores.
The most dangerous failure is a mismatch between what the rubric assumes the model is doing and what the gold model actually does. If a metric is sourced from elsewhere in the workbook rather than calculated at the point of use, a rubric item written around the calculation logic is testing the wrong behavior entirely. An LLM judge presented with that criterion applies it confidently and scores against the wrong standard. The results appear valid. They measure nothing meaningful.
The fix requires rubric writers to work from the gold model outward, not from general financial logic inward. What the model does determines what the criterion should test, not what the model theoretically should do.
These errors compound. A rubric with misscoped items, stacked criteria, and vague formatting checks does not just produce inaccurate scores. It produces inaccurate training signals. AI systems trained against flawed rubrics learn the wrong behavior, and that behavior is difficult to unlearn. The downstream cost is not a bad benchmark number. It is a financial AI system that has been systematically optimized for the wrong standard.
FrontierFinance, a benchmark focused on real-world financial modeling evaluation, addresses this directly by using human experts to define tasks, create rubrics, grade outputs, and perform tasks themselves as human baselines. The core finding is consistent with the broader field: the quality of automated evaluation is entirely a function of the quality of the rubric underneath it. Automated evaluation at scale is only as reliable as the expert judgment used to build the evaluation criteria.
The through-line across prompts, gold models, and rubrics is the same. The quality of the evaluation infrastructure depends entirely on people who have done the underlying financial work at a level where they can recognize failure modes that do not announce themselves.
A prompt written without modeling experience leaves gaps that an AI system fills with plausible but wrong defaults. A gold model built without senior-level review carries hidden errors that propagate through every rubric written against it. A rubric designed without genuine familiarity with how analysts review financial models tests the wrong things at the wrong granularity and misses the errors that matter most.
This is why financial AI evaluation cannot be treated as a data labeling problem. Generic annotation at scale does not produce the judgment required to catch a sign convention error in a working capital schedule, recognize a debt circularity that holds in the base case and breaks under a downside scenario, or identify a valuation bridge that calculates correctly from inputs that were never meant to be hardcoded.
The work requires people who have built DCFs and defended assumptions under scrutiny. People who have seen enough models to know that the most dangerous errors are the ones that look fine. People who understand that surface-valid modeling is not a step toward correct modeling. It is a distinct failure state that standard evaluation frameworks are not built to detect.
In consumer AI, a hallucinated answer is inconvenient. In finance, a structurally flawed model can distort valuation, debt capacity, acquisition pricing, or investment committee decisions. The cost of analytical incorrectness is not a product quality issue. It is a capital allocation issue. The institutions that understand that distinction are building their evaluation infrastructure accordingly. The ones that do not are generating surface-valid financial AI and mistaking it for reliable financial AI.
What surprises most people when they first build financial evaluation taxonomies is that the recurring failures are not dramatic. They are not wrong discount rates or missing schedule tabs. They are sign convention errors in working capital, integration breaks between a summary output and its source schedule, and valuation bridges that calculate correctly from inputs that were never meant to be permanent assumptions. These patterns recur across AI systems and across tasks. Cataloguing them systematically is what converts evaluation work into a feedback loop that makes financial AI progressively more reliable over time.
The long-term competitive advantage in financial AI will not belong solely to model providers. It will belong to firms that build proprietary evaluation infrastructure: domain-specific prompts, expert-constructed gold models, perturbation frameworks, and analyst-designed rubrics precise enough to catch the failures that benchmarks currently miss.
That infrastructure is being built now, by a small number of teams with the right depth of financial expertise behind them. The asset they are accumulating is not data in the conventional sense. It is structured judgment: the documented, auditable, transferable knowledge of what financial analytical correctness actually looks like at the level of the cell.
Three things follow from this.
First, evaluation expertise becomes a durable moat. A team that has built 200 gold models across DCFs, LBOs, and credit analysis, each with perturbation-tested rubrics and documented failure taxonomies, holds something that cannot be replicated quickly. The prompt-rubric-gold model pipeline is not a technical system. It is an accumulated analytical standard, and analytical standards take time to build correctly.
Second, the benchmark numbers will improve, but interpretation will matter more than the numbers themselves. A financial AI system that achieves 80% rubric compliance against a well-designed evaluation framework is meaningfully different from one that achieves 80% against a framework with scope bleed and stacked criteria. As evaluation infrastructure matures, the quality of the rubric will become as important as the score it produces.
Third, the institutions that invest in evaluation infrastructure now will define what financial AI quality means for the next several years. The ones that treat evaluation as a secondary concern, as a checkbox rather than a discipline, will find that the gap between what their models generate and what analysts trust does not close on its own. Surface-valid modeling, left unchecked, does not converge toward structural correctness over time. It gets better at looking correct.
In finance, trust is not generated by fluency. It is generated by analytical reliability. The firms that can measure that reliability systematically, and build evaluation infrastructure precise enough to enforce it, will determine how much of the analyst workflow AI can credibly take on, and how soon.
The ones that cannot will keep producing models that balance.
This piece reflects our ongoing work building financial AI evaluation infrastructure across modeling, credit analysis, and structured research workflows.