Why AI Still Stumbles on Financial Questions

A Patronus AI study found that major language models still struggle to answer questions about SEC filings accurately. Even GPT-4 Turbo, the best performer in the test, reached only 79 percent accuracy, showing why human review remains important in financial workflows.

WTF Index IDIOCRACY
◄ Terminator 1 Idiocracy 3 ►

The story emphasizes AI producing unreliable or fabricated financial answers, eroding truth and quality in professional workflows.

Why AI Still Stumbles on Financial Questions

Large language models are moving into research, customer service, and professional workflows, but finance remains a difficult test. A study from Patronus AI found that even leading models can fail when asked to answer questions based on SEC filings, including cases where the relevant report is largely provided in the prompt.

The Core Finding

Patronus AI tested four language models on questions about corporate financial reports: OpenAI's GPT-4 and GPT-4 Turbo, Anthropics Claude 2, and Metas Llama 2. The strongest result came from GPT-4 Turbo, which reached only 79 percent accuracy in the test.

That matters because the task was not presented as open-ended market forecasting or personal investment advice. The models were being asked to retrieve, interpret, or calculate answers from SEC filings. In many cases, the information needed to answer was available in the material supplied to the model.

The study highlights a practical concern for regulated industries. If a model cannot reliably answer grounded questions from official financial documents, then deploying it directly in customer service or research processes creates obvious risk. In finance, an answer that is nearly right can still be wrong in a consequential way.

What Went Wrong

The models did not fail in just one way. According to the source, they sometimes refused to answer questions. In other cases, they made up facts and figures that were not present in the SEC filings.

That combination is especially difficult for production use. A refusal may slow down a workflow, but a fabricated figure can mislead a user who assumes the system is reading from the supplied document. For financial teams, the problem is not only whether an AI model sounds fluent. The issue is whether it can consistently connect an answer to the right source data.

Anand Kannappan, co-founder of Patronus AI, said the performance is unacceptable and should be much higher for automated and production-ready applications. The point is not that language models have no place in finance. It is that current systems still need careful boundaries when the output depends on precise document understanding.

FinanceBench Raises the Bar

For the test, Patronus AI created FinanceBench, a dataset containing more than 10,000 questions and answers drawn from SEC filings of major public companies. The dataset includes the correct answers and the exact location of those answers in the reports.

The questions are designed to reflect the kind of work financial users might expect an AI assistant to support. Some require finding a fact. Others require basic mathematical or logical reasoning based on line items shown in the filing.

Examples from FinanceBench include:

  • Has CVS Health paid dividends to common shareholders in Q2 of FY2022?
  • Did AMD report customer concentration in FY22?
  • What is Coca Cola’s FY2021 COGS % margin? Calculate what was asked by utilizing the line items clearly shown in the income statement.

These examples show why the benchmark is demanding without being abstract. A useful finance assistant must do more than summarize prose. It must identify relevant information, ignore irrelevant text, perform simple reasoning when needed, and avoid inventing unsupported details.

Why Human Review Still Matters

The researchers believe AI models could still become useful tools for the financial sector as they improve. The study does not frame the technology as hopeless. Instead, it points to a gap between what current models can do and what financial applications require.

For now, that gap means humans still need to remain in the workflow. A person can review whether the model used the right filing, found the right line item, and produced an answer that follows from the source. In regulated settings, that review is not a minor detail. It is part of making the process manageable.

The source also notes that OpenAI's usage guidelines rule out using an OpenAI model to provide individual financial advice without review by a qualified person. That aligns with the broader lesson from the study: language models may assist, but they should not be treated as unchecked financial authorities.

The Long-Context Problem

One technical issue behind these failures is the known difficulty language models have when extracting information from long texts. The source refers to this as the "lost in the middle" problem, where information placed in the middle of a long context can be harder for a model to use reliably.

This raises a bigger question about large context windows. Giving a model more text does not automatically mean it will find and use the right detail. In finance, where long reports contain dense and highly specific information, that limitation becomes more visible.

Improved prompting may help in some cases by making the task clearer and directing the system toward relevant information. The source also notes that Anthropic has developed a method for Claude 2.1 that begins the answer with the sentence "This is the most relevant sentence in the context:". Tests still need to show whether that approach works reliably across many tasks and whether it can produce similar gains for other LLMs such as GPT-4 (Turbo).

The practical takeaway is straightforward. AI may become a stronger financial research assistant, but accuracy, traceability, and review remain central. Until models can handle SEC filings with much higher reliability, financial questions are still too important to leave entirely to automation.