A new cost frontier for financial document parsing

GPT-6 Luna and ten other model configurations, run through Pathway’s production bank-statement harness.

In Jack, Pathway’s financial document harness, GPT-6 Luna parsed a business’s bank statements for under seven cents: 77% less than Gemini 3.7 Flash and 78% less than Gemini 3.8 Flash. It closed every ledger we could check against the statements’ printed balances. Its median time was 10% shorter than Gemini 3.8 Flash and only 9% longer than Gemini 3.7 Flash. And that’s with OpenAI fast mode off.

Cost–latency Pareto frontier

Google GeminiOpenAIPareto frontier
Median completion time (seconds)60s100s140s180s220s260s$10$5$2$1$0.5$0.2$0.1$0.05Mean cost per task (USD) · logarithmic scale3 Flash3.5 Flash3.7 Flash3.8 Flash3.5 Flash LiteGPT-6 LunaGPT-6 SolGPT-6 AstraGPT-5.6 LunaGPT-5.6 TerraGPT-5.6 Sol
Mean inference cost and median completion time per task, for 11 model configurations on the same five tasks. A configuration is on the frontier when no other configuration has both a lower cost and a shorter time.
Model results

Ledgers closed counts ledgers that reconcile with the printed closing balance, out of ledgers with printed opening and closing balances. Time covers the complete task.

ConfigurationCost / taskMedian timeLedgers closed
Gemini 3 Flash$0.245165.3 s33 / 33
Gemini 3.5 Flash$0.752115.8 s26 / 27
Gemini 3.7 Flash$0.29093.7 s27 / 27
Gemini 3.8 Flash$0.305113.3 s27 / 27
Gemini 3.5 Flash Lite$0.195102.3 s27 / 30
GPT-6 Luna$0.068101.9 s33 / 33
GPT-6 Sol$1.254184.5 s29 / 30
GPT-6 Astra$6.497234.8 s31 / 33
GPT-5.6 Luna$0.140158.7 s32 / 32
GPT-5.6 Terra$1.349152.7 s28 / 30
GPT-5.6 Sol$2.634156.1 s33 / 33

Each configuration ran every step of the harness on the same five tasks. A task is one business’s bank statements, parsed from the source PDFs into closed ledgers and classified transactions. The frontier has two points: Gemini 3.7 Flash, the fastest configuration, and GPT-6 Luna, the cheapest. The other configurations cheaper than Gemini 3.7 Flash each gave something up. Gemini 3 Flash and GPT-5.6 Luna took at least 69% longer. Gemini 3.5 Flash Lite was nearly as fast as Gemini 3.7 Flash but left 3 of its 30 checked ledgers open. GPT-6 Luna took the same time as Gemini 3.5 Flash Lite, cost about a third as much, and closed all 33 of its checked ledgers.

GPT-6 Luna also cost 52% less than GPT-5.6 Luna, and its median time was 36% shorter. The two releases used about the same number of tokens on these tasks, so the cost reduction comes from price: the GPT-6 Luna release halved input-token prices and cut output-token prices by about 58%. The shorter time does not. OpenAI trained GPT-6 Luna with methods similar to GPT-6 Astra’s and recommends it for focused, repeated work at scale, such as extracting fields and classifying requests. That is Jack’s workload: each task extracts and classifies hundreds of transactions.

Jack has run in production for two years on the Gemini model family, which we chose for its visual understanding of documents. It has reconciled millions of ledgers. Jack currently runs a task in eight steps of model calls and code. Every call sits behind a provider adapter, so changing models is a configuration change. For this evaluation, we set every call to one model and kept the rest of the harness fixed.

Cost per task is the unit economics of document work like this. Every business a lender underwrites arrives as a stack of statements, and every statement has to become a closed ledger before anyone can make a decision. At under seven cents a task, letting the model read every row is cheaper than engineering around the reading: in our tests, GPT-6 Luna reading the page directly cost less than our code-driven harnesses, in which the model writes code that emits the ledger rows. Checking is cheap as well. GPT-6 Luna’s extraction left more ledgers open than Gemini 3.7 Flash’s; the balance check found every one, and a correction call costing about a third of a cent closed each. The harness keeps only the work that the model should not own: the account IDs that join statements, the arithmetic, and the decision that a ledger is done.

The task

A task’s input is a set of PDFs from a long tail of banks: digital or scanned, a few pages or dozens, with several accounts, overlapping periods and repeated statements. Its output is one ledger per account and statement period: every line item with its date, full description, amount, direction and tags, and the opening and closing balances printed on the statement.

DateDescriptionAmountTags
2026-09-01ACH CREDIT PROCESSOR SETTLEMENT REF 0901+2,500.00payment_processor
2026-09-02ACH DEBIT EXAMPLE BANK LOAN PMT REF 0902−425.00bank_loan
2026-09-03ACH CREDIT PROCESSOR SETTLEMENT REF 0903+1,800.00payment_processor
2026-09-09ACH DEBIT EXAMPLE BANK LOAN PMT REF 0909−425.00bank_loan

Opening balance: $10,000.00 · Closing balance: $13,450.00

A bank statement carries its own check. A ledger closes when its opening balance plus its signed transactions equals the printed closing balance, within $0.05. The example closes: $10,000.00 + $2,500.00 − $425.00 + $1,800.00 − $425.00 = $13,450.00. The statement supplies the expected value and code runs the check, so extraction has an objective that needs no reference answer.

When a ledger does not close, the discrepancy is an error signal. Its size and direction constrain which rows can be wrong: a $1.16 gap is consistent with a $0.58 row recorded in the wrong direction. Reconciliation works from this signal.

Classification then tags each transaction, and related financing transactions are grouped into loan positions. Here, the two $425 payments to Example Bank form one position. Ledgers, tags and positions feed cash-flow and debt-service metrics computed in code.

The five evaluation tasks contain 4 to 12 files and about 140 to 670 transactions each. We selected them by hand to include difficult scans: image-only PDFs without a usable text layer.

The harness

If you do this kind of work with a coding agent, the alternative to a harness is to give the agent the PDFs and ask for ledgers. On a task with a dozen statements across several accounts, three things go wrong. The agent’s context grows with every page it reads. Nothing makes two statements agree on which account is which. Nothing stops the agent from calling a ledger done when the numbers are close. Jack is structured to prevent them. A task runs in eight steps:

  1. Route files. An independent classification step sends each file to the pipeline for its document type.
  2. Resolve accounts. One call reads every statement in the task together and returns the business, its owners and each account, with an integer ID. These IDs are the coordinate system for the rest of the task. Statement and extraction calls receive the list of valid IDs, and code drops records that use any other.
  3. Read printed balances. Each statement is read in parallel for its period and its printed opening and closing balances: the values the check needs.
  4. Remove duplicates. A text-only call compares periods, accounts and balances, and drops statements that another file already covers.
  5. Extract transactions. Each statement is extracted in parallel by a call that sees one statement and the account list. Context is bounded by the statement, not by the task.
  6. Close ledgers. Code checks every ledger. A ledger that does not close goes to correction.
  7. Classify cash flow and loans. Two calls run in parallel over grouped descriptions. Pattern rules in code add tags for checks, wires, peer-to-peer payments and fees.
  8. Group positions. One call groups merchant cash advance activity by funder. Code clusters other loan types by counterparty name.

Between calls, code sorts ledgers, computes running balances, corrects transaction years from the statement period and assigns transaction IDs. Application extraction and a tampering check run beside the sequence. Within a step, statements run in parallel; each step waits for the previous one. A task’s completion time follows its slowest statement at each step, and correction attempts for one ledger run in sequence.

Steps 1 to 4 produce no transactions, and on GPT-6 Luna they are the largest share of the cost. They are also what makes the check possible. Without account IDs, statements cannot be joined into one history per account. Without printed balances, there is nothing to close against. Without deduplication, a statement uploaded twice produces two ledgers that each close on their own, and a month of revenue is counted twice.

Each model call is maintained separately, with its own prompt, Pydantic response type, reasoning setting and model. A provider adapter gives every call one interface: the same request returns the same typed object from Gemini or from the OpenAI Responses API. The reasoning setting is an effort level on OpenAI and a thinking level on Gemini. Extraction, which transcribes rows, runs at low effort. Account resolution, printed balances, classification and correction run at high effort. Telemetry records every call with its step, model, token counts and cost. The chart below is built from that record.

Cost by step

Mean inference cost per task, broken down by step.

$0.290/ task
$0.068/ task
76.7% lesswith GPT-6 Luna

Extract transactions accounts for 68% of the net cost reduction.

Mean inference cost per task, with both models on the same dollar scale.

Step 5, extracting transactions, accounts for the largest saving. Its cost falls by 89% from Gemini 3.7 Flash to GPT-6 Luna, which is 68% of the total reduction. Most of this step’s cost is output, because the model writes every transaction record. GPT-6 Luna’s output price is $0.50 per million tokens, against a promotional $3.75 for Gemini 3.7 Flash that rises to $7.50 on January 1, 2027. GPT-6 Luna also reported fewer output tokens: about 19,400 per task, against 42,000.

The records are the same size from both providers, about 140 characters per transaction in compact JSON. Gemini 3.7 Flash returned indented JSON: whitespace made up 34% of the characters in its extraction responses, against 4% for GPT-6 Luna. It also reported about 6,600 thinking tokens per task for this step. If Gemini 3.7 Flash produced no thinking tokens here and its answer tokens fell by the same 34%, GPT-6 Luna would still cost 69% less per task.

Steps 1 to 4 make up 55% of GPT-6 Luna’s inference cost, against 26% for Gemini 3.7 Flash. Steps 1 to 3 send complete PDFs and return short answers, so most of their cost is input. For the same files, GPT-6 Luna reported about 5.5 times as many input tokens as Gemini 3.7 Flash. The providers count PDF tokens differently, so these counts do not measure the same quantity. At GPT-6 Luna’s input price, steps 1 to 4 still cost about half as much as on Gemini 3.7 Flash. As reading line items becomes cheaper, a larger share of the budget goes to interpreting the documents: which account a statement belongs to, what period it covers, and whether its activity is already included.

Closing ledgers

Reconciliation gives the model a bounded search problem: the current ledger as state, the source PDF as evidence, and the remaining discrepancy, with its direction, as feedback. The model proposes changes through a small edit language, defined as the Pydantic response type BalanceFix:

flip_indices
Reverse the debit or credit direction of specified rows. Their amounts remain unchanged.
remove_indices
Remove specified rows, such as duplicates or transactions assigned to the wrong ledger.
add_transactions
Supply missing transactions as typed records with a date, description, amount and direction from the statement.
give_up
Stop correction when the model cannot identify a supported fix, with an explanation of the unresolved discrepancy.

The model refers to rows by their one-based index. Code applies valid flips, removes rows in reverse index order so that the remaining indices stay valid, adds the new records, then sorts the ledger and recomputes its balances. The edit language cannot change a printed balance. It can change an existing row only by reversing its direction or removing it. A response can satisfy its schema and still propose a wrong edit, so code recomputes the balance after every edit.

Each attempt is a new call that receives the current ledger and the remaining discrepancy. Correction stops when the ledger closes, when the model sets give_up and the ledger is still open, or after three attempts. A ledger that stays open keeps its discrepancy and a reason, and is reported as not closed. The model revises its proposal during inference; the harness owns execution and the stopping condition.

The trace below is the $1.16 case, from the Gemini 3.7 Flash configuration.

Ledger reconciliation

A balance mismatch sends the source PDF and indexed ledger back for correction. One affected row is shown here.

Gemini 3.7 Flash

Input

{
  "transaction": {
    "index": 234,
    "description": "Intérêt versé",
    "type": "credit",
    "amount": 0.58
  },
  "statement_closing_balance": 21348.46,
  "computed_closing_balance": 21349.62,
  "discrepancy": 1.16
}

The full input includes the PDF and remaining ledger rows.

Output

{
  "flip_indices": [234],
  "remove_indices": [],
  "add_transactions": [],
  "give_up": false
}

Correction fields from the recorded response.

Applied by the harness

Code changes the $0.58 credit to a debit, lowering the computed balance by $1.16 to $21,348.46, the statement’s closing balance.

Input14,995Response90Reasoning9,353Call cost $0.0467
A small edit can require substantial reasoning: this call used 9,353 reasoning tokens to produce a 90-token response.

Usage covers the full call. The input is excerpted and reformatted; response and reasoning tokens are reported by Gemini.

Across the five tasks, GPT-6 Luna sent more ledgers to correction than Gemini 3.7 Flash: 8 against 1. Each of the 8 closed on its first correction call. The 8 calls cost $0.003 each on average; the single Gemini 3.7 Flash call cost $0.047. The table shows the result for every configuration.

ConfigurationLedgers checkedClosed after extractionClosed after correctionNot closed
Gemini 3 Flash332670
Gemini 3.5 Flash272601
Gemini 3.7 Flash272610
Gemini 3.8 Flash272700
Gemini 3.5 Flash Lite302163
GPT-6 Luna332580
GPT-6 Sol302811
GPT-6 Astra332832
GPT-5.6 Luna322480
GPT-5.6 Terra302532
GPT-5.6 Sol332850

Totals across the five tasks. A checked ledger has printed opening and closing balances. Ledgers without them cannot be checked: two for GPT-6 Astra, three for GPT-5.6 Luna, one for GPT-5.6 Terra and two for GPT-5.6 Sol. Ledger counts differ because each configuration resolves accounts and statements independently.

A closed ledger satisfies the balance equation. It does not prove that every row is correct: two errors with equal and opposite effects can still close a ledger.

Correction calls are dependent. Each attempt waits for the previous one, and a task waits for its slowest ledger. The median task is the same task for both configurations: GPT-6 Luna made three correction calls on it, and Gemini 3.7 Flash made none. These calls may account for part of GPT-6 Luna’s longer median time. Separating their effect from model speed requires timings for individual calls.

Classifying transactions

Bank descriptions repeat the same patterns with changing reference numbers and formatting noise. Code normalizes that noise into grouping keys and keeps the original records. The model receives one row per group: a representative description with its credit and debit counts and totals. It returns category assignments as lists of group IDs, and code expands each assignment to every transaction in the group. In the trace below, 667 transactions become 116 classification targets, and one group contains 305 credits.

Two classification calls run in parallel over the same groups. One assigns cash-flow categories, such as payment processor, internal transfer and owner transaction; a group can receive more than one. The other assigns a loan type, with a list of known funders in its prompt; a transaction receives at most one loan type.

Transaction classification

667 transactions, grouped into 116 classification targets. Two groups are shown here.

Gemini 3.7 Flash

Input

{
  "groups": [
    {
      "id": 67,
      "description": "ACH-HMS ACH [redacted] HRTLAND PMT SYS ACHDD",
      "credit_count": 1,
      "debit_count": 0
    },
    {
      "id": 78,
      "description": "ACH-TXNS/FEES [redacted] HRTLAND PMT SYS ACHDD",
      "credit_count": 305,
      "debit_count": 0
    }
  ]
}

Two groups from the input table, shown as JSON.

Output

{
  "internal_transfer": [],
  "owner_transaction": [],
  "payment_processor": [67, 78],
  "bank_fee": [1, 86, 88],
  "bank_interest": [],
  "reversal": [2, 34],
  "cash": [82, 83, 84]
}

Response IDs refer to transaction groups.

Applied by the harness

Code applies payment_processor to every member of groups 67 and 78, including all 305 transactions in group 78. The original rows keep their descriptions, dates, amounts and source links.

Input7,181Response112Reasoning1,306Call cost $0.0107
Despite “FEES” in the description, group 78 contains 305 credits. The model assigns it to payment_processor; a fee keyword alone would miss that context.

Usage covers the full call. The input is excerpted and reformatted; response and reasoning tokens are reported by Gemini.

Grouping turns classification into a fixed set of questions, each with a fixed set of answers. OpenAI announced the Decisions API at DevDay on September 29, 2026. It runs GPT-6 Luna on questions the caller defines, each with a fixed list of answers. We are testing it on step 7 in a development environment, with the same grouped input. The results in this article do not use it. We will measure its cost and latency through the same harness.

Document representation

The evaluation tasks include image-only scans. In our document reviews, GPT-6 Luna had more trouble extracting these statements than Gemini. A missed amount or a debit read as a credit leaves a ledger open, which adds a correction call before the ledger is ready for classification. On these tasks, every ledger that GPT-6 Luna’s extraction left open closed in the correction loop.

Document input is part of that comparison. Both OpenAI and Gemini describe PDF inputs that include page images alongside extracted text. Our adapter sends base64-encoded PDF bytes to the Responses API in an input_file item with detail: "high"; Gemini receives the same bytes at medium media resolution. The detail parameter sets the visual detail used for the page images, which matters for small print and dense transaction tables. It does not reconstruct a missing text layer.

Our traces record the calls and their results. They do not show which part of the input path explains the difference on scans: page rendering, visual tokenization, or the interaction between text and images.

Direct PDF input is one of several harness architectures we run against the same tasks. In the most code-driven of them, the model writes no records at all. A coding agent works against the PDF’s text layer, or a layer reconstructed with OCR, and writes code that emits every ledger row and every tag. The model writes the procedure: find the date column, join wrapped descriptions, read separate debit and credit columns. Every architecture’s ledgers go through the same balance check, so each one is held to the same accounting constraint.

These architectures exist to reduce what the model has to write. At GPT-6 Luna’s prices, direct extraction from the page cost less than the code-driven variants in our tests. We will report those results separately. The charts in this article measure direct PDF input.

Measurement notes

The charts compare 11 model configurations on the same five tasks in a local evaluation, with one run per task and configuration. Each configuration runs every step on one model; production assigns steps to different Gemini models. OpenAI fast mode, web research and Gemini model fallback were disabled. Fast mode shortens individual calls, but a task also includes dependent steps and correction passes, so its end-to-end effect has to be measured through the harness. A sixth task is excluded from every configuration because one of its runs was cancelled. With five tasks and one run each, the figures are observed values, not estimates with confidence intervals.

Cost is mean recorded inference spend per task. Time is median end-to-end completion under concurrent load. The Gemini configurations and GPT-6 Luna ran concurrently in one wave; the other OpenAI configurations ran in a second wave. Step costs are per-task means and sum to the full inference cost.

Ledger results come from the recorded calls. We recompute each balance check from the statement-metadata and extraction responses, and apply each recorded correction to the ledger in its prompt. Ledgers without printed opening and closing balances are not checked.

We calculate cost for each recorded call from the provider’s reported token usage and the model’s public API rates at the time of the evaluation, September 29, 2026. We use the published OpenAI pricing and Gemini pricing, including applicable promotional rates and context-length tiers. Input and output are priced separately; reported reasoning tokens are included in output usage. Recorded call costs, including correction passes, are summed for each task and then averaged. Gemini 3.7 and 3.8 Flash were at promotional rates that double on January 1, 2027; at the new rates, their costs in this article would double. These are inference costs before account credits, excluding infrastructure costs. Token counts follow each provider’s accounting.

Loading availability…