Model ML Uses GPT-5.6 Sol to Cut Finance Workflow Time From an Hour to Five Minutes
AI

Model ML Uses GPT-5.6 Sol to Cut Finance Workflow Time From an Hour to Five Minutes

Model ML's finance agent now routes most work through GPT-5.6 Sol, producing editable, traceable PowerPoint decks and Excel workbooks with fewer tokens than Fable 5 per deck and fewer tokens than Opus 5 per workbook, alongside a higher rate of review-ready output.

5 min read
Back to News

Model ML, a startup building AI agents for finance professionals, switched its core workflows to GPT-5.6 Sol on August 10, 2026. The move followed internal benchmarks showing the model produced more review-ready outputs at lower token cost than the alternatives the company had been running.

The practical result: a bespoke tearsheet that previously took an analyst about an hour to assemble now takes about five minutes.

TL;DR
  • -Model ML's finance agent now routes core work through GPT-5.6 Sol, replacing earlier reliance on Opus 4.8 for some workflows and outperforming Opus 5 and Fable 5 on PowerPoint deliverability and Excel token efficiency.
  • -GPT-5.6 Sol completed 100% of PowerPoint test cases versus 76% for Opus 5, hit a 43.3% professional-readiness rate versus 26.7%, and used 36% fewer tokens per Excel workbook than Opus 5 — narrowing the gap between model output and a client-ready file.
  • -Finance and technology leaders evaluating AI for deal-team or investment workflows should test professional-readiness rates and token costs against their own deliverable types before committing to a model or vendor stack.

What Model ML Actually Does

Model ML handles the part of financial analysis that is easy to underestimate. Getting numbers into a model is one problem. Getting a finished, editable, traceable PowerPoint deck or Excel workbook in front of a client or investment committee is another. That last mile — reconciling sources, building formulas, applying formatting, linking every claim back to evidence — is where most of the manual time goes.

Model ML's agents carry an assignment from an initial brief through research, calculations, and a finished file. A core agent plans the work, selects tools, reconciles evidence, and routes each step to the model best suited to it — which the company says is often GPT-5.6 Sol. The company's own document tooling generates native PowerPoint and Excel files with traceable sources, not flat images or locked exports.

Co-founder and CEO Chaz Englander put the improvement directly: "Earlier models could do the work of an analyst, but the user would have to clearly break down the task, specifically what it wanted the output to look like. With GPT-5.6 Sol, we're finding that the agent gets far closer to the final output."

The product is surface-agnostic: a finance professional can start an assignment in email or the Model ML app and continue it in Microsoft Office plug-ins without re-explaining the task.

What the Numbers Show

Model ML runs its own evaluation called Composite, which follows a finance brief from research through a finished deliverable, then scores the numbers, sources, formulas, structure, and visual quality. The benchmarks compared GPT-5.6 Sol against Opus 5, Fable 5, Opus 4.8, GPT-5.6 Terra, and GPT-5.5.

On PowerPoint, the results were decisive on the metrics that matter most to a finance team. GPT-5.6 Sol completed the PowerPoint workflow in 100% of test cases, compared with 76% for Opus 5. It cleared Model ML's professional-readiness gate — meaning the output was ready for substantive review — in 43.3% of cases versus 26.7% for Opus 5, a 16.6 percentage-point lead.

It also led Opus 5 on design coherence, with a Consistency score of 97.8% versus 93.3%. On brief adherence, Sol edged Opus 5 by 0.9 percentage points — 78.8% versus 77.9% — a narrow margin. Sol trailed Opus 5 on visual quality and layout.

The token-efficiency picture is more nuanced. Sol used 1.10 million tokens per deck in the PowerPoint workflow. That is 21% fewer tokens per deck than Fable 5 at 1.40 million, but more than Opus 5 at 953,000 and GPT-5.6 Terra at 720,000. Teams optimizing purely for PowerPoint token cost would look at Terra. Teams prioritizing completed, review-ready decks will find Sol's deliverability rate is the stronger argument.

On Excel, Sol's token advantage over Opus 5 is sharper. GPT-5.6 Sol used 36% fewer tokens per workbook than Opus 5 — 2.44 million versus 3.83 million — and completed workbooks 0.5 minutes per workbook faster than Opus 5 in Model ML's Composite benchmark. The tradeoff: Sol produced fully correct models in 50% of Excel test cases, while Opus 5, Fable 5, and GPT-5.5 each hit 60%.

For key output accuracy, Sol and GPT-5.5 tied at 83.3%.

Model ML used these results to expand GPT-5.6 Sol into production, including some workflows previously handled by Opus 4.8.

How Sol Reaches Those Results

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 family. It costs $5 per million input tokens and $30 per million output tokens through the API and carries a 1,050,000-token context window.

On Agents' Last Exam, a benchmark of long-running professional workflows across 55 fields, Sol scored 53.6 — 13.1 points above Claude Fable 5, with Sol running at max reasoning versus Fable 5 at adaptive reasoning.

The efficiency gains are partly structural. OpenAI trained GPT-5.6 to take a more direct path through a task, optimizing for both task success and efficiency during training. Infrastructure improvements — including kernel optimization, speculative decoding, and smarter context management — compound on top of that. Kernel work reduced end-to-end serving costs by 20%; speculative decoding improvements increased token-generation efficiency by more than 15%.

To the extent those infrastructure gains carry through to API usage, they would contribute to lower per-assignment token costs for workloads like Model ML's. Fewer tokens per deliverable means lower API cost per assignment. At one global asset manager Model ML serves, the tearsheet workflow that took an analyst an hour now runs in about five minutes.

In another workflow, Model ML agents processed virtual data rooms with more than 100,000 rows and hundreds of files in a single pass.

What Finance and Technology Leaders Should Evaluate

Model ML's results are specific to its own agent harness, its document tooling, and its finance-workflow evaluation methodology. The Composite benchmark is Model ML's internal measure, not an independent industry standard. The numbers tell you what Sol did inside one carefully designed system — not what it will do inside yours.

That framing shapes the right question. If your team is producing client-ready financial deliverables at volume — investment memos, tearsheets, Excel models — the relevant metric is not raw benchmark accuracy. It is professional-readiness rate: the share of outputs that can move directly into review without being rebuilt.

Before changing a production stack, test that metric against your own deliverable types and your own source material. Model routing decisions — which model handles which step — affect both quality and cost, and the right balance shifts depending on the workflow. Sol's deliverability advantage on PowerPoint came with a token cost premium over Terra and Opus 5. Sol's Excel token efficiency came with a lower rate of fully correct models than Opus 5.

The infrastructure layer also matters. Model ML routes inference through OpenAI's API. If your organization has data-residency requirements, client confidentiality constraints, or internal policies on what financial data can reach an external model endpoint, those controls should be verified before workflow changes are made. The five-minute tearsheet is a real outcome.

Whether your governance framework treats the underlying data's path to that endpoint as a risk question worth answering is worth confirming before you commit.

Sources and supporting resources
Next
Hoa Sen University Embeds Odoo Enterprise into Its Business Curriculum

Get ERP, Cloud, Data, and AI Updates

News, insights, and practical guidance across ERP, Cloud, Data, AI, digital transformation, and technology projects.

No spam. Unsubscribe anytime.