AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: OpenAI Trains Agents In Software: A Practical Read Of Ironclad’s Terms on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI described training and testing a frontier model on contract-management workflows in Ironclad software, using hosted product environments and synthetic tasks based on public SEC filings. GPT-6 Astra met an average 55% of task rubric criteria, while its reported time estimate was simulated—not a measured customer time saving.

OpenAI said on October 6 that it trained and evaluated its GPT-6 Astra model on multi-step tasks inside contract-management software maker Ironclad’s product, reporting an average of 55% of task criteria met. The work offers an early look at AI models learning business workflows in specialized software, but the results do not establish that the system is ready to handle contract work without human review.

OpenAI said Ironclad staff and OpenAI employees selected 11 tasks spanning legal, commercial and procurement work. Examples included setting up nondisclosure agreements, creating procurement approval processes and updating a reusable contract clause based on a requester’s jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes on each task. Depending on complexity, the evaluation rubric contained 8 to 50 criteria per task.

Ironclad provided hosted copies of its software for the models to use. OpenAI said it generated synthetic training tasks from contracts publicly filed in the U.S. Securities and Exchange Commission’s EDGAR database, with personal information filtered out. It also said it used no OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.

OpenAI reported that GPT-6 Astra met an average 55.0% of rubric criteria, compared with 41.6% for GPT-5.6 Sol in a high-compute setting. An internal model used during Astra’s development reached 63.7%. On one showcase task, Astra met about 94% of criteria. These are shares of evaluation requirements satisfied, not percentages of tasks completed successfully. OpenAI also gave estimated times of 19.2 minutes for Astra and 37 minutes for GPT-5.6 Sol, but said those figures were simulated estimates based on assumed processing and generation speeds, not customer time measurements.

At a glance
reportWhen: Published October 6; further partner wo…
The developmentOpenAI published details of a collaboration with contract-management company Ironclad to train and evaluate an AI model on multi-step workflows in Ironclad’s software.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Contract Workflow Accuracy Matters

The results matter because contract and procurement workflows are judged by whether they preserve all required rules and approvals, not by how many steps an agent gets mostly right. A process may require Finance approval above a spending threshold, Security review for certain requests and Legal review for nonstandard terms. Missing one control can send a purchase through the wrong route, even if the system handled other requirements correctly.

OpenAI’s reported average of 55% of criteria met therefore should not be read as a readiness score for unsupervised use. The evaluation points to progress on a difficult software task, while leaving a wide gap between partial rubric performance and dependable execution. Companies considering agents for consequential work will need to know which requirements were missed, how errors are caught and who remains accountable for approval.

The collaboration also positions software vendors as potential training partners for AI companies. A model that can follow workflows inside a product may make that product more useful. At the same time, the software provider’s lasting value may depend increasingly on its underlying business rules, records, audit trail and controls—not just the screens users operate.

Amazon

contract management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Ironclad Became the Test Environment

OpenAI’s October 6 post described two developments, according to the supplied source material: a release of 722 mathematics manuscripts and a separate account of its work with Ironclad. The Ironclad announcement received less attention, but it describes a different model-development approach: training and evaluating an AI system inside a specialized commercial product, rather than only on general-purpose prompts or a simplified demonstration interface.

OpenAI framed the goal as teaching models to understand business rules, carry out multi-step work in specialized software and check completed work against the original requirements. Its Ironclad evaluation used 11 selected research tasks, not a broad sample of every workflow customers run on the platform. OpenAI’s post also said it is seeking a small number of software-company partners with a concrete example of a task agents cannot reliably complete, experts familiar with the work, a secure test environment and data suitable for research.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Evaluation Leaves Open

The published figures do not show how Astra performed across a larger or independently selected set of Ironclad workflows, nor do they identify the specific criteria missed on each task. The source material describes an evaluation of 11 research tasks; it does not establish how results would translate to live customer work, different contract types or unusual cases.

OpenAI’s simulated time estimates do not measure end-to-end customer time, including review and correction. The available information also does not set out a deployment threshold, error rate for individual high-consequence controls, or a plan for assigning responsibility when an agent makes a mistake. OpenAI and Ironclad have not, in the supplied material, announced a general customer release of this capability.

Amazon

procurement approval workflow software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Next Test Is Reliable Deployment

OpenAI said it is inviting a small number of software companies to work on tasks current agents cannot reliably complete. The proposed partners would supply a specific failure case, knowledgeable staff, a secure environment and research-appropriate data. The supplied material does not name additional partners or give a timetable for the next stage.

For businesses evaluating agents in contract or procurement systems, the practical next step is to ask vendors for task-level results: which criteria failed, how rules are checked, what human approval is required and whether performance has been measured in customer operations. Until those details and broader testing are available, the Ironclad results are evidence of a research direction—not proof of reliable, unsupervised contract automation.

Amazon

NDAs contract templates

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did OpenAI and Ironclad test?

They tested models on 11 legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software. Examples included setting up nondisclosure agreements and building approval processes.

Does a 55% score mean Astra completed 55% of tasks?

No. OpenAI reported the average share of rubric criteria met across the tasks. It is not a task-completion rate, and the published figure does not show that the model completed more than half of the workflows correctly.

Were the time savings measured with customers?

No. OpenAI described the figures—19.2 minutes for Astra and 37 minutes for GPT-5.6 Sol—as simulated estimates based on assumed processing and generation speeds, not measured customer time savings.

What data did OpenAI say it used?

OpenAI said it created synthetic tasks from publicly filed contracts in the SEC’s EDGAR database and filtered out personal information. It said it did not use OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.

Is the system ready to run contract workflows without oversight?

The reported results do not establish that. Astra met an average 55% of rubric criteria, and OpenAI’s own discussion said human oversight remains relevant when agents may lose track of business rules. The supplied material gives no general customer-release plan.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Friendship Ended With Deno, Now Node Is My Best Friend

A developer reports moving a SvelteKit project and static site generator from Deno to Node.js, citing fewer migration changes and 15% faster builds.

Software Testing Tools: A Prime Big Deal Days Guide

Compare software testing tools by purpose, cost, and upkeep, then build a dependable test setup that fits your team.

IBM Expands AI And Quantum Research With IIT Bombay And IISc

IBM is extending research partnerships with IIT Bombay and IISc, spanning Indic-language AI, agentic systems, energy models and quantum-HPC research.

AI Automation Software For Small Businesses: A Prime Big Deal Days Guide

A small-business guide to automation costs and workflow choices, framed around Prime Big Deal Days. The supplied material does not confirm specific event discounts.