🔍 Read the full analysis: OpenAI Trains Agents In Software: A Practical Read Of Ironclad’s Terms on ThorstenMeyerAI.com
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
OpenAI described training and testing a frontier model on contract-management workflows in Ironclad software, using hosted product environments and synthetic tasks based on public SEC filings. GPT-6 Astra met an average 55% of task rubric criteria, while its reported time estimate was simulated—not a measured customer time saving.
OpenAI said on October 6 that it trained and evaluated its GPT-6 Astra model on multi-step tasks inside contract-management software maker Ironclad’s product, reporting an average of 55% of task criteria met. The work offers an early look at AI models learning business workflows in specialized software, but the results do not establish that the system is ready to handle contract work without human review.
OpenAI said Ironclad staff and OpenAI employees selected 11 tasks spanning legal, commercial and procurement work. Examples included setting up nondisclosure agreements, creating procurement approval processes and updating a reusable contract clause based on a requester’s jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes on each task. Depending on complexity, the evaluation rubric contained 8 to 50 criteria per task.
Ironclad provided hosted copies of its software for the models to use. OpenAI said it generated synthetic training tasks from contracts publicly filed in the U.S. Securities and Exchange Commission’s EDGAR database, with personal information filtered out. It also said it used no OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.
OpenAI reported that GPT-6 Astra met an average 55.0% of rubric criteria, compared with 41.6% for GPT-5.6 Sol in a high-compute setting. An internal model used during Astra’s development reached 63.7%. On one showcase task, Astra met about 94% of criteria. These are shares of evaluation requirements satisfied, not percentages of tasks completed successfully. OpenAI also gave estimated times of 19.2 minutes for Astra and 37 minutes for GPT-5.6 Sol, but said those figures were simulated estimates based on assumed processing and generation speeds, not customer time measurements.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Why Contract Workflow Accuracy Matters
The results matter because contract and procurement workflows are judged by whether they preserve all required rules and approvals, not by how many steps an agent gets mostly right. A process may require Finance approval above a spending threshold, Security review for certain requests and Legal review for nonstandard terms. Missing one control can send a purchase through the wrong route, even if the system handled other requirements correctly.
OpenAI’s reported average of 55% of criteria met therefore should not be read as a readiness score for unsupervised use. The evaluation points to progress on a difficult software task, while leaving a wide gap between partial rubric performance and dependable execution. Companies considering agents for consequential work will need to know which requirements were missed, how errors are caught and who remains accountable for approval.
The collaboration also positions software vendors as potential training partners for AI companies. A model that can follow workflows inside a product may make that product more useful. At the same time, the software provider’s lasting value may depend increasingly on its underlying business rules, records, audit trail and controls—not just the screens users operate.
As an affiliate, we earn on qualifying purchases.
How Ironclad Became the Test Environment
OpenAI’s October 6 post described two developments, according to the supplied source material: a release of 722 mathematics manuscripts and a separate account of its work with Ironclad. The Ironclad announcement received less attention, but it describes a different model-development approach: training and evaluating an AI system inside a specialized commercial product, rather than only on general-purpose prompts or a simplified demonstration interface.
OpenAI framed the goal as teaching models to understand business rules, carry out multi-step work in specialized software and check completed work against the original requirements. Its Ironclad evaluation used 11 selected research tasks, not a broad sample of every workflow customers run on the platform. OpenAI’s post also said it is seeking a small number of software-company partners with a concrete example of a task agents cannot reliably complete, experts familiar with the work, a secure test environment and data suitable for research.
AI-powered legal document automation
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Evaluation Leaves Open
The published figures do not show how Astra performed across a larger or independently selected set of Ironclad workflows, nor do they identify the specific criteria missed on each task. The source material describes an evaluation of 11 research tasks; it does not establish how results would translate to live customer work, different contract types or unusual cases.
OpenAI’s simulated time estimates do not measure end-to-end customer time, including review and correction. The available information also does not set out a deployment threshold, error rate for individual high-consequence controls, or a plan for assigning responsibility when an agent makes a mistake. OpenAI and Ironclad have not, in the supplied material, announced a general customer release of this capability.
procurement approval workflow software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Next Test Is Reliable Deployment
OpenAI said it is inviting a small number of software companies to work on tasks current agents cannot reliably complete. The proposed partners would supply a specific failure case, knowledgeable staff, a secure environment and research-appropriate data. The supplied material does not name additional partners or give a timetable for the next stage.
For businesses evaluating agents in contract or procurement systems, the practical next step is to ask vendors for task-level results: which criteria failed, how rules are checked, what human approval is required and whether performance has been measured in customer operations. Until those details and broader testing are available, the Ironclad results are evidence of a research direction—not proof of reliable, unsupervised contract automation.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did OpenAI and Ironclad test?
They tested models on 11 legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software. Examples included setting up nondisclosure agreements and building approval processes.
Does a 55% score mean Astra completed 55% of tasks?
No. OpenAI reported the average share of rubric criteria met across the tasks. It is not a task-completion rate, and the published figure does not show that the model completed more than half of the workflows correctly.
Were the time savings measured with customers?
No. OpenAI described the figures—19.2 minutes for Astra and 37 minutes for GPT-5.6 Sol—as simulated estimates based on assumed processing and generation speeds, not measured customer time savings.
What data did OpenAI say it used?
OpenAI said it created synthetic tasks from publicly filed contracts in the SEC’s EDGAR database and filtered out personal information. It said it did not use OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.
Is the system ready to run contract workflows without oversight?
The reported results do not establish that. Astra met an average 55% of rubric criteria, and OpenAI’s own discussion said human oversight remains relevant when agents may lose track of business rules. The supplied material gives no general customer-release plan.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
