OpenAI Is Training Agents In Software. What Does Ironclad Say?
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: OpenAI Is Training Agents In Software. What Does Ironclad Say? on ThorstenMeyerAI.com

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI described training its GPT-6 Astra model inside hosted copies of Ironclad’s contract-management software, using 11 legal, commercial and procurement tasks. Astra met an average 55% of task criteria, while its estimated completion times were simulated—not measured customer savings—and the results do not establish that the workflows can run without human review.

OpenAI said on October 6 that it trained its GPT-6 Astra model on tasks inside hosted copies of Ironclad’s contract-management software, reporting that Astra met an average 55% of rubric criteria across 11 workflows. The project matters because OpenAI is asking other software companies to help train agents on professional tasks, but the reported results do not show that the system is ready to carry out consequential work without human checks.

OpenAI and Ironclad staff selected 11 legal, commercial and procurement tasks, including creating nondisclosure agreements, setting up approval processes and updating reusable contract clauses based on a requester’s jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes for each task. Each task was assessed against a rubric of 8 to 50 criteria, depending on its complexity; the reported score is the share of those criteria met, not the share of tasks completed successfully.

OpenAI reported that GPT-6 Astra met an average 55.0% of criteria, compared with 41.6% for GPT-5.6 Sol at its high setting. Astra’s estimated time per attempt was 19.2 minutes, compared with 37.0 minutes for Sol. An internal OpenAI model used during Astra’s development reached 63.7%. On one example task, Astra met about 94% of the criteria, but that single result does not describe its average performance.

OpenAI said Ironclad provided hosted copies of its product for model practice. It said synthetic tasks were built from publicly filed contracts in the SEC’s EDGAR database, filtered to remove personal information, and that it used no OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data. OpenAI also said its time figures are simulated estimates based on assumed processing and generation speeds, not measured customer time savings. The figures cover these 11 research tasks, not Ironclad workflows generally.

At a glance
reportWhen: Published October 6; follow-up and part…
The developmentOpenAI published details of a training project with Ironclad that tests frontier models on workflows inside a real business software product.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Contract Workflow Accuracy Matters

The reported average leaves a material gap between carrying out steps and satisfying the rules that govern a business process. A procurement workflow, for example, may need Finance approval above a spending threshold, a Security review for certain requests and Legal review for nonstandard terms. Missing one required approval can make a workflow unusable even if the agent completes other parts correctly. The rubric average does not say which criteria were missed on each task, so it cannot establish whether Astra’s errors were minor or affected essential controls.

OpenAI’s project also points to a possible change in how software agents are developed: vendors may provide environments and expertise so models can be tested on the difficult work their products support. That could make agents more capable inside specialised applications. It also puts focus on what vendors provide beyond their screens: business rules, data structures, audit records and controls that must remain reliable if an agent handles more of the interaction.

For businesses considering agents in contract, finance or customer-record systems, the figures are a research result rather than evidence of a net productivity gain. The reported estimate of about 19 minutes per Astra attempt cannot be compared with a person’s completion time as a measured saving, particularly when the model met only part of the rubric on average and review remains necessary.

Amazon

contract management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Ironclad Test Was Set Up

Ironclad is a contract-management software company, not an agent framework. OpenAI’s October 6 post described work conducted inside the company’s product and presented it as a way to train models to understand business rules, perform multi-step tasks in specialised software and check work against requirements. Some coverage reportedly interpreted “Ironclad” in the title as the name of a hardened agent framework; the source material says it refers to the company.

The selected tasks were designed by Ironclad staff and OpenAI employees who use the product. They covered work in legal, commercial and procurement settings, with task-specific criteria used to grade performance. OpenAI said the training tasks drew on publicly filed contracts, rather than confidential customer contracts. The project therefore provides a limited test of performance on selected workflows; it does not establish how the model performs across all customers, products or contract situations.

OpenAI’s post said the work was intended to show why human oversight and a full contracting platform remain important. Ironclad CTO Sunita Verma emphasized the need for agents to preserve the controls teams rely on. Those qualifications matter when interpreting the scores: passing a subset of criteria is not the same as producing a complete, approved workflow.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Scores Do Not Establish

The post does not provide the full task-by-task results or identify which criteria Astra missed across the 11 workflows. Without that breakdown, readers cannot tell whether failures involved low-impact details or mandatory approvals and other safeguards. The 55% average is not a pass rate, and the one task with about 94% of criteria met does not show how often similar performance occurred.

It is also unclear whether the models were tested on workflows or data beyond the described research tasks, how much human assistance was involved during attempts, or whether the results have been independently evaluated. OpenAI’s time estimates are simulated, so the post does not quantify actual customer time saved, review effort or the rate of errors that would matter in deployment.

The source material does not provide details about partner agreements, commercial terms, a rollout schedule or the number of software companies that may take part. OpenAI’s invitation to a small number of partners indicates a proposed next step, not a confirmed broad program or product launch.

Amazon

document review AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

OpenAI Seeks More Software Partners

OpenAI said it is inviting a small number of software companies to bring tasks current agents cannot reliably complete. It asked prospective partners to provide a concrete example of a failure, people with deep knowledge of the work, a secure test environment and data that can safely be used for research. The source material does not state which companies will participate or when further results will be released.

For any later evaluation, the most useful details will include task-level scores, the criteria missed, the amount of human intervention, and tests showing whether critical business rules were followed. Businesses considering agent use in their own systems will also need to decide what actions require approval and how exceptions are recorded. Until those details are available, the Ironclad work is best read as a reported training experiment and an invitation to collaborate, not proof of autonomous contract operations.

Amazon

contract clause management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Ironclad in this announcement?

Ironclad is a contract-management software company. OpenAI’s post describes training models on tasks inside hosted copies of its product; Ironclad is not the name of a new agent framework in this account.

What does Astra’s 55% score mean?

It is the average share of rubric criteria met across the 11 tasks. It is not the percentage of tasks completed successfully or a claim that 55% of customer workflows can be automated.

Did OpenAI show that Astra saves customers time?

No measured customer time savings were reported. OpenAI described the task durations as simulated estimates based on assumed processing and generation speeds.

Can Astra handle contract workflows without human review?

The reported results do not establish that. The average criteria score was 55%, and the source material says human oversight remains relevant when an agent may lose track of a business rule or required control.

What happens after the Ironclad project?

OpenAI said it is inviting a small number of software companies to contribute difficult tasks, expert input, secure test environments and research-safe data. It has not provided a partner list or schedule in the source material.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Radar That Never Blinks: What SAR Actually Does — for Companies, Institutions, and Governments

Explore how Synthetic Aperture Radar (SAR) works, its applications for companies, governments, and institutions, and why it’s transforming earth monitoring in 2026.

The Delegation Ladder: The Four Agentic Loops, and What Each One Lets You Stop Doing

An analysis of the four agentic loops in AI design, detailing what each allows you to stop doing and how they impact AI workflows.

The citation. Why generative engine optimization rewards the same brand on the least stable ground.

Analysis of how generative engine optimization (GEO) favors established brands, with decay and instability challenges for publishers and small content creators.

Sovereignty Is a Pipe, Not a Passport

Mistral’s European AI firm claims sovereignty through on-premise hosting, but reliance on US cloud infrastructure exposes jurisdictional vulnerabilities under CLOUD Act.