Four models shipped in nineteen days in September, and the useful change for a construction business is not that they are cleverer. It is that they now search and act on their own, in a loop, without being handed the right document first. That moves the bottleneck. The limit on a project is no longer what the model can do. It is whether there is a record worth searching.
Which is a more awkward question than it sounds, because on most jobs the fullest account of what happened is sitting in a group chat on somebody’s phone.
What shipped, and the number that matters
OpenAI released GPT-6 Astra on 3 September, then GPT-6 Sol and GPT-6 Luna on 22 September. Anthropic released Claude Opus 5.5 the same day. The prices tell the story better than the benchmarks do.
| Model | Input | Output | What it is for |
|---|---|---|---|
| GPT-6 Astra | $10 | $50 | The hardest work, computer use, browsing |
| GPT-6 Sol | $2 | $10 | Multistep work, reviewing, analysing |
| GPT-6 Luna | $0.10 | $0.50 | High volume reading, extracting, summarising |
| Claude Opus 5.5 | $4 | $20 | Agentic work, computer use, long research |
Prices are per million tokens, from OpenAI’s model documentation and Anthropic’s release page. OpenAI halved the price of Sol and Luna against what they charged before. Anthropic priced Opus 5.5 twenty per cent below Opus 5 and cut the cost of reading cached tokens from fifty cents to twenty.
The number worth carrying is this one. On OpenAI’s own business workflow test, the cheap model, GPT-6 Sol, scored 33.2 per cent at 27 cents a task. Claude Opus 5 scored 26.9 per cent at 11.1 times that cost. Price and capability used to move together. In September they came apart, which means the routine work on a project no longer needs the expensive tier.
An agent is only as good as the record it can search
The capability that changed is worth being precise about. You give the model a question. It works out what to look up, runs the search itself, reads what comes back, and goes again if the answer is not there. GPT-6 Astra’s documentation lists web search, file search, computer use and the Model Context Protocol as tools it can call, alongside a tool search that lets it pick from its own toolbox rather than being handed a fixed list. Anthropic describe Opus 5.5 as a reliable researcher that found hard to locate sources in testing.
Point that at a construction project and the value is obvious. Nobody wants to read fourteen months of daily records to work out when a sequence changed. Everybody wants the answer, with the entries behind it attached.
The catch is what it needs in order to do that. It needs information it can actually reach: dated entries, named senders, original photographs at full resolution, messages with the time they were sent. Contemporaneous records, in other words, which is the same thing an adjudicator wants and for the same reason. A model is very good at finding the pattern in a record. It cannot recover a decision that was never written down.
So the question this month raises for a contractor is not which model to buy. It is whether the real account of your projects exists anywhere a machine, or an adjudicator, or your own commercial team in eighteen months, can search. On most jobs the honest answer is no, and the reason is covered in our piece on controlling WhatsApp groups: the fastest and most complete record of the job is on personal phones, held by people who may have moved on.
What this looks like on a live job
Four things that are genuinely different now, given a record that exists.
- Assembling a delay position. The hours in an extension of time claim go on finding the entries, not on the argument. An agent that can search a dated site diary, pull the weather, the labour returns and the messages around a given fortnight, and lay them out in order, takes days out of that and leaves the judgement where it belongs.
- Answering the question nobody can answer. When was that detail changed, who approved the substitution, which drawing revision was on site that week. These are search problems over your own record, and search is what got better.
- Finding what is missing. More useful than it sounds. Ask what is absent rather than what is present: which weeks have no record, which instructions were never confirmed, which approvals have no paper behind them. A gap found in month four is a fix. Found in year two it is a dispute.
- Handover and golden thread packs. Collating evidence against an ISO 19650 structure is exactly the tedious, high volume, checkable work the cheap tier is now good enough for.
The two numbers that should keep you honest
The launch pages do not lead with these, so they are worth stating plainly.
The first is 59.3 per cent. That is GPT-6 Astra on Agents’ Last Exam, which tests agents on real professional tasks across 55 sub-industries, against Claude Opus 5 at 55.5 per cent. It is a real capability and it also means around four attempts in ten came back wrong. On a construction project that is the difference between an agent that drafts and a person who signs. Anything heading for a valuation, a notice, a programme or a client is checked by somebody accountable, every time.
The second is not a number but a name. Among the alignment tests OpenAI published with Sol and Luna is one called broken search: the case where the tool fails and the model carries on as though it had not. A model that searches for you can be confidently wrong about what it found. That is precisely the failure mode you cannot afford when the output is going into a delay narrative, so every answer needs the source entry attached and openable. Not a summary. The entry.
It is also worth knowing that the benchmark tables cannot be compared across vendors. Both companies published on 22 September, both cite OSWorld 2.0, and the reported scores range from 60.5 to 81.8 per cent depending on whose page you read, because the test sets, effort settings and cost accounting all differ. Test on ten real tasks from your own week instead. That afternoon beats every comparison chart published this month.
Where the money actually goes
A year ago you sent a question and got an answer, so the price per token was roughly the price of the answer. An agent does not work that way. It plans, looks things up, calls a tool, checks itself and goes round again. The bill is the number of passes multiplied by the tokens each pass burns, and the headline rate is one term in that sum.
Two consequences for anyone budgeting this. Measure cost per completed task, including the retries and the runs you threw away, because that is most of the gap between a pilot that looks cheap and an invoice that is not. And fix the caching before you change the model: cached input on GPT-6 is discounted by ninety per cent, and OpenAI report that GitHub cut the share of prompt tokens needing fresh processing by more than half across billions of requests. Reusing the stable part of a prompt is a bigger saving than most model switches.
How we are using this in Construction Metric
Our position has not changed with the releases, and that is the point of building it this way. Construction Metric exists to turn the traffic a project already generates into a dated, attributed record: the message, the photograph at full resolution, the sender, the time. Capture happens without anyone doing a second job by hand, because capture that needs remembering does not survive a busy Thursday.
What the new models change is what you can then ask of it. The assistant answers questions against your own record rather than the open internet, and returns the entries behind the answer so somebody can check it. The WhatsApp capture is what puts the record there in the first place, and how it works sets out the whole route.
We route to whichever model is right for the job rather than to whatever launched most recently, and most of the work runs on the cheap tier because most of the work is reading and extracting. Being deliberate about that is why a site costs pence per month to run rather than pounds, and it is why a price cut at the cheap end is more interesting to us than a new flagship.
Questions and answers
Can AI actually read our site records and answer questions about the job?
Yes, and that is the part that improved most this month. The newer models decide what to look up, search, read the results and search again, rather than waiting for you to paste the right document in. The condition is that the records exist in a form something can search: dated entries, named senders, the original photographs. If the answer only exists in a WhatsApp group on a foreman’s phone, no model can reach it.
Is it safe to let an agent draft a delay narrative or a valuation?
Draft, yes. Submit, no. On Agents’ Last Exam, which tests agents on real professional tasks, the best score in OpenAI’s own comparison was 59.3 per cent, so roughly four attempts in ten came back wrong. The useful version is an agent that assembles the dated entries, the photographs and the messages behind a claim, and a quantity surveyor who checks it before it goes anywhere. The time saved is in the assembly, which is where the hours actually go.
Which model should a contractor be using?
Almost never the most expensive one. The cheap tier got genuinely capable in September: GPT-6 Sol costs 2 dollars per million input tokens against 10 for GPT-6 Astra, and beat Claude Opus 5 on OpenAI’s own business workflow test at around a ninth of the cost per task. Reading a site diary, extracting fields from a delivery ticket and answering a known question all belong on the cheap tier. Keep the expensive model for the judgement call at the end.
Do we have to give it access to our systems?
Not blanket access, and you should not. The connector standard everyone now uses is the Model Context Protocol, which Anthropic published in November 2024 and handed to a Linux Foundation fund in December 2025. It lets you expose specific things, a project folder, a particular register, rather than opening the whole estate. Decide in writing what an agent may read and what it may change, before you connect anything.
What does this cost to run across a live project?
Less than it did in July, and the figure to watch is cost per finished piece of work rather than the price per million tokens. A cheap model that goes round the loop twenty times can cost more than an expensive one that gets it right first time. Caching is the lever most people miss: cached input on GPT-6 is discounted by 90 per cent, and Opus 5.5 reads cached tokens at a fifth of what Opus 5 charged.
What is the one thing to get right first?
The record, not the model. Every one of these capabilities acts on information you already hold, so the projects that get value are the ones where the day to day traffic is already captured, dated and attributed. If that is not true on your jobs, fixing it is worth more than any model choice, and it keeps paying when the next release lands six weeks later.
Where to start
Not with the model. Take one job, and ask where the real record of it lives. If the answer involves a phone, a spreadsheet somebody maintains out of goodwill, and four people’s memories, then the model is not your constraint and no release will change that.
Fixing it is a smaller piece of work than it looks, and it is the part that keeps paying when the next generation ships in six weeks. We do this with UK contractors and consultants at AI Metric, and Construction Metric is the product that came out of doing it. If you want to talk it through against your own projects and your own contracts, get in touch and we will look at it properly.
