What it costs
Two plans, and nothing metered
A test set built from your own data, held by us, re-run on a schedule and on every change, and a record your customer's risk team can read without asking you first. Unlimited cases, runs, deployments and readers in both plans: if we charged by the run, the cheapest quarter would be the one where we looked least.
Verification
US$1,500per month
Bought so you are not blindsided.
- The test set built from your own corpus, confirmed once by your expert, held and versioned by us
- Runs on any cadence, and within the hour of a change we detect
- Regression detection with the fingerprint diff that says whether it was you or the provider
- Every run sealed, hash-chained and timestamped
Assurance
from US$4,500per month
Bought because your customer asked.
- Everything in Verification
- The evidence pack, mapped to the clauses your customer's risk team cites
- The portal your customer reads directly, under a standing disclosure arrangement rather than per report
- A contracted cadence, so a gap is a breach rather than a quiet quarter
- Candidate models scored against your confirmed cases before you switch
One tells you. The other tells your customer. Which is the only difference worth charging three times as much for, and the reason the second one exists at all.
The date
29 models retire on 23 October 2026
OpenAI publishes the list, and 24 of the 29 have no published replacement at all - so for most of them the migration is not a version bump, it is a choice you have to justify afterwards.
On the list, and probably in your config
- gpt-4
- gpt-4-turbo
- gpt-4o-2024-05-13
- gpt-4.1-nano
- o1
- o3-mini
- o4-mini
Read from the provider's own deprecation page and reconciled nightly, the same way every date on this site is. Every retirement we track.
This is a cadence, not an event
took 4 more. Two of the four were the legacy fine-tuning bases. There are 18 distinct retirement dates across the four providers we track between 10 September 2026 and April 2027.
A forced migration every two to three weeks, on somebody else's roadmap, in perpetuity. Missing this one is not the end of it - it is the largest wave, not the last.
What we do before the date
- 01
Mine a test set from your own documents, tickets and past decisions, and put it in front of one of your experts to confirm. Their side is under four hours.
- 02
Baseline the system you run today, against itself, so we know which of your cases are inherently unstable before we blame a model for them.
- 03
Score the candidates you are considering against that set, before you switch, and tell you which cases each one breaks.
- 04
Keep running it afterwards, on a cadence, sealed - which is the part your customer's risk team is actually asking about.
If we look at your corpus and cannot get a set confirmed before , we will say so before you pay us anything. A date on a page is worth nothing if it is only ever used to hurry you.
The comparable
Whether this is expensive depends entirely on what you put it next to
So here is the ladder, stated rather than left to you. We are not on the first rung, and a page that let you assume we were would be selling you the wrong thing.
An eval platform
$0 to $249 a month, plus metering
Excellent instruments, and one of them ships known-payload replay today. Priced as software, because that is what it is: seats are free and volume is billed, and somebody at your company still has to run it.
Bought by your engineer, for your engineer.
A compliance platform
US$13,000 to US$26,000 a year
Roughly what publicly observed contract data puts a mid-market subscription at. It inventories your systems, holds the policy, and records that the model version changed.
Bought so procurement stops asking. It never measures whether an answer is right.
An assessment engagement
US$13,000 to US$60,000, once
A model-validation contractor runs about US$14,000 a month at the London median day rate. An ISO 42001 first year lands in the same range. Both produce a report, and the fourth one costs what the first did.
Bought because somebody outside asked. Genuinely independent, and it measures your process rather than your answers.
This
US$18,000 a year
An operated service on the third rung's ladder, at a fraction of one engagement, running continuously rather than once. The marginal cost of the second run is near zero, which is why the cadence is not the thing you are billed for.
Bought because somebody outside asked, and it measures the answers.
The honest summary is that we are the cheap option on an expensive ladder, and an absurd one on a cheap ladder. If what you need is a faster way for your own team to measure your own system, the first rung is genuinely better than us and costs a fiftieth as much. This is for the case where your own measurement is the thing that will not be accepted.
Before you spend anything
There is no certificate at the end of this
We are not accredited and we are not an auditor. You are the one paying us, so the issuer-pays conflict applies to us exactly as it does to everyone else who sells assurance. And a fixed set of cases is a claim about those cases: it measures what your expert confirmed, not everything your system does.
The next step
Tell us it is worth a conversation
Not a waitlist. We are asking whether the number above is roughly right, because we would rather find out now than after building the wrong thing. Nothing here commits you to anything, and we will not add you to a mailing list.