What it costs

Two plans, and nothing metered

A test set built from your own data, held by us, re-run on a schedule and on every change, and a record your customer's risk team can read without asking you first. Unlimited cases, runs, deployments and readers in both plans: if we charged by the run, the cheapest quarter would be the one where we looked least.

Verification

US$1,500per month

Bought so you are not blindsided.

  • The test set built from your own corpus, confirmed once by your expert, held and versioned by us
  • Runs on any cadence, and within the hour of a change we detect
  • Regression detection with the fingerprint diff that says whether it was you or the provider
  • Every run sealed, hash-chained and timestamped

Assurance

from US$4,500per month

Bought because your customer asked.

  • Everything in Verification
  • The evidence pack, mapped to the clauses your customer's risk team cites
  • The portal your customer reads directly, under a standing disclosure arrangement rather than per report
  • A contracted cadence, so a gap is a breach rather than a quiet quarter
  • Candidate models scored against your confirmed cases before you switch

One tells you. The other tells your customer. Which is the only difference worth charging three times as much for, and the reason the second one exists at all.

The date

29 models retire on 23 October 2026

OpenAI publishes the list, and 24 of the 29 have no published replacement at all - so for most of them the migration is not a version bump, it is a choice you have to justify afterwards.

On the list, and probably in your config

  • gpt-4
  • gpt-4-turbo
  • gpt-4o-2024-05-13
  • gpt-4.1-nano
  • o1
  • o3-mini
  • o4-mini

Read from the provider's own deprecation page and reconciled nightly, the same way every date on this site is. Every retirement we track.

This is a cadence, not an event

took 4 more. Two of the four were the legacy fine-tuning bases. There are 18 distinct retirement dates across the four providers we track between 10 September 2026 and April 2027.

A forced migration every two to three weeks, on somebody else's roadmap, in perpetuity. Missing this one is not the end of it - it is the largest wave, not the last.

What we do before the date

  1. 01

    Mine a test set from your own documents, tickets and past decisions, and put it in front of one of your experts to confirm. Their side is under four hours.

  2. 02

    Baseline the system you run today, against itself, so we know which of your cases are inherently unstable before we blame a model for them.

  3. 03

    Score the candidates you are considering against that set, before you switch, and tell you which cases each one breaks.

  4. 04

    Keep running it afterwards, on a cadence, sealed - which is the part your customer's risk team is actually asking about.

If we look at your corpus and cannot get a set confirmed before , we will say so before you pay us anything. A date on a page is worth nothing if it is only ever used to hurry you.

The comparable

Whether this is expensive depends entirely on what you put it next to

So here is the ladder, stated rather than left to you. We are not on the first rung, and a page that let you assume we were would be selling you the wrong thing.

  1. An eval platform

    $0 to $249 a month, plus metering

    Excellent instruments, and one of them ships known-payload replay today. Priced as software, because that is what it is: seats are free and volume is billed, and somebody at your company still has to run it.

    Bought by your engineer, for your engineer.

  2. A compliance platform

    US$13,000 to US$26,000 a year

    Roughly what publicly observed contract data puts a mid-market subscription at. It inventories your systems, holds the policy, and records that the model version changed.

    Bought so procurement stops asking. It never measures whether an answer is right.

  3. An assessment engagement

    US$13,000 to US$60,000, once

    A model-validation contractor runs about US$14,000 a month at the London median day rate. An ISO 42001 first year lands in the same range. Both produce a report, and the fourth one costs what the first did.

    Bought because somebody outside asked. Genuinely independent, and it measures your process rather than your answers.

  4. This

    US$18,000 a year

    An operated service on the third rung's ladder, at a fraction of one engagement, running continuously rather than once. The marginal cost of the second run is near zero, which is why the cadence is not the thing you are billed for.

    Bought because somebody outside asked, and it measures the answers.

The honest summary is that we are the cheap option on an expensive ladder, and an absurd one on a cheap ladder. If what you need is a faster way for your own team to measure your own system, the first rung is genuinely better than us and costs a fiftieth as much. This is for the case where your own measurement is the thing that will not be accepted.

Before you spend anything

There is no certificate at the end of this

We are not accredited and we are not an auditor. You are the one paying us, so the issuer-pays conflict applies to us exactly as it does to everyone else who sells assurance. And a fixed set of cases is a claim about those cases: it measures what your expert confirmed, not everything your system does.

The whole of what we do not claim

The next step

Tell us it is worth a conversation

Not a waitlist. We are asking whether the number above is roughly right, because we would rather find out now than after building the wrong thing. Nothing here commits you to anything, and we will not add you to a mailing list.

What would you expect this to cost?

The page says US$1,500 a month. Tell us if that is wrong, and which way.

What you sell, who you sell it to, and which models you are on is the most useful thing you can put here.

Someone will get back to you within 24 hours.