The ChatGPT co-inventor just shipped a model that refuses to talk. It's the AI I actually needed two years ago.
Originally published on LinkedIn.

On September 15, Diogo Almeida came out of two years of stealth with a company called TypeSafe AI, a model called Jev, and a $40 million seed round led by DCVC.
Almeida spent 2018 to 2024 at OpenAI on the InstructGPT and RLHF work that made ChatGPT possible. His launch post on X opened with the question that pushed him out the door: after co-inventing ChatGPT, why have superhuman chat models not led to AGI?
His answer is that the training method is the problem. RLHF teaches a model to please the human reading its answer. That makes it charming and, in his words, unreliable. So he built a different one.
What Jev is
Jev does not generate text. You give it unstructured input and a fixed set of possible answers, and it returns a typed decision with a calibrated probability attached. Almeida's one-line version: "Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out."
TypeSafe calls this a System One model, after the fast, automatic kind of thinking, and trains it with a method they call Reinforcement Learning for Calibrated Decisions. The target is not an answer a rater would prefer. It is a probability that is honest about how sure the model is.
Because the output is a choice from a schema rather than a sentence, the company says it cannot hallucinate or produce a type error. The Register's fair note on that: it is true by construction, since there is no natural language to get wrong.
The numbers, with the label they deserve
TypeSafe reports 70 to 500 milliseconds per decision against 3 to 329 seconds for a frontier chat model, input priced at $0.042 per million tokens, and output free. On one workflow of their own design they report 193.6 times faster and 444.6 times cheaper than a frontier chat model. Their Doom demo runs at ten decisions a second for about $7 an hour.
Those are the company's benchmarks, not independent ones. Remio's write-up makes the right points: the workflow was designed by TypeSafe, the comparison is against a large chat model rather than a small classifier, and the accuracy side of the ledger is not published yet. Early access is a waitlist. I'd read every multiple as a claim to test, not a fact to repeat.
I'm optimistic anyway. Here is why.
What it is for
Look at where the time and money actually go in an automated system. It is not the essay at the end. It is the thousands of small calls in the middle: is this email a reply or a bounce, does this record match that one, is this lead worth a call, close the ticket or escalate it. Every one of those is a decision with a short list of right answers, and today most people make a chat model write a paragraph to get to it.
I run the back office for two jiu-jitsu gyms on AI agents I built. Most of what they do all day is exactly that middle layer: read an inbound message, decide what it is, decide what happens next, and only then draft anything. The drafting is the cheap part. The deciding is where I spent two years building verification gates, because a chat model will tell you it is certain when it is guessing.
A model whose whole job is to return "reply, 0.91" or "bounce, 0.97" and mean it is the piece I kept wishing existed. A calibrated probability is the difference between an agent you supervise and an agent you trust with a threshold.
What I'll be watching
Three things decide whether this matters. Whether the calibration holds up when someone outside the company measures it. Whether the price stays where it is once the waitlist opens. And whether the accuracy on real workloads matches the speed, because a fast wrong answer is the most expensive kind.
If those hold, the interesting consequence is not that chat models get replaced. It is that they get reserved for the work only they can do, and the plumbing underneath gets a model of its own.
Almeida helped build the model that taught the world to talk to computers. Now he has built one that only answers with a decision and a number. For those of us who run software on these things all day, that might be the more useful invention.
What is the first decision in your stack you would hand to a model that cannot write a sentence?