The model can't hallucinate. Your inputs still can.
A new class of AI model guarantees every answer is well-formed. That solves one problem and makes another harder to see.
Last week an AI lab called TypeSafe launched Jev, the first of what it calls System One models. Its founder says he helped develop the instruction-following research behind ChatGPT. The launch post is worth reading in full, for the claims and for how carefully it qualifies them.
The idea is simple. Chat models produce text, and text is a poor interface for software. It has to be parsed and checked. It can go off the rails, return the wrong type or invent a tool call. Jev doesn't generate text at all. You define the possible answers in advance, and it returns a structured decision with a probability attached, much faster and much more cheaply than a conventional model. A malformed answer is ruled out by design. The intended uses are the small decisions inside ordinary software: classify this, route that, score this, extract that. TypeSafe calls them "smart if-statements".
I think this direction is right, and it's where a lot of real business automation is heading. It's also why the problem I work on is about to get more expensive.
A valid answer isn't a correct answer
Type safety guarantees the shape of an answer. It doesn't guarantee the answer is right.
Take a support decision: approve the refund, offer a replacement, or escalate. Every one of those is a valid output, and a type-safe model will always return one of them. What it can't do is account for something it was never told. If the customer has already been promised two replacements that never arrived, and that history wasn't in the information the model received, it can confidently offer a third. The answer will be well-formed and the confidence score may be high. It will still be wrong.
A confidence score tells you how sure the model is, given what it saw. It can't tell you what it didn't see.
Why this makes the input problem worse, not better
With a chat model, a person usually reads the output. A bland or wrong paragraph has a decent chance of being caught before it does damage.
A structured decision inside software isn't read by anyone. It's a branch in the code, and removing the person from that step is the whole point. So an error that used to show up as an odd-sounding paragraph now shows up as a customer routed to the wrong queue, or a claim approved automatically. All that's left behind is a log entry with a confident number next to it.
The better the output guarantees get, the less anyone checks the output. That puts all the remaining risk on what the model was given to decide from.
The lab's own footnotes make the same point
Credit to TypeSafe: their post includes the kind of caveats most launches leave out. Two of them are the argument of this piece.
First, on their side-by-side demo, they note that the input was a short, dense, detailed paragraph, and that this favours their model. Curated input flatters results. That's true of nearly every AI demo, and it's rarely said out loud. It's the same reason pilots look better than rollouts: the pilot runs on material someone prepared, and the rollout runs on whatever the organisation actually has.
Second, their evaluations assume a correct workflow already exists, and every model is tested against the same one. They explain that reliable real-world workflows break a problem into many separate questions, and that getting this right takes careful domain-specific work, done consistently. Deciding what those questions should be, what each one needs to know, and what to leave out is the work that comes before the model. Their benchmark reasonably holds it fixed. In a real business, it's the thing that varies most.
That isn't a criticism. An evaluation has to hold something constant. But it's worth being clear about which part of the problem has been solved.
What this means if you're automating decisions
Before any decision gets wired into your software, three questions are worth answering for each one:
- What does the model need to know to make this decision well?
- Where does that information live today?
- Is it written down, or does it live in someone's head?
The third question is where most of the risk sits. Every business has exceptions, history and judgement calls that were never documented because a person could always just ask. A model can't ask. It will make a valid decision from whatever it was given, and it won't tell you something was missing.
Output guarantees are quickly becoming a solved problem. Input isn't.
If you're planning to put AI decisions inside your workflows, start by finding out what those decisions need to know before they're automated. That's what the audit is for. If you'd rather start on your own, the AI Opportunity Audit template works through the same questions.
Source: Introducing System One Models & Jev, TypeSafe AI, 15 September 2026.
--- date: 2026-09-23 updated: 2026-10-03 title: The model can't hallucinate. Your inputs still can. meta_title: The model can't hallucinate; your inputs can slug: the-model-cant-hallucinate-your-inputs-still-can summary: A new class of AI model guarantees every answer is well-formed. That solves one problem and makes another harder to see. stat: '3' stat_label: questions to answer before any decision gets wired into your software ---
If this sounds like your business, the audit is the quickest way to find out what your AI is missing: one workflow, one 90-minute session, and a written brief on what to fix first.
Book the auditPrefer to talk first? Get in touch — no obligation.