Who Reads the Output
A chat model and a decision model differ in who reads the output. A chatbot writes sentences, and a person reads them. A decision model like TypeSafe’s Jev returns a typed answer with a calibrated probability, and a program reads it. Everything else people argue about, speed and price and “isn’t it just a classifier,” follows from that. I learned it the expensive way, which in this case cost about a fifth of a cent: I pointed Jev at a job where the only reader was me, and it was useless.
Diogo Almeida, TypeSafe’s co-founder and CEO, made the same point on this week’s TWIML AI Podcast. Sam Charrington’s intro frames it as “machine-native intelligence”: models optimized “not for generating strings that people consume, but for making reliable decisions that software can act on.” Diogo spent four and a half years at OpenAI on the team behind InstructGPT and RLHF, the training that turned raw language models into assistants. He helped build the thing he’s now arguing is the wrong tool for most of automation.
Twenty years ago the reader was always a person
Sam and I worked together on corporate portals at Plumtree Software, a company Glenn Kelman co-founded and ran marketing for before he became Redfin’s CEO. From portals to properties. A portal’s whole job was to put the right information on the right person’s screen: the sales numbers, the HR form, the document you couldn’t find on the shared drive. The decision happened after the page loaded, in someone’s head.
The chatbot is the most capable portal ever built. You ask it anything and it puts a well-written answer on your screen. Then you decide. For a person, that’s the right shape.
For software, it’s the wrong one. Code can’t act on a paragraph. Every team that has wired an LLM into a pipeline has written the same glue: ask for JSON, parse it, retry when the parse fails, guess at a threshold because “high confidence” in the prose means nothing. The model was trained to please a reader who isn’t there.
A decision model answers in types
Jev takes unstructured input and a typed question and returns one of three things: a Choice from options you define, a Score against levels you define, or the probability that a statement is true. I covered the mechanics and TypeSafe’s own speed and price claims in The Spec Was Never the Goal, so I won’t repeat them here.
What excites me isn’t the benchmark. It’s that the output is a value. A value goes in an if statement. A threshold on a calibrated probability is a line of code you can test, review and change without retraining or re-prompting anything. The training method is where the difference starts. TypeSafe’s launch post sets it next to the two methods behind today’s LLMs. Reinforcement learning from human feedback (RLHF) rewards answers a person prefers. Reinforcement learning with verifiable rewards (RLVR) rewards outputs a program can check, like a math proof. TypeSafe’s method, Reinforcement Learning for Calibrated Decisions (RLCD), rewards “answers with epistemically honest probabilities”: “higher confidence means higher accuracy.” What a person prefers is the objective that made chat models pleasant, and a program doesn’t care about it. What a program needs is to know how much to trust the answer.
Herbert Simon argued in Administrative Behavior in 1947 that the unit of an organization isn’t the person or the department. It’s the decision, and the premises that feed it. Seventy-nine years later we have a model whose unit is the same thing.
My spike failed because I was the reader
On September 24 I installed the TypeSafe plugin and pointed Jev at my own blog. The idea was to check each post against this site’s voice contract: does the lede state the finding, are the section headers claims or teasers, does the close make a new move or recap the old ones.
On the literal checks it was sharp. On the real lede of my September 16 post it scored about 0.94; on a copy I’d broken on purpose, about 0.03. Eight posts cost about $0.002.
On the relational check it fell over. It rated my real close as a recap at 0.85 and a recap I inserted on purpose at 0.97. Both high, barely apart. Whether a close is a new move depends on everything before it, and a single typed question about one block of text can’t see that. When I deleted a post’s opposing-case section and asked Jev to find it, it picked a different section, because I hadn’t made “none” an option.
None of that is why I scrapped it. I scrapped it because nothing consumed the output. I read every score. Claude already runs the same checks when it edits my drafts, and it explains itself in sentences, which is what I want when I’m the one deciding. A fast, cheap, typed answer is worth something only when code is waiting on it. I had built a portal.
The case that it’s just a classifier
The strongest version of the skeptic’s case goes like this. Typed outputs with calibrated probabilities have existed for decades. Logistic regression gives you a probability. Platt scaling, from 1999, calibrates a classifier’s scores after training. A fine-tuned small model will route support tickets faster and cheaper than anything with an API key. Frontier LLMs now support structured outputs, so you can get a schema-valid Choice from the same model you already pay for. And my own spike shows the limit: the judgments that matter most in writing, and probably in a lot of other work, are relational, and a model that answers one question about one input at a time misses them.
Most of that is right. What it misses is the setup cost. A classifier needs labeled data and a training run for every question. A decision model takes the question in plain language at call time, with your own option list, and returns a calibrated answer the first time you ask. That moves a judgment from “project” to “function call.” The structured-outputs argument is closer, but it inherits the chat model’s latency, its price per token and its confidence, which was tuned to sound right to a person. A schema guarantees the shape of the answer. It doesn’t guarantee that 0.9 means right nine times in ten.
The relational limit is real, and it’s the reason Diogo’s other point lands: AI systems should be engineered from parts, not handed to one model to do everything. A decision model is a part.
The people who use a word decide what it means
“Decision model,” “System One,” and “machine-native” are TypeSafe’s words, and Diogo said on the episode that he doesn’t get the final say on them. The people who build with these models will settle what they’re called, the way they settled “agent” and “prompt.” The skeptics’ “classifier” is in the running too.
That’s a better rule than the one we just watched. On September 29, President Trump signed an executive order directing federal agencies to replace “Artificial Intelligence” and “AI” with “Super Intelligence” and “SI.” AI company leaders signed a separate accord at the same White House summit. I read the rename as a loyalty test more than a vocabulary change: will the companies use the leader’s word? An order can change what goes on agency letterhead. It can’t change what thousands of developers type into a search box. The name that lasts is the one the people doing the work keep using.
Where a decision belongs
The test I now run before putting any model call in a system is one question: who reads the output?
If a person reads it, use a chat model. The sentences are the product, and the person makes the call. If code reads it, the output should be a decision with a type and a calibrated probability, and the threshold belongs in code where it can be tested. Routing a request, gating a risky tool call, flagging a record, picking which of five options a free-text answer meant. Those were never paragraphs.
Twenty years ago Sam and I were building screens so people could make better decisions. Most of the decisions software needs never deserved a screen. They needed an answer a program could trust, and until this year there wasn’t a model built to give one.
Resources
- Sam Charrington and Diogo Almeida, “Why Jev Is Changing How We Build With AI”, The TWIML AI Podcast, episode 779, October 6, 2026.
- TypeSafe AI, “Introducing System One Models & Jev”, September 15, 2026, and documentation (vendor).
- The White House, “Fact Sheet: President Donald J. Trump Inaugurates The Era of Super Intelligence”, September 29, 2026.
- Fox Business, “Trump signs executive order rebranding AI as ‘Super Intelligence’ as tech titans ink separate SI accord”, September 2026.
- Redfin, Management team: Glenn Kelman (Plumtree co-founder biography).
- Long Ouyang et al., “Training language models to follow instructions with human feedback”, OpenAI, 2022 (InstructGPT).
- John Platt, “Probabilistic Outputs for Support Vector Machines and Comparisons to Regularized Likelihood Methods,” 1999.
- Herbert A. Simon, Administrative Behavior, 1947.
About the Author
Kevin P. Davison has over 20 years of experience building websites and figuring out how to make large-scale web projects actually work. He writes about technology, AI, leadership lessons learned the hard way, and whatever else catches his attention—travel stories, weekend adventures in the Pacific Northwest like snorkeling in Puget Sound, or the occasional rabbit hole he couldn't resist.