Evals
Decision models and Drupal: a match made in heaven
AI Validations turns a Drupal field into a yes or no question for an AI model. I compared a decision model, Jev, with GPT-5.6 Sol on 1,500 brand checks to see which one fits that job.
How AI Validations works
Say you run a site where people publish under a brand. There is a brand guideline, and everyone has read it once. Then someone writes "zero emissions", names a competitor or quotes a customer by their full name, and nobody notices until it is live.
AI Validations moves that check to the moment someone presses Save. It adds AI rules to the Field Validation module. You attach a rule to a text, image or audio field, pick an AI provider and model through the AI module, and describe what the field has to satisfy. When the form is saved, the rule sends the value to the model and gets a pass or a fail back. A fail stops the save and tells the editor why, next to the field, like any other validation error.
What happens when an editor presses Save
-
Save
on saveAn editor saves an article. Field Validation runs every rule attached to its fields.
-
Ask
any providerAI Validations sends the field value and your instructions to the provider you picked in the AI module.
-
Answer
yes or noThe model answers one question: does this value pass? A prompt rule asks for XTRUE or XFALSE.
-
Decide
pass or failA pass lets the save through. A fail stops it, with a message on the field.
There are seven rules: a text prompt, the AI module's moderation, text classification, image classification, an image prompt, object detection and audio, which transcribes the file first and then checks the transcript. They sit next to your other validation rules in the same admin screen.
Look at what the model is asked to do. It doesn't write anything. It answers one question with yes or no, every time someone saves. That is a different job from a chatbot, and it is the job decision models are built for.
A yes or no question deserves a yes or no model
A large language model answers by writing. To get a verdict out of it, you ask it to write one word or a bit of JSON and hope it sticks to the format. A decision model skips the writing. TypeSafe's Jev models take text and a question and return a probability: how likely is it that this holds? That gives you a pass or a fail you can put a threshold on, plus a number that says how sure the model is.
Drupal already has the plumbing. Since version 1.6 the AI module has a Decision operation: true or false, choice and score questions, answered with probabilities. It comes with a Decision Explorer to try questions in the browser, decision rules for AI Automators that fill Boolean, list and number fields, and guardrails for decision requests. The TypeSafe AI provider connects Jev to it, and to the text classification and moderation operations that AI Validations' classification and moderation rules use. The Hugging Face provider supports the Decision operation already too, so the same rules can run on models from there. Jev doesn't chat, so the free-text prompt rule still needs a chat model.
The question is whether a decision model is good enough to trust with a Save button. That is what this benchmark tests.
The benchmark
Fernway is a made-up e-bike brand with a brand guideline that reads like a real one. 100 test articles were written against it: 50 that should pass and 50 that should fail, each with a known right answer. Some break a rule in plain sight, some hide the breach in an otherwise good article, and some are near misses that look wrong but are allowed.
| Rule tier | Rules | What happens |
|---|---|---|
| Absolute don'ts | 8 | One breach fails the content |
| Do's | 5 | Missing one is fine, missing two or more fails |
| Preferences | 9 | Never fail content on their own |
| Rule | What it says |
|---|---|
| A1 No competitors | Never name another bike, e-bike, scooter or mobility brand, not even neutrally |
| A2 No safety guarantees | No "100% safe", "crash-proof", "theft-proof" or "unbreakable" |
| A3 No absolute green claims | No "zero-emission", "carbon neutral" or "eco-friendly"; specific facts are fine |
| A4 No prices, discounts or exact dates | Seasons and "later this year" are fine |
| A5 No unsafe or illegal riding | No riding without a helmet, after drinking, on the pavement or with the limiter removed |
| A6 No health claims | Nothing that cures, treats or prevents a condition, not even in a customer quote |
| A7 No personal data | Customers by first name only; staff may appear by full name |
| A8 No profanity or politics | No swearing, masked or not, and no endorsing or attacking politicians or parties |
| D1 Spell the name right | Fernway, never "FernWay" or "Fern Way" |
| D2 Talk to the reader | Use "you" or "your" at least once |
| D3 Qualify range | "Up to", plus a condition such as terrain or load |
| D4 Metric first | km, kg and km/h; imperial may follow in brackets |
| D5 End with a next step | One clear action at the end |
Every article was checked five times by three contenders. All three got the guideline word for word with each article, and a verdict counts as pass when the pass probability is 0.5 or more.
- Jev (jev-latest, version 1.13.0): one yes or no decision.
- GPT-5.6 Sol with reasoning off: a JSON verdict with a confidence.
- GPT-5.6 Sol + reasoning: the same model, told to write its reasoning before the verdict.
1,500
Results
| Measure | Jev | GPT-5.6 Sol | GPT-5.6 Sol + reasoning |
|---|---|---|---|
| Accuracy, mean of 5 runs | 89.2% ± 0.4 | 87.6% ± 1.3 | 99.8% ± 0.4 (best) |
| Breaches caught, of 250 | 235 · 94.0% | 242 · 96.8% | 250 · 100.0% (best) |
| Good articles wrongly blocked, of 250 | 39 | 54 | 1 (best) |
| Articles whose verdict changed between runs | 1 (best) | 6 | 1 (best) |
| Calibration, Brier score (lower is better) | 0.085 | 0.122 | 0.002 (best) |
| Median time per validation | 307 ms (best) | 2,157 ms | 4,405 ms |
| 95th percentile time | 512 ms (best) | 3,420 ms | 5,474 ms |
| Tokens per validation, in / out | 1,554 / 20.0 | 1,277 / 18.8 | 1,321 / 227.7 |
| Cost per 1,000 validations | $0.065 (best) | $1.03 | $5.16 |
GPT-5.6 Sol with written reasoning is close to perfect: one wrong answer in 500. Jev and GPT-5.6 Sol without reasoning land close together, around 88 to 89%, but they go wrong in different places.
Jev
jev-1.13.0, one yes or no decision
446 of 2:49 answers right; 89 of 100 cases right in every run.
GPT-5.6 Sol
reasoning off, JSON verdict
438 of 19:08 answers right; 86 of 100 cases right in every run.
GPT-5.6 Sol + reasoning
writes its reasoning first
499 of 37:31 answers right; 99 of 100 cases right in every run.
Articles in the order of the test set, art-001 to art-100. Each square has five stripes, runs 1 to 5. The grids fill at 50 times the measured speed: every call's own time, run after run.
| Kind of article | Articles | Should | Jev | GPT-5.6 Sol | GPT-5.6 Sol + reasoning |
|---|---|---|---|---|---|
| Compliant | 15 | pass | 100% | 100% | 100% |
| A7 Personal data | 5 | fail | 60% | 84% | 100% |
| One Do missed | 10 | pass | 42% | 8% | 98% (best) |
| 2+ Do's missed | 8 | fail | 87.5% | 90% | 100% |
| A1 Competitor named | 5 | fail | 100% | 100% | 100% |
| A6 Health claim | 5 | fail | 100% | 100% | 100% |
| A3 Green claim | 5 | fail | 100% | 100% | 100% |
| A4 Price or date | 6 | fail | 100% | 100% | 100% |
| A2 Safety guarantee | 5 | fail | 100% | 100% | 100% |
| Preferences only | 12 | pass | 83.3% | 93.3% | 100% |
| Near miss | 13 | pass | 100% | 93.8% | 100% |
| A5 Unsafe riding | 6 | fail | 100% | 100% | 100% |
| A8 Profanity or politics | 5 | fail | 100% | 100% | 100% |
All three catch the obvious breaches every time: competitors, safety guarantees, green claims, prices and dates, unsafe riding, health claims, profanity and politics.
Where Jev and GPT-5.6 Sol without reasoning slip is the rule that needs counting: missing one Do is fine, missing two fails. Of the ten articles that miss exactly one Do and should pass, GPT without reasoning blocked almost every one, with 4 right answers out of 50. Jev got 21 right. Writing the reasoning first fixes it, because the model goes through each rule, marks it met or not, and only then counts.
Jev's other weak spot is personal data. It caught a home address, a phone number and a staff email address, but let a customer's full name through in both articles that had one, every time.
Speed and cost
307 ms
Measured per call over 500 calls for each contender.
| Contender | Median | 95th percentile |
|---|---|---|
| Jev | 307 ms | 512 ms |
| GPT-5.6 Sol | 2156 ms | 3420 ms |
| GPT-5.6 Sol + reasoning | 4404 ms | 5473 ms |
From the average tokens per call. GPT-5.6 Sol at OpenAI's promotional prices: $4 per million input tokens, $0.40 cached and $20 output. Jev at $0.042 per million input tokens from a public price listing, output free; TypeSafe's docs list no prices, so check yours.
| Contender | USD |
|---|---|
| Jev | 0.065 |
| GPT-5.6 Sol | 1.03 |
| GPT-5.6 Sol + reasoning | 5.16 |
Added up, Jev spent 2 minutes 50 seconds of model time on its 500 checks. GPT-5.6 Sol needed 19 minutes, and 37 and a half with reasoning. Most of GPT's input came from OpenAI's prompt cache, because the guideline sits in the system prompt; without that cache it would cost more.
For a Save button, the time matters more than the money. An editor waits for every validation. A third of a second goes unnoticed. Four and a half seconds per rule doesn't, and a form with several AI rules adds them up.
Where each one goes wrong
A few articles tell the story better than the averages.
"Winter commuting, one rider's story" should fail: it quotes a customer by his full name. Jev passed it all five times, at a probability of about 0.75. GPT-5.6 Sol without reasoning caught it once and passed it four times at 0.99. With reasoning it was caught every time.
"E-bike or car for the school run?" should pass: it only compares the bike with a car and a regular cargo bike, and generic comparisons are allowed. Jev passed it every time. GPT-5.6 Sol without reasoning blocked it in four of five runs, each time with a pass probability of 0.02 or less.
"How the Loop changes your commute" gives distances in miles only, so it misses one Do and should pass. Jev sat right on the line: 0.52 in the first run, then between 0.41 and 0.49. It is the only article where Jev's verdict changed between runs. GPT-5.6 Sol blocked it every time without reasoning and passed it every time with it.
"What makes the Haul steer so steadily" has no next step at the end, again one Do missed. It is the reasoning variant's only mistake: one fail in five runs.
Look at the probabilities. Most of Jev's wrong answers sit between 0.2 and 0.6, close to the line, so its doubt is visible. GPT-5.6 Sol without reasoning says 0.99 or 0.01 whether it is right or wrong. A confidence you can't trust is no use for a threshold, and the Brier scores in the table say the same: 0.085 for Jev, 0.122 for GPT without reasoning.
Best of both: let Jev ask for help
Because Jev says how sure it is, it can hand over the cases it is unsure about. I replayed the measured answers: keep Jev's verdict when it is confident, and use the reasoning model's verdict from the same run when Jev's probability falls in a band around 0.5. This is a replay of the results above, not a new run.
| Send to the reasoning model when Jev says | Calls handed over | Accuracy | Average time | Cost per 1,000 |
|---|---|---|---|---|
| Never (Jev alone) | 0% | 89.2% | 339 ms (best) | $0.065 (best) |
| 0.3 to 0.7 | 17% | 96.0% | 1,106 ms | $0.94 |
| 0.2 to 0.8 | 47% | 99.4% | 2,448 ms | $2.50 |
| Always (reasoning alone) | 100% | 99.8% (best) | 4,503 ms | $5.16 |
With a narrow band, Jev handles 83% of the checks on its own and accuracy goes from 89% to 96%. With a wider band you get within half a point of the reasoning model at half its cost and time. AI Validations doesn't do this today, but it would be a small addition to a rule: a threshold band and a second provider for the unsure cases.
Why decision models and Drupal fit
Validation is where decision models shine: short questions, asked on every save, where speed, price and an honest probability matter more than eloquence. Drupal brings the other half. AI Validations turns a field rule into such a question, the AI module lets each rule use the provider that fits it, and Field Validation shows the result where editors already look. You don't write glue code; you configure a rule.
Jev is not the most accurate validator in this test. GPT-5.6 Sol with reasoning is. But Jev is 14 times faster than that and about 80 times cheaper, about as accurate as GPT without reasoning, and more honest about its doubts. For checks that run every time someone presses Save, that trade is often the right one, and with the unsure cases handed over you don't have to choose.
Chapters
15 min read
Other ways to consume this
Listen to this article
11:34Podcast: Decision models and Drupal: a match made in heaven
Podcast episode
5:10Workflows of AI Podcast
Listen in your podcast app
Transcript
- AI disclosure
- This podcast episode is created and spoken by AI based on articles written on workflows-of-ai.com.
- Kwame Kwilson
- Welcome back. Today's post has a big claim in the title: decision models and Drupal, a match made in heaven. Jane, what's it really about?
- Jane Olson
- It's about the Save button. There's a module called AI Validations. It adds AI rules to Field Validation, seven of them, for text, images and even audio. An editor presses Save, the rule asks a model one question about the field, and the model says pass or fail.
- Erin Milleu
- And that's the bit I find interesting. The model isn't writing anything. It's answering yes or no, every single time someone saves. That's quite a different job from a chatbot.
- Jane Olson
- Right. The prompt rule literally tells the model to answer XTRUE or XFALSE. It's a yes or no question wearing a costume.
- Kwame Kwilson
- So where do decision models come in?
- Jane Olson
- Since version one point six, the AI module has a Decision operation. True or false, choice and score questions, answered with probabilities. There's an explorer to try it in the browser, automator rules that fill fields, and guardrails. The TypeSafe AI provider plugs their model, Jev, into it, and the Hugging Face provider supports it too.
- Erin Milleu
- And I'd add the part people skip. A language model gives you a verdict by writing a word and hoping it sticks to the format. A decision model hands you a probability. You can put a threshold on that, and it tells you how sure it is.
- Kwame Kwilson
- Then let's see if it holds up. The benchmark: a made-up e-bike brand called Fernway, a brand guideline, a hundred test articles, half should pass, half should fail. Each checked five times by three contenders. That's fifteen hundred calls. Accuracy: Jev eighty-nine point two percent. GPT five point six Sol, reasoning off, eighty-seven point six. And the same model told to write its reasoning first: ninety-nine point eight.
- Erin Milleu
- One wrong answer in five hundred. That's genuinely impressive.
- Jane Olson
- And Jev, without writing a single word of reasoning, lands right next to GPT without reasoning. Slightly ahead, even.
- Kwame Kwilson
- Now speed and cost, which is where it gets lopsided. Median time per validation: Jev, three hundred and seven milliseconds. GPT without reasoning, about two point two seconds. With reasoning, four point four seconds. Per thousand validations, Jev costs about six and a half cents. GPT, a dollar and three cents. With reasoning, five dollars sixteen.
- Jane Olson
- So the reasoning model is basically Erin. Always right, but you have to sit there while it explains itself.
- Erin Milleu
- I'll take always right. Some of us check our work.
- Jane Olson
- That's fair.
- Kwame Kwilson
- Erin, where do they actually go wrong?
- Erin Milleu
- The counting rule. The guideline says missing one Do is fine, missing two fails. Ten articles miss exactly one, so they should pass. GPT without reasoning got four answers right out of fifty. Jev got twenty-one. With reasoning it goes through each rule, marks it met or not, and then counts. That's what fixes it.
- Jane Olson
- And Jev has one more blind spot. Personal data. It caught a home address, a phone number and a staff email, but it let a customer's full name through, every time.
- Erin Milleu
- Yes, but look at how it was wrong. Most of Jev's wrong answers sit between zero point two and zero point six. Close to the line. GPT without reasoning says ninety-nine percent sure whether it's right or wrong. The Brier score says the same thing: zero point zero eight five for Jev, zero point one two two for GPT.
- Kwame Kwilson
- Which leads to the replay in the post.
- Erin Milleu
- Right. Because Jev says when it's unsure, you can hand only those cases to the reasoning model. If you hand over everything between zero point three and zero point seven, that's seventeen percent of calls, and accuracy goes from eighty-nine to ninety-six. Widen the band to zero point two to zero point eight and you're at ninety-nine point four, for about half the cost of the reasoning model alone.
- Jane Olson
- And in Drupal that's not a big build. AI Validations doesn't do it today, but it'd be a threshold band and a second provider on a rule.
- Kwame Kwilson
- Takeaways. Jane, you first.
- Jane Olson
- For checks that run on every save, speed and price matter more than eloquence. Jev is fourteen times faster than the reasoning model and about eighty times cheaper. And you configure it, you don't code it.
- Erin Milleu
- Mine: trust the model that tells you when it's unsure. That honesty is what makes the hand-over work.
- Kwame Kwilson
- That's the episode. Decision models, a yes or no question on every save, and the numbers to back it up. The full post has every chart. See you next time.
Narrated slides: Decision models and Drupal: a match made in heaven
The post in three sentences
AI Validations lets a Drupal field ask an AI model a yes or no question every time someone saves, and that is exactly the job decision models are built for. In 1,500 brand-guideline checks, the decision model Jev was about as accurate as GPT-5.6 Sol without reasoning, 7 times faster and 16 times cheaper, and it said honestly when it was unsure. GPT-5.6 Sol with written reasoning was nearly perfect but 14 times slower and 80 times more expensive, and sending only Jev's unsure cases to it got close to the best of both.