Skip to main content
Skip to main content

Evals

Decision models and Drupal: a match made in heaven

AI Validations turns a Drupal field into a yes or no question for an AI model. I compared a decision model, Jev, with GPT-5.6 Sol on 1,500 brand checks to see which one fits that job.

How AI Validations works

Say you run a site where people publish under a brand. There is a brand guideline, and everyone has read it once. Then someone writes "zero emissions", names a competitor or quotes a customer by their full name, and nobody notices until it is live.

AI Validations moves that check to the moment someone presses Save. It adds AI rules to the Field Validation module. You attach a rule to a text, image or audio field, pick an AI provider and model through the AI module, and describe what the field has to satisfy. When the form is saved, the rule sends the value to the model and gets a pass or a fail back. A fail stops the save and tells the editor why, next to the field, like any other validation error.

What happens when an editor presses Save

  1. Save

    on save

    An editor saves an article. Field Validation runs every rule attached to its fields.

  2. Ask

    any provider

    AI Validations sends the field value and your instructions to the provider you picked in the AI module.

  3. Answer

    yes or no

    The model answers one question: does this value pass? A prompt rule asks for XTRUE or XFALSE.

  4. Decide

    pass or fail

    A pass lets the save through. A fail stops it, with a message on the field.

There are seven rules: a text prompt, the AI module's moderation, text classification, image classification, an image prompt, object detection and audio, which transcribes the file first and then checks the transcript. They sit next to your other validation rules in the same admin screen.

Look at what the model is asked to do. It doesn't write anything. It answers one question with yes or no, every time someone saves. That is a different job from a chatbot, and it is the job decision models are built for.

A yes or no question deserves a yes or no model

A large language model answers by writing. To get a verdict out of it, you ask it to write one word or a bit of JSON and hope it sticks to the format. A decision model skips the writing. TypeSafe's Jev models take text and a question and return a probability: how likely is it that this holds? That gives you a pass or a fail you can put a threshold on, plus a number that says how sure the model is.

Drupal already has the plumbing. Since version 1.6 the AI module has a Decision operation: true or false, choice and score questions, answered with probabilities. It comes with a Decision Explorer to try questions in the browser, decision rules for AI Automators that fill Boolean, list and number fields, and guardrails for decision requests. The TypeSafe AI provider connects Jev to it, and to the text classification and moderation operations that AI Validations' classification and moderation rules use. The Hugging Face provider supports the Decision operation already too, so the same rules can run on models from there. Jev doesn't chat, so the free-text prompt rule still needs a chat model.

The question is whether a decision model is good enough to trust with a Save button. That is what this benchmark tests.

The benchmark

Fernway is a made-up e-bike brand with a brand guideline that reads like a real one. 100 test articles were written against it: 50 that should pass and 50 that should fail, each with a known right answer. Some break a rule in plain sight, some hide the breach in an otherwise good article, and some are near misses that look wrong but are allowed.

How the Fernway guideline judges content
Rule tier Rules What happens
Absolute don'ts 8 One breach fails the content
Do's 5 Missing one is fine, missing two or more fails
Preferences 9 Never fail content on their own
A rule is broken by what the content claims or encourages, not by touching a topic. A Do that doesn't apply, such as a range figure in an article without one, counts as met.
The absolute don'ts and the do's
Rule What it says
A1 No competitors Never name another bike, e-bike, scooter or mobility brand, not even neutrally
A2 No safety guarantees No "100% safe", "crash-proof", "theft-proof" or "unbreakable"
A3 No absolute green claims No "zero-emission", "carbon neutral" or "eco-friendly"; specific facts are fine
A4 No prices, discounts or exact dates Seasons and "later this year" are fine
A5 No unsafe or illegal riding No riding without a helmet, after drinking, on the pavement or with the limiter removed
A6 No health claims Nothing that cures, treats or prevents a condition, not even in a customer quote
A7 No personal data Customers by first name only; staff may appear by full name
A8 No profanity or politics No swearing, masked or not, and no endorsing or attacking politicians or parties
D1 Spell the name right Fernway, never "FernWay" or "Fern Way"
D2 Talk to the reader Use "you" or "your" at least once
D3 Qualify range "Up to", plus a condition such as terrain or load
D4 Metric first km, kg and km/h; imperial may follow in brackets
D5 End with a next step One clear action at the end

Every article was checked five times by three contenders. All three got the guideline word for word with each article, and a verdict counts as pass when the pass probability is 0.5 or more.

  • Jev (jev-latest, version 1.13.0): one yes or no decision.
  • GPT-5.6 Sol with reasoning off: a JSON verdict with a confidence.
  • GPT-5.6 Sol + reasoning: the same model, told to write its reasoning before the verdict.

1,500

validation calls 100 articles, 5 runs each, 3 contenders. Run on 30 September 2026.

Results

The scoreboard
Measure Jev GPT-5.6 Sol GPT-5.6 Sol + reasoning
Accuracy, mean of 5 runs 89.2% ± 0.4 87.6% ± 1.3 99.8% ± 0.4 (best)
Breaches caught, of 250 235 · 94.0% 242 · 96.8% 250 · 100.0% (best)
Good articles wrongly blocked, of 250 39 54 1 (best)
Articles whose verdict changed between runs 1 (best) 6 1 (best)
Calibration, Brier score (lower is better) 0.085 0.122 0.002 (best)
Median time per validation 307 ms (best) 2,157 ms 4,405 ms
95th percentile time 512 ms (best) 3,420 ms 5,474 ms
Tokens per validation, in / out 1,554 / 20.0 1,277 / 18.8 1,321 / 227.7
Cost per 1,000 validations $0.065 (best) $1.03 $5.16
◆ marks the better value. Breaches caught and wrong blocks are pooled over the five runs.

GPT-5.6 Sol with written reasoning is close to perfect: one wrong answer in 500. Jev and GPT-5.6 Sol without reasoning land close together, around 88 to 89%, but they go wrong in different places.

Every validation, right or wrong

Jev

jev-1.13.0, one yes or no decision

446 of 2:49 answers right; 89 of 100 cases right in every run.

GPT-5.6 Sol

reasoning off, JSON verdict

438 of 19:08 answers right; 86 of 100 cases right in every run.

GPT-5.6 Sol + reasoning

writes its reasoning first

499 of 37:31 answers right; 99 of 100 cases right in every run.

Articles in the order of the test set, art-001 to art-100. Each square has five stripes, runs 1 to 5. The grids fill at 50 times the measured speed: every call's own time, run after run.

Right answers by kind of article
Kind of article Articles Should Jev GPT-5.6 Sol GPT-5.6 Sol + reasoning
Compliant 15 pass 100% 100% 100%
A7 Personal data 5 fail 60% 84% 100%
One Do missed 10 pass 42% 8% 98% (best)
2+ Do's missed 8 fail 87.5% 90% 100%
A1 Competitor named 5 fail 100% 100% 100%
A6 Health claim 5 fail 100% 100% 100%
A3 Green claim 5 fail 100% 100% 100%
A4 Price or date 6 fail 100% 100% 100%
A2 Safety guarantee 5 fail 100% 100% 100%
Preferences only 12 pass 83.3% 93.3% 100%
Near miss 13 pass 100% 93.8% 100%
A5 Unsafe riding 6 fail 100% 100% 100%
A8 Profanity or politics 5 fail 100% 100% 100%
Share of the five runs answered right. A1 to A8 are the absolute don'ts. "One Do missed" and "Preferences only" should pass; "2+ Do's missed" should fail.

All three catch the obvious breaches every time: competitors, safety guarantees, green claims, prices and dates, unsafe riding, health claims, profanity and politics.

Where Jev and GPT-5.6 Sol without reasoning slip is the rule that needs counting: missing one Do is fine, missing two fails. Of the ten articles that miss exactly one Do and should pass, GPT without reasoning blocked almost every one, with 4 right answers out of 50. Jev got 21 right. Writing the reasoning first fixes it, because the model goes through each rule, marks it met or not, and only then counts.

Jev's other weak spot is personal data. It caught a home address, a phone number and a staff email address, but let a customer's full name through in both articles that had one, every time.

Speed and cost

307 ms

median time for a Jev verdict GPT-5.6 Sol took 2,157 ms without reasoning and 4,405 ms with it.
Time per validation

Measured per call over 500 calls for each contender.

Time per validation
ContenderMedian95th percentile
Jev307 ms512 ms
GPT-5.6 Sol2156 ms3420 ms
GPT-5.6 Sol + reasoning4404 ms5473 ms
Cost per 1,000 validations, in US dollars

From the average tokens per call. GPT-5.6 Sol at OpenAI's promotional prices: $4 per million input tokens, $0.40 cached and $20 output. Jev at $0.042 per million input tokens from a public price listing, output free; TypeSafe's docs list no prices, so check yours.

Cost per 1,000 validations, in US dollars
ContenderUSD
Jev0.065
GPT-5.6 Sol1.03
GPT-5.6 Sol + reasoning5.16

Added up, Jev spent 2 minutes 50 seconds of model time on its 500 checks. GPT-5.6 Sol needed 19 minutes, and 37 and a half with reasoning. Most of GPT's input came from OpenAI's prompt cache, because the guideline sits in the system prompt; without that cache it would cost more.

For a Save button, the time matters more than the money. An editor waits for every validation. A third of a second goes unnoticed. Four and a half seconds per rule doesn't, and a form with several AI rules adds them up.

Where each one goes wrong

A few articles tell the story better than the averages.

"Winter commuting, one rider's story" should fail: it quotes a customer by his full name. Jev passed it all five times, at a probability of about 0.75. GPT-5.6 Sol without reasoning caught it once and passed it four times at 0.99. With reasoning it was caught every time.

"E-bike or car for the school run?" should pass: it only compares the bike with a car and a regular cargo bike, and generic comparisons are allowed. Jev passed it every time. GPT-5.6 Sol without reasoning blocked it in four of five runs, each time with a pass probability of 0.02 or less.

"How the Loop changes your commute" gives distances in miles only, so it misses one Do and should pass. Jev sat right on the line: 0.52 in the first run, then between 0.41 and 0.49. It is the only article where Jev's verdict changed between runs. GPT-5.6 Sol blocked it every time without reasoning and passed it every time with it.

"What makes the Haul steer so steadily" has no next step at the end, again one Do missed. It is the reasoning variant's only mistake: one fail in five runs.

Look at the probabilities. Most of Jev's wrong answers sit between 0.2 and 0.6, close to the line, so its doubt is visible. GPT-5.6 Sol without reasoning says 0.99 or 0.01 whether it is right or wrong. A confidence you can't trust is no use for a threshold, and the Brier scores in the table say the same: 0.085 for Jev, 0.122 for GPT without reasoning.

Best of both: let Jev ask for help

Because Jev says how sure it is, it can hand over the cases it is unsure about. I replayed the measured answers: keep Jev's verdict when it is confident, and use the reasoning model's verdict from the same run when Jev's probability falls in a band around 0.5. This is a replay of the results above, not a new run.

Jev first, the reasoning model for the unsure cases
Send to the reasoning model when Jev says Calls handed over Accuracy Average time Cost per 1,000
Never (Jev alone) 0% 89.2% 339 ms (best) $0.065 (best)
0.3 to 0.7 17% 96.0% 1,106 ms $0.94
0.2 to 0.8 47% 99.4% 2,448 ms $2.50
Always (reasoning alone) 100% 99.8% (best) 4,503 ms $5.16
Average time includes Jev's call plus the reasoning call for the cases handed over. Cost is Jev for every call plus the reasoning model for the handed-over share.

With a narrow band, Jev handles 83% of the checks on its own and accuracy goes from 89% to 96%. With a wider band you get within half a point of the reasoning model at half its cost and time. AI Validations doesn't do this today, but it would be a small addition to a rule: a threshold band and a second provider for the unsure cases.

Why decision models and Drupal fit

Validation is where decision models shine: short questions, asked on every save, where speed, price and an honest probability matter more than eloquence. Drupal brings the other half. AI Validations turns a field rule into such a question, the AI module lets each rule use the provider that fits it, and Field Validation shows the result where editors already look. You don't write glue code; you configure a rule.

Jev is not the most accurate validator in this test. GPT-5.6 Sol with reasoning is. But Jev is 14 times faster than that and about 80 times cheaper, about as accurate as GPT without reasoning, and more honest about its doubts. For checks that run every time someone presses Save, that trade is often the right one, and with the unsure cases handed over you don't have to choose.

Chapters

15 min read

Other ways to consume this

Listen to this article

11:34
Articles read out (podcast feed)
Filed under Evals Drupal AI
AI modified

This content was produced by AI and edited by a person.

How was AI used?

AI takes my rambling notes, structures them and writes a first draft. I then prompt it to use interactive components, and at the end I read through it and rewrite it until it says what I want it to say.