🔍 Read the full analysis: 24 AI Decision-Modeling Ideas You Can Try With Jev on ThorstenMeyerAI.com
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Thorsten Meyer of ThorstenMeyerAI.com has catalogued 24 concrete uses for Jev, a decision-modeling AI that returns calibrated answers rather than prose. Three are live in his publishing operation covering roughly 90,000 decisions, 12 are strong fits, 7 need measurement first, and 2 are called poor fits.
A new practitioner report published September 29, 2026 on ThorstenMeyerAI.com maps 24 concrete use cases for Jev, an AI decision-modeling tool that returns calibrated answers instead of generated text. According to author Thorsten Meyer, three of the uses are already live in his own publishing operation, accounting for roughly 90,000 decisions so far, while 12 more meet his four-condition fit test and are ready to build.
Jev, as described in the report, does not write, summarize or extract. Users send it a state — text or JSON — plus a set of typed questions, and it returns calibrated answers with confidence scores that code can branch on without parsing prose. A single call carrying the state and all questions takes about 0.3 to 0.9 seconds and costs about $0.04 per million input tokens, according to Meyer. Three answer types are supported: a probability of yes (0 to 1), a choice with probability per option, and a score on ordered levels.
The report’s central claim concerns confidence. In Meyer’s measurement on a 31-topic classification task, Jev agreed with a frontier LLM 97 to 99% of the time when its confidence was 0.8 or higher, but only 42% of the time below 0.5. The pattern behind most use cases follows from this: act on clear cases with high confidence and route the gray zone to humans or other systems.
Of the 24 mapped uses, the breakdown by Meyer’s own tags is: 3 live, 12 strong fits meeting all four fit conditions, 7 labeled measure first because a visibly failing heuristic is unproven, and 2 labeled poor fits. One live language check scanned 78,889 articles for $2.01, finding 1,576 non-English pieces and fixing 1,553, according to the report. A separate live relevance gate judged about 10,000 story-site pairings in three days, with only 22% clearly on-topic. A classifier fallback agreed with a frontier LLM 89% overall.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
Reported Costs and Methodology
The report documents a production deployment of decision-modeling AI with cost and accuracy figures attributed to a single operator’s measurements rather than vendor claims. The reported figures include an overnight scan of nearly 79,000 articles for about two dollars, which the report compares with the sampled or keyword-based checks it describes as the current norm. The figures have not been independently verified.
The report also sets out a methodology. Meyer’s four-condition fit test — high volume, narrow question, cheap errors, and a visibly failing heuristic — plus a shadow-replay process (300 to 500 past decisions, per-confidence-band comparison, canary rollout at 5 to 10 units) describes one approach for evaluating decision AI systems. The report also documents rejected use cases: two are rejected outright, including one where a canary test found zero duplicates to fix.
How Jev Fits Alongside LLMs
Jev occupies a different niche from general-purpose language models. According to the report, it is designed for thousands of small, narrow judgements — gates, flags, routing, classification and severity scoring — rather than writing or multi-step reasoning. A frontier LLM appears in the report as the reference standard: the live classifier fallback uses Jev as a cheaper layer, falling back to the LLM only on errors, with a keyword rule as the last resort.
The report’s fit test explicitly says to keep an existing keyword rule if it works. Meyer states that condition four — a heuristic that measurably fails — is what disqualifies most of the ‘measure first’ cases, such as a thin-source detector intended to replace a 300-character length rule that cannot distinguish a dense wire item from a teaser. He reports that 88% of the news items he processes start from a bare headline, which motivated that use case.
“Jev is the right tool wherever a system needs thousands of small judgements and can hand the unclear ones to something smarter.”
— Thorsten Meyer, ThorstenMeyerAI.com
Single-Operator Data and Unproven Cases
All performance figures in the report come from one operator’s own measurements and have not been independently verified. The headline agreement figures — 97 to 99% at high confidence — apply to a single 31-topic classification task and may not generalize to other domains or languages. Costs of roughly $0.04 per million input tokens reflect Meyer’s usage and could differ at other volumes.
Seven of the 24 use cases are explicitly tagged measure first because the fourth condition, a visibly failing current heuristic, is unproven. These include the thin-source detector, headline quality scoring and a product-fit check for shopping roundups. The full descriptions of the commerce, software, business-operations and home use cases were only partially detailed in the available source material, so their specific rules and expected error rates are not yet published in full.
From Shadow Tests to Rollout
According to the report’s own methodology, the next step for the seven ‘measure first’ uses is a shadow replay of 300 to 500 past decisions, compared overall and per confidence band, with 20 disagreements read manually to judge who was right. Uses that reach 95% accuracy in the high-confidence band would then be wired in behind an off-by-default flag, canaried on 5 to 10 units, and rolled out gradually. Meyer advises readers to start with the strong fits after applying the four-condition test, and to revisit the rejected dedupe use case only if a duplicate problem is actually measured.
Key Questions
What is Jev, and how does it differ from a chatbot or LLM?
According to the report, Jev does not write or summarize. It takes a state (text or JSON) and typed questions, and returns calibrated answers — probabilities, choices or scores with confidence — that code can branch on directly. It is aimed at thousands of small judgements rather than open-ended generation.
How accurate is Jev according to the report?
In the author’s measurement on a 31-topic classification task, Jev agreed with a frontier LLM 97 to 99% of the time when confidence was 0.8 or higher, and 42% of the time below 0.5. These are single-operator figures and have not been independently verified.
What are the three use cases already running live?
A relevance gate judging story-site fit (about 10,000 pairings in three days), an English-language check (78,889 articles scanned for $2.01, with 1,553 non-English pieces fixed), and a classifier fallback that agreed with a frontier LLM 89% overall.
When should you not use Jev?
The report lists four conditions that must all hold: high volume, a narrow question, cheap errors, and a current heuristic that measurably fails. Two mapped use cases failed the test, including same-event dedupe, where a canary found zero duplicates to fix.
How does the report recommend rolling out a new use case?
Replay 300 to 500 real past decisions, compare results overall and per confidence band, and manually review 20 disagreements. Wire the use in only where the high-confidence band reaches 95%, keep an off-by-default flag, canary on 5 to 10 units, then roll out.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
