🔍 Read the full analysis: 24 Ways To Explore Decision Modeling With Jev on ThorstenMeyerAI.com
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
TL;DR
In a September 29, 2026 article, Thorsten Meyer mapped 24 possible uses for Jev, a tool that returns typed judgments for software to act on. Meyer says three uses are live in his publishing operation, 12 are strong fits, seven need measurement and two are poor fits; these are his assessments, not independently verified results.
Thorsten Meyer published a map of 24 ways to use Jev, a tool that answers typed questions about text or structured data so software can make decisions. Meyer says three applications are already running in his publishing operation, while 12 other uses meet his criteria for a strong fit; the figures and performance results in the article are the author’s own reports.
Meyer describes Jev as a system for returning structured answers rather than writing or summarizing text. Depending on the question type, an answer can be a yes-or-no probability, a choice among options with probabilities and confidence, or a score on an ordered scale. The application code decides what to do with those answers. Meyer reports that one call takes about 0.3 to 0.9 seconds and costs about $0.04 per million input tokens; the article does not provide an independent measurement of those figures.
Three uses are marked live: checking whether stories fit a publication, checking whether article text is in English, and classifying headlines into topics when another system fails. Meyer says the language check scanned 78,889 articles for $2.01, found 1,576 non-English items and fixed 1,553. He also reports about 10,000 story-and-site pairings judged in three days for the relevance gate, with 22% clearly on-topic, and 89% agreement with a frontier large language model for the topic classifier.
Meyer’s proposed fit test has four parts: high volume, a narrow question, low-cost errors or a path for uncertain cases to receive more capable review, and evidence that an existing heuristic fails. He advises replaying 300 to 500 past decisions, checking results by confidence band and reviewing disagreements before deployment. His suggested rollout uses a separate feature flag, starts off, then goes to 5% to 10% of units as a canary.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
A Test Before Adding Automation
The map outlines criteria teams can use to assess automation proposals before integrating them into production. Meyer says a tool should address a measured problem and that a working keyword rule may not need to be replaced. He also describes using a confidence threshold to keep uncertain cases on an existing path or send them for human review.
The 24 examples cover different risk levels. Meyer labels disclosure checks and comment moderation strong fits, while he marks headline quality and product matching as requiring measurement first. For disclosure checks, he proposes sending uncertain or missed cases for human review rather than allowing automatic publication. These are recommendations; the article does not report that every proposed application has been deployed or validated.
The reported scan figures concern one operation and are attributed to Meyer. They do not establish that other publishers would achieve similar accuracy, costs or savings. Meyer proposes replaying local decisions and reviewing disagreements before enabling a use case.
Three Publishing Uses Are Live
The article describes three live examples in Meyer’s publishing workflow: Jev supplies an answer, and software acts according to a defined rule. For relevance screening, Meyer says clearly low-fit stories can be dropped when confidence is high, while ambiguous cases continue through the prior publishing process. The English-language check rewrites items below a stated probability threshold while keeping the same URL.
Meyer says a 31-topic classification test found 97% to 99% agreement with a frontier language model when Jev’s confidence was at least 0.8, compared with 42% below 0.5. The article attributes these figures to Meyer’s measurement; it does not describe the full evaluation setup, the size or composition of the test set, or an independent replication. These missing details limit what can be concluded about how the result applies elsewhere.
The map distinguishes proposed uses from demonstrated ones. Meyer says a thin-source detector could flag stories that lack verifiable facts, but marks it “measure first.” He describes same-event deduplication as a poor fit for now because his canary found no duplicates. The examples connect the proposed criteria to whether a problem has been observed.
“Use Jev only when all four conditions hold: High volume. Narrow question. Cheap errors. A heuristic fails visibly. Measured, not assumed.”
— Thorsten Meyer
Results Need Independent Checks
The article does not provide an external audit or enough methodological detail to independently assess the reported accuracy, cost and throughput figures. It also does not establish whether the 12 “strong fit” applications have been implemented; the labels are Meyer’s assessment under his own four-condition test. The supplied source text ends as it begins the commerce and customer-operations section, so the remaining examples and the full count breakdown across categories are not available here.
It is also unclear how results would vary across publishers, languages, content types or different error costs. Meyer advises reviewing disagreements and validating confidence bands on a team’s own past decisions, but the article does not report results from such tests for the proposed uses marked “measure first.”
Measure Before Wider Rollout
Meyer’s recommended next step for a prospective user is to replay 300 to 500 real past decisions, compare results overall and by confidence band, and review 20 disagreements to determine which answer was correct. He says a use should be wired in only where the high-confidence band reaches 95%. For rollout, he proposes keeping the feature flag off by default, testing on 5% to 10% of units, and expanding from there.
The article does not announce a product launch, outside evaluation or deployment schedule. Any broader adoption will depend on teams measuring their own error rates and deciding whether the remaining uncertainty is acceptable for the task.
Key Questions
What is Jev, according to the article?
Meyer describes Jev as a tool that takes text or JSON plus typed questions and returns structured judgments, such as probabilities, classifications or scores, for application code to use.
How many Jev applications does Meyer say are already live?
Meyer marks three uses as live in his publishing operation: relevance screening, English-language checks and a fallback topic classifier.
What does Meyer’s fit test require?
He calls for high volume, a narrow question, inexpensive errors or a review path for uncertain cases, and evidence that the current heuristic fails.
Are the performance figures independently verified?
The article attributes its accuracy, cost and processing figures to Meyer’s measurements. It does not provide an independent audit or replication.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
