JEV
JEV is a decision model from TypeSafe AI. You give it a situation and a list of possible answers, and it returns a decision with a confidence score.
This page has three parts: what JEV is, how it differs from the LLMs you already use, and how we run it on our own videos and content before anything gets filmed. Case studies from other teams are on their own pages, linked at the end.
TL;DR

The whole argument on one page: not a chatbot, a decision engine. Everything below expands it.
▶ Watch: What Is JEV AI? How I Use It for YouTube & Shorts (7 min). The full walkthrough of the loop we run below.
The quick version
The LLMs we know and love answer by writing an essay, word by word. That is great for complex open-ended tasks, but it means slow and expensive output.
JEV does not write anything. You give it a situation and a list of possible answers. It picks one and returns a confidence score.
That is why it is best for large systems that need a ton of classification and decision making on every input.
An LLM generates language, JEV evaluates choices.
See the difference
| typical LLM (GPT, Claude, general models) | JEV | |
|---|---|---|
| you give it | a question | input plus fixed options |
| it does | writes the answer word by word | scores options in parallel |
| you get back | free-form text | a typed choice with probabilities |
| best for | open-ended work: writing, explaining, brainstorming, rewriting | repeated judgments: yes/no, scores, labels, rankings, confidence |
A normal LLM generates an answer. JEV makes a decision. They are complementary, not competitors. The LLM keeps writing. The decision layer moves to a model built for deciding.
How we use it on our own videos
This is our own version, not a demo. Every short and every long-form video we publish goes through the same loop before it gets filmed.
What Is JEV AI? How I Use It for YouTube & Shorts (7 min)
The video follows this exact loop: what JEV is, JEV against an LLM, scoring many questions in parallel, the API setup, scoring Shorts before filming, and testing titles and thumbnails.
- 0:00 Can AI predict video performance?
- 0:30 JEV vs LLM
- 1:15 Parallel video scoring
- 2:15 Setting up the JEV API
- 2:33 Scoring Shorts before filming
- 4:37 Testing YouTube titles and thumbnails
- 6:45 My final take on JEV
Why this matters to us: everyone can generate ten videos a day now, so the bottleneck moved from making to choosing. The skill is no longer producing. It is knowing what not to publish. Before JEV, our signal scoring layer was one of the most expensive layers in the stack, because every candidate script needed judgment across a whole library of transcripts. Now that judgment costs cents, so we can score every script instead of only the one we were already sure about.
Before you start: the API key
JEV runs through an API. Join the waitlist at TypeSafe AI, generate an API key, and keep it in your environment file. Everything below assumes that key exists, because the loop is only useful when scoring is cheap enough to run on every idea.
1. We built a library from our own niche
Instead of judging an idea in isolation, we started with what already performs. We collected 248 Shorts and 415 titles from the AI and no-code niche with their reported performance, so there is a baseline to score against.
I run a [your niche] content channel for [audience].
Here is a dataset of [N] posts from my niche with captions and reported
performance (views, likes, comments, shares where available).
Clean the data, then tell me what performance baseline I should calibrate
against. Flag any rows with missing or unreliable numbers.
2. We wrote the questions down once
JEV only answers what you ask, so the questions are the product. Ours cover hook strength, structure, content type, niche fit, and confidence on every judgment. Same scale, every item, so the scores stay comparable.
Define a scoring prompt for my niche where each post gets:
- hook strength (1-10)
- structure from setup to payoff (1-10)
- content type: tutorial, story, proof, listicle, contrarian take, or rant
- niche fit (1-10)
- confidence for each judgment
Return typed fields, not prose. Score all items consistently so the scores are
comparable across the library.
3. We score the draft before we film it
This is the step that saves the six hours.
Score this draft short using the same questions and the same scale as my library.
State where it beats my baseline and where it falls below it. If the format is a
mismatch for my niche, say which format performs better and why. Do not rewrite
anything yet.
The absolute score matters less than the mismatch. A strong opening delivered as a rant, in a niche where tutorials consistently win, is a format problem, and we would rather find that in a score than in the edit.
4. We fix the weakest part and score it again
JEV judges. An LLM rewrites. Then JEV judges the rewrite.
Here is the draft and its scores. Rewrite only the lowest-scoring section,
keeping my voice and the rest unchanged. Then rescore the revision with the same
questions and show both score sets side by side.
We publish when the idea beats our baseline. Then it goes back into the library, so the next round starts from a better one.
5. We score the packaging too
For long-form, the decision point moves earlier. Before anyone evaluates the content, they evaluate the title and thumbnail.
Score these title options for hook strength, title format, click worthiness,
niche fit, and confidence, against my library of 415 titles from AI channels.
Then evaluate my thumbnail concept. Tell me whether it adds curiosity or repeats
the same promise the title already makes.
Five title variations in ninety seconds, instead of publishing one and hoping.
What JEV cannot do
Virality depends on far more than text: audience trust, distribution, timing, delivery, editing, visual quality, topic demand, and platform behavior.
A confidence score is not the probability of virality. It describes confidence in the model's judgment under the questions you defined. This is decision support, not prediction.
Our datasets use captions, not full transcripts, so JEV sees the written opening and structure but not delivery, pacing, editing, visuals, or sound.
Case studies
Other teams running JEV in production. Each one has its own page, credited to the author, with a link to the original post.
- JEV maxxing for marketers — 7 marketing use cases at 30x for under $3
- JEV cut our SEO and GEO costs by 90% — the same client run: ~$250 to ~$25
- Instant compaction instead of summarization prompts — score every tool call, drop what is irrelevant
- Browser Use plus JEV: flights in 7 seconds for $0.0039
- Scoring 700 leads before outreach for $0.09
- Lead finding on Reddit — 96 real leads where Opus 5 found 72, and 4 junk kept against 50
- Live sales call copilot — scores the deal and flags objections while the call is still happening
What stays human
You still own the offer, the argument, the proof, the taste, and the final call. JEV answers the questions you define. It cannot decide what your audience should understand or associate with your brand.
Quick checklist
- Collect a niche library with real performance data
- Define consistent scoring questions and a scale
- Score one draft against your baseline
- Identify the weakest section or the format mismatch
- Rewrite only that part, then rescore
- Score titles and thumbnail together before publishing
Related resources
More from us on JEV:
Want this set up for your business?
We turn these workflows into working marketing and sales systems.