Most of the AI in marketing is pointed at the part everyone can see: the headline, the email, the ad variant. The part nobody sees is the thousand small decisions behind them. Is this lead real? Which team gets this message? Does this search term belong in the account? Will this product title get rejected? Those decisions are usually an if statement that is too brittle, or a person who is too expensive.
On September 15, TypeSafe AI introduced a model built for exactly that layer. It is called Jev, it is in early access, and its trick is what it gives up: it does not generate text.
I read the announcement and the docs. What follows about the model comes from them. The marketing uses are designs worth testing, not results.
What Jev is#
TypeSafe calls Jev the first System One model, a class built to make fast, structured decisions that software can use directly. The company’s own phrase is “a frontier-intelligence function call”: you send in some text or JSON, called the state, plus typed questions, and you get typed answers back. No prose to parse, and no reply that wanders off into an essay.
There are three kinds of question:
| Question | Asks | Returns |
|---|
| Choice | Which of these options? | The choice, probabilities, confidence |
| Score | Where on this rubric? | A score, probabilities, confidence |
| Noul | Is this true? | A number from 0 to 1 |
Every question in a request is evaluated in parallel and independently, against the same state. That is why the docs tell you to ask small, specific questions and combine the answers in your own code. “Rate this lead” becomes a handful of gut checks, and the weighting lives in a function you can read and change, not in a prompt you have to rewrite.
The name is a nod to the economist William Stanley Jevons, who noticed that more efficient coal use led to more coal use. The bet is that cheap decisions lead to a lot more decisions.
The pitch is speed and price. TypeSafe lists Jev at $0.042 per million input tokens, with output free, and answers in roughly 70 to 500 milliseconds. Its big comparisons against large language models come from workflow evals the company built itself, and it says so: the reference answers are the average of two other labs’ flagship models, the figures on its home page are on the high end of what to expect, and it cannot yet prove the pricing is sustainable. I read the numbers as a reason to run a test, not as a result.
Where it fits in a marketing stack#
The docs list advertising, lead generation, customer support and e-commerce marketplaces among the places to try it. They also describe the shapes of decision it suits: classifying, scoring, routing, detecting, ranking, verifying, and pulling known fields out of messy text. This is how I would map that onto marketing work.
- Lead and inbox triage. The state is the message from a contact form or chat, plus whatever you know about the company. One Choice asks what the person wants: buy, learn, get support, sell to us, something else. One Score asks how time-sensitive it is. One Noul asks whether a competitor is named. Your code decides who gets it. The messages nobody has time to read carefully are where this earns its keep.
- Search terms and site search. Classify every query as brand, competitor, generic or problem-aware, and flag the ones that have nothing to do with what the store sells. That is a job for millions of rows, which is the volume where a price per input token starts to matter. The output is a list of negative keyword candidates and a cleaner split between brand and non-brand, which is where reported ROAS tends to flatter itself.
- Campaign naming and attribution cleanup. Map thousands of messy campaign names and UTM values onto a fixed taxonomy: channel, objective, funnel stage. A Choice can take up to 255 options, so a real taxonomy fits. The result is a clean field in the warehouse instead of a regular expression that only its author understands.
- Product feed quality. Classify products into your category tree, pull out attributes such as material, audience and size system, and flag titles that carry promotional text a channel may reject. Jev returns typed fields, so there is nothing to parse. It cannot rewrite a bad title, though. A good split is that Jev finds the titles that need work and a writing model fixes only those.
- Ad comments and reviews. Sort comments under ads and product reviews into questions, complaints, praise, buying signals, spam and brand-safety risks. Questions go to support, buying signals go to someone who can answer quickly, spam gets hidden, and the review themes feed the next creative brief. The model reads text only, so for video and image ads the state is the transcript, the copy and a description of the creative, not the pixels.
- Lifecycle triggers. TypeSafe’s own side-by-side demo includes a churn likelihood question, and the docs name churn signals and purchase intent as features to extract. I would build the state from recent orders, support notes and a browsing summary, ask a few atomic questions, combine them in code, and start a win-back flow only above a threshold. When the model is unsure, the right move is to send nothing. Fewer emails to the wrong people is a result in itself.
- A checker for AI-written copy. The announcement’s “verify everything” use case is the one I would reach for first. Let a writing model produce the ad or the email, then have Jev check each draft: does it promise a discount we are not running, make a health claim, name a competitor, break a platform rule, miss the brand voice? Reject the drafts that fail and regenerate. The writer writes and the judge judges, and neither pretends to be the other.
What the code looks like#
Here is lead triage with the Python SDK. The questions use the shapes from the docs, the model version is pinned, and a low confidence sends the lead to a person instead of guessing.
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient
# Pin the version, so tuned thresholds keep their meaning.
client = TypeSafeClient(model="jev-1.13.0")
def route(message: str, company: str) -> str:
response = client.system_one(
state=f"Company: {company}\nMessage: {message}",
questions={
"intent": Choice(
instructions="What does this person want?",
criteria={
"buy": "Asks about price, scope or timing",
"learn": "Researching, no sign of buying yet",
"support": "Is a customer and needs help",
"vendor": "Is trying to sell something to us",
"other": "Spam, job seekers, anything else",
},
),
"urgency": Score(
instructions="How time-sensitive is the request?",
criteria=["No deadline", "Soon", "States a deadline"],
),
"competitor": Noul(
instructions="The message names a competing company",
),
},
)
answers = response.answers
intent = answers["intent"]
if intent.confidence < 0.6:
return "human_review" # the model says it is not sure
if intent.choice == "buy":
urgent = answers["urgency"].score >= 1
rival = answers["competitor"].noul > 0.5
return "sales_now" if urgent or rival else "sales_queue"
routes = {"support": "support", "vendor": "ignore"}
return routes.get(intent.choice, "nurture")
The 0.6 is a placeholder. The docs say to start with conservative thresholds, test on your own data and adjust, and I agree.
Let it say “I’m not sure”#
The feature I like most is not the speed. It is that every Choice and Score comes with a confidence. The docs suggest three paths: act automatically when confidence is high, ask for a check when it is medium, and do not act when it is low. The thresholds should scale with the stakes. Tagging a campaign can be automatic at a modest threshold. Sending a discount code to someone who may not want it should need a much higher one. Noul answers carry no confidence, so for those you set a threshold on the probability itself.
Marketing automation has a long history of acting on a guess. A model that can decline to answer is a small design change with a large effect on how many wrong emails go out.
What to check before you trust it#
TypeSafe publishes a page of known weak spots for the current version, which is the kind of page I wish more vendors wrote. These are the ones that matter for marketing.
- It is not a calculator. It does not count reliably, and it reads dates as text. Keep ROAS, CAC, budgets, flight dates and send windows in code, and use Jev only for the judgment.
- It reads literally. It answers the question you wrote, not the one you meant. Write the exact condition, and put the edge cases in the criteria.
- Inbound text is hostile until proven otherwise. The docs say plainly that content written to steer the model, such as an instruction hidden in a message, can move the answer. A lead form is public input. Never let one answer trigger something you cannot undo, keep the confidence gates, and test with messages that try to game the score.
- Valid is not the same as right. As I read the announcement, the claim that it cannot hallucinate is about the shape of the answer: it always fits the schema you defined. A well-formed, confident, wrong label is still possible, which is why you measure it.
- Keep the state small. Accuracy falls when the state is full of unrelated detail, so filter in code and send only what the question needs.
- English first. English is where accuracy is best, and other languages are handled less well. If your market speaks Swedish, German or Finnish, test on your own content before you route anything, and watch the confidence.
- Text only. No images, audio or video. Turn them into text first.
- Pin the version. The alias
jev-latest moves when a new release ships, so answers can change without any change on your side. If you tune thresholds, pin the versioned ID and log which model answered. - It is early. It is in early access, the rate limits can change without notice, the service is currently based on the West Coast, and the evals are the company’s own. TypeSafe says it does not train on customer requests. If you handle EU customer data, read the data processing agreement and check what your consent covers before any personal data goes into the state.
A first project that cannot hurt#
Pick a decision that a person makes today and that you can check afterwards. Campaign taxonomy is a good one, and so is routing inbound messages. Then:
- Collect a few hundred past examples, each with the answer a person gave.
- Run them through Jev in shadow mode, so it answers and nothing acts on it.
- Compare its answers with the people’s, and read the disagreements first. Sometimes the model is wrong. Sometimes the person was.
- Check calibration: of the answers at high confidence, how many were right? Of the ones at low confidence, how many did you want to see anyway?
- Turn on the automatic path for the safest decision only, keep a person on the rest, and keep logging.
It is a small project, and it ends with a number from your own data instead of one from a blog post, this one included.
Where this leaves me#
Marketing is full of decisions that are too fuzzy for an if statement and too small for a person. Jev is a bet that this gap is model-shaped. I like the bet, mostly because of the limits the vendor volunteers: it does not write, it does not count reliably, and it says when it is not sure. That is a short list, and it is the right shape for something that sits in the middle of a pipeline.
Treat everything here, my enthusiasm included, as a prior, and let your own data do the updating.
Sources: TypeSafe’s announcement (September 15, 2026) and its docs on use cases, confidence, models and known weak spots.