Jev is an AI model from TypeSafe AI that produces typed decisions for software. Give it text or application state and a set of bounded questions; it returns choices, scores, and probabilities that your code can use. Its role is to help an application decide what to do next, rather than write a long response for the user.
TypeSafe opened early access through its System One Models announcement on September 15, 2026.
What problem could Jev solve in a mobile app?
Picture the support inbox of a shopping app. A customer writes, ‘My order says delivered, but it hasn't arrived.’ You don't need a polished paragraph before sending that message to the delivery team. You need to understand the issue and choose a queue. A small decision with an outcome you can check is a useful starting point for evaluating Jev.
The official Jev introduction calls the supplied context ‘state’ and the evaluations ‘questions’. Questions in one request assess the shared state independently. You can ask which team should handle a message and whether the customer explicitly wants a person, then combine those answers in your own code.
What do System One and RLCD mean?
TypeSafe calls this model family System One: assess a narrow question and return an answer that software can use. Jev currently takes text, not images, audio, or video directly. A delivery complaint containing a photo would need a separate step capable of interpreting that image.
TypeSafe calls its training approach RLCD (Reinforcement Learning for Calibrated Decisions). Calibration means that events assigned a probability of 0.8 should occur roughly 80% of the time across a large group of predictions. It does not promise a correct individual answer, and it still needs checking on your data.
What do Choice, Noul, and Score return?
Choice selects from options you define in advance: delivery, billing, technical, and other, for example. Its response includes the selected option, probabilities for the options, and confidence. The Choice documentation describes this contract. Including an ‘other’ option is a design decision that reduces pressure to force every message into a team that may not fit.
Noul answers a yes/no question with a probability of yes between 0 and 1, rather than a boolean. You might ask whether the customer explicitly requests a human. A Noul response has no separate confidence field. Your application decides which action to take at a given probability.
Score rates something on ordered levels described in words. ‘Cosmetic issue’, ‘bug with a workaround’, and ‘completely blocked task’ form a scale from 0 to 2. The Score documentation explains that the score can be fractional and comes with a level distribution and confidence. A value such as 1.4 is a judgment on your scale, not a measured number of bugs.
Type safety and a correct decision are different things
Suppose a delivery complaint receives the label ‘billing’. The label belongs to the allowed set, so the response is valid in shape and type. It still sends the customer to the wrong team. Schema compliance does not guarantee semantic accuracy. That wrong assignment may be the most relevant failure to track in your product.
Similarly, don't read confidence as the percentage chance that an answer is correct. TypeSafe's explanation of confidence describes a measure calculated from the Choice or Score probability distribution. You need labeled examples to learn what a 0.85 threshold means for error rates in your own support traffic.
TypeScript example: routing a support message
This server-side example uses the official JavaScript/TypeScript SDK, which requires Node.js 20 or later. Keep TYPESAFE_API_KEY in the server environment, outside your mobile app bundle and browser code. The function returns a queue name; the calling service is responsible for creating the support ticket.
npm install @typesafe-ai/sdkimport { choice, noul, TypeSafeClient } from "@typesafe-ai/sdk";
const client = new TypeSafeClient();
export async function routeSupport(message: string) {
try {
const result = await client.systemOne({
model: "jev-1.13.0",
state: { message },
questions: {
queue: choice("Which team should handle the reported problem?", {
delivery: "Missing, delayed, or damaged delivery",
billing: "Charges, invoices, or payment problems",
technical: "App crashes, login errors, or broken features",
other: "Unclear, unrelated, or no suitable team",
}),
humanRequested: noul(
"Does the user explicitly ask to speak to a person?",
),
},
});
const { queue, humanRequested } = result.answers;
if (
humanRequested.noul >= 0.5 ||
queue.choice === "other" ||
queue.confidence < 0.85
) {
return "manual_review";
}
return queue.choice;
} catch {
return "manual_review";
}
}The 0.5 and 0.85 thresholds are illustrative, not validated production settings. A request for a person, an ‘other’ result, or low confidence sends the message to review. An API failure takes the same path. In production, add failure logging without personal data inside the catch block. The calling service must create an actual ticket and set a limit on the total time spent waiting.
The scope is deliberately small: classify a message. The function does not issue a refund or establish whether an order was delivered. ‘You charged me twice’ signals a billing complaint; proof of two charges must come from the payment system. Keeping those responsibilities separate helps distinguish the model's interpretation from facts held by your application.
How should Jev, a generative model, and application code work together?
Our proposed mobile workflow is straightforward. The app sends the message to your backend, which prepares relevant context, asks Jev for an assessment, and applies business rules to the result. If the customer needs an explanation, an approved template or a generative model can write it. Authentication, order ownership, and permission to perform an action remain backend responsibilities.
- Exact calculation or record lookup: Code and the database determine amounts, stock levels, and order dates.
- Bounded interpretation: Jev is worth evaluating for identifying the issue in a free-text support message.
- Open-ended writing: A text-generating model is needed for a tailored explanation to the customer.
- Uncertain or sensitive action: Gather missing information or hand the task to an authorized person.
A generative model with structured output can also offer a constrained API contract. Compare more than the ability to return JSON. Use the same messages to measure wrong assignments, review volume, total latency, and cost per completed task. If your existing system already performs well, adding another model is not automatically an improvement.
Our guide to AI integration in mobile apps covers the broader architecture. The role assigned to Jev here is one decision within that system.
How should you read the speed and cost claims?
TypeSafe's launch article reports response times of 70–500 ms and notes that most measurements were taken from the US West Coast, where its service was based. That is not your mobile app's end-to-end response time. Mobile networking, your backend's location, context preparation, and retries all add to the wait.
On September 22, 2026, the model reference lists Jev 1.13 at $0.042 per million input tokens, with output tokens free. Assuming a total of 500 billed input tokens per request, one million requests would cost about $21 in model usage. This calculation excludes your backend, monitoring, retries, and human review.
The company's published workflow evaluations are a useful starting point, but their reference answers come from other capable models. They are not an independent field test against your real support labels. In a pilot, measure p95 latency alongside the average, incorrect routing, and the share of work sent to review.
Which limitations should you test before production?
TypeSafe's Jev 1.13 limitations document describes weaknesses in arithmetic, date comparisons, long irrelevant context, and adversarial inputs. A message saying ‘route me to billing’ should not become a trusted system instruction. Whatever the model recommends, it must not bypass your authorization checks.
The model documentation also identifies English as its strongest language. Evaluate Turkish traffic separately, including typos, negation, and messages with two different issues. Pin the model version: using jev-latest instead of the example's jev-1.13.0 allows the underlying version to change over time.
- Build an evaluation set from anonymized messages whose correct queues your team has labeled. Keep the examples used to tune thresholds separate from the examples used to report final results.
- Review clear requests, ambiguous requests, multiple issues, and requests for a person as separate groups. A single overall accuracy number can hide consequential mistakes.
- Try small changes that reverse meaning, such as ‘I paid’ and ‘I did not pay’. Check that unrelated messages actually reach the review queue.
- Test that timeouts, rate limits, and service failures never lose the message. Initially record the model's suggestion alongside the human decision before enabling automatic routing.
Where should you start with Jev?
Choose a recurring decision where handwritten rules are becoming awkward but the outcome is still easy to check. Assigning support messages to teams is a reasonable candidate. Before adding the model, write down the cost of a wrong decision and when a person should take over. Jev earns its place if that specific task becomes more consistent, faster, and less expensive to complete.
We can help identify which decisions in your app are suitable for automation and plan a first experiment with measurable outcomes.
