Jev-Based Metric
A categorical metric judged by Jev, a decision model that picks one category with calibrated probabilities instead of generating text.
Overview
A Jev-based metric is a categorical metric judged by Jev, TypeSafe AI's "System One" decision model, instead of an LLM. Jev never generates text. It answers typed questions about a state with a fixed answer and a calibrated probability for each option, which makes it fast and cheap for verdicts that are one of a few known outcomes.
Jev is a model provider in Rhesis, not a metric type. Any categorical metric can use it by selecting Jev as its model.
How It Works
When a categorical metric runs on Jev, Rhesis turns it into a single Choice question:
- State: the test input and the AI response
- Instructions: the metric's evaluation prompt and evaluation steps
- Options: the metric's categories
- Result: Jev picks one category and returns a probability for each; Rhesis marks the test passed if the pick is one of the passing categories
The result's reason gives the chosen category and its probability. There is no written reasoning.
Jev vs. LLM Judge
LLM judge: Reads a rubric in plain language and returns a score or category with written reasoning. Works for every metric type, but each call takes seconds and output tokens are billed.
Jev: Takes a question with fixed answer options and returns one option with probabilities. Calls are fast, input tokens are cheap and output is free, which matters most for long inputs such as multi-turn conversations.
Use Jev when the verdict is one of a few known outcomes. Use an LLM judge when you need a numeric score or an explanation.
Using Jev in the Platform
- Open Models and choose the Jev tile
- Enter your TypeSafe API key
- Click Test Connection
- Select Jev as the model of a categorical metric, or make it the default model for evaluation
Using Jev with SDK
Limitations
- Categorical metrics only: Decision models work only with , in both single-turn and multi-turn scope. A pass/fail check works when written as categories. Numeric judges and conversational judges (such as goal achievement) raise
- Evaluation only: Jev can't be used for test generation or test execution
- No rationale: Results carry a category and probabilities, not an explanation
- Literal reading: Negations are taken at face value, and Jev can't count or do arithmetic
Best Practices
- Self-describing categories: Name each category so it stands on its own, e.g. rather than
- Positive instructions: Phrase the evaluation prompt so the meaning doesn't hinge on a "not"
- Few outcomes: Keep categories mutually exclusive and few in number
- Use the probabilities: Review low-confidence verdicts by hand or re-check them with an LLM judge