Overview
The Rhesis SDK evaluates LLM applications with built-in metrics from several frameworks, custom metrics you define, and metrics managed on the platform.
Metric types
- Single-turn metrics evaluate an individual exchange between a user input and a system output — RAG quality, response quality, safety, and custom checks.
- Conversational metrics evaluate interactions across multiple turns — coherence, goal achievement, role adherence, knowledge retention, and tool usage.
Metric Scopes
Every metric has a metric_scope that controls where and when it runs. The three scope values are:
Single-Turn— runs during single-turn test evaluation and per-turn trace evaluationMulti-Turn— runs during multi-turn test evaluation and conversation-level trace evaluationTrace— enables automatic evaluation against live traces (see below)
A metric can have any combination of these scopes. Scopes are additive: including more values makes the metric eligible in more contexts.
Default Scopes by Metric Class
| Metric Class | Default Scope | Notes |
|---|---|---|
NumericJudge | Single-Turn, Multi-Turn | See note below on multi-turn behavior |
CategoricalJudge | Single-Turn, Multi-Turn | See note below on multi-turn behavior |
ConversationalJudge | Single-Turn, Multi-Turn | Receives structured ConversationHistory |
GoalAchievementJudge | Single-Turn, Multi-Turn | Receives structured ConversationHistory |
GarakDetectorMetric | Single-Turn only | Operates on individual prompt/response pairs |
How single-turn metrics work in multi-turn tests: When a NumericJudge or CategoricalJudge is used in a multi-turn evaluation, the full conversation is serialized to plain text and passed as the output parameter. The metric evaluates the conversation as a single text blob. It is also passed the conversation_history, but only to return relevant_turns, the turns its verdict rests on. This means the evaluation quality depends entirely on the evaluation_prompt you write. For turn-aware evaluation (e.g., analyzing coherence between specific turns), use a ConversationalJudge instead, which receives the full structured conversation with individual turns.
You can override the default scope when creating a metric:
Trace Evaluation
Any metric can also evaluate live production traces. There is no separate scope for it: add the metric to a project’s Trace Metrics in the project settings, and that project’s traces are evaluated with it. Single-Turn metrics run after each turn, and Multi-Turn metrics run on the whole conversation once it goes quiet.
For full details on how trace metrics evaluation works — including the two-phase pipeline, debounce timing, project configuration, and first-turn handling — see the Trace Metrics documentation.
Framework Integration
Rhesis integrates with the following open-source evaluation frameworks:
- DeepEval - Apache License 2.0 The LLM Evaluation Framework by Confident AI
- DeepTeam - Apache License 2.0 The LLM Red Teaming Framework by Confident AI
- Garak - Apache License 2.0 LLM Vulnerability Scanner by NVIDIA
These tools are used through their public APIs. The original licenses and copyright notices can be found in their respective repositories. Rhesis is not affiliated with these projects.
Custom metrics
Beyond framework-provided metrics, Rhesis provides custom judges you define with a prompt and scoring rules:
NumericJudge,CategoricalJudge— single-turn scoring and classification (see Single-Turn Metrics)ConversationalJudge,GoalAchievementJudge— multi-turn evaluation (see Conversational Metrics)
Generate and improve metrics with MetricSynthesizer
MetricSynthesizer creates metric definitions from natural-language instructions. Use it when
you know what you want to evaluate but want the SDK to draft the metric fields needed by the
platform.
The synthesizer returns a dictionary suitable for a Metric entity or the backend metric create
payload.
| Method | Input | Output |
|---|---|---|
generate(prompt) | Natural-language metric description | New metric fields such as name, evaluation_prompt, score_type, and thresholds |
improve(existing_metric, prompt) | Current metric dictionary plus edit instructions | Updated metric fields for the same metric |
The generated fields follow the same metric schema used by custom judges:
| Field | Notes |
|---|---|
score_type | Must be numeric or categorical. |
threshold_operator | Used for numeric metrics; valid values include =, <, >, <=, >=, and !=. |
metric_scope | Can include Single-Turn and Multi-Turn depending on the intended evaluation context. |
Platform Integration
Metrics can be managed both in the platform and in the SDK. The SDK provides push and pull methods to synchronize metrics with the platform.
get_annotations returns both kinds of judgement about the metric: where someone disagreed with it on a test result, and the metric tuning judgements about whether the metric itself is any good. See Annotations.
Next steps
If a metric or provider you need is missing, open an issue on GitHub .