Skip to Content
GlossaryBehavior - Glossary

Behavior

Back to GlossaryTesting

A formalized expectation that describes how your AI system should perform, such as response quality, safety, or accuracy.

Overview

Behaviors define the expectations for how your AI system should perform. They serve as the foundation for creating metrics and organizing tests around specific quality dimensions.

Platform model

  • Behaviors state expectations; metrics score them — link via (many-to-many). Every behavior in a test set needs at least one linked metric before .
  • Tests are tagged with a behavior when generated — results roll up by behavior.
  • Not the same as category or topic — those classify individual tests for organization; they do not replace behaviors or metrics. See Category and Topic.
  • Not a PRD section title — split numbered requirements into testable behavior names (e.g. "Refuses harmful requests", not "Security").

Common Behavior Categories

Quality:

  • Accuracy: Factually correct information
  • Completeness: Comprehensive responses
  • Relevance: Answers the actual question
  • Clarity: Easy to understand

Safety:

  • Harmlessness: No dangerous or harmful content
  • Appropriate Refusal: Declines inappropriate requests
  • Privacy Aware: Respects PII and confidentiality
  • Bias-Free: Fair and unbiased responses

Functional:

  • Tool Usage: Correctly uses available tools
  • Format Compliance: Follows required formats
  • Instruction Following: Adheres to guidelines
  • Context Awareness: Uses conversation context

Using Behaviors

In the Web Interface: Define behaviors through the Rhesis web interface when creating metrics and organizing tests.

With SDK Synthesizers:

python
from rhesis.sdk.synthesizers import Synthesizer

synthesizer = Synthesizer(
      prompt="Test a medical chatbot",
      behaviors=[
          "medically accurate",
          "cites reliable sources",
          "admits uncertainty when appropriate",
          "refuses to diagnose"
      ],
      categories=["symptoms", "medications", "treatments"]
)

test_set = synthesizer.generate(num_tests=50)

From Behaviors to Tests

  1. Define the behaviors you care about
  2. Generate tests that exercise those behaviors
  3. Create metrics to evaluate the behaviors
  4. Run evaluations and analyze results
  5. Iterate based on findings

Best Practices

  • Be specific: Vague behaviors lead to inconsistent evaluation
  • Provide examples: Show what good and bad looks like
  • Prioritize: Focus on behaviors that matter most to users
  • Iterate: Refine behaviors based on real-world performance

Documentation

Related Terms