Skip to main content
This tutorial shows the specification-driven workflow — the recommended way to evaluate your product in Galtea. Instead of manually configuring datasets and metrics, you define specifications (behavioral expectations), and Galtea derives everything else.

Overview

The specification-driven flow works like this:
  1. Define specifications — describe what your product should do, cannot do, and must follow
  2. Generate or link metrics — AI generates judge prompts from your specs, or you link existing metrics
  3. Create datasets from specs — dataset type is auto-derived from the specification

Prerequisites

  • A product with a description (created via dashboard or SDK)
  • A version to evaluate
  • The Galtea SDK installed and configured

Step 1: Define Specifications

Specifications represent testable behavioral expectations. There are three types:
  • Capability — what the product can do (e.g., “Can explain investment concepts”)
  • Inability — what the product cannot do due to hard technical limits (e.g., “Cannot execute transactions”)
  • Policy — rules the product must follow (e.g., “Must refuse personalized investment advice”)
Policy specifications require a dataset_type (ACCURACY, SECURITY, or BEHAVIOR) that determines how the spec is evaluated. Capability specifications always get the BEHAVIOR dataset type, assigned automatically — you never choose it. Inability specs do not have a dataset type.
You can also create specifications from the dashboard with AI assistance, and you are never asked to pick a type up front. Capabilities and policies are created from the product’s Specifications section; inabilities from its Product section. Both are sidebar entries of the open product. When you create one, Fill with AI opens automatically: describe the behavior in a rough note and the AI rewrites it into a properly written description, then classifies it behind the scenes. Complete with AI does the same classification for a description you have already written, without rewriting the text. If the description does not match where you are creating it, the form says so and will not save it. See Specifications for the full flow.
Metrics define how each specification is scored. You have two options: From the dashboard, open your product, pick Specifications in the sidebar, open the dropdown on a specification, and click Generate Metrics. The AI creates judge prompts and evaluation parameters tailored to each spec. See AI Metric Generation for the full workflow.

Option B: Manual Metric Creation and Linking

Create a metric with a custom judge prompt and link it to a specification:

Step 3: Create Datasets from Specifications

Datasets can be created from specifications in two ways:

Option A: AI-Generated Dataset Configurations (Dashboard)

From the dashboard, open your product, pick Specifications in the sidebar, select the Policy and Capability specifications you want to generate datasets for, and use the bulk actions menu to click Create Datasets. The system will suggest dataset configurations — including name, type, variants, strategies, and max test cases — all auto-derived from your specifications. Review each candidate, edit if needed, and save.
The bulk Create Datasets action covers Policy specifications with a SECURITY or BEHAVIOR dataset type, and every Capability specification (always BEHAVIOR). An ACCURACY specification is generated one at a time instead, because its questions and their known-correct answers are built from a knowledge base file you provide: use Setup for evaluation on the specification’s row, which asks you for that file. The system uses the specification’s description as context — for Security datasets it becomes the custom_variant_description, and for Behavior datasets it shapes the scenario generation.
Setup for evaluation does Step 2 and Step 3 at once. Saving a Capability or Policy by hand lands on its page, which opens the dialog on its own when there is something to generate; it stays available afterwards from that specification’s row menu. Saving an Inability never opens it, since an inability needs no dataset and no metric. You set how many test cases to generate, attach the knowledge base file if the specification is ACCURACY, and confirm. It creates whatever that specification is still missing — the dataset, its metrics, or both — then tells you what it did. The dataset’s test cases keep generating in the background. If it generated metrics, click Review to open the review page and accept or edit them there; closing the dialog instead discards them, and you would have to generate them again.

Option B: SDK — Create Dataset with Specification ID

Pass specification_id instead of type — the dataset type and variant are auto-derived:

Step 4: Run the Evaluation

With specifications, their metrics, and their datasets in place, galtea.evaluations.run() runs everything in one call. Pass your agent and it resolves each specification’s datasets and metrics, runs the agent on every test case, and submits the results for scoring — no manual per-test-case loop.
To see which datasets a specification resolved to, use specifications.get_datasets():

Next Steps

Per-Test-Case Control

Drive each test case yourself when you need to choose the output and metrics per case.

AI Metric Generation

Automatically generate metrics from your specifications using AI.