Evaluators
Create evaluator
POST
Create evaluator
Creates a new evaluator for your organization. You must specify
Single/Multi Select:
Available LLM config fields (all optional except
For complete details, see the Filters API Reference.
type and score_value_type. The eval_class field is optional and only used for pre-built templates.
Authentication
All endpoints require API key authentication:Evaluator Types and Score Value Types
Evaluator Types (Required)
Important: The evaluator
type field now represents the primary interface/use case, but automation can be added independently via llm_config or code_config. This decouples the annotation method from the evaluator type.llm: Primarily LLM-based evaluators (can also have code automation)human: Primarily human annotation-based (can have LLM or code automation for assistance)code: Primarily code-based evaluators (can also have LLM automation as fallback)
Score Value Types (Required)
numerical: Numeric scores (e.g., 1-5, 0.0-1.0)boolean: True/false or pass/fail evaluationspercentage: 0-100 percentage scores (use decimals; 0.0–100.0)single_select: Choose exactly one option from predefined choicesmulti_select: Choose one or more options from predefined choicesjson: Structured JSON data for complex evaluationstext: Text-based feedback and comments- (Legacy)
categoricalandcommentremain readable for older evaluators
Pre-built Templates (Optional)
You can optionally use pre-built templates by specifyingeval_class:
keywordsai_custom_llm: LLM-based evaluator with standard configurationcustom_code: Code-based evaluator template
Unified Evaluator Inputs
All evaluator runs now receive a single unifiedinputs object. This applies to all evaluator types (llm, human, code).
Structure:
input(any JSON): The request/input to be evaluated.output(any JSON): The response/output being evaluated.metrics(object, optional): System-captured metrics (e.g., tokens, latency, cost).metadata(object, optional): Context and custom properties you pass; also logged.llm_inputandllm_output(string, optional): Legacy convenience aliases.
Required Fields
name(string): Display name for the evaluatortype(string): Evaluator type -"llm","human", or"code"score_value_type(string): Score format -"numerical","boolean","categorical", or"comment"
Optional Fields
evaluator_slug(string): Unique identifier (auto-generated if not provided)description(string): Description of the evaluatoreval_class(string): Pre-built template to use (optional)configurations(object): Custom configuration based on evaluator typecategorical_choices(array): Required whenscore_value_typeis"categorical"
New Format (Recommended)
The new evaluator format uses clean, flat configuration fields instead of nested
configurations. This format allows you to add both LLM and code automation to any evaluator type, decoupling the annotation method from the evaluator type.New Top-Level Fields (All Optional)
Score Config Shapes
Numerical/Percentage:LLM Config
model and evaluator_definition):
- Core:
model,stream - Sampling:
temperature,top_p,max_tokens,max_completion_tokens - Penalties:
frequency_penalty,presence_penalty,stop - Formatting:
response_format,verbosity - Tools:
tools,tool_choice,parallel_tool_calls
Code Config
Passing Conditions
Uses the universal filter format. Example:Legacy Format (Still Supported)
The legacyconfigurations format remains fully functional for backward compatibility.
Configuration Fields by Type
Fortype: "llm" evaluators:
evaluator_definition(string): The evaluation prompt/instruction. Must include{{input}}and{{output}}template variables. Legacy{{llm_input}}and{{llm_output}}are also supported for backward compatibility.scoring_rubric(string): Description of the scoring criteriallm_engine(string): LLM model to use (e.g., “gpt-4o-mini”, “gpt-4o”)model_options(object, optional): LLM parameters like temperature, max_tokensmin_score(number, optional): Minimum possible scoremax_score(number, optional): Maximum possible scorepassing_score(number, optional): Score threshold for passing
type: "code" evaluators:
eval_code_snippet(string): Python code withmain(eval_inputs)function that returns the score
type: "human" evaluators:
- No specific configuration fields required
- Use
categorical_choicesfield whenscore_value_typeis"single_select"or"multi_select"
score_value_type: "single_select" | "multi_select":
categorical_choices(array): List of choice objects withnameandvalueproperties
Examples
New Format Examples
LLM Evaluator with Automation (Numerical)
Human Evaluator with LLM Assistance
This shows how a human evaluator can have LLM automation for suggested scoring, decoupling annotation method from evaluator type.
Code Evaluator (Boolean)
Single Select Evaluator with LLM
Legacy Format Examples
Custom LLM Evaluator (Numerical)
Human Categorical Evaluator
Code-based Boolean Evaluator
LLM Boolean Evaluator
Using Pre-built Template
Response
Status: 201 CreatedError Responses
400 Bad Request
401 Unauthorized
Create evaluator