Datasets

A dataset is the set of test cases that drives an offline evaluation. In the Evaluation SDK, a dataset is either a list of Python dicts, where each dict is an input record and you define the keys to match what your task function expects, or a Studio dataset that the SDK fetches for you.

Record structure

Record structure

Each record is a plain dict. The keys are up to you:

dataset = [
    {"sentence": "Hello, how are you?", "groundtruth": "English"},
    {"sentence": "Bonjour, comment ça va?", "groundtruth": "French"},
    {"sentence": "Hola, ¿cómo estás?", "groundtruth": "Spanish"},
]

In your task and scorer, access the record via ctx.input_record:

from mistralai.evaluations import TaskContext, ScorerContext

async def task(ctx: TaskContext) -> str:
    return ctx.input_record["sentence"]  # access any key you defined

def scorer(ctx: ScorerContext) -> int:
    return 1 if ctx.input_record["groundtruth"].lower() == str(ctx.output).lower() else 0
Use a Studio dataset

Use a Studio dataset

If you curate records in Studio, pass the dataset to evaluation.run() with Dataset, by slug or by ID. The SDK fetches every record before the run starts:

from mistralai.evaluations import Dataset, Evaluation, Evaluator, Project

run = await client.evaluation.run(
    project=Project(name="Language Detection"),
    evaluation=Evaluation(name="Managed Dataset Eval"),
    dataset=Dataset(slug="language-detection-golden-set"),
    task=task,
    evaluators=[Evaluator(name="accuracy", scorer=scorer)],
)
SelectorUse it when
Dataset(slug="language-detection-golden-set")You want a readable, stable script.
Dataset(id="018f879d-20cd-7e9f-a1bc-2f4f08b6f170")You already have the dataset UUID.

Pass exactly one of id and slug. The dataset must already exist: selecting it by slug never creates it.

Each record's payload becomes the input record, as is: ctx.input_record holds the same dict an equivalent inline record would. For the language detection example, use Studio records shaped like this:

payload = {"sentence": "Bonjour, comment ça va?", "groundtruth": "French"}

When importing those records from JSONL, wrap each payload in the Studio Dataset import format:

{"payload":{"sentence":"Bonjour, comment ça va?","groundtruth":"French"},"properties":{}}

Record properties and source information aren't added to the input record. On runs saved to Studio, the run keeps the dataset ID, and each input record keeps the ID of the Studio record it came from.

i
Information

Studio datasets aren't snapshots

Each run reads the records that exist when it starts, up to 10,000. Later edits to the dataset affect future runs only. If records are added or removed while the SDK fetches them, the run fails before it starts instead of running on an inconsistent set.

Dataset is accepted by evaluation.run(). Optimization and Retry failed records take a list of records.

Transform records before the run

Transform records before the run

When your Studio records don't match the shape your task expects, or when you need a list of records, for example to optimize, fetch and transform them yourself.

Fetch the records with the Mistral SDK and adapt each record's payload yourself. The list endpoint is paginated. This helper fetches every page and validates records into a LangItem shape:

from typing import Any, TypedDict

from mistralai.evaluations import Mistral

class LangItem(TypedDict):
    sentence: str
    groundtruth: str

def as_dict(value: Any) -> dict[str, Any]:
    if isinstance(value, dict):
        return value
    if hasattr(value, "model_dump"):
        return value.model_dump(mode="json")
    return dict(value)

def require_string(value: Any, *, record_id: str, field: str) -> str:
    if isinstance(value, str) and value:
        return value
    raise ValueError(f"Dataset record {record_id} is missing {field}")

def to_lang_item(record: Any) -> LangItem:
    payload = as_dict(record.payload)

    return {
        "sentence": require_string(payload.get("sentence"), record_id=record.id, field="payload.sentence"),
        "groundtruth": require_string(
            payload.get("groundtruth"),
            record_id=record.id,
            field="payload.groundtruth",
        ),
    }

async def fetch_eval_dataset(client: Mistral, dataset_id: str) -> list[LangItem]:
    dataset: list[LangItem] = []
    page = 1
    page_size = 100

    while True:
        result = await client.beta.observability.datasets.list_records_async(
            dataset_id=dataset_id,
            page_size=page_size,
            page=page,
        )
        records = result.records.results

        dataset.extend(to_lang_item(record) for record in records)

        if not result.records.next:
            break
        page += 1

    return dataset

Then pass the resulting list to dataset:

import asyncio
import os

from mistralai.evaluations import (
    Evaluation, Evaluator, Mistral, Project, ScorerContext, TaskContext,
)

client = Mistral(api_key=os.environ["MISTRAL_API_KEY"])

async def task(ctx: TaskContext) -> str:
    response = await client.chat.complete_async(
        model="mistral-small-latest",
        messages=[
            {
                "role": "user",
                "content": (
                    "What language is this sentence in? Reply with ONLY the language name. "
                    f"Sentence: {ctx.input_record['sentence']}"
                ),
            }
        ],
    )
    return str(response.choices[0].message.content)

def scorer(ctx: ScorerContext) -> int:
    groundtruth = ctx.input_record["groundtruth"].lower()
    predicted = str(ctx.output).lower()
    return 1 if groundtruth == predicted else 0

async def main():
    dataset_id = os.environ["MISTRAL_DATASET_ID"]
    dataset = await fetch_eval_dataset(client, dataset_id=dataset_id)

    run = await client.evaluation.run(
        project=Project(name="Language Detection"),
        evaluation=Evaluation(name="Managed Dataset Eval"),
        dataset=dataset,
        task=task,
        evaluators=[
            Evaluator(
                name="accuracy",
                description="1 if the detected language matches the groundtruth.",
                scorer=scorer,
            ),
        ],
        metadata={
            "studio_dataset_id": dataset_id,
            "record_count": len(dataset),
        },
    )
    run.show(level="records")

asyncio.run(main())

Validate the fields you need from record.payload, then return the exact dict shape your task and scorers expect.

Type safety with TypedDict

Type safety with TypedDict

Use TypedDict to make record schemas explicit and get IDE autocompletion:

from typing import TypedDict

class LanguageRecord(TypedDict):
    sentence: str
    groundtruth: str

dataset: list[LanguageRecord] = [
    {"sentence": "Hello, how are you?", "groundtruth": "English"},
    {"sentence": "Bonjour, comment ça va?", "groundtruth": "French"},
]
What to include in records

What to include in records

Records can contain anything your task or scorer needs:

Field typePurposeExample
Task inputsWhat the task processesprompt, context, text, question
Ground truthReference output for scoringexpected, groundtruth, reference_answer
MetadataExtra context for scorers or LLM judgescategory, difficulty, grading_guidance

Include ground truth in records when you want to compare the task output against a known-good answer:

dataset = [
    {
        "prompt": "What is the capital of France?",
        "expected": "Paris",
        "difficulty": "easy",
    },
    {
        "prompt": "Explain the difference between precision and recall.",
        "expected": "Precision measures true positives over predicted positives; recall measures true positives over actual positives.",
        "difficulty": "medium",
    },
]

def accuracy_scorer(ctx: ScorerContext) -> int:
    return 1 if ctx.input_record["expected"].lower() in str(ctx.output).lower() else 0
Best practices

Best practices

Keep datasets focused

A dataset built around a single task or capability produces clearer signals than a broad, mixed-topic collection. Maintain separate datasets for distinct evaluation goals (for example, language_detection, qa_factual, code_generation).

Curate for representativeness

  • Include edge cases and failure modes, not easy examples alone.
  • Balance your dataset: if 90% of records are easy cases, the evaluation won't reveal real problems.
  • Remove records where even a human couldn't reliably score the response (ambiguous inputs add noise).

Version your datasets

Freeze your dataset between runs if you want to track performance over time. Even small changes to records can make runs incomparable. Studio datasets are mutable: each run reads their current records. Use meaningful names like qa_baseline_2025_06 rather than test_data.

Ground truth quality matters

Inaccurate or ambiguous ground truth produces noisy scores. If you use an LLM judge (see Evaluators), include a grading_guidance field to give the judge explicit scoring instructions per record.

Organizing in Studio

Organizing in Studio

The Evaluation SDK organizes results using Projects, Evaluations, and Runs in Studio, not by the dataset itself. Pass your dataset to evaluation.run():

from mistralai.evaluations import Evaluation, Project

run = await client.evaluation.run(
    project=Project(name="Language Detection"),
    evaluation=Evaluation(name="Accuracy Eval"),
    dataset=dataset,  # your list of dicts, or a Dataset
    task=task,
    evaluators=[...],
)

Tags and metadata on the run help you trace back which dataset version was used:

run = await client.evaluation.run(
    ...
    tags=["dataset:qa_baseline_2025_06", "model:mistral-small"],
    metadata={"dataset_version": "2025-06", "record_count": len(dataset)},
)
FAQ

FAQ