Story 01 · AI Engineer

Evaluate an AI application before release

Turn a proposed support-assistant change into evaluation evidence and a reviewable governance decision before a separate release system acts.

The situation

A support team has updated its retrieval and prompt configuration. Before shipping, the AI engineer needs clear evidence that the candidate meets the team’s quality criteria.

The question

Does this candidate have enough evidence to support a release review?

The governance workflow

From question to governed outcome

  1. 01

    Select the governed inputs

    Start with the dataset, prompt, model, and candidate configuration that will be evaluated.

  2. 02

    Run the evaluation

    Use the configured evaluation provider to run the candidate and persist item-level results.

  3. 03

    Inspect the result

    Review progress, metrics, failures, and the comparison with the relevant baseline.

  4. 04

    Test the policy

    Simulate the quality policy against the available evidence before relying on it.

  5. 05

    Review the decision

    Inspect the persisted governance outcome, its explanation, evidence graph, lineage, and audit context.

In Studio

The screens behind the story

Evaluation progress and results stay attached to the governed candidate.

A policy evaluates available evidence; it does not manufacture a passing result.

The final outcome is reviewable with the evidence that supported it.

The outcome

The team has a durable, explainable basis for its release review instead of a spreadsheet of scores or an informal approval.

OSS capabilities used

What makes this workflow possible

Evaluation DatasetsEvaluator ProvidersEvaluation Jobs and ResultsPolicy EvaluationGovernance Decisions