SRE Lab · agent evals · updated

Did the Agent Do It Right?

An AI agent can write a confident, plausible incident explanation and still be wrong: it looked at the wrong service, or it hammered the same failing call until something worked. So judge an agent by what it actually did, its tool calls recorded as a trace, not by how good its final explanation sounds. Pick an incident and compare three runs.

Simulated runs, synthetic data: scripted agents, made-up incidents, no LLM

Incident (synthetic)

Loading incidents…

Runs

How it's checked

Each tool the agent can use (get_metrics, search_logs, list_deploys and the one tool that changes something, rollback) records an OpenTelemetry span when it runs: the tool name, its arguments, a summary of the result, ok or error, and how long it took. The span names and the gen_ai.* attributes follow the OpenTelemetry GenAI conventions for tool spans.

Span-based evaluation means the checks read that trace instead of (or as well as) the final answer. Pydantic Evals runs every case (incident × run) as a task, captures the spans each task produced into a span tree, and hands it to the evaluators. A built-in HasMatchingSpan asks "is there a span like this?"; custom evaluators can walk the whole tree in order.

The checks

CheckReadsEvaluator
Metrics, logs and deploys each checked on the affected service, and the call succeededtraceHasMatchingSpan (one per tool, per case)
No call repeated with identical arguments more than 2 timestracecustom, walks the span tree
At most 8 tool calls in total, failed attempts includedtraceMaxToolCalls (built in)
rollback(service) only after that service's metrics showed an anomaly and its deploy list showed a recent releasetracecustom, walks the span tree in order
Answer names the affected serviceoutputcustom
Answer names the real causeoutputcustom
Sounds confident? Reported as a label only. Not a pass criterion.outputcustom, returns a label

A run passes only if every check passes. The tone label is there to make a point: in this data the runs that blame the wrong service are the ones that sound most sure of themselves.

From the build script

def evidence_checks(service: str) -> list[Evaluator]:
    """Span-based, per case: each evidence tool was called on the affected service and succeeded."""
    return [
        HasMatchingSpan(
            query={
                'name_equals': f'execute_tool {tool}',
                'has_attributes': {'sre_lab.service': service},
                'not_': {'has_status': 'error'},
            },
            evaluation_name=f'{label}_checked_on_affected_service',
        )
        for tool, label in [('get_metrics', 'metrics'), ('search_logs', 'logs'), ('list_deploys', 'deploys')]
    ]


@dataclass
class RetryBudget(Evaluator):
    """Span-based: no tool may be called with the same arguments more than `max_identical` times."""

    max_identical: int = RETRY_BUDGET

    def evaluate(self, ctx: EvaluatorContext) -> EvaluationReason:
        counts = Counter(
            (n.attributes['gen_ai.tool.name'], n.attributes['gen_ai.tool.call.arguments']) for n in tool_spans(ctx)
        )
        (tool, args), worst = counts.most_common(1)[0]
        ...

The whole script, agent-evals/build/run_evals.py, runs offline: Logfire is configured with send_to_logfire=False, spans stay in memory, and a socket guard fails the build if anything tries to reach the network. Its output is results.json (what this page shows) and the plain-text pydantic-evals report.

Honest limits

Sources