AI Observability
Artificial intelligence (AI) observability collects and connects evidence about how an AI application runs, such as model calls, retrieved context, tool actions, errors, latency and cost. It helps explain the path taken by a request or agent task.
AI ObservabilityAI Evaluation & ReliabilityAI Engineering
An agent says a refund succeeded. Its trace shows that the payment tool timed out even though the model returned a fluent answer. The application produced a response, but the requested action was not confirmed. Inspecting the connected operations helps locate that mismatch.
OpenTelemetry describes a trace as the path of a request through operations, with individual operations recorded as spans. Applied to an AI application, a trace can connect retrieval, model generation and tool calls rather than leaving unrelated logs to be assembled by hand.
This matters because a successful response code is not proof that an AI task succeeded. Evaluation checks the outcome against the goal; observability supplies evidence for diagnosis. Record enough context to investigate failures, while limiting access to private prompts, retrieved records and tool results. In episode 31, Giovanna Carofiglio discusses collecting data for observability and evaluation in agent systems; the refund scenario here is an illustration.
Sources
- OpenTelemetry: Traces — Defines traces and spans for linked operations; this page applies that model to an AI task.
Go deeper
- What is AI observability, and why do you need it in production? AI, decoded · What Is AI Observability
- What's the difference between AI observability, evaluation, and benchmarking? AI, decoded · Observability vs. Evaluation vs. Benchmarking
- OpenTelemetry: Demo docs
Explore a distributed application to see connected traces, metrics and logs in practice.