Field note

AI Agent Observability: See Every Trace, Tool Call and Cost

AI agent observability explained: what to trace, the OpenTelemetry GenAI conventions, the metrics that matter, evals vs tracing, privacy, and how teams debug agents in production.

Contents

A long engine control room lined with gauges and indicator panels

Image: Dmitrijs Safrans on Unsplash.

"Your AI agent just took 45 seconds to answer a simple question. Was it the model? A slow tool call? A retry loop?" That is how the OpenTelemetry project opens its guide to GenAI observability, and its answer is blunt: "without observability, you are guessing" (OpenTelemetry, 2026).

Agents are harder to observe than ordinary software because the interesting part, the decision, happens inside a model. The same prompt can take a different path on the next run. A traditional log tells you that a step failed; an agent trace has to tell you why the agent chose that step at all.

Practitioners have noticed. In LangChain's 2026 survey of 1,340 practitioners, 89% had implemented observability for their agents and 62% had detailed tracing of individual steps and tool calls. Among teams with agents in production, those figures rose to 94% and 71.5% (LangChain). Tracing has become table stakes.

What a good agent trace contains

A trace is a tree. The root is the whole run; the children are each model call and each tool call, in order.

Level Capture
Run Goal/prompt, agent version, start/end time, outcome, total tokens and cost
Model call Model name, input and output tokens, latency, finish reason, and optionally the messages
Tool call Tool name, arguments, result or error, latency
Decision points Why the agent stopped: final answer, turn limit, error, or human handoff
Human steps Approval requested, who approved, when

The finish reason is underrated. It tells you whether the model stopped because it was done (stop) or because it wanted a tool (tool_calls). A run that ends on a turn limit is a different failure from one that ends on a wrong final answer.

The emerging standard: OpenTelemetry GenAI conventions

OpenTelemetry, the open standard most backend observability already uses, now has semantic conventions for generative AI. They standardise "how GenAI operations are recorded — the model being called, input and output token counts, and when opted in, the full content of prompts, completions, tool calls, and tool results" (OpenTelemetry, 2026).

A trace built on them has an invoke_agent span at the top, with child chat spans for each model call and execute_tool spans for each tool invocation. Key attributes:

  • gen_ai.request.model: which model ran
  • gen_ai.usage.input_tokens / gen_ai.usage.output_tokens: token counts per call
  • gen_ai.response.finish_reasons: why generation stopped

And two standard metrics: gen_ai.client.operation.duration for latency and gen_ai.client.token.usage for consumption. Coding agents already emit them. The same post notes VS Code Copilot, OpenAI Codex and Claude Code all export OpenTelemetry data. Because it is a standard, you can send it to any OTLP-compatible backend instead of locking into one vendor.

The metrics that matter

  1. Task success rate: judged against an outcome, not "the run didn't crash."
  2. Cost per successful task, including the cost of failed attempts. Anthropic measured agents using about 4× the tokens of a chat, and multi-agent systems about 15× (Anthropic, 2025).
  3. Latency: the second-biggest production barrier after quality (20% of LangChain respondents).
  4. Tool error rate by tool: one flaky tool can derail every run that touches it.
  5. Turns per task: a rising trend often means the agent is flailing.
  6. Human override rate: how often a reviewer changes the agent's output. It is the most honest quality signal you have.

Tracing is not evaluation

Observability tells you what happened. Evaluation tells you whether it was good. You need both, and most teams have more of the first: 89% have observability, but only 52.4% run offline evaluations and 37.3% online ones (LangChain).

Anthropic's approach to closing the gap:

  • Start small, now. About 20 real queries were enough to see the impact of early changes, when "a prompt tweak might boost success rates from 30% to 80%."
  • Judge outcomes, not paths. Agents "might take completely different valid paths to reach their goal," so check the end state rather than whether it followed your expected steps.
  • Use an LLM judge with a rubric: factual accuracy, citation accuracy, completeness, source quality, tool efficiency. A single call scoring 0.0–1.0 with pass/fail was the most consistent with human judgment.
  • Keep humans in the loop. Human testers found that early agents preferred "SEO-optimized content farms over authoritative but less highly-ranked sources," a bias automated evals missed.

How traces change debugging

Anthropic's account is the clearest case study. Users reported agents "not finding obvious information," and the team "couldn't see why. Were the agents using bad search queries? Choosing poor sources? Hitting tool failures? Adding full production tracing let us diagnose why agents failed and fix issues systematically" (Anthropic, 2025).

Traces also improve tools. Anthropic built an agent that used a flawed tool dozens of times and rewrote its description, which led to a 40% decrease in task completion time for later agents using the improved description. That is only possible when you can see every tool call and its failure.

Privacy: observe decisions, not necessarily content

Prompts and tool results often contain personal or confidential data, so capturing them has a cost. OpenTelemetry's defaults reflect this: "By default, no prompt content or tool arguments are captured with GenAI telemetry," only metadata such as model, token counts and durations, and content capture is opt-in (OpenTelemetry, 2026).

Anthropic takes the same line in production: it monitors "agent decision patterns and interaction structures — all without monitoring the contents of individual conversations." A sensible default:

  • Always capture metadata: models, tools called, tokens, latency, errors, finish reasons.
  • Capture content in development and for opted-in or internal workloads.
  • Keep traces where the data is allowed to live. For sensitive work that may mean on the same machine or network.

A practical setup checklist

  • Every run has an ID, a version and an outcome
  • Every model call records model, tokens, latency and finish reason
  • Every tool call records name, arguments, result or error, and latency
  • Cost is computed per run and per successful task
  • A small eval set (≈20 real cases) runs on every prompt or tool change
  • Human overrides are logged and reviewed weekly
  • Content capture is a deliberate, documented choice
  • Traces are exportable in a standard format (OpenTelemetry)

How AGNT records runs

Disclosure: AGNT is our product. Every AGNT run writes an execution receipt to a local SQLite database: the exact prompt that started the run, every tool call with its input, output and any error, timestamps, token usage and cost per step, and the final response stored verbatim. Workflows also record each node's input, output and duration. Traces are full-text searchable, success rate is tracked per agent, and runs can be reconstructed step by step months later. Because receipts stay on your machine, content capture does not mean sending your data to a third-party observability service. See agent verification and the execution and streaming API.

FAQ

What is AI agent observability?

The ability to see what an agent did on each run: every model call, tool call, result, decision, latency and cost, so you can debug failures and measure quality.

What is the difference between LLM observability and agent observability?

LLM observability tracks individual model calls. Agent observability tracks whole runs: sequences of model calls and tool actions, the decisions between them and the outcome.

What should I log for an AI agent?

Run ID, version and outcome; per model call the model, tokens, latency and finish reason; per tool call the name, arguments, result or error and latency; plus cost and any human approvals.

Is there a standard for tracing AI agents?

Yes. The OpenTelemetry semantic conventions for generative AI define standard spans such as invoke_agent, chat and execute_tool, and attributes like gen_ai.usage.input_tokens.

Sources

AI Agent ObservabilityAgent TracingOpenTelemetryLLM ObservabilityEvaluation