VICKY TECH JOURNAL
Vicky Tech Journal
AI Agents

AI Agent Observability: How to Monitor AI Agents in 2026

AI agent observability dashboard showing agent traces, tool calls, and performance metrics.

You built an AI agent. It answers most questions correctly, so you shipped it. Then, three weeks later, someone tells you it's calling the wrong tool half the time, taking 20 seconds to reply, or that your OpenAI bill tripled overnight — and you have no idea why, because all you can see is the final message it sent back.

That gap between "it works" and "I know why it works" is exactly what AI agent observability solves. This guide explains what it actually means, why an agent needs it more than a normal app does, which tools people use in 2026, and then walks you through a complete, working Python example so you can trace, inspect, and debug your own agent by the end of this article.

What Is AI Agent Observability?

AI agent observability is the practice of collecting the traces, logs, metrics, and evaluation results that show everything an agent did while completing a task — not just its final answer.

AI agent observability explained with a simple example

Diagram showing an AI agent trace connecting a user request, model calls, tool calls, and final response.

Imagine a delivery driver. If you only get a text saying "Delivered," that's a chatbot-style answer — you know the outcome, nothing else. Now imagine you get a live map showing every turn the driver took, how long they waited at each stop, and a note explaining why they skipped one address. That second version is observability: the full route, not just the destination.

An AI agent's "route" usually includes several model calls, one or more tool calls (a weather API, a database query, a calculator), and the reasoning steps connecting them. Observability records that entire route as a trace, made up of smaller units called spans — one span per model call, tool call, or logical step.

Observability vs monitoring vs evaluation

Comparison of AI agent tracing, monitoring, and evaluation showing execution details, aggregate metrics, and output quality.

Beginners often mix up three related ideas:

  • Tracing captures one agent run in detail — every span, in order, with inputs and outputs.
  • Monitoring aggregates many runs into dashboard metrics: average latency, error rate, daily cost.
  • Evaluation judges output quality — was the answer correct, safe, and helpful — separately from whether it ran successfully.

You can have a trace showing an agent ran fast and error-free, and still have a wrong answer. Observability tells you what happened; evaluation tells you whether it was good.

Why Do AI Agents Need Observability?

A traditional app follows the same code path every time. An agent doesn't — the same prompt can trigger a different tool, a different number of steps, or a different answer on two separate runs, because the model is deciding what to do at each step. That non-determinism is exactly why agents need deeper visibility than a typical web service, and it's part of what separates what AI agents are from a simple chatbot in the first place.

Debugging incorrect answers and tool calls

Without a trace, "the agent gave a wrong answer" is a dead end. With a trace, you can see exactly which tool it called, what arguments it passed, and what that tool returned — often revealing that the tool call itself was fine, but the agent misread the result (or vice versa).

Detecting loops, latency, and failures

Agents can get stuck calling the same tool repeatedly, waiting on a slow API, or retrying after a silent failure. A trace makes these patterns visible as a timeline instead of a mystery.

Understanding and controlling costs

Every model call inside an agent's run consumes tokens, and a multi-step agent can rack up several calls per user request. Without per-run visibility, a single expensive edge case can hide inside an average that still looks fine.

What Should You Monitor in an AI Agent?

AI agent monitoring dashboard displaying latency, error rate, token usage, cost, and task success.

Traces, spans, and logs

A trace is the full record of one agent run. Spans are its building blocks — each model call, tool call, or sub-step gets its own span, nested to show what triggered what. Logs are timestamped text events that can sit alongside spans for extra context (a validation warning, a retry message).

Tool calls and model calls

For each tool call, record the tool name, the arguments the model passed in, and what came back — including errors. For each model call, record the model name, the prompt, the response, and token counts.

Latency, errors, and task success

Track how long each span took (and the run as a whole), whether any span ended in an error, and — separately — whether the overall task was actually completed successfully.

Token usage and cost

Input tokens, output tokens, and (where applicable) cached tokens all affect cost differently. Multi-step agents can also make calls from sub-agents, so cost tracking needs to account for the whole run, not just the first model call.

Output quality and safety

This needs a dedicated evaluation step — automated checks, an LLM-as-judge, or human review — layered on top of the raw execution data.

MetricMeaningWhy it mattersExample
TraceFull record of one agent runGives you the complete picture, not just the final answerOne user request → one trace
SpanA single step inside a traceLets you isolate exactly which step failed or was slowexecute_tool: get_weather
LatencyTime taken per span / per runSlow steps hurt user experience and may signal a stuck loopA tool call taking 8s instead of 200ms
Error rateShare of spans/runs ending in failureFlags unreliable tools or brittle promptsA tool call failing on unusual input
Token usageInput/output tokens per model callDirectly drives cost and can flag runaway promptsA prompt growing every retry
Task successWhether the goal was actually achievedThe metric users actually care aboutCorrect final answer, delivered once
AI agent observability architecture connecting application telemetry through OpenTelemetry to observability platforms.

Best AI Agent Observability Tools in 2026

It's worth separating the telemetry standard from the platforms built on top of it. OpenTelemetry defines how trace data should be structured; Langfuse, LangSmith, and Arize Phoenix are products that collect, store, and display that data (and, in Langfuse's and Phoenix's case, are themselves built on OpenTelemetry).

OpenTelemetry

OpenTelemetry is the vendor-neutral, open source standard for traces, metrics, and logs, supported by dozens of observability vendors. Its GenAI semantic conventions — maintained in a dedicated repository, since this part of the spec is still evolving — define standard span types for agent work, including create_agent, invoke_agent, and execute_tool, along with attributes like gen_ai.usage.input_tokens and gen_ai.usage.output_tokens. It's worth knowing this standard exists even if you never touch it directly, because most platforms below speak it under the hood.

Langfuse

Langfuse is an open source (MIT core) AI engineering platform covering tracing, prompt management, and evaluation. It's self-hostable for free, or available as a hosted "Langfuse Cloud" with a genuinely usable free Hobby tier (50,000 units/month at the time of writing). Langfuse was acquired by ClickHouse in January 2026, which now backs it with ClickHouse's analytics database. It has native Python and JS SDKs and can also ingest plain OpenTelemetry data.

LangSmith

LangSmith is LangChain's proprietary observability and evaluation platform. Despite the name, it works with many providers and frameworks beyond LangChain, including OpenAI, Anthropic, CrewAI, and the Vercel AI SDK, and it accepts standard OTLP data too. Its free Developer tier and paid Plus tier are hosted-only; self-hosted or hybrid deployment is available only on its custom-priced Enterprise plan.

Arize Phoenix

Arize Phoenix is an open source observability and evaluation tool built directly on OpenTelemetry and the OpenInference instrumentation project. It can run with a single local command for experimentation, or be self-hosted for production. Arize also sells a managed enterprise version, Arize AX, built on the same open standards.

LangfuseLangSmithArize Phoenix
LicenseOpen source (MIT core)ProprietaryOpen source
Self-hostingFree, full core featuresEnterprise plan onlyFree
Free hosted tierYes (Hobby)Yes (Developer)N/A (self-host or Arize AX)
Built on OpenTelemetryYesAccepts OTLPYes
Best fitTeams wanting open-source tracing + evals + prompt management in one placeTeams already on LangChain/LangGraph, or wanting a fully managed platformTeams that want a lightweight, OpenTelemetry-native, self-hostable option

No single tool here is an unqualified "winner" — the right pick depends on whether you need self-hosting, how your budget looks, and which framework you're already using. Pricing and free-tier limits shift regularly, so check each tool's current pricing page before committing.

How to Monitor an AI Agent Using Python: Step-by-Step

This section builds one small, coherent "weather agent" and instruments it end-to-end with Langfuse, since it's free, open source, and has a well-documented Python SDK — a good first stop if you're following our Python roadmap for beginners and want a practical AI project to add to it. If you haven't built an agent before, our guide on how to build an AI agent from scratch is a good companion piece.

Step 1 — Understand the example and prerequisites

You'll need Python 3.10+, pip, a free Langfuse account (Hobby tier or your own self-hosted instance), and an OpenAI API key. Using the OpenAI API will incur a small charge per run — this example makes exactly one short model call, so cost is minimal, but never skip that step.

This example demonstrates one full agent turn: a tool call plus a model call, both captured in a single trace. It does not demonstrate multi-agent handoffs, streaming responses, or production-scale sampling — those need more advanced instrumentation than a beginner tutorial should cover in one pass.
Python weather agent workflow instrumented with Langfuse to trace tool calls and OpenAI model responses.

Step 2 — Set up the Python environment

mkdir weather-agent-observability
cd weather-agent-observability
python -m venv venv
source venv/bin/activate      # Windows: venv\Scripts\activate
pip install langfuse openai python-dotenv

Step 3 — Create a minimal agent or agent-like workflow

Create these files:

weather-agent-observability/
├── .env
├── requirements.txt
└── agent.py

.env — never commit this file or hardcode these values in your code:

LANGFUSE_PUBLIC_KEY=pk-lf-xxxxxxxxxxxxxxxx
LANGFUSE_SECRET_KEY=sk-lf-xxxxxxxxxxxxxxxx
LANGFUSE_BASE_URL=https://cloud.langfuse.com
OPENAI_API_KEY=sk-xxxxxxxxxxxxxxxxxxxxxxxx

requirements.txt:

langfuse
openai
python-dotenv

Step 4 — Add tracing and capture tool calls

agent.py — read the comments, they explain each traced step:

import os
from dotenv import load_dotenv
from langfuse import get_client
from langfuse.openai import openai  # Langfuse's traced drop-in for the OpenAI SDK

load_dotenv()
langfuse = get_client()

# Swap this for whichever current chat model your OpenAI account has access
# to — model names change often, so check OpenAI's model list before running.
MODEL_NAME = "gpt-4o-mini"

# A tiny mock "tool." In a real agent this would call a live weather API.
MOCK_WEATHER_DB = {
    "chennai": "32°C, humid, light rain expected",
    "hyderabad": "29°C, clear skies",
    "mumbai": "31°C, cloudy",
}


def get_weather(city: str) -> str:
    city_key = city.strip().lower()
    if city_key not in MOCK_WEATHER_DB:
        raise ValueError(f"No weather data available for '{city}'")
    return MOCK_WEATHER_DB[city_key]


def run_weather_agent(user_query: str, city: str) -> str:
    with langfuse.start_as_current_observation(
        as_type="span", name="weather-agent-run"
    ) as root_span:
        root_span.update(input={"query": user_query, "city": city})

        # --- Tool call span ---
        with langfuse.start_as_current_observation(
            as_type="span", name="tool-get_weather"
        ) as tool_span:
            tool_span.update(input={"city": city})
            try:
                weather = get_weather(city)
                tool_span.update(output={"weather": weather})
            except ValueError as e:
                tool_span.update(output={"error": str(e)}, level="ERROR")
                raise  # stop the run — the agent has nothing to answer with

        # --- Model call: Langfuse's OpenAI wrapper auto-traces this
        #     as a nested "generation" observation ---
        completion = openai.chat.completions.create(
            name="generate-answer",
            model=MODEL_NAME,
            messages=[
                {"role": "system", "content": "You are a concise weather assistant."},
                {
                    "role": "user",
                    "content": (
                        f"The weather in {city} is: {weather}. "
                        f"Answer the user's question: {user_query}"
                    ),
                },
            ],
        )
        answer = completion.choices[0].message.content
        root_span.update(output={"answer": answer})

    langfuse.flush()  # short-lived scripts must flush before exiting
    return answer


if __name__ == "__main__":
    print(run_weather_agent("Should I carry an umbrella today?", "Chennai"))

Step 5 — Run the example and inspect a trace

python agent.py

Open your Langfuse project and go to Traces. You should see a new trace named weather-agent-run containing two nested items: the tool-get_weather span and a generate-answer generation span with the model's input, output, and token counts attached. This nested structure — one root span containing a tool span and a generation span — is exactly what makes an agent's trace different from a single log line: you can see the tool result and the reasoning that used it, side by side.

AI agent trace timeline showing a root span, nested weather tool span, model generation, and a tool error.

Step 6 — Identify a failure and debug it

Re-run the script with a city that isn't in the mock database:

run_weather_agent("Should I carry an umbrella today?", "Delhi")

This time, get_weather() raises a ValueError. The tool-get_weather span records the error and is marked at ERROR level before the exception stops the run — so in your trace list you'll see a failed run where the tool span, not the model span, is the one flagged. That distinction is the whole point of tracing a failure instead of just logging "something went wrong": it tells you where to look first.

Step 7 — Track latency, token usage, and cost

Because the example uses Langfuse's OpenAI wrapper, token usage (input and output tokens) is captured automatically from the API response, and Langfuse estimates cost using its maintained per-model pricing data. Treat that estimate as a starting point, not a bill — always confirm actual charges against OpenAI's own pricing page, since per-token rates and any pricing tiers can change independent of Langfuse's cached copy of them.

How to Debug Common AI Agent Problems

Incorrect tool selection

Check the trace for which tool the model actually called versus which one the task needed. Often the fix is a clearer tool description, not a code change — the model chooses tools based on how they're described to it.

Repeated tool calls or loops

A trace showing the same tool span appearing many times in one run is the signature of a loop. Usually the model isn't getting confirmation that a step succeeded, so it retries. Check what the tool is actually returning on success.

Slow responses

Compare span durations inside the trace. A slow model call and a slow tool call need completely different fixes (a smaller/faster model versus a faster API or caching), so don't guess — look at which span is actually eating the time.

Unexpected token costs

Look at the gen_ai.usage.input_tokens value on the generation span. A growing input size across retries — a full conversation history being re-sent every time, for example — is a common, invisible-until-you-trace-it cost driver.

Incorrect or low-quality answers

If the trace shows every span succeeded and the answer is still wrong, that's an evaluation problem, not an observability one. Add a scoring step — a code check, an LLM-as-judge, or human review — to your trace so quality gets tracked over time.

AI Agent Observability Best Practices

AI agent observability best practices including sensitive data redaction, alerts, evaluation datasets, and production monitoring.

Protect sensitive data in traces

Prompts, tool arguments, and model outputs can contain real user data — OpenTelemetry's own GenAI conventions flag message-content attributes as "Opt-In" specifically because of this. Redact or mask personal information before it's traced, limit who can access your observability dashboard, and set a data retention period instead of keeping traces forever. Our guide on how to protect your data covers the general principles that apply here too.

Set meaningful alerts

Alert on things that actually need a human — a spike in error rate or cost, not every single failed run. Too many alerts trains people to ignore them.

Create evaluation datasets

Save real (anonymized) traces where the agent got something wrong, and turn them into a small test set you re-run whenever you change a prompt or swap a model. This turns "I think it's better now" into something you can actually verify.

Monitor after deployment

An agent that worked in testing can still drift in production as real users send inputs you didn't anticipate. Observability isn't a one-time setup step — check dashboards on a schedule, not just when something breaks.

Frequently Asked Questions

What is AI agent observability?

It's the practice of capturing traces, logs, metrics, and evaluation scores from an AI agent so you can see every model call, tool call, and decision it made, not just its final answer.

What is the difference between AI agent monitoring and tracing?

Tracing is the detailed record of a single agent run. Monitoring aggregates many runs into dashboard-level metrics like average latency or daily cost.

Which AI agent observability tools have free or self-hosted options?

Langfuse and Arize Phoenix are open source and free to self-host. LangSmith has a free hosted tier but only offers self-hosting on its Enterprise plan.

Can I monitor an AI agent built with Python?

Yes — Python has strong SDK support from Langfuse, LangSmith, Arize Phoenix, and OpenTelemetry itself, as shown in the tutorial above.

How do I track AI agent token usage?

Tools that wrap your model provider's SDK (like Langfuse's OpenAI integration used in this tutorial) capture token counts automatically from each API response.

Does observability detect hallucinations?

No, not on its own. It shows you what happened during execution. Catching incorrect or made-up answers needs a separate evaluation step.

Is OpenTelemetry useful for AI agents?

Yes — it's the open standard most observability platforms are built on, and its GenAI conventions define how agent and tool spans should be structured, even though that part of the spec is still evolving.

Conclusion

AI agent observability comes down to one habit: stop looking only at the final answer, and start looking at the trace that produced it. Once you can see every model call, tool call, error, and token count in one place, debugging stops being guesswork. Start small — instrument one agent the way this tutorial did, get comfortable reading a trace, then add evaluation on top once execution visibility feels routine. If you're building toward this as a broader skill set, our AI/ML engineer roadmap and list of free developer tools are good next stops.

Comments