OpenTelemetry Tracing for Production AI Agents: A Step-by-Step Guide

Key Takeaways

  • Set service.name explicitly on every service before you write a single span. The SDK's fallback, unknown_service plus the executable name, can make unrelated services look identical in your backend.
  • Nest spans in the order work actually happens, with the agent span wrapping the model call and the tool span sitting under whichever span issued it, then carry that structure across process and queue boundaries by propagating the traceparent header on every call.
  • Once you run more than one service, export through a batch processor to a Collector rather than straight to the backend. Centralizing there lets you add sampling, redaction, and backend changes without touching agent code.
  • Keep a ParentBased head sampler so sub-agents follow the orchestrator's decision, and add Collector tail sampling once volume nears roughly a thousand traces per second and you need every error trace kept.
  • Change three defaults before production traffic: cap attribute value length, size and monitor the batch queue so dropped spans are visible, and stop relying on head-only sampling to catch failures.

Your orchestrator logs a call to its research sub-agent, and a few hundred milliseconds later the tool service that sub-agent depends on logs an inbound request, but nothing in either log proves the two records belong to the same user session. OpenTelemetry tracing resolves that blind spot by carrying one trace ID from the orchestrator through the sub-agent to the tool span. The whole request then reads as one joined record, with the failed span and its parent in the same view. Building that record starts with how traces, spans, and the context passed between services fit together.

How Tracing Connects a Multi-Agent Request

A trace is the record of one request end to end, and a span is one timed operation inside it, such as the orchestrator's planning call or the model inference. A tool lookup is another span. Each span carries a span ID, start and end times, attributes, and the ID of its parent span. The orchestrator's invoke_agent span is the root. The sub-agent span is its child, and the tool span is a child of whichever span issued the call. Your backend uses the parent IDs to rebuild that tree and draw it as a waterfall.

Spans never leave the process that created them. They go to the exporter and from there to your Collector or backend. What travels to the next service is the SpanContext, a small identity bundle holding the trace ID and span ID, along with trace flags and optional trace state. Over HTTP and gRPC, OpenTelemetry's default propagator serializes it as the W3C traceparent header [1]:

traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01

Read it field by field:

  • version: 00, two hex characters. The value ff is invalid.
  • trace-id: 4bf92f3577b34da6a3ce929d0e0e4736, 32 hex characters identifying the whole request. All zeros is invalid.
  • parent-id: 00f067aa0ba902b7, 16 hex characters identifying the calling span. The receiving service uses it as the parent of its first span.
  • trace-flags: 01, two hex characters. The rightmost bit is the sampled flag; 01 means the caller recorded the trace and 00 means it did not [1].

The receiving service extracts this header, starts its first span under that parent, and injects a fresh traceparent when it calls the next hop.

Prerequisites for Instrumenting Your Agent

Have four things in place before Step 1.

  • Language SDK: Install the OpenTelemetry API and SDK for each service from the Python releases, Node.js releases, Java releases, or Go releases page. Tracing is stable in all four.
  • An OTLP endpoint: Run a Collector reachable from every service (gRPC on 4317 and HTTP on 4318 by default), or use a backend that accepts OTLP directly. Step 7 covers the choice.
  • A trace backend: You need a store that can search by trace ID and draw a waterfall.
  • A service name per service: Decide the service.name for the orchestrator, each sub-agent, and each tool service before writing code. Every span inherits it, and the backend groups by it.

How to Trace a Production AI Agent With OpenTelemetry: Step by Step

You configure OTel tracing for a multi-agent request by starting with a tracer provider and correct resource, then adding nested agent, model, and tool spans. You carry the context across a traceparent hop, export through a batch processor, and choose sampling and Collector settings you can defend on cost. In Step 8 you confirm one request appears as a single trace spanning two services.

Step 1: Configure the Tracer Provider and Tracer

Set service.name before you create a single span, because the resource is attached when the provider is built. If you skip it, the SDK falls back to unknown_service: followed by the process executable name, so every Python service in your system can report as unknown_service:python [2].‍

import os
from opentelemetry import trace
from opentelemetry.sdk.resources import Resource, SERVICE_NAME
from opentelemetry.sdk.trace import TracerProvider

resource = Resource.create({
    SERVICE_NAME: os.environ.get("OTEL_SERVICE_NAME", "orchestrator"),
    "service.namespace": "support-agent",
})
provider = TracerProvider(resource=resource)
trace.set_tracer_provider(provider)
tracer = trace.get_tracer("support_agent.orchestrator")

‍Prefer the environment variable in deployment manifests and keep the code fallback only for local runs. OTEL_SERVICE_NAME takes precedence over a service.name entry in OTEL_RESOURCE_ATTRIBUTES [3], and in the Python SDK attributes passed directly to Resource.create outrank both environment variables. Pick one source of truth per service.

Check the name first when traces look merged or missing. Two services that both report unknown_service:python share the same service identity. The conventions also forbid different names across replicas of a horizontally scaled service [2], because that splits one logical service across several entries.

Step 2: Choose Auto-Instrumentation or Manual Spans

Use auto-instrumentation for the plumbing your agent shares with every other service, including network clients and servers and LLM client libraries. Write manual spans for the agent boundaries no library can see, including which agent ran and the tool and session associated with it. The two work together because your manual code calls trace.get_tracer() against the same global provider. A manual tool span started inside an auto-instrumented HTTP handler nests under that handler's span without extra wiring.

The zero-code entry points by language:

LanguagePackageZero-code entry point
Pythonopentelemetry-distro plus opentelemetry-bootstrap -a installopentelemetry-instrument python agent.py
Node.js@opentelemetry/auto-instrumentations-nodenode --require @opentelemetry/auto-instrumentations-node/register app.js
Javaopentelemetry-javaagent.jarjava -javaagent:path/to/opentelemetry-javaagent.jar -jar myapp.jar
Goopentelemetry-go-compile-instrumentationotelc go build in place of go build

The Python contrib repository labels its instrumentation packages beta and advises against general production use [4], so treat the Python zero-code path as a fast start and expect to add manual spans before you depend on the output. Framework packages fill the remaining gap in the split above. Library instrumentation covers shared plumbing and your manual spans cover agent boundaries, while a framework package covers the framework's own calls. If you build on LangChain or LangGraph, the official opentelemetry-instrumentation-genai-langchain package traces them after LangChainInstrumentor().instrument(). If you build on AWS Strands Agents, it emits OpenTelemetry spans when you install it as strands-agents[otel].

Step 3: Create and Nest Agent, Model, and Tool Spans

Name and attribute your spans with the GenAI semantic conventions so your backend and any evaluator reading the traces later can distinguish the span types [5]. The conventions are still marked Development, which means attribute names can change between releases, but they are the names the ecosystem is converging on. Agent spans are named invoke_agent {gen_ai.agent.name}, model spans {gen_ai.operation.name} {gen_ai.request.model}, and tool spans execute_tool {gen_ai.tool.name} [5].

from opentelemetry import trace
from opentelemetry.trace import Status, StatusCode

tracer = trace.get_tracer("support_agent.orchestrator")

def handle_request(session_id: str, user_message: str) -> dict:
    with tracer.start_as_current_span(
        "invoke_agent orchestrator",
        attributes={
            "gen_ai.operation.name": "invoke_agent",
            "gen_ai.agent.name": "orchestrator",
            "gen_ai.provider.name": "openai",
            "gen_ai.conversation.id": session_id,
        },
    ):
        with tracer.start_as_current_span(
            "chat planner-model",
            kind=trace.SpanKind.CLIENT,
            attributes={
                "gen_ai.operation.name": "chat",
                "gen_ai.provider.name": "openai",
                "gen_ai.request.model": "planner-model",
            },
        ) as model_span:
            plan = call_model(user_message)
            model_span.set_attribute("gen_ai.usage.input_tokens", plan.input_tokens)
            model_span.set_attribute("gen_ai.usage.output_tokens", plan.output_tokens)

        with tracer.start_as_current_span(
            "execute_tool lookup_order",
            attributes={
                "gen_ai.operation.name": "execute_tool",
                "gen_ai.tool.name": "lookup_order",
                "gen_ai.tool.call.id": plan.tool_call_id,
            },
        ) as tool_span:
            try:
                return call_tool_service(plan.arguments)
            except ToolServiceError as exc:
                tool_span.set_attribute("error.type", type(exc).__name__)
                tool_span.set_status(Status(StatusCode.ERROR, str(exc)))
                tool_span.record_exception(exc)
                raise
  • ‍gen_ai.conversation.id on the agent span ties separate traces from the same user session together, so set it on every root span. Token counts belong on the model span as integers. The conventions say gen_ai.usage.input_tokens includes cached tokens, with gen_ai.usage.cache_read.input_tokens reported alongside when the provider exposes it [5].
  • Call set_status as well as record_exception, because record_exception only adds an exception event and leaves the status unchanged. The with block already sets ERROR and records any uncaught exception, so the explicit pair above matters when you catch an error, log it, and still want the span marked failed. Leave status Unset on success; Ok marks success that an application developer or operator has explicitly validated.

Prompt and completion text are opt-in attributes under the conventions. By default, the official GenAI instrumentations capture no message content because it can contain PII [5]. Keep that default until Step 7, where the Collector can redact what you do capture.

Step 4: Propagate traceparent Across HTTP and gRPC

Inject the current SpanContext into outgoing headers and extract it on the receiving side. That keeps the trace joined across a service boundary. The default propagator set is tracecontext,baggage (OTEL_PROPAGATORS), so traceparent is written automatically wherever you call inject [3].

import httpx
from opentelemetry import trace
from opentelemetry.propagate import inject, extract

# Orchestrator side: the client call made inside the execute_tool span
def call_tool_service(arguments: dict) -> dict:
    headers: dict[str, str] = {}
    inject(headers)  # writes traceparent and baggage from the current span
    response = httpx.post(
        "https://tools.internal/lookup_order", json=arguments, headers=headers
    )
    response.raise_for_status()
    return response.json()

# Tool service side: continue the trace from the incoming request
tool_tracer = trace.get_tracer("support_agent.tool_service")

def handle_lookup_order(request) -> dict:
    ctx = extract(request.headers)
    with tool_tracer.start_as_current_span(
        "lookup_order", context=ctx, kind=trace.SpanKind.SERVER
    ):
        return query_orders(request.json())

If you instrument the HTTP library instead, HTTPXClientInstrumentor().instrument() or RequestsInstrumentor().instrument() performs the inject call for you and creates the CLIENT span. For gRPC, GrpcInstrumentorClient().instrument() injects into call metadata and GrpcInstrumentorServer().instrument() extracts it and attaches the context on the server, with GrpcAioInstrumentorClient and GrpcAioInstrumentorServer for grpc.aio.

Three in-process and queued hops need their own handling.

  • asyncio: Tasks copy context at creation, so a sub-agent launched with asyncio.create_task inherits the parent span.
  • Thread pools: Submit work as executor.submit(contextvars.copy_context().run, fn) or install opentelemetry-instrumentation-threading.
  • Queues: Inject into message headers on publish, call extract(message.headers) in the consumer, and use span links for batch consumers, since a span has one parent.

Step 5: Add a Batch Processor and OTLP Exporter

Export through a BatchSpanProcessor feeding an OTLPSpanExporter, and point it at your Collector or backend with OTEL_EXPORTER_OTLP_ENDPOINT. The generic endpoint variable is a base URL to which the SDK appends /v1/traces for OTLP/HTTP; the per-signal OTEL_EXPORTER_OTLP_TRACES_ENDPOINT is used exactly as written [6].

export OTEL_SERVICE_NAME=orchestrator
export OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector:4318
export OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf
export OTEL_EXPORTER_OTLP_HEADERS="api-key=<your-key>"
from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter
from opentelemetry.sdk.trace.export import BatchSpanProcessor

exporter = OTLPSpanExporter()  # reads the OTEL_EXPORTER_OTLP_* variables
processor = BatchSpanProcessor(exporter)
provider.add_span_processor(processor)

The Python SDK and distro pick gRPC when OTEL_EXPORTER_OTLP_PROTOCOL is unset, so a service pointed at the HTTP port 4318 shown above can end up sending gRPC to it. This happens even though the specification defaults to http/protobuf [6]. Setting the protocol explicitly prevents the mismatch.

An agent that fans out to several sub-agents and tools can emit many spans per request, and each span may carry large attributes. That volume puts pressure on the batch queue, so you need to know its limits before production traffic arrives.

The processor defaults are a 2048-span queue, 512 spans per export, a 5-second schedule delay, and a 30-second export timeout [3]. You change them with OTEL_BSP_MAX_QUEUE_SIZE, OTEL_BSP_MAX_EXPORT_BATCH_SIZE, OTEL_BSP_SCHEDULE_DELAY, and OTEL_BSP_EXPORT_TIMEOUT. When the queue is full, the Python SDK logs Queue full, dropping Span., and appending the new span evicts the oldest queued span. The SDK suppresses duplicate warnings within 20-second buckets, so the log alone understates how many spans were lost. To count the drops, set OTEL_PYTHON_SDK_INTERNAL_METRICS_ENABLED=true, which exposes the otel.sdk.processor.span.processed counter with error.type=queue_full.

Step 6: Choose a Sampling Strategy

  • Head sampling: The SDK makes the decision at the root span, which keeps the process cheap. It cannot use the completed trace's error or latency because neither exists when the decision occurs.
  • Tail sampling: A Collector waits for the trace to complete before applying a policy. This lets you keep every error trace while sampling the rest, and high-volume agent systems often combine it with head sampling.

The SDK default is parentbased_always_on, which records everything and lets child services follow the root's decision [3]. To head-sample, set OTEL_TRACES_SAMPLER=parentbased_traceidratio and OTEL_TRACES_SAMPLER_ARG=0.25 on every service. Keep the parentbased variant. The root sampler decides once, and the sampled flag in traceparent carries the decision to your sub-agents and tool services. A plain traceidratio sampler ignores the parent's flag, so services with different ratios produce partial traces.

Head sampling decides at the root span, before an error or latency value exists, so it cannot select traces on either outcome [7]. That matters for agents, because a single failure often hides inside a long, mostly healthy run. Michele Mancioppi's analysis for The New Stack explains that head sampling at single-digit rates tends to miss localized problems [8]. A rare bad tool call is one example, since the trace that contains it is unlikely to be among the few you keep.

Tail sampling answers this by holding spans in the Collector until the trace completes, so the decision can use what actually happened. You can then apply policies such as status_code, latency, and probabilistic [9]:

processors:
  tail_sampling:
    decision_wait: 30s
    num_traces: 50000
    policies:
      - name: keep-slow-traces
        type: latency
        latency:
          threshold_ms: 5000
      - name: keep-errors
        type: status_code
        status_code:
          status_codes: [ERROR]
      - name: baseline
        type: probabilistic
        probabilistic:
          sampling_percentage: 10

Tail sampling is beta and stateful [9]. Because the processor holds spans in memory until it decides, every span of a trace must reach the same Collector instance. If you run multiple replicas, you need the load-balancing exporter in front to route by traceID. The config above sets num_traces: 50000 and decision_wait: 30s, so the buffer holds at most 50,000 traces while each waits up to 30 seconds for a decision. If the Collector evicts a trace from its circular buffer before decision_wait, it drops that trace without a decision [9]. Raising num_traces or lowering decision_wait reduces that risk but increases memory.

The authors of the OpenTelemetry sampling guidance suggest that sampling becomes worth the complexity at roughly a thousand traces per second and is not worth it at tens per second [7]. Below that line, run parentbased_always_on and skip tail sampling. You pay more for storage, but you avoid losing the trace needed to reconstruct a failed session.

Step 7: Decide Whether You Need a Collector

Export directly to the backend while you have one service in development, and add a Collector as soon as you run the orchestrator and tool service as separate deployments. Direct export is simple and has no extra moving parts, but it couples your application code to the backend. A change in redaction, sampling, or destination then becomes a code change in every agent. A Collector moves those decisions into one configuration that your agents never see. It also lets each service hand spans off quickly while the Collector handles retries, batching, and sensitive-data filtering.

Order the processors as the Collector documentation recommends:

  1. memory_limiter
  2. sampling or filtering
  3. context-dependent enrichment
  4. transforms
  5. batch

Merge the Step 6 tail_sampling block into this file before you start the Collector.

receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317
      http:
        endpoint: 0.0.0.0:4318

processors:
  memory_limiter:
    check_interval: 1s
    limit_mib: 4000
    spike_limit_mib: 800
  # paste the tail_sampling block from Step 6 here
  attributes/strip-prompts:
    actions:
      - key: gen_ai.input.messages
        action: delete
  batch: {}

exporters:
  otlp:
    endpoint: tracing-backend:4317

service:
  pipelines:
    traces:
      receivers: [otlp]
      processors: [memory_limiter, tail_sampling, attributes/strip-prompts, batch]
      exporters: [otlp]

The attributes/strip-prompts block in the config above is the Step 7 payoff for the no-content default you kept in Step 3. It deletes gen_ai.input.messages, so prompt text never leaves your network. If you do need to capture content, use the attributes processor's delete and hash actions or the redaction processor's blocked_values patterns to mask captured PII in prompts and tool returns before anything leaves your network. Because the Collector configuration now holds your backend destination, protect it as a secret whenever it contains a backend API key or another credential.

Keeping OTLP between your services and the Collector lets you switch destinations by editing the Collector's exporters block instead of changing the agent code written in Steps 1 through 5. The exporter you choose then determines whether a translation step sits between the Collector and the backend.

Jaeger accepts OTLP natively on 4317 and 4318, so it needs no translation. Tempo also takes OTLP, but it accepts gRPC on 4317 by default and needs its HTTP receiver enabled for 4318. Zipkin does not, so it needs a translation layer such as the zipkin-otel server module or the Collector's Zipkin exporter. Commercial observability backends also ingest OTLP, so they follow the same pattern as Jaeger and Tempo.

Step 8: Verify One Request Across Two Services

Send one request through the orchestrator, capture its trace ID, and confirm the backend shows a single trace containing spans from both service.name values. Log the ID from inside the root span so you can search for it:

ctx = trace.get_current_span().get_span_context()
print(f"trace_id={ctx.trace_id:032x} span_id={ctx.span_id:016x}")

Then test the tool service's extraction on its own by sending the W3C example header directly:

curl -H "traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01" \
  https://tools.internal/lookup_order

Search the backend for 4bf92f3577b34da6a3ce929d0e0e4736 and check four things:

  • The lookup_order SERVER span's parent ID should equal the span ID of the orchestrator's execute_tool INTERNAL span; if it shows no parent, extract ran against the wrong carrier.
  • The trace should list exactly two services; one service means Step 1 is wrong, and three means a replica has a different name.
  • The longest bar tells you whether the model call or the tool call dominates latency.
  • A span with status ERROR should carry the exception event and the error.type you set in Step 3; if the tool service failed but the trace shows Unset, the handler caught the exception without calling set_status.

Defaults to Change Before Production

  • Attribute value length: The default has no length limit and a count limit of 128 [3], so a 50 KB prompt or tool return rides along in full. Set OTEL_ATTRIBUTE_VALUE_LENGTH_LIMIT to a value sized for your largest expected tool return, and keep full message content out of span attributes.
  • Dropped spans on a full queue: Raise OTEL_BSP_MAX_QUEUE_SIZE, lower OTEL_BSP_SCHEDULE_DELAY, and turn on the internal metrics from Step 5.
  • Head-only sampling: A traceidratio sampler below 1.0 with no tail policy discards a fixed share of error traces by construction. Either run at 1.0 or pair head sampling with a status_code tail policy in the Collector.

Tracing Agents With the Fiddler Platform

Fiddler AI Observability and Security platform ingests the OTLP traces produced above. Agentic Observability organizes them across the agentic hierarchy: application, session, agent, trace, and span. You can drill from an application-level alert to the agent and trace, then inspect the LLM or tool span that caused it.

Ingest is OpenTelemetry-native through fiddler-otel for any Python agent, so the spans and service names from the earlier steps feed the agentic hierarchy described above. For the frameworks covered in Step 2, the fiddler-langgraph, fiddler-langchain, and fiddler-strands SDKs add framework-aware spans for LangGraph, LangChain, and AWS Strands Agents.

Once the traces show where a failure happened, the next step is to control and score behavior at that point. Fiddler Guardrails checks the request before the model is invoked and the response before it reaches the user or a downstream system. It allows, blocks, or redacts per policy, with under 80ms response time.

Batteries-included Fiddler Centor Models score prompts and responses in-environment, with no external LLM call and no per-evaluation cost. Continuous Evaluations run the same evaluators before deployment and in production. You can see how trace, evaluation, and policy views line up on the Fiddler platform.

Bringing It Together

After the eight steps you have one trace per request that names the orchestrator and every sub-agent. It also identifies the model calls and tool service, with error status and token counts on the spans that carry them. The sampling and Collector choices in Steps 6 and 7 determine whether the rare failing trace survives for investigation. The next step is to put the traces in front of evaluators, so a faithfulness or tool-selection problem is scored as it happens rather than found during a replay.

Request a demo to review how OTLP traces from your orchestrator and tool services would map onto the agentic hierarchy and what the first-day integration path looks like for your framework.

References

[1] W3C, Trace Context Recommendation https://www.w3.org/TR/trace-context/

[2] OpenTelemetry Semantic Conventions, Service Resource Attributes https://opentelemetry.io/docs/specs/semconv/resource/service/

[3] OpenTelemetry Specification, SDK Environment Variables https://opentelemetry.io/docs/specs/otel/configuration/sdk-environment-variables/

[4] OpenTelemetry Python Contrib Repository https://github.com/open-telemetry/opentelemetry-python-contrib

[5] OpenTelemetry Semantic Conventions for Generative AI Repository https://github.com/open-telemetry/semantic-conventions-genai

[6] OpenTelemetry Specification, OTLP Exporter Configuration https://opentelemetry.io/docs/specs/otel/protocol/exporter/

[7] OpenTelemetry Documentation, Sampling Concepts https://opentelemetry.io/docs/concepts/sampling/

[8] Michele Mancioppi, The New Stack, Distributed Tracing Sampling With OpenTelemetry https://thenewstack.io/distributed-tracing-sampling-opentelemetry/

[9] OpenTelemetry Collector Contrib, Tail Sampling Processor README https://github.com/open-telemetry/opentelemetry-collector-contrib/blob/main/processor/tailsamplingprocessor/README.md

[10] OpenTelemetry Collector, Internal Telemetry https://opentelemetry.io/docs/collector/internal-telemetry/

Frequently Asked Questions

Is OTel the Same Thing as OpenTelemetry?

Yes. OTel is the official short form of OpenTelemetry, the vendor-neutral open source framework for generating, collecting, and exporting traces, metrics, and logs, hosted by the Cloud Native Computing Foundation (CNCF). The two names refer to one project and one set of SDKs.

What Do the Four Fields in a traceparent Header Mean?

The header holds a version and a 32-character trace ID shared by every span in the request. It also contains a 16-character parent ID naming the calling span and trace flags whose low bit records the sampling decision [1].

How Do I Monitor the Collector Itself?

Once spans pass through the Collector, the SDK queue metrics from Step 5 no longer cover losses that happen inside the Collector, so you need to watch the Collector itself. It exposes its own metrics on a Prometheus endpoint, by default http://127.0.0.1:8888/metrics, with the level set under service::telemetry::metrics::level [10]. When otelcol_exporter_queue_size rises toward otelcol_exporter_queue_capacity, the export queue is filling. Sustained otelcol_receiver_refused_* counts indicate possible client data loss. A climbing otelcol_exporter_send_failed_spans_total shows failed sends, which are not necessarily data loss because retries may occur. The signal that marks actual drops is the log line “Dropping data because sending_queue is full” [10], so alert on it. Finally, the health_check extension on localhost:13133 gives you a liveness and readiness endpoint for Kubernetes.

When Should I Export Directly Instead of Running a Collector?

Direct export fits a single service in development or test, where you want traces in minutes and no other service or team depends on the configuration. Once two services share a trace, or you need redaction or tail sampling, move to a Collector.

Should I Use Auto-Instrumentation or Write Spans by Hand?

Use both. Auto-instrumentation covers HTTP, gRPC, and LLM client libraries with no code changes and gives you context propagation for free. Since the Python contrib packages are still labeled beta [4], plan on manual spans at the agent boundaries from the start.