OpenTelemetry GenAI Semantic Conventions: Tracing LLM Calls and AI Agents
An agent that takes forty seconds to answer a question and costs eleven cents is not a model problem until you can see which of its nine model calls, four tool calls and two retrievals ate the time and the tokens. Most teams discover this the hard way, because the first generation of LLM tooling shipped proprietary trace formats that stop at the vendor’s dashboard. The OpenTelemetry GenAI semantic conventions are the vendor-neutral answer: a shared vocabulary of span names, attributes and metrics so that a chat call, a tool execution and an agent invocation look the same in any backend.
The vocabulary is still moving. As of 6 October 2026 every GenAI signal is at Development stability, the specification has left the main semantic-conventions repository for a dedicated one, and the token metric most tutorials still cite has been replaced. This post reads the current text, shows what a correct trace looks like, gives a runnable Python example and a Collector pipeline, and works through the cost and cardinality consequences.
What this covers: the spec’s current status and location, span and metric design, agent and tool spans, content capture and privacy, a Python example, a Collector configuration, and the trade-offs that decide whether this telemetry is affordable.
Context and Background
Classic application performance monitoring assumes a request is a short, deterministic path through code. An LLM application breaks both assumptions. A single user request fans out into model calls whose latency varies by an order of magnitude, tool calls chosen at runtime by the model, retrievals against vector stores, and sometimes recursive calls into other agents. The unit of cost is the token, not the CPU second, and the most useful debugging artifact, the prompt and the completion, is also the most sensitive and the largest.
OpenTelemetry (OTel) is the Cloud Native Computing Foundation project that standardizes traces, metrics and logs, and its semantic conventions define the attribute names that give those signals meaning. The GenAI conventions sit on top of that: they say that a model call is a span named chat gpt-4o-mini with gen_ai.operation.name set to chat, that token counts go in gen_ai.usage.input_tokens, and that the prompt text is opt-in. Because the names are shared, an instrumentation library for the OpenAI SDK and one for Anthropic emit comparable data, and a backend can build one latency-by-model view for both.
Before this work, teams chose between vendor SDK wrappers, framework-specific callbacks and ad hoc logging. Each produced its own schema. If you have read our piece on eBPF observability and kernel tracing for APM, the contrast is useful: eBPF observes from below the application with no cooperation, while GenAI telemetry is semantic and has to be emitted by code that understands what a prompt, a tool call and a finish reason are. The two are complementary, since kernel-level data tells you the pod was throttled and GenAI spans tell you the third model call was the slow one.
The stakes grew when agents arrived. Work such as Claude computer use and desktop agent architecture and dynamic skill loading in Claude agents produces loops of perception, planning and action where one user-visible answer is dozens of operations. Without a trace tree you cannot answer basic questions such as how many inference calls a run made, or which tool failed before the model gave up. The official project documentation lives at the OpenTelemetry site, and the normative text is now maintained in the semantic-conventions-genai repository.
Where the specification lives and what its status means
The first fact to internalize about the OpenTelemetry GenAI semantic conventions is a change of address. The pages under opentelemetry.io/docs/specs/semconv/gen-ai/ now carry a notice that the GenAI conventions have moved to the dedicated semantic-conventions-genai repository, and the same notice sits on the old pages in the main repository. The new repository’s model/manifest.yaml declares the registry as semantic-conventions-genai, with stability development, a schema URL of https://opentelemetry.io/schemas/gen-ai-dev/1.42.0-dev, and a dependency on the core semantic conventions at version 1.44.0. Its documentation is generated from YAML models with the Weaver tool, and the repository was receiving commits as recently as 5 October 2026.
Development stability has concrete meaning. Attribute names, metric names, and requirement levels may change in any release, and in the last year some did. Every attribute table in the current spec marks the gen_ai.* keys as Development, while the borrowed general keys such as error.type, server.address and server.port are marked Stable. That split is a useful mental model: the plumbing is stable, the GenAI vocabulary is not. You should pin instrumentation versions, and you should write dashboards against a thin layer of recording rules or views that you can retarget when a name changes.
The Reference Architecture for GenAI Telemetry
The architecture has four layers: the application and agent framework, a GenAI instrumentation library that applies the conventions, the OTel SDK that batches and exports, and a Collector that governs what leaves your network. The conventions are the contract between the second and third layers and between the pipeline and every backend.

Figure 1: Reference architecture for GenAI telemetry, with the semantic conventions as the contract between instrumentation and SDK.
Figure 1 shows the data path. Instrumentation wraps the model SDK or the agent framework and creates spans and metrics following the gen_ai.* names. The SDK exports over OTLP, the OpenTelemetry Protocol. The Collector is where you enforce policy, which for GenAI means redacting content, sampling traces and shaping metrics before they reach storage that bills by volume.
Direct answer: what the conventions define
The conventions define six things: model-call spans (inference, embeddings, retrieval, memory, fetch response), agent spans (create agent, invoke agent, invoke workflow, plan), a tool execution span, client and server metrics for duration and tokens, an opt-in event for full request details, and provider-specific rules for Anthropic, OpenAI, AWS Bedrock and Azure AI Inference, plus the Model Context Protocol. Content such as prompts is opt-in everywhere.
Spans describe logical operations as the caller sees them
The spec states that GenAI spans represent logical operations as observed by the caller. A span should cover the operation from initiation until the response is fully received or the operation is terminated by an error or cancellation. If a transient failure triggered an automatic retry, the span should cover the logical operation with all retries.
That retry rule matters more than it looks. Many SDKs retry silently on rate limits. If you instrument at the HTTP layer you see three attempts and a confusing waterfall; if you instrument at the logical operation you see one span with a duration that includes backoff, which is what the user experienced. The trade-off is that the span hides how many attempts happened, so an instrumentation that wants that detail has to add its own attribute or child spans.
The inference span
The inference span is the workhorse. Its name should be {gen_ai.operation.name} {gen_ai.request.model}, so chat claude-sonnet-4 or generate_content gemini-2.5-flash style names, and its kind should be CLIENT. The spec allows INTERNAL for models that run in the same process as the caller, and recommends CLIENT when the system normally runs in a separate process or the call crosses an instrumented protocol such as HTTP.
Two attributes are Required: gen_ai.operation.name and gen_ai.provider.name. The operation name has well-known values in the registry, among them chat, generate_content, text_completion, embeddings, retrieval, execute_tool, invoke_agent, invoke_workflow, create_agent, plan, fetch_response and a family of memory operations such as search_memory and upsert_memory. The provider name is a discriminator for the telemetry flavor: well-known values include openai, anthropic, aws.bedrock, azure.ai.openai, gcp.gemini, gcp.vertex_ai, gcp.gen_ai, mistral_ai, cohere and x_ai. The spec says it should reflect the instrumentation’s best knowledge and may differ from the real upstream, for example when a client talks to a proxy that relays to another vendor.
Older tutorials and older library versions used gen_ai.system for this role. The current registry uses gen_ai.provider.name, and the deprecated OpenAI instrumentation package still documents a gen_ai.system column. If a backend shows both, you are looking at mixed library vintages.
The conditionally required attributes include gen_ai.request.model, error.type when the call failed, gen_ai.conversation.id when the framework has one, gen_ai.output.type when the request asks for a format such as json, and server.port when server.address is set. Recommended attributes carry the sampling parameters (gen_ai.request.temperature, top_p, max_tokens, stop_sequences), the response (gen_ai.response.id, gen_ai.response.model, gen_ai.response.finish_reasons) and usage.
One detail deserves attention. The spec says that for gen_ai.conversation.id, instrumentation should not invent a value when none exists: a new UUID, a trace ID or a hash of the request content is explicitly named as a bad fallback. The reasoning is that a fabricated ID looks like a real session in a backend and corrupts any per-conversation analysis.
Usage attributes grew a modality and cache vocabulary
Token usage on spans is no longer just input and output. The current text defines gen_ai.usage.input_tokens and gen_ai.usage.output_tokens, plus gen_ai.usage.cache_read.input_tokens, gen_ai.usage.cache_write.input_tokens and gen_ai.usage.reasoning.output_tokens, and per-modality variants for text, image and audio such as gen_ai.usage.image.input_tokens. This tracks how providers actually bill: a cached prefix costs less than fresh input, and reasoning tokens are billed as output but are invisible in the completion text.
The practical point is that input token counts include cached tokens. The metric description says so explicitly for the counters, and the cache attributes are subsets. If you compute cost as input times price you will overcharge cached prompts unless you subtract the cache-read portion and price it separately. Our context-engineering work, covered in context engineering for production LLM agents, argues that prompt prefix stability is the largest cost lever in long-running agents, and these attributes are what let you measure whether your prefix actually hits the cache.
Agent Tracing with OpenTelemetry GenAI Semantic Conventions
A single agent run is a tree, and the conventions give each node a type. The tree has a workflow at the top when the application has a user-facing entry point, one or more agent invocations beneath it, optional plan spans, and under those the leaf operations: model calls, tool executions and retrievals. Getting the nesting right is what turns a flat list of spans into an explanation.

Figure 2: A span tree for a support agent that calls a knowledge-base tool, then delegates a refund to a second agent.
Figure 2 is a worked example. The root is invoke_workflow, the triage agent runs as invoke_agent triage, and its children are a plan span, a first model call that decides to use a tool, the tool execution with a retrieval nested beneath it, a second model call that reads the result, and a delegation to a refund agent that makes its own model call and tool call. Reading this tree answers questions a log file cannot, such as whether the refund agent was invoked before or after the knowledge-base lookup completed.
The span types and their names
Agent spans extend and override the model-span conventions. The agent-span document defines these.
| Operation | gen_ai.operation.name |
Span name format | Span kind |
|---|---|---|---|
| Create an agent (remote services) | create_agent |
create_agent {gen_ai.agent.name} |
CLIENT |
| Invoke a remote agent | invoke_agent |
invoke_agent {gen_ai.agent.name} |
CLIENT |
| Invoke an in-process agent | invoke_agent |
invoke_agent {gen_ai.agent.name} |
INTERNAL |
| Invoke a workflow | invoke_workflow |
workflow-specific, uses gen_ai.workflow.name |
not asserted here |
| Plan or task decomposition | plan |
plan {gen_ai.agent.name} |
not asserted here |
| Execute a tool | execute_tool |
execute_tool {gen_ai.tool.name} |
INTERNAL |
| Retrieval | retrieval |
retrieval {gen_ai.data_source.id} |
CLIENT |
When the agent name is not available, the invoke and plan names fall back to the bare operation name. Where I wrote “not asserted here”, the text I read for this post did not pin a kind for that span in the part I verified, so check the current document before encoding a rule.
The split between a CLIENT and an INTERNAL agent span is the one people miss. A remote agent, for example an agent hosted by a cloud service that you call over the network, is a client span, and the service on the other side may produce its own server span if it is instrumented. An agent that is just a function in your process is INTERNAL. Choosing wrongly does not break anything, but it distorts service maps, which draw edges between services for CLIENT and SERVER pairs.
Agent identity and the main agent entity
The agent attributes are gen_ai.agent.id, gen_ai.agent.name, gen_ai.agent.description and gen_ai.agent.version. For hosted agents the spec says the ID should be the provider-assigned stable identifier of the agent resource, with an AWS Bedrock agent ARN given as an example. The version attribute accepts schemes such as 1.0.0 or a date, which gives you a way to correlate behavior changes with agent releases, the same way you would tag a microservice version.
The registry also defines a gen_ai.main_agent entity with matching id, name and description fields, and the tool and agent spans are expected to be associated with it. The idea is that in a multi-agent system you want to know which top-level agent a deep span belongs to, even when a sub-agent made the call. That is an entity on the resource side rather than another span attribute, which keeps the attribute cardinality of individual spans lower.
Tool execution spans
The tool span is where agent debugging usually pays off. Its name is execute_tool {gen_ai.tool.name}, its kind is INTERNAL, and gen_ai.operation.name is execute_tool. Required attributes are the operation name and gen_ai.tool.name. Recommended attributes include gen_ai.tool.call.id, which should match the identifier the model assigned to the call, gen_ai.tool.description and gen_ai.tool.type, with examples function, extension and datastore.
Two attributes are opt-in: gen_ai.tool.call.arguments and gen_ai.tool.call.result. Both can carry user data, since a tool such as a flight search takes a traveler’s location and dates, and a database tool returns rows. They follow the same opt-in discipline as prompts, which is covered in the content-capture section below.
The spec is also clear about who instruments tools. GenAI instrumentation that can wrap tool calls should do so, unless another instrumentation reliably covers all tool types, and application developers are encouraged to hand-instrument tools their own code invokes. If a tool is a Model Context Protocol call, the MCP convention applies: its client span is named {mcp.method.name} {target}, where the target matches the tool name or prompt name when one exists, and the kind is CLIENT. The MCP document asks that gen_ai.operation.name also be set to execute_tool for tool calls so that consumers can treat MCP tool calls like other tools. It also warns instrumentations not to record two different spans for one call, so if your framework instrumentation and an MCP instrumentation both see the call, one of them should defer.
Agent skills are specialized tools
The agent-span document now includes refinements for Agent Skills: a load-skill span, a read-skill-resource span and a command-execution span, all as specializations of the execute-tool span. Instrumentations are told to distinguish generic tools from skills using framework tool names or heuristics and record the applicable refinement, again without creating two spans for one call. For teams building on the pattern described in Claude skills architecture, this means progressive disclosure becomes visible in traces: you can see which skills were loaded, in what order, and how many tokens the loading cost.
Plan spans and the limits of agent tracing
The plan span, with gen_ai.operation.name set to plan, represents task decomposition. The spec says model calls and tool spans produced by the plan should be children of the plan span, and that instrumentations should report it when they can reliably determine the planning phase. The qualifier matters. Frameworks differ in whether planning is an explicit step or an emergent property of repeated model calls, and a trace that invents a plan boundary is worse than one that omits it.
Agent tracing has a hard limit that the conventions cannot remove: spans show what happened, not why. A model chose the wrong tool because of the prompt, the tool descriptions and the context, none of which are in the span unless you opted in to capture them. The attribute gen_ai.tool.definitions is opt-in and the spec says it is not recommended to populate its non-required properties by default because it can be large. The instrumentation can therefore tell you that search_kb was chosen but, by default, not what the model saw when it chose.
Metrics: Duration, Tokens and the Counter Redesign
Spans give you per-request detail, and the OpenTelemetry GenAI semantic conventions define metrics that give you the rates and distributions you alert on. The GenAI metrics are small in number and worth knowing exactly, because the token metrics changed shape this year.
Latency metrics
The client metrics are histograms in seconds. gen_ai.client.operation.duration records operation duration with the Required gen_ai.operation.name attribute and the usual provider, model, error and server attributes. gen_ai.client.operation.time_to_first_chunk and gen_ai.client.operation.time_per_output_chunk cover streaming. Each has recommended explicit bucket boundaries from 0.01 seconds doubling to 81.92 seconds, which matches the long tail of LLM latency far better than the default HTTP-oriented buckets.
On the serving side, gen_ai.server.request.duration, gen_ai.server.time_to_first_token and gen_ai.server.time_per_output_token are for people running model servers rather than calling them. Agent-level metrics are new: gen_ai.invoke_agent.duration, plus gen_ai.invoke_agent.inference_calls and gen_ai.invoke_agent.tool_calls which are histograms of how many model calls and tool calls each agent invocation made, and gen_ai.execute_tool.duration for tools. The workflow metric gen_ai.invoke_workflow.duration uses bucket boundaries from 1 second to 7200 seconds, a nod to long-running workflows.
The two counts-per-invocation histograms are underrated. A rising p95 of gen_ai.invoke_agent.inference_calls is the earliest signal that an agent is looping, since the average run can look healthy while a tail of runs takes thirty iterations. You can alert on it without reading a single prompt.
The token metrics changed
If you learned GenAI metrics from a 2025 tutorial, you learned gen_ai.client.token.usage: one histogram with a Required gen_ai.token.type attribute whose values were input and output. I confirmed that this metric appears in the main semantic-conventions repository at versions 1.37.0 and 1.40.0, and that it does not appear in the current text in the dedicated repository.
The replacement splits the job in two. Counters named gen_ai.client.inference.usage.input_tokens, gen_ai.client.inference.usage.output_tokens, gen_ai.client.inference.usage.cache_read.input_tokens, gen_ai.client.inference.usage.cache_write.input_tokens and gen_ai.client.inference.usage.reasoning.output_tokens measure cumulative consumption, broken down by a gen_ai.token.modality attribute whose values are text, image, audio and unknown. Separate histograms, gen_ai.client.inference.operation.input_tokens and gen_ai.client.inference.operation.output_tokens, describe the distribution of tokens per operation and carry no modality breakdown.
The design note in the repository explains why, with an example worth reproducing. If one histogram records a measurement per modality, a workload of 100 identical operations with 100 text tokens and 200 image tokens each produces 200 measurements, half at 100 and half at 200. A 95th percentile query returns 200, but every actual operation consumed 300 tokens. No operation was ever that size. Percentiles over a partitioned measurement are meaningless, so per-operation distributions had to lose the modality dimension, and totals had to move to counters where each token lands in exactly one bucket and sums stay valid.
Cache and reasoning tokens became separate counters instead of attributes for a related reason. Most providers do not report how many cached tokens were text versus image, so a combined attribute would force instrumentation to guess. The cache_read, cache_write and reasoning counters report subsets of the input and output counters, and the cache-write counter can carry a provider-specific TTL attribute because Anthropic prices 5-minute and 1-hour cache writes differently. The commit that made the change landed on 22 September 2026, under the message that it fixes usage metrics to provide meaningful aggregation and break down by modality, reasoning and cache usage.
The consequence for you is a migration. Dashboards that sum gen_ai.client.token.usage by gen_ai.token.type will go quiet when an instrumentation library upgrades to the new names. I could not verify from the library changelogs which released versions emit which metric names, so treat that as something to test on a staging collector before you upgrade production.
Content Capture, Privacy and Hand-Instrumenting the GenAI Conventions
The most valuable telemetry in an LLM system is the text, and it is also the telemetry most likely to violate a policy. The conventions take a clear position: model instructions, user messages and model outputs are sensitive and often large, so instrumentations should not capture them by default, and should offer an opt-in.

Figure 3: The content capture decision, from the default of nothing through span attributes and events to external storage.
Figure 3 summarizes the three usage patterns the spec describes. The default is to record nothing. The second pattern records gen_ai.system_instructions, gen_ai.input.messages and gen_ai.output.messages on the spans, which the spec positions for situations where telemetry volume is manageable and privacy rules either do not apply or the storage complies with them, with pre-production named as the example. The third stores content externally and records references on the span, which the spec recommends for production where volume and sensitivity are concerns, because external storage allows separate access controls.
Structured content and the span attribute caveat
The input and output attributes follow JSON schemas published in the repository, with messages made of a role and a list of parts, such as text, tool calls and tool call responses. On log events they must be recorded in structured form. On spans they may be recorded as a JSON string where structured attributes are unsupported, which is the common case today because OpenTelemetry attribute values have been limited to primitives and arrays of primitives. The spec points to an OpenTelemetry enhancement proposal on extending attributes to support complex values, so this limitation is expected to relax.
Practically, that means a span-attribute capture gives you a JSON blob in a string field. It is searchable in some backends and opaque in others, and it counts against any attribute-length limit. The spec acknowledges this, noting that content may be larger than backend limits for telemetry envelopes or attribute values, and it permits instrumentations to offer truncation that preserves JSON structure. Truncation protects the backend but removes exactly the tail of a long prompt where tool results usually sit, so decide per environment.
The event option and the upload hook
The details event is named gen_ai.client.inference.operation.details. It is Opt-In, describes the request including chat history and parameters, and the spec says instrumentations should set its severity to DEBUG, severity number 5. A second event, gen_ai.evaluation.result, exists for recording evaluation outcomes. Events are logs under the hood, so you can route them through a different pipeline and a different retention policy than traces, which is the main reason to prefer them to span attributes in production.
The upload hook is the most operationally interesting. The spec says an instrumentation may support in-process hooks that receive the instructions, inputs and outputs objects, in the formats defined by the conventions, before they are serialized, together with the span. The hook is independent of the capture flags, is invoked regardless of the sampling decision, and can modify the content or enrich the span. The spec leaves one item open: a TODO notes that a common approach for recording references to externally stored content is not yet documented, so reference attributes are implementation specific today.
In the new Python packages the control surface is concrete. The opentelemetry-instrumentation-genai-openai package, at version 1.2b0 released 24 September 2026, reads the environment variable OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT with the values span_only, event_only, span_and_event and no_content, the last being the default. It also supports OTEL_INSTRUMENTATION_GENAI_COMPLETION_HOOK=upload together with OTEL_INSTRUMENTATION_GENAI_UPLOAD_BASE_PATH pointing to an fsspec-compatible location such as a local path or a gs:// bucket, with the upload extra of opentelemetry-util-genai installed. The package states that it uses the latest experimental conventions unconditionally, with no opt-in variable. The older OTEL_SEMCONV_STABILITY_OPT_IN=gen_ai_latest_experimental mechanism still appears on the OpenTelemetry website’s GenAI page, so you may meet it in other stacks.
Package status is also worth recording. The README of the earlier opentelemetry-instrumentation-openai-v2 package, last released as 2.4b0 on 1 May 2026, now says it is deprecated, only receives security patches, and points to the genai-openai package, warning that the replacement contains breaking changes. An Anthropic package, opentelemetry-instrumentation-genai-anthropic, is published at the same 1.2b0 version. Both depend on opentelemetry-util-genai and opentelemetry-api~=1.43, and all are beta releases.
A runnable Python example
Library instrumentation is the right default, but hand-instrumenting once teaches you the model, and you need it for your own tools anyway. The example below uses only the OpenTelemetry API and SDK. It builds the exact tree in Figure 2 in miniature, with a stub in place of a real model call so it runs without credentials, and it records the new counter-style token metrics. Install opentelemetry-sdk and opentelemetry-exporter-otlp; the latest SDK release at the time of writing is 1.45.0.
import os, time, json, random
from opentelemetry import trace, metrics
from opentelemetry.trace import SpanKind, Status, StatusCode
from opentelemetry.sdk.resources import Resource
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.sdk.metrics import MeterProvider
from opentelemetry.sdk.metrics.export import PeriodicExportingMetricReader
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
from opentelemetry.exporter.otlp.proto.grpc.metric_exporter import OTLPMetricExporter
res = Resource.create({"service.name": "support-agent", "service.version": "0.3.1"})
tp = TracerProvider(resource=res)
tp.add_span_processor(BatchSpanProcessor(OTLPSpanExporter())) # OTEL_EXPORTER_OTLP_ENDPOINT
trace.set_tracer_provider(tp)
mp = MeterProvider(resource=res, metric_readers=[
PeriodicExportingMetricReader(OTLPMetricExporter(), export_interval_millis=15000)])
metrics.set_meter_provider(mp)
tracer = trace.get_tracer("support-agent.genai")
meter = metrics.get_meter("support-agent.genai")
# Metric names follow the current GenAI registry (Development stability).
duration = meter.create_histogram("gen_ai.client.operation.duration", unit="s",
description="GenAI operation duration")
in_tokens = meter.create_counter("gen_ai.client.inference.usage.input_tokens", unit="{token}")
out_tokens = meter.create_counter("gen_ai.client.inference.usage.output_tokens", unit="{token}")
inf_calls = meter.create_histogram("gen_ai.invoke_agent.inference_calls", unit="{inference_call}")
CAPTURE = os.getenv("OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT", "no_content")
def fake_model(model, messages):
"""Stand-in for a provider SDK call. Returns text and usage numbers."""
time.sleep(random.uniform(0.05, 0.2))
return {"id": "resp_demo", "model": model, "text": "ok", "finish": "stop",
"in": 420, "out": 85}
def chat(model, messages, agent_name):
attrs = {"gen_ai.operation.name": "chat", "gen_ai.provider.name": "openai",
"gen_ai.request.model": model, "gen_ai.request.temperature": 0.2,
"gen_ai.request.max_tokens": 512, "gen_ai.agent.name": agent_name}
t0 = time.perf_counter()
with tracer.start_as_current_span(f"chat {model}", kind=SpanKind.CLIENT,
attributes=attrs) as span:
try:
r = fake_model(model, messages)
except Exception as exc:
span.set_attribute("error.type", type(exc).__name__)
span.set_status(Status(StatusCode.ERROR))
raise
span.set_attribute("gen_ai.response.id", r["id"])
span.set_attribute("gen_ai.response.model", r["model"])
span.set_attribute("gen_ai.response.finish_reasons", [r["finish"]])
span.set_attribute("gen_ai.usage.input_tokens", r["in"])
span.set_attribute("gen_ai.usage.output_tokens", r["out"])
if CAPTURE in ("span_only", "span_and_event"): # opt-in, JSON string on spans
span.set_attribute("gen_ai.input.messages", json.dumps(
[{"role": "user", "parts": [{"type": "text", "content": messages[-1]}]}]))
common = {"gen_ai.operation.name": "chat", "gen_ai.provider.name": "openai",
"gen_ai.request.model": model}
duration.record(time.perf_counter() - t0, common)
in_tokens.add(r["in"], {**common, "gen_ai.token.modality": "text"})
out_tokens.add(r["out"], {**common, "gen_ai.token.modality": "text"})
return r
def run_tool(name, call_id, fn):
with tracer.start_as_current_span(f"execute_tool {name}", kind=SpanKind.INTERNAL,
attributes={"gen_ai.operation.name": "execute_tool", "gen_ai.tool.name": name,
"gen_ai.tool.type": "function", "gen_ai.tool.call.id": call_id}):
return fn()
def invoke_agent(name, question):
calls = 0
with tracer.start_as_current_span(f"invoke_agent {name}", kind=SpanKind.INTERNAL,
attributes={"gen_ai.operation.name": "invoke_agent", "gen_ai.provider.name": "openai",
"gen_ai.agent.name": name, "gen_ai.agent.version": "0.3.1"}):
chat("gpt-4o-mini", [question], name); calls += 1
run_tool("search_kb", "call_001", lambda: time.sleep(0.1))
chat("gpt-4o-mini", [question, "tool result"], name); calls += 1
inf_calls.record(calls, {"gen_ai.agent.name": name})
with tracer.start_as_current_span("invoke_workflow support_flow",
attributes={"gen_ai.operation.name": "invoke_workflow",
"gen_ai.workflow.name": "support_flow"}):
invoke_agent("triage", "Where is my refund?")
tp.shutdown(); mp.shutdown()
Run it with OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4317 pointing at a Collector. The span names, kinds and attribute keys match the current text, with two deliberate simplifications. First, I set the span name for the workflow to include the workflow name, because the workflow span name format was one of the items I did not pin down from the spec; confirm it against the live document. Second, the example marks span-attribute capture with the same environment variable the Python packages use, so switching to a real instrumentation library does not change your operating procedure.
Notice what the example does not do. It never puts conversation.id on a metric, it records modality only on the counters, and it keeps content off by default. Those are the three habits that keep GenAI telemetry cheap and safe, and the following sections explain why.
Collector Design: Redaction, Sampling and Cost Control
The Collector is where telemetry built on the OpenTelemetry GenAI semantic conventions becomes governable. SDK-side configuration is per service and per language, while a Collector pipeline is one file that enforces the same policy over every service that sends through it. As of this writing the Collector release train is at v0.162.0, published on 29 September 2026.

Figure 4: A Collector pipeline that strips content, tail-samples traces and batches both traces and metrics.
Figure 4 shows the shape. The trace pipeline protects the Collector from overload first, removes sensitive content second, decides which traces to keep third, and batches last. The order is deliberate. Tail sampling holds whole traces in memory until a decision is made, so you want to strip large attributes before they occupy that memory, and you want the memory limiter ahead of everything.
receivers:
otlp:
protocols:
grpc: { endpoint: 0.0.0.0:4317 }
http: { endpoint: 0.0.0.0:4318 }
processors:
memory_limiter:
check_interval: 1s
limit_percentage: 75
spike_limit_percentage: 20
# Production policy: content never leaves the trust boundary on spans.
transform/drop_genai_content:
error_mode: ignore
trace_statements:
- context: span
statements:
- delete_key(attributes, "gen_ai.input.messages")
- delete_key(attributes, "gen_ai.output.messages")
- delete_key(attributes, "gen_ai.system_instructions")
- delete_key(attributes, "gen_ai.tool.definitions")
- delete_key(attributes, "gen_ai.tool.call.arguments")
- delete_key(attributes, "gen_ai.tool.call.result")
tail_sampling:
decision_wait: 30s # agent runs are long; wait for the root to finish
num_traces: 50000
policies:
- name: keep-errors
type: status_code
status_code: { status_codes: [ERROR] }
- name: keep-slow-runs
type: latency
latency: { threshold_ms: 20000 }
- name: sample-the-rest
type: probabilistic
probabilistic: { sampling_percentage: 5 }
batch:
send_batch_size: 2048
timeout: 5s
exporters:
otlphttp/traces:
endpoint: https://traces.example.internal:4318
otlphttp/metrics:
endpoint: https://metrics.example.internal:4318
service:
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, transform/drop_genai_content, tail_sampling, batch]
exporters: [otlphttp/traces]
metrics:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [otlphttp/metrics]
The endpoints are placeholders, and the transform and tail-sampling processors ship in the contrib distribution, not the core one. The redaction processor is a second option: its README describes it as deleting attributes that are not on an allow-list and masking values that match a blocked pattern, which suits the case where you want to keep content but mask card numbers or emails inside it. Its listed stability is beta for traces and alpha for logs and metrics, and the tail sampling processor is listed as beta for traces. Confirm the component status in the version you deploy.
Two details in the configuration carry real weight. decision_wait: 30s exists because agent runs are long. The default wait suits a web request that finishes in a second, and with a default window a slow agent run is judged before its slowest spans arrive, so the latency policy never fires. Second, tail sampling requires that all spans of a trace reach the same Collector instance. When you run several replicas you need load balancing by trace ID in front of them, which is a standard but easily forgotten requirement.
Where content should go instead
If your policy allows content capture for debugging, do not route it through the trace pipeline at all. Send the content events through a separate logs pipeline to a store with its own retention, access control and deletion tooling, and keep only a reference on the span. This follows the spec’s own recommendation for production, and it means a request to delete a user’s data is a deletion in one content store, not a scrub of every trace backend and its backups.
Cost and Cardinality: What GenAI Telemetry Actually Costs
Observability bills for GenAI tracing grow with two numbers: the volume of span data and the number of distinct metric time series. GenAI workloads inflate both, and the conventions give you the levers if you use them deliberately. The figures below are illustrative arithmetic, not measurements; replace them with your own traffic.
Metric cardinality is a product of attributes
Each metric’s series count is the product of the distinct values of its attributes. Take gen_ai.client.operation.duration with attributes for operation name, provider, request model, response model and error type. Suppose you run 3 operations, 2 providers, 6 request models, 6 response models (mostly identical to the request but with version suffixes) and 5 error types. The upper bound is 3 x 2 x 6 x 6 x 5 = 1,080 series, and each histogram series multiplies again by its bucket count, which for the recommended 14 boundaries gives 15 buckets plus sum and count. That is manageable. Now add gen_ai.conversation.id as a metric attribute with 100,000 active conversations a day and the bound becomes 108 million series. That is the classic way to produce a surprise invoice.
The rule that follows is simple: identifiers belong on spans and events, dimensions with a bounded set of values belong on metrics. gen_ai.conversation.id, gen_ai.response.id, gen_ai.tool.call.id and prompt variables are unbounded and should never be metric dimensions. Even gen_ai.response.model deserves scrutiny when providers return dated model snapshots, since each new snapshot creates a new series. The spec’s own metric attribute lists keep to a short bounded set, and a Collector or SDK view that drops everything else is cheap insurance.
Token counters are cost telemetry, if you compute them correctly
The new counters are designed to serve as a proxy for spend. Because gen_ai.client.inference.usage.input_tokens includes cached tokens and the cache-read counter reports a subset of it, an estimate of input cost uses three terms: uncached input at the base rate, cache reads at the cache-read rate, and cache writes at the cache-write rate. The rates are provider and model specific and change often, so keep them in a lookup table outside the telemetry, and join at query time rather than baking prices into spans.
A worked illustration with made-up rates: suppose input costs 1.00 per million tokens, cache reads cost 0.10, cache writes cost 1.25, and in an hour your counters show 50 million input tokens, of which 30 million were cache reads and 2 million were cache writes. Uncached input is 50 minus 30 minus 2 = 18 million, assuming the write counter also falls within the input total as its description says. The cost is 18 x 1.00 + 30 x 0.10 + 2 x 1.25 = 18 + 3 + 2.5 = 23.5 currency units, versus 50 if you naively multiplied. The numbers are invented, but the structure is the point: a cache hit rate of 60 percent more than halves the bill, and you can only see that rate if the cache attributes are captured.
For outlier detection use the per-operation histograms instead. The spec says they should not be used for totals or cost calculations, only for percentiles. A p99 of input tokens per operation that doubles after a release tells you a prompt template grew, long before the monthly invoice does.
Trace volume and sampling arithmetic
A single agent run with 4 model calls, 3 tool calls, 1 retrieval, 1 agent span and 1 workflow span is about 10 spans. At 1 million runs a day that is 10 million spans, before child spans from HTTP, database and framework instrumentation, which commonly double or triple it. If content capture is on and each model call attaches 8 KB of messages, the content alone adds roughly 32 KB per run, so 32 GB a day at that volume. These are illustrative, but the ratio is what matters: content can dwarf the structural data by an order of magnitude, which is why the conventions push it off the span by default.
Tail sampling is the standard mitigation. The configuration above keeps every error, every run slower than 20 seconds, and 5 percent of the rest. For agents you often want a fourth policy that keeps runs with an unusually high count of inference calls, using a span attribute or a numeric attribute policy, because looping runs are where the cost sits. The caveat is that tail sampling is stateful and memory-hungry: 50,000 traces held for 30 seconds is a real working set, so size num_traces against your arrival rate times the decision wait, and watch the processor’s own metrics for dropped traces.
Trade-offs, Gotchas, and What Goes Wrong
The first trade-off is stability against usefulness. The OpenTelemetry GenAI semantic conventions are Development stability and the metric rename shows breaking changes do land. Adopt them anyway, because proprietary alternatives lock you in harder, but isolate the dependency: pin library versions, keep dashboards behind recording rules, and read the repository’s changelog fragments before upgrading. I could not find a published release or version tag for the dedicated repository when I checked, only a manifest schema URL of gen-ai-dev/1.42.0-dev, so there is no version number to pin for the spec itself.
The second is mixed vintages. A fleet with several services will run different instrumentation versions at once, so you will see gen_ai.system next to gen_ai.provider.name, and the old token histogram next to the new counters. Normalize in the Collector with the transform processor, or accept two queries during the migration. Check what each language ecosystem ships: this post verified Python packages only, and I did not verify the JavaScript, Java or Go GenAI instrumentations’ current convention versions.
The third is sampling that hides the failure you care about. Head sampling at span creation cannot know that a run will loop or fail, so it throws away the traces you need. The spec acknowledges that some attributes matter for sampling decisions and should be provided at span creation time, but outcome-based decisions need tail sampling, with the memory cost described above.
The fourth is content leakage through places you did not think about. Tool arguments and results, tool definitions, system instructions and prompt template variables are all flagged as potentially sensitive in the spec, not just user messages. A privacy review that only covers gen_ai.input.messages misses the others. Exception messages are another route: error text from a provider can echo a prompt fragment, so check what your instrumentation puts in exception events before shipping them.
The fifth is double counting. If an agent framework instrumentation and a model SDK instrumentation both create inference spans for one call, you get duplicate tokens in every metric. The spec repeatedly tells instrumentations to avoid two spans for one operation, but it cannot stop two libraries from both being installed. Verify with a single request and count the spans.
The last is semantic drift in what a span means. Whether a span covers retries, whether streaming time is included in duration, and whether a tool span includes argument validation all vary between libraries. Compare a trace against a stopwatch before you trust a latency number for an SLO.
Practical Recommendations
Start with the narrowest useful slice of the OpenTelemetry GenAI semantic conventions. Instrument model calls and tool executions first, because they carry the cost and failure signal, then add agent and workflow spans once the leaves are right. Prefer an instrumentation library for the model SDK you use, hand-instrument your own tools, and keep content capture off in production.
Treat the Collector as the policy point. Strip or redact content there, run tail sampling with a decision wait that matches your longest agent run, and keep the metrics pipeline separate so that sampling traces never affects counters. Build cost views from the usage counters joined to an external price table, and build quality-of-behavior alerts from the per-invocation histograms.
A short checklist to work through before you call this production-ready:
- Pin
opentelemetry-sdk, the GenAI instrumentation packages and the Collector distribution, and record the versions in your runbook. - Confirm the metric names your libraries emit, then write dashboards against recording rules so a rename is a one-line change.
- Set content capture to
no_contentby default, and enable it per environment, not per developer laptop. - Drop or redact
gen_ai.input.messages,gen_ai.output.messages,gen_ai.system_instructions, tool definitions and tool arguments and results in the Collector. - Keep identifiers off metrics; limit metric attributes to operation, provider, model, modality and error type.
- Add a tail sampling policy for errors, slow runs and high inference-call counts.
- Alert on the p95 of
gen_ai.invoke_agent.inference_callsand on the per-operation input token p99. - Run a duplicate-span check with one request after every instrumentation change.
Frequently Asked Questions
What are the OpenTelemetry GenAI semantic conventions?
They are a set of standard names for telemetry produced by generative AI applications: span names and attributes for model calls, embeddings, retrievals, agents and tool executions, plus metrics for duration and token usage and an opt-in event for request details. They let different SDKs and frameworks emit comparable data so any OpenTelemetry-compatible backend can analyze it. All of the GenAI signals are currently at Development stability.
Are the GenAI semantic conventions stable?
No. Every gen_ai.* attribute and metric in the current specification is marked Development, which allows breaking changes between releases. The general attributes they reuse, such as error.type and server.address, are Stable. The specification has also moved from the main semantic-conventions repository to a dedicated GenAI repository, and the token metrics were redesigned in September 2026, so pin versions and isolate dashboards.
Does OpenTelemetry capture prompts and completions by default?
No. The specification says instrumentations should not capture instructions, inputs or outputs by default, and should provide an opt-in. The Python GenAI instrumentation packages use OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT with the values span_only, event_only, span_and_event and the default no_content. For production, the spec recommends storing content externally and keeping only references on spans.
What happened to the gen_ai.client.token.usage metric?
It existed as a single histogram with a required gen_ai.token.type attribute in the main semantic-conventions repository through at least version 1.40.0. The current GenAI repository replaces it with counters named gen_ai.client.inference.usage.* that carry a gen_ai.token.modality attribute, and per-operation histograms named gen_ai.client.inference.operation.*. Check which names your instrumentation emits before upgrading dashboards.
How do I trace an AI agent with OpenTelemetry?
Create an invoke_agent span for each agent run, with child chat spans for model calls and execute_tool spans for tools, nested under an invoke_workflow span when there is a user-facing entry point. Set gen_ai.agent.name, gen_ai.operation.name and gen_ai.provider.name, and propagate context so sub-agents and tools stay in one trace. Use a library where one exists and hand-instrument your own tools.
Which OTel Collector processors should I use for LLM observability?
Use memory_limiter first, a transform processor to delete sensitive content attributes, tail_sampling to keep errors and slow runs, and batch last. These ship in the contrib distribution. The redaction processor suits cases where you keep content but mask values by pattern. Set the tail sampling decision wait longer than your longest agent run, and balance by trace ID across replicas.
Further Reading
- Claude computer use architecture for LLM desktop agents: how perception-action loops produce the long span trees discussed here.
- Claude skills architecture and dynamic LLM agents: the skill-loading pattern that the agent-span refinements now model.
- Context engineering for LLM agents in production: why cache hit rate and prompt size, which these metrics expose, drive agent cost.
- eBPF observability and kernel tracing for APM: the infrastructure-level complement to semantic GenAI telemetry.
- OpenTelemetry GenAI semantic conventions repository: the normative source for spans, metrics, events and the attribute registry.
- OpenTelemetry semantic conventions documentation: the project site entry point, which now redirects to the repository above.
By Riju — about
