Skip to content
This repository was archived by the owner on Oct 1, 2026. It is now read-only.

AAPP-1159: Add OTLP traces backend with dual-write to DO Traces API - #45

Merged
kevinli-do merged 2 commits into
mainfrom
kevinli/aapp-1159-otlp-traces-adk
May 8, 2026
Merged

kevinli-do merged 2 commits into
mainfrom
kevinli/aapp-1159-otlp-traces-adk

Conversation

@kevinli-do

@kevinli-do kevinli-do commented May 7, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Wire OpenTelemetry GenAI-conventioned spans out of gradient-adk to a local OTLP collector while preserving the existing decorator API (@entrypoint, @trace_llm/@trace_tool/@trace_retriever, add_*_span) and the X-Gradient-Trace-Id header. Traces are emitted via OTLP/HTTP-proto to OTEL_EXPORTER_OTLP_ENDPOINT (default http://localhost:4318), with a MultiTracker fan-out so we can dual-write to the legacy DO Traces API behind two env-var flippers during the Galileo decommissioning window.

This is the library-side migration for the AAPP-1071 Tracing & Observability epic. The collector deployment for ADK pods is AAPP-1160, and customer pods that don't pull a fresh ADK release migrate transparently via the AAPP-1171 Galileo proxy — so this lands behind DIGITALOCEAN_OTLP_WRITE_ENABLED=false by default and does not force any rollout on its own.

What's in the box

  • runtime/otel_setup.py — one-time TracerProvider with ParentBased(root=ALWAYS_ON) sampler, BatchSpanProcessor (5s timeout, queue 2048, drop on full), and OTLPSpanExporter. Resource attributes: service.name=adk-agent, service.version, do.trace.visibility=customer, plus agent_workspace_name/agent_deployment_name as a dev/test fallback (the AAPP-1160 collector resource processor is the canonical source in prod).
  • runtime/otlp_tracker.py — OTLPTracesTracker mirroring the legacy tracker surface. Maps existing NodeExecution.metadata flags to GenAI semconv:
    • is_llm_call → chat span with gen_ai.system, gen_ai.request.model, gen_ai.request.temperature, gen_ai.usage.{input,output,total}_tokens, gen_ai.server.time_to_first_token
    • is_tool_call → execute_tool span with gen_ai.tool.name, gen_ai.tool.call.id
    • is_retriever_call → retrieval span with gen_ai.retrieval.query
    • is_workflow / is_agent_call → invoke_agent span (sub-spans become real OTel children via active context)
    • error path → Status.ERROR + record_exception(original_exception) with traceback preserved
    • inputs/outputs as span events (gen_ai.input.messages, gen_ai.output.messages), redaction-aware and byte-capped at 256 KB
  • runtime/multi_tracker.py — fan-out wrapper. Returns the legacy DO trace_id for X-Gradient-Trace-Id during dual-write and falls back to the OTel root trace_id once the legacy tracker is removed (Phase 4 of AAPP-1170).
  • decorator.py — extracts W3C traceparent from request headers via TraceContextTextMapPropagator and passes it as the parent context to on_request_start; forwards evaluation-id as evaluation_run_uuid (cross-ticket requirement from AAPP-1165).
  • network_interceptor.py — registers a RequestHook that injects the current OTel context as traceparent on outbound httpx calls to inference / KBaaS / external LLM provider URLs, so Serverless and Dedicated Inference can join the same trace per AAPP-1172. INFERENCE_URL_PATTERNS is extended to cover api.openai.com, api.anthropic.com, generativelanguage.googleapis.com, api.cohere.ai, api.x.ai for accurate gen_ai.system derivation.
  • runtime/helpers.py — _ensure_tracker() constructs DigitalOceanTracesTracker and/or OTLPTracesTracker based on DIGITALOCEAN_GALILEO_WRITE_ENABLED (default true) and DIGITALOCEAN_OTLP_WRITE_ENABLED (default false), wrapping in MultiTracker when both are enabled. OTLP-only mode no longer requires a DO API token.

Conversation log redaction (Galileo parity)

DIGITALOCEAN_CONVERSATION_LOGS_ENABLED is snapshotted onto the per-request state at on_request_start (so a mid-request env flip cannot partially redact a single trace). When false, gen_ai.input.messages and gen_ai.output.messages events are omitted and gen_ai.input.redacted=true / gen_ai.output.redacted=true markers are set. Token counts, model, temperature, durations, tool names, and error status are always preserved — parity with check_and_remove_sensitive_data() in genai-agent/logger/galileo_logger.py.

Env-var contract (per-agent, materialized by cluster-api per AAPP-1167)

Var Default Purpose
DIGITALOCEAN_GALILEO_WRITE_ENABLED true Per-agent toggle for legacy DO Traces / Galileo writer
DIGITALOCEAN_OTLP_WRITE_ENABLED false Per-agent toggle for OTLP writer
DIGITALOCEAN_CONVERSATION_LOGS_ENABLED true Per-agent redaction toggle (maps to Agent.ConversationLogsEnabled)
OTEL_EXPORTER_OTLP_ENDPOINT http://localhost:4318 OTLP collector base URL
OTEL_EXPORTER_OTLP_TRACES_ENDPOINT derived from above OTLP traces endpoint override
DIGITALOCEAN_TEAM_ID / DIGITALOCEAN_ORG_ID unset Tenant attribution fallback for local/dev runs (collector is canonical in prod)
DISABLE_TRACES unset Disable all tracing globally (existing behavior)

Test plan

  • tests/runtime/test_otlp_tracker.py — GenAI semconv attributes for chat / execute_tool / retrieval spans; root-span gen_ai.conversation.id and gen_ai.evaluation.run_uuid; resource attributes (service.name=adk-agent, do.trace.visibility=customer, team_id, org_id); error path sets Status.ERROR.
  • tests/runtime/test_otlp_tracker_lifecycle.py — __init__ does not call client.list_agent_workspaces; record_exception receives the original BaseException (class + message preserved through to the exception event); redaction flag is snapshotted at on_request_start (mid-request env flip is ignored).
  • tests/runtime/test_redaction.py — DIGITALOCEAN_CONVERSATION_LOGS_ENABLED=false omits gen_ai.input.messages / gen_ai.output.messages events and sets gen_ai.{input,output}.redacted=true; token counts and model still present.
  • tests/runtime/test_payload_cap.py — 1 MB synthetic payload truncates to ≤ 256 KB with ...<truncated> marker.
  • tests/runtime/test_multi_tracker.py — fan-out hits both backends; submit_and_get_trace_id prefers legacy trace_id and falls back to OTel trace_id when legacy is None.
  • tests/runtime/test_sampling.py — ParentBased(ALWAYS_ON) honors remote unsampled parent (no spans recorded; outbound traceparent carries flags=00); records when remote parent is sampled (outbound flags=01).
  • tests/decorator_test.py — extended for traceparent extraction into a parent context and evaluation-id → evaluation_run_uuid. All existing decorator tests still pass.
  • tests/runtime/helpers_test.py — extended for OTLP-only mode without DO API token.
  • Full broad regression suite: 242 passed, 4 skipped (optional pydantic_ai / crewai deps), 0 failed (tests/test_a2a skipped due to unrelated a2a-sdk packaging mismatch on main).
  • Manual: OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4318 DIGITALOCEAN_OTLP_WRITE_ENABLED=true python main.py against otelcol-contrib --config debug-exporter — to be done by reviewer / before enable-flag flip.

Out of scope

  • Streaming TTFT capture mid-stream (plan defers; time_to_first_token_ns is only populated for non-streamed responses today — separate follow-up against AAPP-1175).
  • Collector deployment / Helm chart wiring → AAPP-1160.
  • Galileo domain proxy for in-flight pods → AAPP-1171.
  • Galileo click-out retirement → AAPP-1168.

Ref: AAPP-1159 (epic AAPP-1071)

Wire OpenTelemetry GenAI-conventioned spans out of gradient-adk to a
local OTLP collector while preserving the existing decorator API and
the X-Gradient-Trace-Id header. Traces are emitted via OTLP/HTTP-proto
to OTEL_EXPORTER_OTLP_ENDPOINT (default localhost:4318). Dual-write to
the legacy DO Traces API runs behind two env-var flippers
(DIGITALOCEAN_GALILEO_WRITE_ENABLED / DIGITALOCEAN_OTLP_WRITE_ENABLED)
so the migration from Galileo can be staged per-agent.

What's new:
- runtime/otel_setup.py: TracerProvider with ParentBased(ALWAYS_ON)
  sampler, BatchSpanProcessor, OTLPSpanExporter, and resource
  attributes (service.name=adk-agent, do.trace.visibility=customer,
  workspace/deployment fallback for dev).
- runtime/otlp_tracker.py: OTLPTracesTracker mirroring the legacy
  tracker surface and mapping NodeExecution metadata to GenAI
  semconv (gen_ai.operation.name, gen_ai.system, gen_ai.request.*,
  gen_ai.usage.*, gen_ai.server.time_to_first_token,
  gen_ai.tool.{name,call.id}, gen_ai.retrieval.query). Root span
  carries gen_ai.conversation.id (from session-id header) and
  gen_ai.evaluation.run_uuid (from evaluation-id header) for
  AAPP-1163/AAPP-1164/AAPP-1165 trace-stream and conversation-log UX.
- runtime/multi_tracker.py: fan-out wrapper that prefers the legacy
  trace_id for X-Gradient-Trace-Id during dual-write and falls back
  to the OTel root trace_id once the legacy tracker is removed
  (Phase 4 of AAPP-1170).
- decorator.py: extract W3C traceparent from request headers via
  TraceContextTextMapPropagator and pass it as the parent context
  to on_request_start; forward evaluation-id as evaluation_run_uuid.
- network_interceptor.py: register a RequestHook that injects the
  current OTel context as a traceparent header on outbound httpx
  calls to inference / KBaaS / external LLM provider URLs, so
  Serverless and Dedicated Inference can join the same trace per
  AAPP-1172. Extended INFERENCE_URL_PATTERNS to cover OpenAI,
  Anthropic, Google, Cohere, x.ai for accurate gen_ai.system.
- helpers.py: _ensure_tracker dynamically constructs a single
  tracker, two trackers wrapped in MultiTracker, or no tracker at
  all based on the env-var flippers.

Conversation log redaction (parity with galileo_logger.py):
DIGITALOCEAN_CONVERSATION_LOGS_ENABLED is snapshotted on the
per-request state at on_request_start; when false we omit
gen_ai.input.messages / gen_ai.output.messages events and set
gen_ai.input.redacted / gen_ai.output.redacted markers, while
preserving token counts, model, durations, and error info.
Stringified payload events are byte-capped at 256 KB.

Tenant attribution (team_id / org_id) is sourced from env vars
only at SDK level; the AAPP-1160 collector resource processor is
the canonical source of truth in production. The SDK does not
perform any blocking DigitalOcean API lookups during tracker
construction.

Tests cover: GenAI semconv mapping for LLM/tool/retriever, root
span per-request attributes, redaction with payload omission,
256 KB payload cap, MultiTracker fan-out and trace_id fallback,
ParentBased sampling honoring upstream sampling decisions,
outbound traceparent injection, ingress traceparent extraction,
no DO API call at tracker init, and original exception type
preservation through record_exception.

Ref: AAPP-1159 (epic AAPP-1071)
@kevinli-do kevinli-do self-assigned this May 8, 2026
Suppress optional LangGraph import warnings so CLI JSON mode stays machine-readable, and update the A2A integration to work with the current a2a-sdk APIs used in CI. This keeps the OTLP tracing PR green without changing its user-facing tracing behavior.
@kevinli-do
kevinli-do merged commit a83189f into main May 8, 2026
5 checks passed
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants