>_ Analyst Engineering

OpenTelemetry Traces for Analysts: Investigate a Payment Without Grepping Logs

Written by Ahmed at Analyst Engineering, a Senior Technical Business Analyst with 10+ years in banking and payments delivery.

Cover for a guide to OpenTelemetry traces for analysts, showing a W3C traceparent header, a span waterfall, and a SEPA Instant payment timing out at the sanctions screening hop.

Key takeaways

  • A log line records that something happened; a span records how long it took, what called it, and whether it failed. A trace is the tree of spans for one request, so the waterfall answers in one screen what a log search answers in twenty minutes.
  • The W3C traceparent header carries four fields: version, a 32 hex character trace-id, a 16 hex character parent-id, and the trace-flags byte. Every hop that does not forward it, including a Kafka producer, starts a new and disconnected trace.
  • A trace_id is not a payment identifier. One payment produces several traces (initiation, retries, status updates, returns) and one batch trace can carry hundreds of payments, so the UETR must be recorded as a span attribute you can search on.
  • In a waterfall, compare each client span with the server span it called. When the client gave up at two seconds and the server finished at six, the customer saw a failure for a payment that downstream systems completed, which is where duplicates come from.
  • Traces are only useful if the requirement says so: propagate W3C trace context over HTTP and Kafka headers, record the internal payment id and UETR as attributes, keep every error and slow trace when sampling, and never record names, IBANs, or remittance text.

An OpenTelemetry trace is the tree of spans one request produced across every service it touched, each span carrying a start time, a duration, a parent, attributes, and a status. For an analyst, that means a single waterfall answers the two questions a log search answers slowly: where did the time go, and which hop failed. The catch is that a trace is a technical identifier, not a payment identifier, so the investigation only works if the UETR travels with it as a span attribute and the trace context survives every hop, Kafka included.

I spent years tracing payments the way reading production logs describes: take the UETR, search every index, sort by time, rebuild the story by hand. It works, slowly, and it falls apart at the hop where a service logged its own id instead. Traces fix most of that, provided somebody wrote the requirement that makes them useful. This guide is for the developer analyst who has been handed a Jaeger link on an incident bridge, and for the one writing the story that decides whether next year’s traces are any good.

What is a distributed trace, and how is it different from a log?

A log line is a sentence a service wrote at a point in time. A span is a measured unit of work with a beginning, an end, and a known caller. A trace is every span that shares one trace_id, assembled into a tree.

Log lineSpan
RecordsThat something happenedThat an operation ran, how long, and its outcome
TimeOne timestampStart and end timestamps
CausalityNone, you infer it from timestampsExplicit parent span id
CorrelationOnly if a developer logged the idAutomatic, via trace_id
Best forThe detail of what a service decidedWhere the time went and which hop failed

The parent pointer changes the job. With logs, you infer that the orchestrator called screening because the timestamps line up. With spans, the screening span names its parent, so the tool draws the call tree for you. I still read logs, because the span says screening took 6.4 seconds and the log says why. But I open the trace first.

OpenTelemetry (OTel) is the vendor-neutral standard for producing traces: SDKs per language, auto-instrumentation agents, and the OpenTelemetry Protocol (OTLP) for shipping data to a backend such as Jaeger, Grafana Tempo, Datadog APM, or Honeycomb.

What is inside a span?

Every span carries the same core fields, and knowing them is most of the skill.

{
  "traceId": "4bf92f3577b34da6a3ce929d0e0e4736",
  "spanId": "a3ce929d0e0e4736",
  "parentSpanId": "00f067aa0ba902b7",
  "name": "POST /screen",
  "kind": "CLIENT",
  "startTime": "2026-10-05T09:14:02.118Z",
  "endTime": "2026-10-05T09:14:04.121Z",
  "status": { "code": "ERROR" },
  "attributes": {
    "http.request.method": "POST",
    "server.address": "sanctions-gateway.internal",
    "error.type": "timeout",
    "payment.id": "PAY-20261005-004417",
    "payment.uetr": "7f4c1a20-9e6b-4d31-8a55-2c9d10bb4e77",
    "payment.scheme": "SCT_INST"
  },
  "resource": { "service.name": "payment-orchestrator" }
}

Read it field by field:

  • Kind. SERVER (inbound call), CLIENT (outbound call), PRODUCER (creating a message), CONSUMER (processing it), or INTERNAL. A client span and the server span it called are two records of one call, seen from each side.
  • Status. Unset by default, Error when the operation failed, Ok only when a developer explicitly marks it. Under the HTTP conventions, a client span is marked Error on a 4xx, while a server span leaves a 4xx unset, because a 400 is the caller’s mistake, not the server’s failure. Both mark 5xx as errors.
  • Attributes. Key and value pairs. Some are standard semantic convention names; the custom payment.* ones are what make a trace searchable by business identifier.
  • Resource. Who emitted it: service.name, version, environment.

How does trace context travel between services?

Through a header. The W3C Trace Context specification, a W3C Recommendation since November 2021, defines traceparent:

traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01
             │  │                                │                │
             │  trace-id (32 hex, 16 bytes)      parent-id        trace-flags
             version (00)                        (16 hex)         (01 = sampled)

The trace-id is identical on every span in the trace. The parent-id is the caller’s span id, so the receiver can attach its server span as a child. Neither may be all zeros. The lowest bit of trace-flags is the sampled flag. A companion tracestate header carries vendor-specific key and value pairs.

Trace Context Level 2, still a Candidate Recommendation Draft as of its March 2024 version, adds a flag bit signalling a random trace-id, so 03 starts appearing in the flags field. Any proxy that validates the flags strictly will drop it and start a fresh trace, which makes it a cheap test case.

Mixed estates are normal in banks. Datadog SDKs inject and extract both their own x-datadog-* headers and W3C traceparent by default, which is what lets a Datadog-instrumented channel and an OTel-instrumented payments core produce one trace.

Kafka is where traces break. HTTP auto-instrumentation propagates context for free. For messaging, the producer has to write traceparent into the Kafka record headers and the consumer has to read it. The OpenTelemetry Java agent does this for the standard Kafka clients, but a hand-rolled producer, a Kafka Connect pipeline, or a batch consumer frequently does not. The symptom is a trace that ends at publish payment.received and a second, unrelated trace that starts at the consumer. If your Kafka tests consume a test message, assert the header is there:

kcat -C -b localhost:9092 -t payment.received -o -1 -e \
  -f 'headers: %h\nvalue: %s\n'
# headers: traceparent=00-4bf92f35...-7a1d9c03e2b4f611-01

How does trace_id relate to the correlation id, UETR, and EndToEndId?

IdentifierAssigned byScopeSurvives
trace_idFirst instrumented serviceOne technical request treeOnly as far as context propagates, rarely outside the bank
Correlation id (X-Request-ID)Gateway or clientOne API call, sometimes a sessionWherever developers chose to forward it
Internal payment idYour payment engineOne payment inside the bankYour own systems
UETRFirst bank in the chainOne payment, globallyEvery bank and scheme hop, by rule
EndToEndIdThe originating customerThe customer’s referencePassed unchanged, unique by luck

The trap: one payment produces several traces. Initiation is one. A retry after a timeout is another. The asynchronous status update when the pacs.002 arrives is a third. A return three days later is a fourth. Meanwhile one batch job trace can carry four hundred payments. So the trace_id is the wrong search key for “what happened to this payment”. The UETR is the right one, as the ISO 20022 identifier guide argues for logs too, and the only way to search traces by UETR is to have recorded it as an attribute.

My rule in a specification: record payment.id and payment.uetr on every span that handles a payment. Record the EndToEndId only if your data policy allows it, because the customer writes it, and I have seen invoice numbers, names, and once a phone number in that field.

How do you read a trace waterfall?

The waterfall draws each span as a bar on a shared time axis, indented under its parent. Jaeger calls it the trace timeline; Datadog offers flame graph, waterfall, and span list views of the same data. Read it in this order:

  1. The root span. Its duration is what the customer experienced. For an instant payment, compare it with the scheme budget before anything else.
  2. The longest path down. Follow the widest child, then its widest child. Jaeger highlights the critical path, the chain of spans that actually determines total duration, which saves you from chasing a slow span that ran in parallel and did not matter.
  3. Client against server. For each outbound call, compare the client span with the server span underneath. Client 2,003 ms against server 60 ms means the time went in the network, a proxy, or a connection pool. Client 2,003 ms with a server span that ends at 6,400 ms means the caller gave up and the callee carried on.
  4. Gaps. White space between a parent’s start and its first child, or between siblings, is work nobody instrumented or a consumer that took 900 ms to pick up a message. Both are findings.
  5. Repeated siblings. Three identical POST /screen spans in a row are retries. Check whether the downstream operation is idempotent; idempotency testing covers what to assert.
  6. Status and attributes. Click the red span. Read error.type, the status code, and the payment attributes.

The same tree is exactly what sequence diagrams from logs reconstructs by hand. With traces, the export already has the parent pointers.

How do you find the slow or failing span across many traces?

One trace explains one payment. An incident needs the population. Each backend has a query language for that.

Grafana Tempo uses TraceQL:

{ resource.service.name = "sanctions-gateway" && span:duration > 1500ms }

{ span.payment.scheme = "SCT_INST" && status = error }

{ resource.service.name = "payment-orchestrator" } >> { status = error }

The >> operator selects error spans that are descendants of orchestrator spans, which is how you ask “which downstream hop failed for these payments” in one line.

In Datadog’s Trace Explorer, span attributes take an @ prefix: service:sanctions-gateway @payment.scheme:SCT_INST status:error. In Honeycomb, BubbleUp compares for you: select the slow region on a latency heatmap and it ranks which attribute values differ most from the baseline. When 97 percent of the slow spans carry screening.list_version=2026-10-05.2, you have your lead.

Which semantic conventions should you expect, and are they stable?

Semantic conventions are OTel’s agreed attribute names. They matter to an analyst because they are what you type into the query box.

HTTP is stable. Expect http.request.method, http.response.status_code, http.route on server spans, url.full on client spans, server.address, and error.type. Span names follow {method} {route}, such as POST /payments/{id}/refunds. Older instrumentation still emits pre-stable names like http.method and http.status_code, and OTel provides the OTEL_SEMCONV_STABILITY_OPT_IN setting for migrating, so check which generation your estate emits before you write a query that returns nothing.

Messaging is still in development as of October 2026, which means names can change. The current ones: messaging.system (kafka), messaging.destination.name (the topic), messaging.operation.type (create, send, receive, process, settle), messaging.destination.partition.id, messaging.consumer.group.name, messaging.kafka.message.key, and messaging.kafka.offset. The conventions use span links as the default way to connect consumer to producer, because one batch can hold messages from many traces, so some tools show the consumer as a linked trace rather than a child.

Also new: in March 2026 OTel announced the deprecation of the Span Events API in favour of log-based events correlated with the current span. Existing span events still display, but new requirements should ask for structured logs carrying trace_id and span_id.

A worked incident: a SEPA Instant payment timing out at sanctions screening

Monday, 09:20. Operations raise a P2: since about 09:00, outbound SEPA Instant Credit Transfers (SCT Inst) from the mobile app are failing for some customers with “payment could not be completed”, and two customers have complained that they retried and were charged twice.

The budget context first. The Instant Payments Regulation requires the payee’s bank to make funds available and confirm within 10 seconds of the payer’s bank receiving the order, and since the 2025 SCT Inst rulebook took effect on 5 October 2025, the scheme’s target maximum execution time is five seconds. The bank’s own share is a fraction of that, so the orchestrator gives its pre-send checks a hard two second client timeout.

Screening also needs context. Since 9 January 2025, eurozone payment service providers verify at least daily whether their own customers are subject to EU targeted financial sanctions, instead of screening each in-scope instant transfer against EU lists. Transaction screening against other regimes the bank applies, such as OFAC or the UK list, can still run per payment, and in this bank it is a synchronous call. The sanctions screening guide covers what that call reads.

Step 1: one failing payment. The complaint gives me two internal payment ids, one per attempt. In Tempo:

{ span.payment.id = "PAY-20261005-004417" }

One trace comes back for the first attempt. The second id, PAY-20261005-004431, has its own UETR and a clean, successful trace. That alone tells me the retry was a brand new payment, not a replay of the first, so no duplicate check could ever have caught it.

Step 2: the waterfall of the first trace.

POST /instant-payments             channel-api           2,340 ms  ERROR 503
└─ process payment                 payment-orchestrator  2,295 ms  ERROR
   ├─ POST /fraud/score            fraud-service            41 ms
   ├─ POST /screen   (CLIENT)      payment-orchestrator  2,003 ms  ERROR error.type=timeout
   │  └─ POST /screen (SERVER)     sanctions-gateway     6,412 ms  screening.result=CLEAR
   └─ publish payment.rejected     payment-orchestrator      4 ms

The client span stopped at 2,003 ms with error.type=timeout. The server span under it ran for 6.4 seconds and finished with screening.result=CLEAR. The customer was told the payment failed for a payment that screening cleared. Nothing was sent to the clearing mechanism, so no pacs.008 left the bank; the “double charge” is a funds reservation from attempt one that was not released when the orchestrator rejected, followed by a successful attempt two.

Step 3: the population.

{ resource.service.name = "sanctions-gateway" && span:duration > 2s }

Grouped by screening.list_version, every slow span carries the list version loaded at 08:58. The screening vendor’s index rebuild after a list update was running on the same nodes serving live requests.

Step 4: what the trace could not tell me. Why the reservation was not released is not in any span, because the orchestrator records no span for the reservation release on the rejection path. That gap is a finding, and the log line I eventually found said release skipped: state=SCREENING_PENDING. The payment state machine had no transition from that state on timeout.

The incident report had four findings: the list update process degrades live latency, the timeout path skips the reservation release, the rejection path is invisible in traces, and the app offers retry with no idempotency key. If you want to practise this chain of trigger, defect, and condition, the free duplicate refund mission is the same shape of problem: a gateway gives up, the backend finishes, and the client retries.

The domain background for incidents like this, scheme rules, screening, and the payment lifecycle, is in Break Into Banking, and the technical progression from logs to traces is in The Technical Skills Guide for BAs.

What should an NFR or story say so traces are actually useful?

Traces are only as good as the requirement that shaped them. Auto-instrumentation gives you HTTP spans with no business identifiers, broken at the first Kafka hop, sampled at random. These are the lines I put in a non-functional requirements section for any payment flow; the companion piece on API observability requirements covers metrics, SLOs, and alerts.

IdRequirement
OBS-T1Every service propagates W3C Trace Context (traceparent, tracestate) on every outbound HTTP call and accepts it on every inbound call, including trace-flags 03.
OBS-T2Every Kafka producer writes traceparent into the record headers; every consumer extracts it and links or parents its processing span to it.
OBS-T3Every span that handles a payment carries payment.id, payment.uetr, and payment.scheme. Status transitions carry payment.status and, on rejection, payment.reason_code (for example AB05, AM04).
OBS-T4No span attribute, span name, or log line contains a name, IBAN, address, remittance text, or full card number. The OTel Collector applies an allow-list redaction processor as a backstop, not as the primary control.
OBS-T5Sampling retains 100 percent of traces containing an error span and every trace whose root exceeds the latency SLO threshold; other traffic may be sampled.
OBS-T6Every log line carries trace_id and span_id, so a span links to its logs and back.
OBS-T7Every timeout, retry, and compensation (reservation release, reversal) is its own span, so the failure path is as visible as the happy path.
OBS-T8A business rejection (sanctions hit, insufficient funds) sets an attribute, not span status Error; a technical failure sets Error.

OBS-T5 is the one that bites. Head sampling at 10 percent, decided at the first service, means nine in ten of the failing payments you need have no trace. Keeping every error and slow trace needs tail-based sampling in the Collector, an infrastructure cost someone must approve, which is why it belongs in a requirement.

OBS-T8 is a decision the business has to make with you. If a sanctions rejection marks the span as an error, the error rate dashboard spikes every time compliance does its job, and people learn to ignore it.

The acceptance test fits in the story: initiate one SCT Inst payment in SIT, search the tracing backend by its UETR, and assert one connected trace from the channel API through the Kafka consumer to the instant gateway, with payment.uetr on every span and no personal data on any of them.

The takeaway

A trace is the tree of spans one request produced, and reading its waterfall answers where the time went and which hop failed faster than any log search: root duration first, then the critical path, client spans against server spans, gaps, and repeated siblings. Context travels in the traceparent header, and it dies at any hop that does not forward it, Kafka most often. The trace_id is a technical identifier, so record the UETR and payment id as attributes and search on those.

Then write the requirement that makes it work: propagate context everywhere, carry the business identifiers, keep every error and slow trace, instrument the failure paths, and never put personal data on a span. When you are back in the logs, reading logs during a major incident and the wider production support skills are the companions to this one.

Ahmed is a Senior Technical Business Analyst with 10+ years in banking and payments. He builds practical guides and tools for analysts at The Tech BA Toolkit.

Tags: Observability, OpenTelemetry, Distributed Tracing, Production Support, Payments

About the author

Analyst Engineering is written by Ahmed, a Senior Technical Business Analyst with 10+ years of banking and payments delivery experience: ISO 20022 and SWIFT messaging, payments API integration, Kafka event validation, and production support. Every article comes from real delivery work, and each one is reviewed and updated as tools and standards change.

Go deeper on this

Not ready to buy? The free downloads are a no-cost place to start, and every article here stays free.

Free account

Practice on the Labs, keep your progress

A free account, no password: an email link signs you in. It saves your steps and self-assessments on the Labs, shows your missions on a dashboard, unlocks the solutions, and, if you tick the box, sends you new missions and articles when they ship.

Your email is used to sign you in. Nothing else, unless you ask. Privacy.