OpenTelemetry Traces for Analysts: Investigate a Payment Without Grepping Logs
Written by Ahmed at Analyst Engineering, a Senior Technical Business Analyst with 10+ years in banking and payments delivery.
Key takeaways
- A log line records that something happened; a span records how long it took, what called it, and whether it failed. A trace is the tree of spans for one request, so the waterfall answers in one screen what a log search answers in twenty minutes.
- The W3C traceparent header carries four fields: version, a 32 hex character trace-id, a 16 hex character parent-id, and the trace-flags byte. Every hop that does not forward it, including a Kafka producer, starts a new and disconnected trace.
- A trace_id is not a payment identifier. One payment produces several traces (initiation, retries, status updates, returns) and one batch trace can carry hundreds of payments, so the UETR must be recorded as a span attribute you can search on.
- In a waterfall, compare each client span with the server span it called. When the client gave up at two seconds and the server finished at six, the customer saw a failure for a payment that downstream systems completed, which is where duplicates come from.
- Traces are only useful if the requirement says so: propagate W3C trace context over HTTP and Kafka headers, record the internal payment id and UETR as attributes, keep every error and slow trace when sampling, and never record names, IBANs, or remittance text.
An OpenTelemetry trace is the tree of spans one request produced across every service it touched, each span carrying a start time, a duration, a parent, attributes, and a status. For an analyst, that means a single waterfall answers the two questions a log search answers slowly: where did the time go, and which hop failed. The catch is that a trace is a technical identifier, not a payment identifier, so the investigation only works if the UETR travels with it as a span attribute and the trace context survives every hop, Kafka included.
I spent years tracing payments the way reading production logs describes: take the UETR, search every index, sort by time, rebuild the story by hand. It works, slowly, and it falls apart at the hop where a service logged its own id instead. Traces fix most of that, provided somebody wrote the requirement that makes them useful. This guide is for the developer analyst who has been handed a Jaeger link on an incident bridge, and for the one writing the story that decides whether next year’s traces are any good.
What is a distributed trace, and how is it different from a log?
A log line is a sentence a service wrote at a point in time. A span is a measured unit of work with a beginning, an end, and a known caller. A trace is every span that shares one trace_id, assembled into a tree.
| Log line | Span | |
|---|---|---|
| Records | That something happened | That an operation ran, how long, and its outcome |
| Time | One timestamp | Start and end timestamps |
| Causality | None, you infer it from timestamps | Explicit parent span id |
| Correlation | Only if a developer logged the id | Automatic, via trace_id |
| Best for | The detail of what a service decided | Where the time went and which hop failed |
The parent pointer changes the job. With logs, you infer that the orchestrator called screening because the timestamps line up. With spans, the screening span names its parent, so the tool draws the call tree for you. I still read logs, because the span says screening took 6.4 seconds and the log says why. But I open the trace first.
OpenTelemetry (OTel) is the vendor-neutral standard for producing traces: SDKs per language, auto-instrumentation agents, and the OpenTelemetry Protocol (OTLP) for shipping data to a backend such as Jaeger, Grafana Tempo, Datadog APM, or Honeycomb.
What is inside a span?
Every span carries the same core fields, and knowing them is most of the skill.
{
"traceId": "4bf92f3577b34da6a3ce929d0e0e4736",
"spanId": "a3ce929d0e0e4736",
"parentSpanId": "00f067aa0ba902b7",
"name": "POST /screen",
"kind": "CLIENT",
"startTime": "2026-10-05T09:14:02.118Z",
"endTime": "2026-10-05T09:14:04.121Z",
"status": { "code": "ERROR" },
"attributes": {
"http.request.method": "POST",
"server.address": "sanctions-gateway.internal",
"error.type": "timeout",
"payment.id": "PAY-20261005-004417",
"payment.uetr": "7f4c1a20-9e6b-4d31-8a55-2c9d10bb4e77",
"payment.scheme": "SCT_INST"
},
"resource": { "service.name": "payment-orchestrator" }
}
Read it field by field:
- Kind.
SERVER(inbound call),CLIENT(outbound call),PRODUCER(creating a message),CONSUMER(processing it), orINTERNAL. A client span and the server span it called are two records of one call, seen from each side. - Status.
Unsetby default,Errorwhen the operation failed,Okonly when a developer explicitly marks it. Under the HTTP conventions, a client span is markedErroron a 4xx, while a server span leaves a 4xx unset, because a 400 is the caller’s mistake, not the server’s failure. Both mark 5xx as errors. - Attributes. Key and value pairs. Some are standard semantic convention names; the custom
payment.*ones are what make a trace searchable by business identifier. - Resource. Who emitted it:
service.name, version, environment.
How does trace context travel between services?
Through a header. The W3C Trace Context specification, a W3C Recommendation since November 2021, defines traceparent:
traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01
│ │ │ │
│ trace-id (32 hex, 16 bytes) parent-id trace-flags
version (00) (16 hex) (01 = sampled)
The trace-id is identical on every span in the trace. The parent-id is the caller’s span id, so the receiver can attach its server span as a child. Neither may be all zeros. The lowest bit of trace-flags is the sampled flag. A companion tracestate header carries vendor-specific key and value pairs.
Trace Context Level 2, still a Candidate Recommendation Draft as of its March 2024 version, adds a flag bit signalling a random trace-id, so 03 starts appearing in the flags field. Any proxy that validates the flags strictly will drop it and start a fresh trace, which makes it a cheap test case.
Mixed estates are normal in banks. Datadog SDKs inject and extract both their own x-datadog-* headers and W3C traceparent by default, which is what lets a Datadog-instrumented channel and an OTel-instrumented payments core produce one trace.
Kafka is where traces break. HTTP auto-instrumentation propagates context for free. For messaging, the producer has to write traceparent into the Kafka record headers and the consumer has to read it. The OpenTelemetry Java agent does this for the standard Kafka clients, but a hand-rolled producer, a Kafka Connect pipeline, or a batch consumer frequently does not. The symptom is a trace that ends at publish payment.received and a second, unrelated trace that starts at the consumer. If your Kafka tests consume a test message, assert the header is there:
kcat -C -b localhost:9092 -t payment.received -o -1 -e \
-f 'headers: %h\nvalue: %s\n'
# headers: traceparent=00-4bf92f35...-7a1d9c03e2b4f611-01
How does trace_id relate to the correlation id, UETR, and EndToEndId?
| Identifier | Assigned by | Scope | Survives |
|---|---|---|---|
trace_id | First instrumented service | One technical request tree | Only as far as context propagates, rarely outside the bank |
Correlation id (X-Request-ID) | Gateway or client | One API call, sometimes a session | Wherever developers chose to forward it |
| Internal payment id | Your payment engine | One payment inside the bank | Your own systems |
| UETR | First bank in the chain | One payment, globally | Every bank and scheme hop, by rule |
| EndToEndId | The originating customer | The customer’s reference | Passed unchanged, unique by luck |
The trap: one payment produces several traces. Initiation is one. A retry after a timeout is another. The asynchronous status update when the pacs.002 arrives is a third. A return three days later is a fourth. Meanwhile one batch job trace can carry four hundred payments. So the trace_id is the wrong search key for “what happened to this payment”. The UETR is the right one, as the ISO 20022 identifier guide argues for logs too, and the only way to search traces by UETR is to have recorded it as an attribute.
My rule in a specification: record payment.id and payment.uetr on every span that handles a payment. Record the EndToEndId only if your data policy allows it, because the customer writes it, and I have seen invoice numbers, names, and once a phone number in that field.
How do you read a trace waterfall?
The waterfall draws each span as a bar on a shared time axis, indented under its parent. Jaeger calls it the trace timeline; Datadog offers flame graph, waterfall, and span list views of the same data. Read it in this order:
- The root span. Its duration is what the customer experienced. For an instant payment, compare it with the scheme budget before anything else.
- The longest path down. Follow the widest child, then its widest child. Jaeger highlights the critical path, the chain of spans that actually determines total duration, which saves you from chasing a slow span that ran in parallel and did not matter.
- Client against server. For each outbound call, compare the client span with the server span underneath. Client 2,003 ms against server 60 ms means the time went in the network, a proxy, or a connection pool. Client 2,003 ms with a server span that ends at 6,400 ms means the caller gave up and the callee carried on.
- Gaps. White space between a parent’s start and its first child, or between siblings, is work nobody instrumented or a consumer that took 900 ms to pick up a message. Both are findings.
- Repeated siblings. Three identical
POST /screenspans in a row are retries. Check whether the downstream operation is idempotent; idempotency testing covers what to assert. - Status and attributes. Click the red span. Read
error.type, the status code, and the payment attributes.
The same tree is exactly what sequence diagrams from logs reconstructs by hand. With traces, the export already has the parent pointers.
How do you find the slow or failing span across many traces?
One trace explains one payment. An incident needs the population. Each backend has a query language for that.
Grafana Tempo uses TraceQL:
{ resource.service.name = "sanctions-gateway" && span:duration > 1500ms }
{ span.payment.scheme = "SCT_INST" && status = error }
{ resource.service.name = "payment-orchestrator" } >> { status = error }
The >> operator selects error spans that are descendants of orchestrator spans, which is how you ask “which downstream hop failed for these payments” in one line.
In Datadog’s Trace Explorer, span attributes take an @ prefix: service:sanctions-gateway @payment.scheme:SCT_INST status:error. In Honeycomb, BubbleUp compares for you: select the slow region on a latency heatmap and it ranks which attribute values differ most from the baseline. When 97 percent of the slow spans carry screening.list_version=2026-10-05.2, you have your lead.
Which semantic conventions should you expect, and are they stable?
Semantic conventions are OTel’s agreed attribute names. They matter to an analyst because they are what you type into the query box.
HTTP is stable. Expect http.request.method, http.response.status_code, http.route on server spans, url.full on client spans, server.address, and error.type. Span names follow {method} {route}, such as POST /payments/{id}/refunds. Older instrumentation still emits pre-stable names like http.method and http.status_code, and OTel provides the OTEL_SEMCONV_STABILITY_OPT_IN setting for migrating, so check which generation your estate emits before you write a query that returns nothing.
Messaging is still in development as of October 2026, which means names can change. The current ones: messaging.system (kafka), messaging.destination.name (the topic), messaging.operation.type (create, send, receive, process, settle), messaging.destination.partition.id, messaging.consumer.group.name, messaging.kafka.message.key, and messaging.kafka.offset. The conventions use span links as the default way to connect consumer to producer, because one batch can hold messages from many traces, so some tools show the consumer as a linked trace rather than a child.
Also new: in March 2026 OTel announced the deprecation of the Span Events API in favour of log-based events correlated with the current span. Existing span events still display, but new requirements should ask for structured logs carrying trace_id and span_id.
A worked incident: a SEPA Instant payment timing out at sanctions screening
Monday, 09:20. Operations raise a P2: since about 09:00, outbound SEPA Instant Credit Transfers (SCT Inst) from the mobile app are failing for some customers with “payment could not be completed”, and two customers have complained that they retried and were charged twice.
The budget context first. The Instant Payments Regulation requires the payee’s bank to make funds available and confirm within 10 seconds of the payer’s bank receiving the order, and since the 2025 SCT Inst rulebook took effect on 5 October 2025, the scheme’s target maximum execution time is five seconds. The bank’s own share is a fraction of that, so the orchestrator gives its pre-send checks a hard two second client timeout.
Screening also needs context. Since 9 January 2025, eurozone payment service providers verify at least daily whether their own customers are subject to EU targeted financial sanctions, instead of screening each in-scope instant transfer against EU lists. Transaction screening against other regimes the bank applies, such as OFAC or the UK list, can still run per payment, and in this bank it is a synchronous call. The sanctions screening guide covers what that call reads.
Step 1: one failing payment. The complaint gives me two internal payment ids, one per attempt. In Tempo:
{ span.payment.id = "PAY-20261005-004417" }
One trace comes back for the first attempt. The second id, PAY-20261005-004431, has its own UETR and a clean, successful trace. That alone tells me the retry was a brand new payment, not a replay of the first, so no duplicate check could ever have caught it.
Step 2: the waterfall of the first trace.
POST /instant-payments channel-api 2,340 ms ERROR 503
└─ process payment payment-orchestrator 2,295 ms ERROR
├─ POST /fraud/score fraud-service 41 ms
├─ POST /screen (CLIENT) payment-orchestrator 2,003 ms ERROR error.type=timeout
│ └─ POST /screen (SERVER) sanctions-gateway 6,412 ms screening.result=CLEAR
└─ publish payment.rejected payment-orchestrator 4 ms
The client span stopped at 2,003 ms with error.type=timeout. The server span under it ran for 6.4 seconds and finished with screening.result=CLEAR. The customer was told the payment failed for a payment that screening cleared. Nothing was sent to the clearing mechanism, so no pacs.008 left the bank; the “double charge” is a funds reservation from attempt one that was not released when the orchestrator rejected, followed by a successful attempt two.
Step 3: the population.
{ resource.service.name = "sanctions-gateway" && span:duration > 2s }
Grouped by screening.list_version, every slow span carries the list version loaded at 08:58. The screening vendor’s index rebuild after a list update was running on the same nodes serving live requests.
Step 4: what the trace could not tell me. Why the reservation was not released is not in any span, because the orchestrator records no span for the reservation release on the rejection path. That gap is a finding, and the log line I eventually found said release skipped: state=SCREENING_PENDING. The payment state machine had no transition from that state on timeout.
The incident report had four findings: the list update process degrades live latency, the timeout path skips the reservation release, the rejection path is invisible in traces, and the app offers retry with no idempotency key. If you want to practise this chain of trigger, defect, and condition, the free duplicate refund mission is the same shape of problem: a gateway gives up, the backend finishes, and the client retries.
The domain background for incidents like this, scheme rules, screening, and the payment lifecycle, is in Break Into Banking, and the technical progression from logs to traces is in The Technical Skills Guide for BAs.
What should an NFR or story say so traces are actually useful?
Traces are only as good as the requirement that shaped them. Auto-instrumentation gives you HTTP spans with no business identifiers, broken at the first Kafka hop, sampled at random. These are the lines I put in a non-functional requirements section for any payment flow; the companion piece on API observability requirements covers metrics, SLOs, and alerts.
| Id | Requirement |
|---|---|
| OBS-T1 | Every service propagates W3C Trace Context (traceparent, tracestate) on every outbound HTTP call and accepts it on every inbound call, including trace-flags 03. |
| OBS-T2 | Every Kafka producer writes traceparent into the record headers; every consumer extracts it and links or parents its processing span to it. |
| OBS-T3 | Every span that handles a payment carries payment.id, payment.uetr, and payment.scheme. Status transitions carry payment.status and, on rejection, payment.reason_code (for example AB05, AM04). |
| OBS-T4 | No span attribute, span name, or log line contains a name, IBAN, address, remittance text, or full card number. The OTel Collector applies an allow-list redaction processor as a backstop, not as the primary control. |
| OBS-T5 | Sampling retains 100 percent of traces containing an error span and every trace whose root exceeds the latency SLO threshold; other traffic may be sampled. |
| OBS-T6 | Every log line carries trace_id and span_id, so a span links to its logs and back. |
| OBS-T7 | Every timeout, retry, and compensation (reservation release, reversal) is its own span, so the failure path is as visible as the happy path. |
| OBS-T8 | A business rejection (sanctions hit, insufficient funds) sets an attribute, not span status Error; a technical failure sets Error. |
OBS-T5 is the one that bites. Head sampling at 10 percent, decided at the first service, means nine in ten of the failing payments you need have no trace. Keeping every error and slow trace needs tail-based sampling in the Collector, an infrastructure cost someone must approve, which is why it belongs in a requirement.
OBS-T8 is a decision the business has to make with you. If a sanctions rejection marks the span as an error, the error rate dashboard spikes every time compliance does its job, and people learn to ignore it.
The acceptance test fits in the story: initiate one SCT Inst payment in SIT, search the tracing backend by its UETR, and assert one connected trace from the channel API through the Kafka consumer to the instant gateway, with payment.uetr on every span and no personal data on any of them.
The takeaway
A trace is the tree of spans one request produced, and reading its waterfall answers where the time went and which hop failed faster than any log search: root duration first, then the critical path, client spans against server spans, gaps, and repeated siblings. Context travels in the traceparent header, and it dies at any hop that does not forward it, Kafka most often. The trace_id is a technical identifier, so record the UETR and payment id as attributes and search on those.
Then write the requirement that makes it work: propagate context everywhere, carry the business identifiers, keep every error and slow trace, instrument the failure paths, and never put personal data on a span. When you are back in the logs, reading logs during a major incident and the wider production support skills are the companions to this one.
Ahmed is a Senior Technical Business Analyst with 10+ years in banking and payments. He builds practical guides and tools for analysts at The Tech BA Toolkit.
Tags: Observability, OpenTelemetry, Distributed Tracing, Production Support, Payments
About the author
Analyst Engineering is written by Ahmed, a Senior Technical Business Analyst with 10+ years of banking and payments delivery experience: ISO 20022 and SWIFT messaging, payments API integration, Kafka event validation, and production support. Every article comes from real delivery work, and each one is reviewed and updated as tools and standards change.
Related articles
- Reading Production Logs: Trace One Transaction's Trail How an analyst reads production logs to debug a system: correlation ids, log levels, searching by transaction, and following one request across services.
- Draw Sequence Diagrams from Splunk and Datadog Logs: The Flow as It Actually Ran Turn correlated Splunk or Datadog logs into an accurate Mermaid sequence diagram with AI. The queries, the export shape, the prompt, and the verification step.
- What I Learned Reading Logs During a Major Incident Field notes from a major payments incident: reading logs under pressure, what the trail revealed, and the observability lessons that became requirements.
- The Production Support Skills Nobody Teaches Analysts Production support skills for technical analysts: triage, tracing transactions, reading logs, staying calm, and turning incidents into requirements.
Go deeper on this
Not ready to buy? The free downloads are a no-cost place to start, and every article here stays free.
Free account
Practice on the Labs, keep your progress
A free account, no password: an email link signs you in. It saves your steps and self-assessments on the Labs, shows your missions on a dashboard, unlocks the solutions, and, if you tick the box, sends you new missions and articles when they ship.
Your email is used to sign you in. Nothing else, unless you ask. Privacy.