>_ Analyst Engineering

API Observability Requirements: SLOs, Log Fields, Metrics, and Alerts an Analyst Writes

Written by Ahmed at Analyst Engineering, a Senior Technical Business Analyst with 10+ years in banking and payments delivery.

Cover for a guide to API observability requirements, showing an SLO target, an error budget, required log fields, and an alert with a runbook link.

Key takeaways

  • Observability is a requirement, not a developer's preference. If the specification does not name the log fields, metrics, SLOs, and alerts, the API ships with whatever the framework emits by default, and the first incident discovers the gap.
  • Write an SLO as a ratio of good events over valid events in a fixed window, such as 99.9 percent of POST /payments requests answered without a 5xx over 30 days. That target implies an error budget of 43.2 minutes of total failure per 30 days.
  • Every log line for a payment needs the same core fields: timestamp, service, trace_id, correlation id, internal payment id, UETR, event, status, and reason code. Names, IBANs, and remittance text never go in a log.
  • Metric labels must be low cardinality. Scheme, channel, status, and reason code are good labels; payment id and UETR are not, because every distinct label value creates a new time series.
  • An alert without a runbook link is a notification, not a control. Each alert definition needs a condition, a threshold, a window, a severity, an owner, and a runbook URL, and it should fire on a customer symptom rather than an internal cause.

API observability requirements are the non-functional requirements that make a payment API diagnosable in production: a service level objective (SLO) with its error budget, the structured fields on every log line, metrics with fixed names and low-cardinality labels, alerts with thresholds and runbook links, dashboards that show business outcomes, and Gherkin acceptance criteria that prove all of it in a test environment. If the specification does not say these things, the API ships with whatever the framework emits by default.

Every major incident I have worked ended with the same post-mortem action: “improve logging and monitoring”. That action is a requirement written too late. The analyst who specifies the API is the person best placed to write it early, because observability is mostly about business identifiers and business outcomes, and those are in the analyst’s head, not the framework’s. This is a systems analyst deliverable that sits next to the contract in the APIs for Analysts path, and it extends the general method in non-functional requirements.

What are the four golden signals for a payment API?

Google’s Site Reliability Engineering (SRE) book names four signals to monitor on any user-facing system: latency, traffic, errors, and saturation. Translated to a payment initiation API:

SignalSRE definitionPayment API version
LatencyTime to service a request, successful and failed measured separatelyp50, p95, p99 of POST /payments, split by outcome
TrafficDemand on the systemPayments submitted per minute, by scheme and channel
ErrorsRequests failing explicitly, implicitly, or by policy5xx rate, plus business rejections by reason code
SaturationHow full the most constrained resource isKafka consumer lag, connection pool usage, screening queue depth

Two details matter for payments. Split latency by outcome, because a fast rejection averaged with slow successes produces a healthy-looking number while customers wait. And treat “implicit” errors seriously: an HTTP 202 for a payment that then waits an hour for a pacs.002 status report is a failure the status code never shows.

How do you write an SLO and an error budget?

An SLO is a target ratio of good events to valid events over a fixed window. Write it so a developer can compute it from a metric without asking you anything.

IdSLOGood eventWindowTarget
SLO-1AvailabilityPOST /payments returns a non-5xx response30 days rolling99.9%
SLO-2LatencyPOST /payments responds within 800 ms30 days rolling99%
SLO-3Instant outcomeSCT Inst payment gets a positive (ACCP) or negative (RJCT) confirmation within 5 s of receipt30 days rolling99.5%

SLO-3 is the one most specifications miss, and it is the one the business cares about. The 2025 SCT Inst rulebook, effective 5 October 2025, sets a target maximum execution time of five seconds, so an outcome SLO aligned to that number turns a scheme rule into something you can monitor.

The error budget is the complement. A 99.9 percent target over 30 days allows 0.1 percent failure: 43.2 minutes of total outage, or 1,000 failed requests in every million. Write down what happens when it is spent, because that is the part that changes behaviour: for example, “when SLO-1’s budget is exhausted, non-urgent releases to the payment API pause until the 30 day ratio recovers”. Without that sentence, an SLO is a chart.

Which structured log fields does the specification require?

Logs must be structured (JSON, one object per line) so they can be searched by field rather than by regular expression. These are the fields I require on every log line that concerns a payment:

FieldExampleWhy
timestamp2026-10-09T09:14:02.118ZUTC, milliseconds, ISO 8601
levelINFOERROR and WARN must mean something
service.namepayment-orchestratorWho wrote it
trace_id, span_id4bf92f35...Jump from log to trace and back
correlation_idreq-8d1f...The API request that started it
payment.idPAY-20261009-000731Internal key
payment.uetr7f4c1a20-9e6b-...The key every bank shares
eventpayment.status_changedA fixed vocabulary, not free text
payment.statusRJCTISO 20022 status code
payment.reason_codeAM04Mandatory when status is RJCT
payment.schemeSCT_INSTSplits every query
{"timestamp":"2026-10-09T09:14:02.118Z","level":"INFO","service.name":"payment-orchestrator","trace_id":"4bf92f3577b34da6a3ce929d0e0e4736","span_id":"a3ce929d0e0e4736","correlation_id":"req-8d1f0c2a","event":"payment.status_changed","payment.id":"PAY-20261009-000731","payment.uetr":"7f4c1a20-9e6b-4d31-8a55-2c9d10bb4e77","payment.scheme":"SCT_INST","payment.status":"RJCT","payment.reason_code":"AM04"}

And the negative requirement, which is the one that ends up in an audit finding when omitted: no log line contains a customer name, IBAN, address, remittance text, card number, or credential. The UETR is the reason you never need the IBAN to find a payment; the ISO 20022 identifier guide explains why it is the right key.

Which metrics, with which names and labels?

Name metrics once in the specification so dashboards and alerts do not each invent their own. Where OpenTelemetry already defines a stable metric, use it: http.server.request.duration is a histogram in seconds with http.request.method, http.route, and http.response.status_code attributes. A Prometheus exporter typically exposes it as http_server_request_duration_seconds. For business metrics, follow Prometheus naming: one unit, base units such as seconds, and _total on counters.

MetricTypeLabels
payments_submitted_totalCounterscheme, channel
payments_completed_totalCounterscheme, status (ACCP, ACSC, RJCT)
payments_rejected_totalCounterscheme, reason_code
payment_outcome_duration_secondsHistogramscheme
screening_request_duration_secondsHistogramoutcome
kafka_consumer_lag_messagesGaugetopic, consumer_group

The rule that saves the platform team: labels must be low cardinality. Prometheus warns that every unique combination of label values is a new time series, so a payment_id or uetr label turns one metric into millions of series. Identifiers belong in logs and traces. Metrics carry the categories you aggregate by.

The API contract and the metric catalogue are both part of a complete interface specification, and API Documentation from Scratch walks through writing one end to end.

What does an alert definition need?

Google’s SRE guidance is that pages should fire on symptoms users feel, require a human to act, and be rare enough not to be ignored. An alert specification needs seven things: name, condition, threshold, window, severity, owner, and runbook link.

For SLO-based paging, the SRE Workbook recommends multiwindow, multi-burn-rate alerts. For a 99.9 percent SLO, its starting point is to page when the burn rate exceeds 14.4 over one hour and over five minutes (2 percent of the 30 day budget spent), and to page at a burn rate of 6 over six hours and 30 minutes (5 percent spent). As a Prometheus rule:

groups:
  - name: payments-api-slo
    rules:
      - alert: PaymentsApiErrorBudgetFastBurn
        expr: |
          (
            sum(rate(http_server_request_duration_seconds_count{http_route="/payments",http_response_status_code=~"5.."}[1h]))
            / sum(rate(http_server_request_duration_seconds_count{http_route="/payments"}[1h]))
          ) > (14.4 * 0.001)
          and
          (
            sum(rate(http_server_request_duration_seconds_count{http_route="/payments",http_response_status_code=~"5.."}[5m]))
            / sum(rate(http_server_request_duration_seconds_count{http_route="/payments"}[5m]))
          ) > (14.4 * 0.001)
        labels:
          severity: page
          team: payments-core
        annotations:
          summary: "POST /payments is burning the 30 day error budget at 14x"
          runbook_url: "https://runbooks.bank.example/payments-api/error-budget-burn"

Business alerts sit alongside the technical ones, and they are the analyst’s to define:

AlertConditionSeverityRunbook
Instant timeouts risingAB05 or AB06 rejections above 1% of SCT Inst for 10 minutesPageCheck CSM and beneficiary reachability
Screening slowp99 screening above 1.5 s for 10 minutesPageScreening degradation
Rejection mix shiftAny reason code above 3x its 7 day baseline for 30 minutesTicketRejection triage
Payments stuckAny payment awaiting its pacs.002 for more than 15 minutesPageStuck payment recovery

The last one catches the implicit failure the status code never shows. And every runbook link must resolve before go-live; a 404 runbook at 3 a.m. is its own incident. The production support skills piece covers what the person holding the pager actually does with it.

Which dashboards does each business outcome need?

Build dashboards around outcomes, not hosts. One payment API dashboard, top to bottom:

  1. Payments accepted and rejected per minute, by scheme. The first thing anyone on a bridge asks.
  2. Rejected by reason code, top ten, stacked. A sudden AC04 (closed account) spike means a data problem upstream; an AB05 spike means a reachability problem downstream.
  3. Latency against the SLO line, p50, p95, p99, split by outcome.
  4. Error budget remaining for each SLO, as a percentage of the window.
  5. Dependencies: screening, fraud, and clearing latency and error rate.
  6. Saturation: consumer lag, dead letter queue depth, connection pool usage.

Each panel names its metric in the specification. A dashboard nobody specified is a dashboard whose numbers nobody can explain in an incident review.

How do you write observability acceptance criteria in Gherkin?

Observability is behaviour, so it gets acceptance criteria and gets tested in SIT like anything else, using the same style as how to write API requirements.

Feature: Payment API observability

  Scenario: A rejected payment is logged, counted, and traceable
    Given a SCT Inst payment is submitted with insufficient funds
    When the payment is rejected
    Then one log line has event "payment.status_changed", status "RJCT", reason_code "AM04"
    And that log line carries payment.id, payment.uetr, trace_id, and correlation_id
    And payments_rejected_total{scheme="SCT_INST",reason_code="AM04"} increases by 1
    And no log line for the payment contains the debtor IBAN or name

  Scenario: Trace context survives the Kafka hop
    Given a payment is submitted with a valid traceparent header
    When the payment is consumed from the payment.received topic
    Then the consumer span belongs to the same trace or links to the producer span
    And every span for the payment carries payment.uetr

  Scenario: The fast-burn alert fires and points to a runbook
    Given a fault injection returns 5xx for 2% of POST /payments requests in SIT
    When the condition has persisted for one hour
    Then PaymentsApiErrorBudgetFastBurn fires with severity "page"
    And its runbook_url returns HTTP 200

The second scenario is where most estates fail, and the reason is in OpenTelemetry traces for analysts: a producer that does not write traceparent into the Kafka record headers splits every payment into two unrelated traces.

The takeaway

Observability requirements are ordinary requirements with unusual subjects. Write the SLOs as ratios with a window and a target, and state what happens when the error budget is spent. Name the log fields, make the UETR and payment id mandatory, and ban personal data explicitly. Name the metrics with low-cardinality labels, define each alert with a threshold, window, severity, owner, and working runbook link, and lay the dashboard out by business outcome. Then prove all of it in SIT with Gherkin, the same way you prove the payment itself works.

Ahmed is a Senior Technical Business Analyst with 10+ years in banking and payments. He builds practical guides and tools for analysts at The Tech BA Toolkit.

Tags: Observability, Non-Functional Requirements, SLOs, APIs, Payments

About the author

Analyst Engineering is written by Ahmed, a Senior Technical Business Analyst with 10+ years of banking and payments delivery experience: ISO 20022 and SWIFT messaging, payments API integration, Kafka event validation, and production support. Every article comes from real delivery work, and each one is reviewed and updated as tools and standards change.

Go deeper on this

Not ready to buy? The free downloads are a no-cost place to start, and every article here stays free.

Free account

Practice on the Labs, keep your progress

A free account, no password: an email link signs you in. It saves your steps and self-assessments on the Labs, shows your missions on a dashboard, unlocks the solutions, and, if you tick the box, sends you new missions and articles when they ship.

Your email is used to sign you in. Nothing else, unless you ask. Privacy.