Onboarding as a Developer Analyst: Repository, Logs, and SQL in Week One
Written by Ahmed at Analyst Engineering, a Senior Technical Business Analyst with 10+ years in banking and payments delivery.
Key takeaways
- A developer analyst is onboarded when they can take a production question, find the code path that answers it, prove the answer with logs and a query, and hand the developers a root cause note they do not need to redo.
- Request five accesses on day one, in writing, because each one takes days to approve: read access to the repository, the log platform, a read-only database replica, the test environments, and the message viewer that shows raw payment messages.
- An ISO 20022 file that fails XSD validation never becomes a payment, so the payment table is empty for it. Your identifier is the inbound file and its MsgId, not the UETR or the EndToEndId.
- pain.001.001.09 changed element shapes that pain.001.001.03 files got away with: BIC became BICFI, and RequestedExecutionDate became a choice that wraps the date in a Dt element. A channel that half-upgraded its generator produces exactly these FF01 rejections.
- The first deliverable for a developer analyst is a one-page root cause note: symptom, evidence with file names and log lines, cause, fix owner, and a regression test that stops the same file shape shipping twice.
A developer analyst is onboarded when they can take a production question, find the code that answers it, prove the answer with logs and a query, and write it up so the developers do not have to redo the work. In week one that means five accesses requested on day one, the repository read with an AI assistant, one real message followed by its identifiers, and a first root cause note.
On my first Monday with a payments engineering team, the onboarding plan said “week one: environment setup and reading.” The board said something else. The first ticket assigned to me, created the Friday before I arrived, read in full: “Investigate FF01 rejections from new corporate channel.” A large corporate client had just switched their treasury system to send payment files over a new host-to-host connection, and some of their pain.001 files were coming back rejected with a pain.002 carrying group status RJCT and reason code FF01, invalid file format. Not all of them. Some. The client’s treasury team was calling the relationship manager every morning, and the relationship manager was calling the team lead.
I did not know the codebase, the log platform, the database, or a single colleague’s name. This article is how I worked that ticket, and how the same week became the onboarding plan I now use for every developer analyst seat. It is part 9 of The First 90 Days, the role angle for analysts who read code, logs, and data for a living. If you want the general plan first, start with part 1, the first 90 days.
What does “onboarded” mean for a developer analyst?
For a business analyst, onboarded means stakeholders trust your summary of what they need. For a developer analyst, it is narrower and easier to test: you can answer “why did the system do that?” with evidence, without borrowing a developer’s afternoon.
That breaks into four capabilities, and each one maps to an access you need:
| Capability | What it looks like | Access it depends on |
|---|---|---|
| Find the code path | ”Inbound files are validated in XsdValidator, against the schema version configured per channel” | Repository read access |
| Follow a message | ”File X arrived 07:42:13, failed validation 07:42:14, pain.002 sent 07:42:15” | Log platform |
| Count the pattern | ”41 files rejected with FF01 this month, all from one channel, mostly on Mondays” | Read-only database replica |
| Reproduce it | ”This file fails in SIT with the same error, and passes with one element changed” | Test environments and a message viewer |
When you can do all four on a system you joined last month, you are onboarded. Everything below is about getting there inside a fortnight instead of a quarter.
Which accesses should you request on day one?
Request all five on day one, in writing, in one message to your manager, with the ticket you need them for. Access requests in banks go through an identity and access management workflow with approvers, and every approver adds a day. If you request them one at a time as you discover you need them, you lose your first two weeks to waiting.
The message I sent at 09:30 on day one:
Hi [manager],
To work PAY-4812 (FF01 rejections, new corporate channel) I need the
following. All read-only. Could you approve or point me to the right
request form?
1. Repository: read access to the payments intake and mapping repos
2. Logs: the log platform, payments index, production read
3. Database: read-only replica of the payments schema (no PII
unmasking needed)
4. Environments: SIT and UAT, with the ability to submit a test file
through the host-to-host test channel
5. Message viewer: read access to raw inbound and outbound messages
(pain.001, pain.002), production
I will not need write access to anything this month.
Two details in there do more than they look. “Read-only” and “no PII unmasking” get approved faster because the approver’s risk is lower, and in a bank the approver is usually weighing exactly that. And naming the ticket gives the approver a reason they can write into the access log, which is what audit asks them for.
While you wait, you are not idle. Repository access usually arrives first, and that is where the next step happens.
How do you read an unfamiliar payments repository in one day?
You do not read it. You interrogate it with three questions, and you use an approved AI assistant to point you at the files, then you open the files yourself.
On my team, the assistant was the organization’s approved coding assistant, running inside the IDE against the cloned repository, with the policy that source code was allowed and production data was not. Check your own policy before you open a single file in an assistant; part 3 of this series covers how to find out what is approved without making it awkward.
The three questions that unlock any intake system:
I am a technical analyst investigating why some incoming pain.001
files are rejected with FF01 (invalid file format).
In this repository, answer with file paths and line numbers only.
Do not summarize behavior you have not located in code.
1. Where do inbound files from the host-to-host channel enter the
application? Name the entry point class or handler.
2. Where is XSD validation performed, and which schema file or files
are loaded? Is the schema version chosen by the file's namespace,
by configuration per channel, or fixed?
3. Where is the pain.002 rejection built, and how is the reason code
chosen? Is FF01 hardcoded, mapped from an exception type, or
configured?
If you cannot find something, say NOT FOUND rather than guessing.
The answers came back in under a minute. Then I spent forty minutes verifying them, which is the right ratio. Two were correct. One was half right: validation happened where the assistant said, but the schema was not chosen by namespace. It was configured per channel, in a YAML file the assistant had not mentioned because it was in a separate configuration repository I did not have yet.
That gap was the most important thing I learned all week, and the assistant could not have told me, because it could only see what I had cloned. The lesson generalizes: an AI reading a repository knows only that repository, and in a bank the behavior that matters is often in configuration that lives elsewhere. Ask a developer “where does per-channel configuration live?” on day two. It is the cheapest question you will ask all year.
For the full method of reading code you did not write, including how to ask an assistant to trace a call path, see AI in the codebase for analysts. If cloning and branches are new to you, Git for analysts covers the twenty minutes of Git you actually need.
What to write down while you read
Keep a file called system-notes.md from the first hour. Not a polished document, a log of facts with their source:
## Intake (verified 2026-09-29)
- Entry: h2h/InboundFileListener.java:41 (polls SFTP drop every 30s)
- Validation: intake/XsdValidator.java:88, schema chosen per channel
from config repo channels/<channel>.yaml key `schemaVersion`
- Reject builder: status/Pain002Builder.java:120, FF01 when
XsdValidationException is thrown (hardcoded mapping, line 133)
- QUESTION: who owns the config repo? (asked Priya, waiting)
Every line has a file and a line number, or it is marked as a question. That habit is what lets you hand this to a developer later and have them trust it at a glance.
How do you follow one rejected file through the logs?
Here is the part that catches people who come from the payment side rather than the file side. A file that fails XSD validation never becomes a payment. It never gets a payment id, its transactions are never written to the payment table, and the UETR and EndToEndId inside it may never be parsed. Your identifier is the file: its name, its hash, and the MsgId in the group header if the parser got that far.
So the identifier chain for an intake rejection is:
| Identifier | Where it lives | Available when the file fails XSD? |
|---|---|---|
| File name | SFTP drop, intake log | Yes, always |
| File hash (SHA-256) | Intake log, archive | Usually, if the team logs it |
MsgId (GrpHdr/MsgId) | Inside the XML | Only if the logger extracts it before validation |
| PmtInfId | Inside the XML | Rarely |
| EndToEndId, UETR | Inside each transaction | No, not in the payment tables |
| OrgnlMsgId in the pain.002 | Outbound rejection | Yes, it echoes the MsgId |
That last row is the trick. The outbound pain.002 echoes the original MsgId in OrgnlGrpInfAndSts/OrgnlMsgId, so even when the inbound log only has the file name, the rejection message ties the file to its MsgId.
My first log query, on the team’s log platform, filtered on the channel and the rejection:
index=payments-intake channel="H2H-CORP-07" level=ERROR
"XsdValidationException"
| table _time, file_name, msg_id, error_line, error_message
| sort _time
The error_message column was the gift. The validator logged the parser’s own message, and in every failing row it was one of two lines:
cvc-complex-type.2.4.a: Invalid content was found starting with
element 'BIC'. One of '{BICFI, ClrSysMmbId, LEI, Nm, PstlAdr,
Othr}' is expected.
cvc-complex-type.2.3: Element 'ReqdExctnDt' cannot have character
[children], because the type's content type is element-only.
Two errors, both on elements that hold perfectly ordinary values: a bank identifier and a date. That is the signature of a version mismatch, not of bad data, and I will come back to it. For the general method of reading log lines like these, see reading production logs.
What SQL does a developer analyst run in the first week?
Two queries: one to size the pattern, and one to check that nothing else is hiding behind it. The read-only replica arrived on day three.
The intake schema had three tables that matter here: inbound_file (one row per received file), payment (one row per transaction, only for files that passed validation), and payment_status (every status a payment has had, with reason code and timestamp).
Query 1: size the rejection pattern by day, channel, and reason.
SELECT
CAST(f.received_at AS DATE) AS received_day,
f.channel_id,
f.schema_namespace,
f.reject_reason_code,
COUNT(*) AS files
FROM inbound_file AS f
WHERE f.received_at >= DATE '2026-09-01'
AND f.status = 'RJCT'
GROUP BY CAST(f.received_at AS DATE), f.channel_id,
f.schema_namespace, f.reject_reason_code
ORDER BY received_day, files DESC;
The result answered the “some, not all” mystery in one screen. Every FF01 came from channel H2H-CORP-07. Every rejected file declared the namespace urn:iso:std:iso:20022:tech:xsd:pain.001.001.09. And the channel had also sent dozens of files on the same namespace that were accepted. So the namespace was not the problem by itself. Something inside some version 09 files was.
Query 2: for the files that passed, what is the latest status of each payment? This is the check most people skip. If the channel has a format problem severe enough to fail XSD sometimes, it may have subtler problems that pass XSD and fail later, at business validation or at screening.
WITH latest AS (
SELECT
p.payment_id,
p.end_to_end_id,
p.uetr,
s.status,
s.reason_code,
s.status_at,
ROW_NUMBER() OVER (
PARTITION BY p.payment_id
ORDER BY s.status_at DESC, s.status_seq DESC
) AS rn
FROM payment AS p
JOIN inbound_file AS f ON f.file_id = p.file_id
JOIN payment_status AS s ON s.payment_id = p.payment_id
WHERE f.channel_id = 'H2H-CORP-07'
AND f.received_at >= DATE '2026-09-01'
)
SELECT status, reason_code, COUNT(*) AS payments
FROM latest
WHERE rn = 1
GROUP BY status, reason_code
ORDER BY payments DESC;
The status_seq tiebreaker in the ORDER BY matters: two statuses written in the same millisecond are common in payment systems, and without a tiebreaker ROW_NUMBER picks one arbitrarily, so the same query can return different answers on different runs. If window functions are new, SQL window functions for analysts explains ROW_NUMBER with exactly this latest-status pattern, and SQL for analysts covers the joins.
Query 2 came back clean: accepted payments from the channel were settling normally, with a handful of AC01 (incorrect account number) rejections that were genuine data issues on the client side and nothing to do with my ticket. That ruled out a wider format problem and kept the ticket small.
How do you reproduce the failure safely?
You take the shape of the failing file, never the file itself, into a lower environment.
Production payment files contain real names, account numbers, and amounts. They do not go into SIT, into a ticket, or into an assistant. What you need is the structure of the failing element, so you build a synthetic file that reproduces it. I pulled one rejected file in the message viewer, read the failing region, and wrote a minimal pain.001.001.09 with test data that had the same structure.
Then validated locally before submitting anything, with a ten-line script against the exact XSD the configuration pointed to:
from lxml import etree
schema = etree.XMLSchema(etree.parse("pain.001.001.09.xsd"))
doc = etree.parse("synthetic_failing.xml")
if schema.validate(doc):
print("VALID")
else:
for err in schema.error_log:
print(f"line {err.line}: {err.message}")
It reproduced both errors. Now the cause was visible, because the synthetic file made it obvious. The client’s treasury system had been upgraded to declare version 09, but one of its payment templates, the one used for urgent same-day runs, still generated two elements in the version 03 shape:
What the urgent-run template produced, in the version 03 shape:
<ReqdExctnDt>2026-09-28</ReqdExctnDt>
<CdtrAgt><FinInstnId><BIC>EXMPGB2LXXX</BIC></FinInstnId></CdtrAgt>
What pain.001.001.09 requires:
<ReqdExctnDt><Dt>2026-09-28</Dt></ReqdExctnDt>
<CdtrAgt><FinInstnId><BICFI>EXMPGB2LXXX</BICFI></FinInstnId></CdtrAgt>
In pain.001.001.03, ReqdExctnDt is a plain date and the agent identifier element is BIC. In pain.001.001.09, ReqdExctnDt is a choice that wraps the value in Dt (or DtTm), and the agent identifier is BICFI. The standard payment runs used the updated template and passed. The urgent runs, rarer and concentrated at the start of the week when treasury cleared the weekend backlog, used the old template and failed. That was the Monday pattern in Query 1.
Submitting the synthetic file through the SIT host-to-host channel produced the same pain.002 with FF01 at group level, which closed the loop: same input shape, same system behavior, in an environment where nothing real was at risk. Scripting checks in Python takes this validation script further into a reusable check, and ISO 20022 XML traps catalogues the other version and namespace mistakes that produce the same symptom.
What is a developer analyst’s first deliverable?
A root cause note. One page. It is the developer analyst equivalent of a business analyst’s first workshop summary: it shows the team how you think, and if it is good, it becomes the template they use.
Mine, for PAY-4812:
# PAY-4812: FF01 rejections from H2H-CORP-07
## Symptom
Some pain.001 files from channel H2H-CORP-07 rejected at intake with
pain.002 GrpSts RJCT, reason FF01. First seen 2026-09-15.
## Evidence
- Query 1 (attached): all FF01 since 09-01 are from H2H-CORP-07,
all declare pain.001.001.09. Accepted files on the same namespace
exist, so the namespace alone is not the cause.
- Intake log: every failure is cvc-complex-type on <BIC> or
cvc-type on <ReqdExctnDt>. Sample: file H2HC07_20260928_0712.xml,
log entry 07:12:44.
- Reproduced in SIT with a synthetic file (attached, test data only).
## Cause
Client's urgent same-day payment template generates two version 03
element shapes inside a version 09 file: <ReqdExctnDt> as a bare date
and <BIC> instead of <BICFI>. Standard template is correct.
## Not the cause
- Our XSD and channel config (schemaVersion: 09) are correct per the
agreed onboarding spec.
- Accepted payments from the channel settle normally (Query 2).
## Fix owner
Client treasury: correct the urgent template. Relationship manager
to relay; the two element corrections are below for their vendor.
## Proposed regression test
Add both synthetic files (bad BIC, bad ReqdExctnDt) to the intake
negative suite, asserting pain.002 GrpSts=RJCT, reason FF01, and
OrgnlMsgId echoed. Prevents a future config change silently
accepting malformed version 09 files.
## Open question
Should FF01 rejections include the parser line in the pain.002
AddtlInf? Would have saved the client two weeks. Needs product view.
Three things make this note work. “Not the cause” protects the team: the first question from management was going to be “is our system wrong?”, and the note answers it before it is asked. The fix owner is outside the team, which is common and is exactly what a developer analyst is for: proving with evidence that the change belongs to someone else, politely, so the relationship manager can say it with confidence. And the open question is the first piece of feedback you give as the new person, framed as a product question rather than a criticism. Part 5, giving feedback as a new analyst, is about exactly that framing.
The note took me a day to write and the team lead forwarded it to the relationship manager unchanged. That is the moment a developer analyst stops being new.
For the end-to-end investigation method behind this, from ticket to root cause on a payment that did get past intake, see how a technical BA investigates a failed payment.
How do the five onboarding levers apply to a developer analyst?
Every role in this series uses the same five levers. For a developer analyst they look like this.
AI. Use it to locate, never to conclude. Ask for file paths and line numbers, demand NOT FOUND over guesses, and remember it only knows what you cloned. The configuration repository is where it will fail you first. Part 3 covers building a personal knowledge pack; for a developer analyst, that pack is your system-notes.md plus the schema files plus the log query library you build in week one.
Documentation. Trust the code over the wiki, and the configuration over the code. In my case the Confluence page for the host-to-host channel said “schema version auto-detected from namespace,” which had been true two releases earlier. Part 2 explains how to keep a trust register so you know which pages to believe.
Questions. A developer analyst’s questions are cheap for developers to answer if they come with a file and a line. “In XsdValidator.java:88 the schema comes from the config repo; who owns that repo, and is per-channel config reviewed?” gets an answer in a chat message. “How does validation work?” gets a meeting invite next week. Part 4 has the full question format.
Feedback. Your first feedback is almost always about observability: the log line that did not include the MsgId, the rejection that did not say which element failed. These are uncontroversial, they make everyone’s life easier, and they show you have been in the evidence.
First deliverable. The root cause note above. Pick a ticket that is real, small, and has a visible customer impact. Do not pick a refactoring or a “documentation improvement”; pick something where your evidence changes what someone outside the team does next.
A 90-day plan for a developer analyst
| Weeks | Focus | Output you can show |
|---|---|---|
| 1 | Five accesses requested on day one; repository interrogated with AI; system-notes.md started | Notes file with entry points, validation, mapping, and status code paths, each with a file and line |
| 2 | Follow one message by its identifiers through logs and database, end to end | A trace of one real payment from inbound file to final status, with every hop timestamped |
| 3 | First real ticket worked with evidence | Root cause note with queries, log lines, and a reproduction |
| 4 | Query library | Saved queries for latest status per payment, rejections by reason per day, and files per channel |
| 5 to 6 | Lower environments | Can submit synthetic files to SIT and UAT; knows the refresh schedule and who owns test data |
| 7 to 8 | Second system boundary | Same trace as week 2, through the next system: screening, the gateway, or the ledger |
| 9 to 10 | Regression contribution | Two or more negative tests added to the suite from your investigations |
| 11 to 12 | Observability feedback | A short list of log and message improvements, each tied to an incident it would have shortened |
| 13 | Review | The 90-day review with your manager, with this table as the evidence |
The rhythm is deliberate: one new system boundary every two weeks. By week 13 you can follow a payment from the corporate’s file to the ledger posting without asking anyone, and that is what the team hired you for.
What does “onboarded” look like at day 90?
Five tests. If you can pass them, you are no longer the new developer analyst, regardless of what the calendar says.
- Given a UETR, an EndToEndId, or a file name, you can produce the full timeline of that payment in under fifteen minutes, from logs and the database, without asking anyone.
- You know where behavior lives: which rules are in code, which in configuration, which in reference data tables, and who owns each.
- Your root cause notes are forwarded unchanged. Nobody has to redo your query or re-check your log lines.
- You have added tests from your own investigations, so at least two failures that happened on your watch cannot happen silently again.
- Developers ask you questions. Usually “what does the data say about X?” That is the clearest sign you have crossed over.
Practise before your next first week
If you want to rehearse this before your next start date, the Become a Technical Analyst track on the Labs drops you into Northline Pay, a fictional payments platform with an API contract, Kafka events, logs, and a database, and asks you to do exactly this: follow evidence through a system you have never seen and hand over a finding. It is the closest thing I know to a simulated first week, and a free account saves your progress.
The takeaway
Developer analyst onboarding compresses into one sentence: get read access to everything on day one, interrogate the repository with AI but verify every answer with a file and a line, follow one real message by its identifiers until you can tell its whole story, and hand over a root cause note that nobody has to redo. Remember that a file rejected at XSD validation never becomes a payment, so investigate the file, not the transaction. Do one new system boundary every two weeks and you will know the platform better than most of the team by day 90.
The SQL in this article is the part most analysts are weakest at in week one, and it is the part that turns a hunch into evidence. SQL for Business Analysts takes you from your first SELECT to queries you can defend against production data. For the wider set of technical skills behind this seat, reading code, logs, and APIs, see The Technical Skills Guide for BAs. If the channel you inherit is an API rather than a file drop, API Fundamentals for Analysts covers requests, status codes, and auth with APIs you can call while you read. And once your query library repeats every week, The BA Automation Guide shows how to schedule it. Reading the repository with AI on a team whose policy is still undecided? The AI-Powered Analyst covers how to do it without putting anything at risk. The free downloads are a good place to start.
Next in the series: part 10, onboarding as a systems analyst, where the same payment is followed across every system in the bank rather than through one.
Ahmed is a Senior Technical Business Analyst with 10+ years in banking and payments. He builds practical guides and tools for analysts at The Tech BA Toolkit.
Tags: Developer Analyst, Onboarding, ISO 20022, SQL, Payments
About the author
Analyst Engineering is written by Ahmed, a Senior Technical Business Analyst with 10+ years of banking and payments delivery experience: ISO 20022 and SWIFT messaging, payments API integration, Kafka event validation, and production support. Every article comes from real delivery work, and each one is reviewed and updated as tools and standards change.
Related articles
- The Analyst's First 90 Days: An Onboarding Plan for a New Company or Team A 90 day onboarding plan for technical analysts in four phases: orient, map, contribute, own. What to produce each week, with a banking ISO 20022 example.
- AI in the Codebase: How Analysts Read a Repository They Did Not Write Point AI at the repo and answer questions no document can: where a rule really lives, what a status actually means, what a pull request changes for the business.
- How a Technical BA Investigates a Failed Payment A walkthrough of how a technical business analyst actually investigates a failed payment: the questions, the tools, and following one transaction from the complaint to the cause.
- The ISO 20022 XML Traps: Why a Schema-Valid Message Still Gets Rejected Namespaces, element order, the business header, empty vs absent, amount precision, and code choices. The XML layer that fails messages your test tool accepts.
Go deeper on this
Not ready to buy? The free downloads are a no-cost place to start, and every article here stays free.
Free account
Practice on the Labs, keep your progress
A free account, no password: an email link signs you in. It saves your steps and self-assessments on the Labs, shows your missions on a dashboard, unlocks the solutions, and, if you tick the box, sends you new missions and articles when they ship.
Your email is used to sign you in. Nothing else, unless you ask. Privacy.