A provider's refusal is not our defect, and reporting it as one costs days
Affects: backend/src/integrations/http.ts, backend/src/middlewares/error-handler.ts, backend/src/lib/errors.ts, backend/src/schemas/common.schema.ts, public/openapi.json, backend/test/error-handler.test.ts (six new assertions), app/api/billing/checkout-session/route.ts.
What was wrong
Every ProviderHttpError fell through to the error handler's final branch and came back as 500 internal_error — "The request could not be completed." — with a request id and nothing else. That branch is written to be strict on purpose: its comment explains, correctly, that a raw Postgres message names tables and columns and that a unique-violation text can confirm an email exists. The strictness was right. Applying it to a Stripe refusal was not.
The two failures are not the same event and they do not want the same audience. A 500 says a defect was shipped, the response is deliberately uninformative, and the log is where the answer is. A provider declining a parameter says the request was wrong, the provider already explained why, and nobody needs to open a dashboard. Collapsing the second into the first threw away an explanation that had already been written and paid for.
What it cost, measured rather than asserted
The billing runner (D-073) failed on deploy and returned upstream_refused with upstreamStatus: 500. That was enough to establish that the frontend was fine and the backend was not — which is the whole reason upstream_refused and upstream_unreachable are separate codes, and it is the second time that separation has paid for itself. It was not enough to establish anything else.
The reason then existed in exactly one place: a hosting dashboard's log viewer. Across two working sessions that viewer could not be read — the table is virtualised, so the rows never enter innerText, and the only visible trace was a filter facet reading "Error 4". A failure that was live, reproducible on demand, and completely undiagnosable from outside is not a logging problem. It is a design problem in what the API chooses to say.
The decision
ProviderHttpError answers 502 provider_error, carrying the provider's name, the status it answered with, and its own identifier fields — Stripe's code/param/type, QuickBooks' fault code, Salesforce's errorCode. The frontend proxy forwards the same identifiers under upstream, behind an allowlist. A reader now gets param: "line_items[0][price_data][unit_amount]" in the response body instead of a request id and a trip to a dashboard.
The message is not carried, at any layer. That is the line, and it is drawn where it is for a specific reason: identifiers are bounded vocabularies, and prose is not. Stripe quotes the rejected value back into its message, and on this API a rejected value is routinely a customer's name or email address — the errorHandler test asserts exactly that, using a body containing both, and fails if either reaches the wire. So the useful half travels and the dangerous half stays in the log, rather than the previous arrangement where both stayed in the log and the endpoint said nothing at all.
Three smaller choices inside that:
The fault is extracted in the `ProviderHttpError` constructor, not by whoever catches it. The constructor is the only place the untruncated body still exists — body is cut to 400 characters, and a cut JSON document does not parse. A caller doing this afterwards would succeed on short refusals and silently return nothing on long ones, which is the wrong way round: the long ones are the interesting ones.
The keys are found by walking the document, not by a path per provider. The three providers nest them differently and a walk finds all three without this file encoding three schemas that only the providers may change.
The frontend forwards an allowlist, not the error object. The upstream may add a field at any time, and the next field it adds should not reach a public marketing page merely because nobody thought about it.
The alternative that was rejected
Keep the 500 and read the logs. That is what was in place, and it is what failed twice. It also assumes the reader has console access — true for the operator, false for every integrator who will ever call this API, and the API is a public artifact with a published contract. Requiring a hosting account to learn that a parameter was rejected is not a contract.
What would make this wrong
A provider that puts customer data in code or param. None of the three does — they are enumerations and parameter paths — but the walk is deliberately narrow (three key names, 120-character ceiling, no free text) so that a provider getting this wrong leaks an identifier-shaped string rather than a paragraph. If one ever does, the fix is to drop that provider's key from FAULT_KEYS, not to widen what is forwarded.
The other way it goes wrong is scope creep: someone adds message to FORWARDED_FAULT_FIELDS because a particular error was hard to read. The test that greps a customer's name out of the response body exists to make that a failing test rather than a quiet regression.
What this does not do
It does not fix the checkout. The runner still fails; this decision is about making the reason legible rather than about the reason itself. Recording it separately is the point — the diagnosis tooling was a real defect on its own terms, and it would have been one even if the checkout had worked.