Architecture & Operations
When payment failover is actually safe
Failover is a finality problem, not only a routing problem. Where it is safe, where it is not, and what runs today.
The rule to remember: before sending an operation to another provider, ask whether the previous attempt can still move money. If it cannot, choosing the next provider is a routing decision. If it can, or if you do not know, a second attempt risks a second economic effect.
The job
A payments team adds a second provider for the same rail. The first failover proposal fits on one line: if Provider A returns an error, send the operation to Provider B. That line is correct for some errors and costly for others. This guide sorts out which, for pay-ins and for payouts, and states what Orangepill does today.
Before you start
- Written for CTOs, heads of payments and engineers deciding failover behaviour across providers.
- Assumes more than one provider can handle the same product, currency and destination.
- Assumes you can say, per provider, what a final outcome looks like: which statuses, which errors, and when a settled transfer can still be returned.
- Maturity: pay-in routing by priority and failover for payment requests are in production. Payout failover and Resolve are in development.
Flow
- Attempt 1 Provider A
- Error or silence What kind of failure?
- Is attempt 1 final? Can it still move money?
- Final, nothing moved Attempt 2 with Provider B may be safe
- Completed Stop. Record the outcome
- Unknown Wait, query, reconcile. No failover
The reasoning
1. A final decline: another attempt may be safe
Provider A rejects a payout because your account with it is not enabled for the destination bank. It returns a 4xx response with a reason, and no transfer reference. Nothing was accepted, so nothing can settle. Sending the payout through Provider B is a routing decision.
Two checks come first. Is the decline final for this provider, or can the same request still be processed later? And did the provider create anything (a pending record, a reference, a reserved amount) before rejecting? Most declines are clean. Some are not, and the difference sits in provider documentation and behaviour, not in the HTTP status code.
This is why Orangepill's automatic failover stops on a provider rejection today. A rejection is only treated as safe once it is known, for that provider, not to leave anything behind.
2. A timeout or an unknown outcome: another attempt can be dangerous
- 10:00:00 A payout is sent to Provider A.
- 10:00:30 The request times out. Provider A may or may not have accepted it.
- 10:00:31 Health checks mark Provider A as degraded. Provider B is healthy and eligible.
- 10:00:32 The failover rule sends the payout to Provider B. It settles in seconds.
- 10:04:10 Provider A, which had accepted the first request, settles it too.
The health check was right: Provider A was slow. It said nothing about the payout already inside Provider A. Health tells you where to send the next operation. It does not change the state of an attempt already in flight. See A timeout is not a decline for the evidence you need before a second attempt.
3. Pay-ins and payouts carry different risks
| Pay-in | Payout | |
|---|---|---|
| Who acts | The payer authorises or completes the payment. The provider collects. | Your platform instructs. The provider sends, with no further step from anyone. |
| What an attempt often creates | A request for the payer: a redirect, a QR code, a key. Money moves only when the payer acts. | A transfer instruction that can settle on its own. |
| Cost of a duplicate | An unused payment request, or a payer charged twice who needs a refund. | Value sent twice. Getting it back usually needs the beneficiary's cooperation. |
| When the outcome is unknown | Confirm with the provider before asking the payer to pay again. | Hold the funds, query, reconcile. Send again only on evidence. |
This difference explains a choice in Orangepill's pay-in failover. When creating a payment request (a Bre-B dynamic QR code or key) fails with a network error, including a timeout, Orangepill asks the next eligible provider. At that step no money has moved: the payer has not seen a QR code yet. If the first provider created a request but the response was lost, that request is never shown to the payer. The same rule applied to a payout would be unsafe, which is why payouts do not fail over today.
4. Idempotency is not finality
Idempotency deduplicates requests sent to you. Orangepill's idempotency_key makes sure one key
produces one payout record, however many times your client retries. It cannot tell Provider B that Provider A
already paid, and a provider's own idempotency reference means nothing to a different provider.
Finality is a property of the money: the point after which an outcome can no longer change, except through a new and visible event such as a return or a reversal. Each rail reaches it differently. Some are final within seconds of acceptance; others only at settlement; some pay-ins can be disputed long after they succeeded. A failover rule needs finality per provider. Idempotency alone cannot supply it.
5. Decision table
| What you know about attempt 1 | Pay-in | Payout | Orangepill today |
|---|---|---|---|
| Configuration error before any provider call (bad credentials, inactive integration) | Safe to try the next provider | Safe to try the next provider | Payment requests fail over automatically. Payouts are not retried elsewhere. |
| Connection refused or DNS failure: the request never reached the provider | Safe to try the next provider | Safe, once you know the request was not delivered | Payment requests fail over. A payout is marked failed (INITIATE_FAILED). |
| Explicit rejection with a reason, no reference returned | Usually safe, if this rejection is known to leave nothing behind | Usually safe, after confirming no transfer exists | Payment requests stop. A payout is marked failed (INITIATE_FAILED). |
| Timeout, connection reset or 5xx after the request was sent | Acceptable only if the step creates a payer-facing request and nothing was collected | Not safe | Payment requests fail over (no money has moved at that step). A payout is held for operator attention and not resent. |
| Provider reports a duplicate reference | Not safe: the provider already holds something | Not safe | Payment requests stop. |
| Provider status is pending or processing | Wait | Wait | Payouts keep querying the provider for a final status. |
| Provider status is a final failure | A new attempt may be safe | A new attempt may be safe after reconciliation | Payout failed, reservation released. A new attempt is your decision, as a new payout. |
| Provider status is completed | Never fail over | Never fail over | Payout completed, journal posted. |
| Error you cannot classify | Treat as unknown | Treat as unknown | Payment requests stop. |
A payout that ends failed with RETRY_EXHAUSTED belongs in the "unknown" rows, not the
"final failure" row. It means Orangepill stopped checking without a final answer from the provider, so the
transfer may still settle. Do not send it again, through any provider, until the provider or a statement
confirms it did not.
6. What runs in Orangepill today
- Pay-in routing by priority Production. Each product has routes to
provider integrations, ordered by priority and configured per tenant. On
POST /v4/checkout/paymentsand payment intents, one provider handles each request: the one you name, or the first active route by priority. There is no automatic failover on these paths. If a payment endsfailed, creating a new payment through another route is your decision.POST /v4/payments/explainshows how routing evaluates your configuration; it is advisory. - Automatic failover for payment requests Production. For payment requests (Bre-B dynamic QR and key) and checkout-session attempts, Orangepill moves to the next eligible provider only on configuration or network errors. It stops on a provider rejection, a duplicate reference or an unknown outcome. The failover trace is not exposed to clients.
- Payout failover In Development. A payout that fails or times out is not sent to another provider.
- Resolve In Development. The operations that cannot safely continue, the "unknown" rows above, are what Resolve is being built to govern: Investigate, Establish, Decide, Verify, ending in Continue, Wait, Remediate or Escalate. It has no API today.
What can go wrong
- "Any error means try the next provider." A timeout on a payout becomes two transfers when the first provider settles late.
- Health checks drive failover for attempts already sent. A provider marked unhealthy can still complete what it accepted.
- A new idempotency key for the failover attempt hides the link. Two payouts exist with nothing
connecting them. Link attempts with your own reference in
metadata. - All 4xx responses are treated as clean declines. Some providers create a record before rejecting. Classify per provider.
- A pay-in is retried while the payer is still completing the first one. The payer finishes both.
- A late reversal after a "final" success. Finality on some rails is not the end. Plan for returns as new events, not as reasons to fail over.
What Orangepill guarantees today
- Pay-in routing selects one provider per request by priority, or the provider you name.
- Payment-request failover moves on only after configuration or network errors, and stops on rejection, duplicate reference or unknown outcome.
- A payout is never sent to a second provider automatically.
- A timeout, reset or 5xx when sending a payout to the provider does not mark it failed or trigger a resend. It is flagged for operator attention.
- One payout per
idempotency_keyin your tenant.
What Orangepill does not guarantee
- Payout failover across providers. It is in development.
- Automatic failover on
POST /v4/checkout/paymentsor payment intents. - That the provider shown by
POST /v4/payments/explainis the one that executes. - Exactly-once execution across providers. Idempotency applies inside Orangepill.
- Visibility of the payment-request failover trace through the API.
- Automated handling of unknown outcomes. Resolve is in development.
Production considerations
- For every provider, write down what "final" means: which statuses, which error codes, and how long a settled transfer can still be returned.
- Classify provider errors into four groups: configuration, not delivered, rejected, ambiguous. Anything you cannot classify goes into ambiguous.
- Configure routes and priorities per product in Console, and check them with
POST /v4/payments/explainbefore launch. - Make any payout resend through another provider a human decision today, backed by a provider status, a statement line or a reconciliation result.
- Read
GET /v4/payments/{paymentId}/attemptsto see which provider handled each attempt. - Alert on attempts that stay non-terminal past the rail's normal time.
- In sandbox, test a decline, a network error and a timeout separately. They should produce three different behaviours.
To set up a second provider for pay-ins, follow Add a second payment provider without rewriting your product. The architecture is described on Multi-Rail Orchestration.
Run your failover rule against a timeout.
Bring the rule your stack uses to try another provider. We will walk it through a decline, a timeout and a late success.