Back to Blog

Automatic Provider Failover: Handling Payment Provider Downtime Without Disrupting Users

Abstract network diagram showing failover path switching

Payment providers go down. Not often, and rarely for long, but it happens to every provider eventually. How your platform handles those minutes determines whether the outage is invisible to your users or becomes an incident you are explaining to customers the next morning.

The gap between a good outcome and a bad one almost never comes down to how fast your engineering team reacts. By the time someone notices the alerts, investigates the cause, confirms it is the provider and not your system, and pushes a fix, the window is already long. The outcome is determined before the outage starts, by whether automatic failover logic exists and where it lives.

What a provider outage actually looks like

Provider degradation rarely presents as a clean binary: working or not. The more common pattern is a gradual increase in errors or latency. API calls start timing out at a rate that looks like noise at first. Then the error rate climbs. Then calls stop completing. Then, typically, the provider's status page updates.

If your platform sends all payment traffic to a single provider, this gradual degradation hits your users in a specific way: some transactions fail, some succeed depending on timing, and the experience becomes unpredictable. Users who hit a failure typically retry, which adds load to a system that is already struggling. The provider's recovery gets slower because more retry traffic is hitting it. You end up compounding the problem.

A platform with failover logic detects the degradation signal early, routes new transactions to a secondary provider before the primary provider is fully down, and absorbs the outage without user-visible failures. The detection, decision, and rerouting all need to happen faster than a human can respond.

Detection: the signal that triggers failover

The detection problem is harder than it looks. A single timeout is not a failover signal. Error rates fluctuate naturally. The challenge is distinguishing between normal variance and genuine provider degradation in real time, without triggering unnecessary failovers that cause their own disruption.

The approach that works in practice uses a sliding window of recent responses from each provider. You track error rate (5xx responses, timeouts, and connection failures) over a short window, typically 60 to 120 seconds. When the error rate crosses a threshold, you enter a "degraded" state for that provider. In degraded state, new routing decisions deprioritize or exclude that provider. You continue sampling traffic to that provider at low volume to monitor recovery, rather than completely cutting it off.

The thresholds matter. A threshold that is too sensitive causes false failovers during normal traffic spikes. A threshold that is too loose means you wait too long before rerouting. In our internal benchmarks, a 15% error rate over a 90-second window has performed well as an initial trigger, with recovery confirmed after error rate drops below 5% for at least 60 consecutive seconds. Those numbers are not universal. They should be calibrated to your traffic volume and your providers' typical behavior.

One thing to get right early: distinguish between a provider error that means "this transaction failed and should be retried elsewhere" and a provider error that means "this transaction may have processed and you should not retry it." A 5xx response or timeout often falls into the second category. Blindly retrying a payment that may have already been submitted to the provider is how you end up with duplicate charges. The failover logic has to check idempotency before rerouting.

The routing decision during failover

When the primary provider is in degraded state, the routing layer needs to select an alternative. This is not as simple as "use the backup provider." A few things to consider.

Not all providers handle all transaction types equally. Your fallback for PayNow-style instant transfers may not be the same as your fallback for cross-border payments. The routing logic needs to understand which providers are capable of processing each transaction type, not just which ones are currently healthy.

Fee structures differ between providers. A failover to a more expensive provider is acceptable during an outage. What is not acceptable is using the outage as an implicit policy to route traffic to the cheapest provider regardless of health. The routing decision during failover should be: find the healthiest provider that can handle this transaction type, subject to user-configured preferences. The cost optimization logic takes a back seat to availability during degradation.

In-flight transactions need special handling. A payment that was submitted to the degraded provider before failover triggered should be tracked separately. If the provider eventually processes it, you need to know not to retry it. If the provider times out definitively (not just slowly), you need to know it is safe to resubmit. Maintaining this state per transaction, not just per provider, is what prevents both lost payments and duplicates.

Recovery and traffic rebalancing

Recovery is often less-thought-about than the initial failover. Once a provider recovers, you do not want to immediately send 100% of traffic back to it. The pattern that works is gradual traffic reintroduction: start with a small percentage (say, 10%), monitor error rates for a few minutes, increase the percentage if the provider looks stable, reach full traffic over 10 to 20 minutes.

This matters because providers often recover unevenly. The system may be up but processing slowly. Sending a full traffic load immediately can cause a secondary degradation event as the recovering provider gets overwhelmed before it has fully stabilized.

There is also the question of what to do with the secondary provider once the primary recovers. If you have routed significant volume to the secondary during the outage, you may want to keep a portion of that traffic there rather than snapping back immediately. Outages are useful forcing functions: they often reveal that the secondary provider is performing well and that a permanent traffic split is worth considering.

Where the logic should live

The critical design question is where this detection and routing logic runs. The two options are: in your application code, or in a dedicated payment routing layer.

In application code, failover logic tends to be incomplete. It handles the case the engineer thought of when building it, often the total failure case, but misses the partial degradation case, the idempotency case, and the gradual recovery case. Maintaining that logic across multiple services that touch payments is also expensive. When a provider changes its error format or latency characteristics, you have to update every service.

In a dedicated routing layer, failover logic runs once, is tested against the actual behavior of each provider, and applies consistently to every payment. Application code makes a single call to the routing layer. The routing layer handles degradation detection, provider selection, idempotency enforcement, and recovery. Your application does not need to know any of it happened.

The practical test for whether your failover logic is good: can a junior engineer on call during a provider outage at 2 AM do nothing and have it resolve on its own? If yes, the logic is in the right place. If the answer is "it depends on which engineer is on call and whether they know which configuration to update," the logic belongs in the infrastructure layer, not in application code.

Reconciliation after failover

One operational detail that gets overlooked: when you fail over to a secondary provider, the reconciliation state for transactions that were in flight becomes complex. You have transactions that were submitted to Provider A, timed out, and then submitted to Provider B. You need to reconcile against both providers' settlement records for those transactions, identify which one actually settled, and ensure the other one did not also settle.

This is not rare. It is exactly what happens during failover. Your reconciliation infrastructure needs to handle the multi-provider, multi-submission case without requiring manual intervention. The reconciliation record for each transaction should track all submission attempts, not just the final settled one.

Getting both the failover and the reconciliation right is what turns a provider outage from an incident into a background event that your users never see and your finance team handles with a routine end-of-day review.

Build on Checker

Payment routing, reconciliation, and compliance for Southeast Asian fintech platforms. Start free, scale as you grow.

Get API Key Free