mksim.pro
Back to all posts
IT 9 min read

When the provider is the incident: hybrid infrastructure and somebody else's maintenance

Azure had an infrastructure incident that hit hybrid scenarios, VPN and cloud VMware services. According to The Register, the trouble came out of internal maintenance work. Why this is the nastiest class of outage to own, and what actually helps.

Microsoft Azure had an infrastructure incident that affected hybrid scenarios, VPN and cloud VMware services. According to The Register, the problems arose during internal maintenance work on the provider's own side. The report carries no detail on scale or duration, and I am not going to invent any.

What matters here is the class of failure rather than this particular one. It is an outage you did not cause, did not schedule and cannot call off. Your load did not spike, nobody attacked you, you shipped nothing. Something changed inside your supplier, and it reached you as broken connectivity. For an owner or a CTO this is worse than any traffic peak, because the usual levers do not move anything. The lever is in another company.

Why somebody else's maintenance beats your own peak

A traffic peak is a predictable enemy. You know your seasons and your reporting periods. You can scale up in advance, run a load test, put people on call, hold releases. A peak has a calendar, and the calendar is yours.

Provider maintenance works differently, for three reasons.

  • You do not know the timing. Work that the provider classes as internal and safe may not appear in its public schedule at all. Your change calendar and theirs live separate lives.
  • You cannot postpone it. Your own maintenance window moves with one decision from a director. Theirs moves with nothing, and a support ticket raised mid-incident does not speed up somebody else's repair.
  • You do not control the rollback. You roll back your own release and you know roughly how long that takes. When their change breaks, your recovery time equals their recovery time.

On top of that sits an information asymmetry. The provider knows exactly what it changed and sees its own telemetry end to end. You see consequences, plus a public wording that arrives late and stays careful. While the status page is silent or talks about degradation in general terms, your team looks for the cause at home by default: rechecking the last release, the firewall rules, the certificates. Those hours go into disproving theories that were wrong from the start.

Which leads to an uncomfortable planning conclusion. An availability model where every failure is assumed to be yours understates risk in a systematic way. Part of your downtime comes from a zone where you have neither a button nor information. What the contract actually says about that I covered separately: public cloud SLA.

Why the hybrid seam tears first

The incident report names hybrid scenarios, VPN and cloud VMware services specifically. That is not a random set. It is the thinnest part of connectivity, and it depends on two parties at once.

Inside a single data centre, traffic runs over a network owned by one operator, who has both the telemetry and the ability to fix things. The hybrid seam is built from different material: tunnels and gateways, routes, virtualisation mirrors, agreed versions and timeouts, credentials and certificates. At every one of those points, state is held by both sides simultaneously, and both have to read it the same way.

The practical effect is this. A broken tunnel does not feel like a service going down, it feels like weirdness. Replication lags. Authentication reaches for a domain that now sits behind a severed link, and the application on the far side simply thinks longer than usual. Half the system works, half does not, and the team spends the expensive first minutes working out what is even happening. If an incident opens with "is this us at all?", time is already gone.

The return trip is its own trap. Connectivity comes back and the systems stay out of sync: queues have piled up, replicas have drifted, background jobs fired twice, some operations ran more than once. The moment everyone treats as the end of the incident actually opens its second half, and that half often runs longer than the first. Hybrid architecture is sensible in itself and usually unavoidable, which I wrote about here: on-premise vs cloud. The bill for it arrives in moments like this, and it is paid in recovery complexity.

Monitoring goes blind exactly when you need it

The most expensive part of the story is observability. Monitoring usually lives where the workload lives: same subscriptions, same network, same gateways. It is cheaper and simpler to collect that way. The saving has a price, and the price is charged once, in full.

When the shared layer breaks, the dashboard stays green. Not because things are fine, but because the bad news never arrived. Alerts stay quiet for the same reason: the delivery channel sits inside the same blast radius as the fault. The on-call engineer watches a calm screen while the first real signal comes from a customer phone call.

That is the nastiest class of outage. The cause is at the supplier, your metrics insist you are healthy, and the business is stopped. Worse, the green dashboard does active harm: it grants false confidence and pushes the start of your response out to the point where complaints arrive from customers rather than from systems. How to build observability that answers questions instead of generating noise I covered here: from logs to metrics and back. For today, one specific point decides everything: the vantage point has to sit outside the thing it watches.

What to actually do

  1. Put an independent vantage point outside. The check should live with a different supplier and reach your services the same way a customer does: external synthetic monitoring of key journeys, plus an alerting channel that does not depend on corporate mail or the VPN. It is cheap insurance that pays for itself in the first outage, because it answers the question that dominates the first minutes - does the service work from outside.

  2. Monitor the hybrid seam as an object in its own right. Tunnel state, latency and loss between sites, replication lag, authentication response time. These break before your business metrics do and buy you time.

  3. Sort workloads by how much failure they must survive. Make the decision once, honestly: which services must survive the loss of a region, which can be down for an hour, which can be down for a day. The list is usually shorter than it feels in the moment, and protecting everything else spends money on resilience nobody needed.

  4. Verify that the move is actually executable. A plan to shift a critical workload to another region or another environment is worth exactly what its last rehearsal was worth. That means drills with a stopwatch, on real data, with a real cutover. Where to start: cloud exit plan.

  5. Subscribe to the provider's status pages and notifications in advance. It sounds trivial, but mid-incident the gap between "we spent two hours looking at ourselves" and "we saw the supplier notice in five minutes" is measured in money. Agree internally who watches those channels in the first minutes.

  6. Write down the protocol for a supplier outage. Who declares the incident, who says what to customers, when degraded mode goes on. The call that the cause is external should follow a rule rather than the instinct of whoever is on shift.

  7. Rehearse the return, not only the failure. How you reconcile divergence once the link is back, what happens to work that ran twice, in which order systems come up. This part is almost always untested and almost always the longest.

Honest caveats

  • Multi-region and multi-cloud cost money and add complexity. Bills grow, running costs grow, and you acquire your own failure modes in replication and failover. Architecture built to survive a rare outage sometimes breaks more often than the outage it was meant to survive.

  • Deciding what deserves that protection is a business call. The CTO brings the options and the price, the owner picks the level of risk. Trying to protect everything equally ends with a budget spread so thin that nothing is genuinely protected.

  • SLA credits do not cover your loss. At best you get service credits proportional to the unavailable time. Lost orders, idle staff and the conversations with customers do not enter that calculation, and the contract is a poor substitute for insurance against consequences.

  • The independence of your vantage point is relative. A second monitoring supplier may sit on the same backbone or the same registrar. Full independence does not exist, but partial independence still beats watching from inside the outage.

  • A provider failure does not transfer your responsibility to the customer. The customer sees your service. "Our supplier had an outage" explains the cause and cancels no obligation. Who decides what in that moment I covered here: the first 30 minutes of an incident.

In short

  • The Azure incident, per The Register, traces to internal maintenance work and hit hybrid scenarios, VPN and cloud VMware services.
  • Somebody else's maintenance is worse than your own peak: you do not know the timing, cannot postpone it and do not control the rollback.
  • The hybrid seam tears first because both sides hold its state at once, and restored connectivity leaves systems out of sync.
  • Monitoring that lives in the same cloud shows green precisely when you need it most. An external vantage point and a separate alerting channel close that gap cheaply.
  • Surviving the loss of a region costs money. Which workloads deserve it is a business decision, and it is validated by drills rather than by a document.

If you run hybrid infrastructure and do not know what your dashboards will show on the day the fault sits with your supplier, that is worth finding out in advance. Discuss your case through the form on the home page, and a first conversation commits you to nothing.

Back to all posts
Contact

If this resonated, write to me. I reply personally.

WhatsApp