Incident Report

Live stream delivered no strikes on 1 September

A routine server update caused an unplanned database failover. The service stayed online and answered requests normally, which is exactly why the problem was not obvious, to you or to us.

Resolved1 September 2026Historical data restored

Affected window

06:46-21:45

UTC, 1 September 2026

Status

Resolved

Service fully restored

Historical record

Complete

Affected window fully recovered

What happened

At 06:46 UTC on 1 September 2026, a routine automated maintenance update on our primary database server restarted a component that our high availability system treats as a signal that the server has failed. The cluster responded exactly as designed and promoted the standby server.

The failover itself completed cleanly in seconds, and the API came back up on the second server without interruption. What did not resume was new strike data. A configuration difference between the two servers meant that part of the platform did not start automatically on the promoted machine, so the API continued to answer requests normally while its data stopped advancing.

The result was a service that looked entirely healthy from the outside and had nothing new to serve.

06:45:59 UTC → 21:45:09 UTC

No new strike data was available during this window. The most recent strike before the interruption was timestamped 06:45:59 UTC.

How this affected you

Because the API never went offline, the failure was silent rather than loud. Depending on how you use the service:

  • If you use the WebSocket stream: your connection was established and authenticated normally, and you received the usual confirmation frame with your tier and filters. After that you received no strike messages. No error was sent, because from the service’s perspective there was genuinely nothing new to send.
  • If you use the REST endpoints: requests continued to return 200 OK with valid responses. However, no data newer than 06:45:59 UTC was available, so recent time windows returned empty or unchanged results.
  • If you use zone alerts: no proximity alerts were evaluated or delivered during the window.

No stored data was lost or corrupted. Everything from before the interruption remains intact and accurate. Nothing was deleted: the affected window is an absence of new data rather than the removal of existing data, and we have since restored the missing period in full. Historical queries covering 1 September now return a complete record.

The one thing that cannot be repaired is the live delivery itself. Strikes that would have been pushed to you in real time during the window were not, and no backfill can undo that. If you acted on the absence of strikes during those hours, treat that period as having had no coverage rather than no activity.

If you rely on historical completeness

The 06:46 to 21:45 UTC window has been fully recovered and is available through the regular historical endpoints. If you pulled and stored data covering 1 September before the recovery completed, re-run that query to pick up the full record. If the window matters for your reporting and you want confirmation of its state, contact us and we will tell you where recovery stands for that period.

Timeline

All times UTC on 1 September 2026

Automatic server update triggers an unplanned database failover.

Standby server promoted. API remains online throughout.

New strike data stops.

A configuration difference between the two servers left part of the platform inactive after the promotion.

Issue reported by a customer and investigation begins.

Our monitoring did not cover this condition. This is addressed below.

New strike data resumes. Live stream begins delivering strikes again.

Restored without any further interruption to connected clients.

Service returned to the original primary server.

Brief planned interruption of under 90 seconds. Streaming clients reconnected normally.

All systems verified healthy. Incident closed.

Strike volumes confirmed at normal levels.

Missing window fully recovered.

Strikes from the affected period restored and verified complete.

What we have changed

Completed

Automatic updates can no longer trigger a failover

The component whose restart caused this is now excluded from automatic update handling on both servers. This has been applied and verified.

Completed

The affected services now run on both servers

Everything involved in this incident, covering the strike data path and alerting, is now present and verified on both servers. A future failover brings those up automatically rather than a subset of them.

Completed

Failover verified in both directions

We returned service to the original server and confirmed the standby rejoined correctly, with data replicating and no lag between the two.

Completed

Missing window restored

Strike data for 06:46 to 21:45 UTC on 1 September has been recovered and verified. Recovered strikes are identical in content and accuracy to those from any other period. Historical queries covering that window are now complete.

In progress

Alerting on data staleness

The most important gap this exposed is that our monitoring covered the health of the service but not the freshness of its data. We are adding a check that alarms when the most recent strike is older than a few minutes. It would have caught this within minutes of it starting, and it catches this whole category of problem regardless of the underlying cause.

In progress

Enforced consistency between servers

The length of this outage came down to a configuration difference between our two servers. We are putting an automated check in place so any such difference is reported immediately rather than surfacing during an incident.

What we take from this

The failover machinery worked correctly. Your data was never at risk, replication never fell behind, and nothing already stored was lost or damaged.

What failed was quieter. A configuration difference between our two servers, and a gap in what our monitoring covered: the health of the service rather than the freshness of its data. This reached us through a customer report rather than our own alerting, and the changes above are how we make sure the next one does not.

We are sorry for the disruption, and specifically for the fact that the service gave you no signal that anything was wrong. A stream that goes quiet without saying so is worse than one that reports an error, and improving that behaviour is part of the follow up work above.

Published 1 September 2026. Current service health is on the status page. If you have questions about this incident, or need to know whether the affected window matters for your account, get in touch.

Coverage areas

Lightning data provided as-is; not for safety-critical use. Commercial use is permitted on every current plan. Read the EULA →