Some time shortly after 7 a.m. Monday, Oct. 4, one of the major players in the social media space suffered what turned out to be a major service breakdown.
Facebook isn’t just a bit player. Through its various properties it is a social media behemoth, and the lengthy interruption to service for its various platforms will have ripple effects for some time to come. It may lead to regulatory oversight, possibly even a mandate to separate some of those intertwined services.
Initially, users on one of the Facebook-owned services that Monday morning noticed minor hiccups. For instance, although I could read Facebook posts and comments, I was unable to reply or comment. Initially I wrote that off as simply being a post or comment having been deleted by the author, even though such content remains visible until a refresh.
However, when I encountered a similar problem on my Instagram feed, I suspected Facebook was experiencing a problem of sorts. Exactly what it was I didn’t know, and it was not for another hour or two that the Facebook tab on a couple of my devices finally went into a “site not found” mode.
At that point I checked on another Facebook-owned service, WhatsApp, and found it too could not post content.
By this time the service loss period was stretching into a third hour. Certainly the social media giants have outages, but stretching beyond three hours is very uncommon nowadays. By 11 a.m. the first articles from cybersecurity experts were beginning to appear online.
At his KrebsOnSecurity site, respected reporter and analyst Brian Krebs was one of the first to post a “what we know” description of the outage. His opening paragraph summarized the situation succinctly: “Facebook and its sister properties Instagram and WhatsApp are suffering from ongoing, global outages. We don’t yet know why this happened, but the how is clear: Earlier this morning, something inside Facebook caused the company to revoke key digital records that tell computers and other Internet-enabled devices how to find these destinations online.”
Meanwhile, other platforms were filling up with posts expressing everything from anger to anguish, not just from people unable to post cat photos, but also from users who’ve built businesses around a Facebook and/or an Instagram presence. The stock market took note of the latter and Facebook stock dropped around five points, although some of that was due to a whistleblower’s appearance on 60 Minutes the night before to describe some of the company’s unsavoury business practices.
At the core of the breakdown in service was a software error in an audit command that essentially made all Facebook servers invisible, and in turn unreachable from the internet. The company was actually forced to provide initial statements on the outage through a Twitter account.
As if the six- to eight-hour Monday outage wasn’t damaging enough, the California-based company suffered a shorter loss of service on Friday, rounding out a work week that may leave lasting implications. In a terse press release Facebook stated “We know how much you depend on us to communicate with one another. We fixed the issue – thanks again for your patience this week.”
To its credit the company did its best to keep users of its services somewhat apprised of developments. Relatively early on its VP of infrastructure, Santosh Janardhan released a statement, that while short on specifics, at least set a tone of contrition.
“People and businesses around the world rely on us every day to stay connected. We understand the impact that outages like these have on people’s lives, as well as our responsibility to keep people informed about disruptions to our services. We apologize to all those affected, and we’re working to understand more about what happened today so we can continue to make our infrastructure more resilient.”
By the next day that same VP was ready to provide the details that some were demanding. He didn’t disappoint and essentially confirmed that an update had gone awry and that the company’s own security measures actually kept employees at bay for a while as they attempted to right the Facebook ship.
Santosh gave a lengthy summary of the service outage, summing it up this way: “We’ve done extensive work hardening our systems to prevent unauthorized access, and it was interesting to see how that hardening slowed us down as we tried to recover from an outage caused not by malicious activity, but an error of our own making. I believe a tradeoff like this is worth it — greatly increased day-to-day security vs. a slower recovery from a hopefully rare event like this. From here on out, our job is to strengthen our testing, drills, and overall resilience to make sure events like this happen as rarely as possible.”
Follow me on Facebook (facebook.com/PeterVogelCA), on Twitter (@PeterVogel), or on Instagram (@plvogel).
