Tag: disaster recovery

  • Big Power Issues

    Recently British Airways had a massive computer failure and had to cancel hundreds of flights across a couple days. This was a large disruption for travel as tens of thousands of passengers were stranded around the world. In addition to the concerns many of us have about security, I also worry about large scale systems that continue to grow in scale, with more users depending on them, and not necessarily getting updated with modern technologies. While I know many of the customer facing systems have been enhanced for companies like airlines, I’m not sure the core infrastructure that backs these systems has changed.

    British Airways has blamed the cause of the outage on a power failure, with some finger pointing between the airline and their outsourcing company. BA appears to say a power surge caused problems with it’s UPS’s and batteries, which resulted in a situation where “the controlled contingency migration to other facilities could not be applied.”

    I’m sure some of the information being published is carefully vetted by lawyers and doesn’t necessarily reflect the actual technical issue, but it doesn’t matter. A well designed HA solution shouldn’t care if a primary system drops off line, falters, or anything else. In the worst case, any IT system failover can be forced, understanding a potential loss of data. In the SQL world, we could always remove a primary and just go with whatever data is on the secondary. Certainly within a few hours we could be up and running.

    I suspect that the architecture of the BA system is not well designed for large failures. I know that dealing with power is tricky. While working for a large company (10k+ employees), our data center went down one day, which affected our public presence, customer support, and plenty of internal systems. We had UPS service underway, which had disconnected main power for some reason and we couldn’t quickly switch to the public grid when a system failed. It took a few hours to reroute power, during which our CTO paced the data center floor in an irate manner. Certainly that didn’t speed things along.

    As noted in the piece I’ve linked, there are some strange inconsistencies with the explanation. It’s entirely likely that a data quality issue caused problems that took hours to rectify. I’ve been on the wrong side of those, working through queries and attempting to piece together some version of the correct data from various sources. If that’s the case, then I’d really like to know what the faliure of their systems allowed bad data to disrupt operations. There are learning opportunities here for us data professionals.

    I do hope that some of the details of how BA architected their system gets shared among technical staff. At the least, I’d like to have large companies like UPS, Wal-Mart, British Airways, and other companies publish and share information about technical details and architectures. Some of the tech companies do this already (Amazon, Google, etc.), but it’s good for other industries to share their information as well. We owe it to the industry to learn what works well and what fails as we grow systems to larger scales and more complex interactions. Certainly I’m proud that many SQL Server experts share their experiences, helping others learn and make fewer mistakes in the future.

    Steve Jones

    The Voice of the DBA Podcast

    Listen to the MP3 Audio ( 5.4MB) podcast or subscribe to the feed at iTunes and Libsyn.

  • The Quiet Zone

    I’ve been in a data center when most servers turned off. I’ve actually heard dozens of systems powered off quickly, and it’s a strange sound. You become so used to the white noise of numerous fans that having them turned off is a little unnerving. It’s neat when it’s a scheduled patch day and all servers cleanly shut down together. It’s an altogether different experience when there’s an unexpected issue and management sees their expensive hardware not working.

    However, imagine losing your servers because of a loud noise. That’s what happened to ING Bank when a fire extinguishing test caused a number of hard drives to fail. To be fair, the loud noise was north of 130db, which is very loud. Since sound is really vibration, the impact to read/write heads caused numerous failures. The bank needed 10 hours to restart systems in their DR center, and managed to do so. While that might not have been what the bank officials wanted, this is a good DR test, and I hope they learned a few things that might help to fail over much quicker in the future.

    This might be a good reason to think about SSDs, which are less susceptible to vibration than the older, spinning rust drives. I’d guess that there are other issues that could affect SSDs and someone is going to discover them at an inopportune time. Already we’ve seen dramatic improvement in SSD technology, driven by numerous early issues relate to writes and reliability.

    Engineering facilities is hard, and there can be many unexpected issues. I’m sure the people that designed the fire suppression system weren’t concerned about the noise; they were concerned about shutting down flames quickly. I’m sure that the people filling the system didn’t think a little extra pressure would matter. These seemingly innocuous decisions can cascade, which is why we practice and preach DR preparation. Not just backups, but restores and quick fail over.

    If it’s not your organization, it might be humorous rather than stressful, but you never know what design flaws might lurk inside your facilities. I once worked in a data center that had only about half the cooling that we expected. Why? The engineers assumed that since we worked an 8 hour day, so did the computers, which we’d turn off at night. Luckily they had built a pad into their calculations so we were only short half the capacity rather than two thirds.

    Steve Jones

    The Voice of the DBA Podcast

    Listen to the MP3 Audio ( 3.7MB) podcast or subscribe to the feed at iTunes and Libsyn.

  • Data Loss or Downtime

    This editorial was originally published on Jan 7, 2011. It is being re-run as Steve is on vacation.

    I was watching Kimberly Tripp of SQL Skills talk recently about VLDB disasters and how to recover from them. One of the first things she said in the session was getting a damaged database back online, even without all of the data, was important. Often her clients need to keep working, and it is important that they get the system back online, even without all the data. This allows business applications and business people to get back to work.

    That is interesting. I had always thought of my production OLTP databases as needing to be online, but also needing all the critical data. After Mrs. Tripp’s talk, I had to rethink that a bit and consider that a little data loss might be acceptable.

    To me this is a topic that is worth understanding. At the very least, it will help you make decisions in the event of some disaster for how you will proceed. So for this Friday’s poll:

    What is more important to you: downtime or data loss?

    My feeling is that most of the people would really rather have the database online, even without all the data so they can continue to work. I realized that most of the time, getting the site back up, having lookup and other types of ancillary data (like products, prices, etc), was the most important thing. Recovering other data such as older orders, was secondary.

    Once the database is up, you can then work on getting other data back and merging it into the production system.

    Let us know what you think this Friday and what’s more important to your business (and why).

    Steve Jones

  • Two Days Off

    I almost couldn’t believe this when I saw the article. The Verizon Cloud is shutting down for 48 hours. Apparently they have maintenance scheduled for this weekend and notified their customers that their virtual machines will be shut down early Saturday morning. There are some legacy Verizon cloud-type services that will be available, but the platform they’ve been pushing to customers will be down.

    This isn’t good news for Verizon or their customers, but it also doesn’t help the cloud overall as a service. This outage reinforces the idea that reliability isn’t necessarily better for vendors than individuals. If costs for the cloud are anywhere near that of on-premise hosting, this event would certainly make me think twice about moving anything really important in my organization to a single cloud vendor.

    I suspect that most cloud vendors have outages like this, but they don’t shut down their entire clouds. When a large amount of maintenance is needed at a data center, most vendors would migrate customers to a separate data center or another part of their cloud while they perform their work. Either this mainenance is a major change to Verizon’s entire infrastructure that can’t be staged on just a part of their system, or they’ve poorly planned their architecture and maintenance.

    Either way, this weekend will certainly be a good DR test for enterprises that might have important applications hosted with Verizon. It will also be a chance to test how these clients notify their own customers of potential issues or how they respond to problems. I wouldn’t want to have an application hosted with Verizon as it would be a lot of work for me and likely a weekend away from family, but I know I’d learn a lot about how well I’ve prepared my own systems for fault tolerance.

    Steve Jones

    The Voice of the DBA Podcast

    Listen to the MP3 Audio ( 2.2MB) podcast or subscribe to the feed at iTunes and LibSyn.