Tag: disaster recovery

  • Deleting a Database

    Who among us has deleted a production database?

    I’d hope it’s very few of you that have done this in your career. I’m sure a few of you have deleted (or truncated or updated all rows for) a table in production. I’ve done that a few times, but fortunately, I’ve been able to recover the data quickly. I had this happen in SQL 6.5 and was grateful I could start a single-table restore before my phone rang.

    Here’s another question: which of you has had a storage admin delete or remove some remote storage and cause you database problems? Has anyone had that happen in their environment? I haven’t had this in production, but I have had this happen to test systems, and I was very irate with the storage people when it did. After that, I’m sure they were very cautious about changing any configuration for database servers. I’m also sure that also contributed to my struggles in getting more space promptly as well, so I’m not sure I came out ahead in that situation.

    A hospital system had issues after this happened to them: engineers deleted critical storage that connected to a database system. Fortunately, no services have stopped, and no patients are missing services, at least as far as we know. Everything has to be moving slower, and that might mean that staff is spending more time on “downtime procedures”, i.e. paper, than focusing on patients. Knowing a few medical professionals, this means they’re more stressed and working harder to be sure patients aren’t affected. Sucks for them, and I doubt the hospital compensates them for an engineer’s mistake.

    I haven’t heard about this happening in a long time, and I’m surprised by that. Almost all storage these days for server systems is remote, especially in the cloud, and it would be easy to click the wrong button or select the wrong disk when re-configuring a system and remove critical storage. Maybe we’ve gotten better at popping up warnings that slow people down and prevent mistakes.

    Or maybe we don’t delete disks and only add them to database systems 😉

    In any case, I hope they can recover things quickly and easily. If you’ve seen this, let me know. If you haven’t, here’s a reminder that this could happen. You should be sure your backups are running AND you can perform a test restore.

    Steve Jones

    Listen to the podcast at Libsyn, Spotify, or iTunes.

    Note, podcasts are only available for a limited time online.

  • Responding to a Disaster

    Ryan had a planes, trains, and automobiles situation a few weeks ago. On the same day, I didn’t. A delay for a few hours, but an easy trip home. Lots of other people didn’t have a smooth trip, and I had a few friends who spent an extra night somewhere or had flights canceled and decided not to take a trip. The Crowdstrike outage hit the entire world, causing problems everywhere. Just before that happened, Azure had an outage in the central region.

    If you were affected by these, you have my sympathies. If your IT job included responding to these, you get all the virtual hugs from me (and Brent). I know what it’s like when someone calls you to handle an outage, and I know it can turn your life upside down. Hopefully, it hasn’t been too stressful to you.

    In Brent’s post, he talks about reviewing your DR plans and presenting your findings to your boss. Good advice, and worth reviewing yearly as most of our environments will experience some amount of change. If you follow this advice, make sure you present this in a logical, dispassionate way. I see plenty of technology professionals who get upset when DR isn’t given enough priority. That doesn’t help the situation. Channel your inner Spock and present your results, allowing someone else to decide what to do and accepting their priorities. You can advocate for change, but keep in mind that there is often no shortage of work and limited time. If someone else decides you can accept the risk of poor plans, document that and move on.

    Plenty of us will experience a disaster at some point. I’ve had more than a few occur at various positions, and I know many colleagues who have gone through a wide variety of failures. Even for companies that have extensive DR plans, there will be challenges when a disaster occurs. We’re certainly not at the point where any AIs can handle reading our DR document and implementing fixes. Instead, those of us responding will have a guide, maybe a great guide, for how to respond, but we will need to adapt our plans to the current situation. Having the mindset that our plans might not be perfect (as Mike says), will help you deal with any challenges that arise.

    Aside from the technical situation, there are also mental challenges in a disaster. You will feel pressure from your employer, but also pressure at home. Your partner might not understand the extra hours. If you miss commitments made to your family, they will be disappointed, and I hope you are as well. You might have expected your schedule to include fun events, and you likely didn’t expect to get less sleep. Your diet might suffer in times of crisis and long hours.

    All of these situations can be managed, but it helps to think about them when you’re calm and make plans. Warn friends and family. Think about how to limit unhealthy foods or situations, set expectations with yourself that help you manage stress, and most of all, treat yourself with kindness. Many of us want to help and support our co-workers but ensure that you (and them) get breaks. Imagine you’re in a Crowdstrike situation that lasts for multiple days. Talk about how that will look, even with just a few people and you’ll be better prepared when a disaster event occurs.

    Hopefully, it won’t be a large-scale event, but better to prepare for that and experience a minor situation than the reverse.

    Steve Jones

    Listen to the podcast at Libsyn, Spotify, or iTunes.

    Note, podcasts are only available for a limited time online.

  • A Staffing Disaster

    There was a failure recently at an Azure data center in Australia when a utility power sag caused equipment to trip offline at one of the Azure data centers in Australia. You can read about it here, but essentially the headline is that there were only three people on site when the incident occurred, and that caused them to be unable to restart the equipment in time before an outage occurred.

    In a little more detail, there weren’t enough people to quickly restart the equipment chillers after the incident. The staff had to access the equipment on a roof when 13 of the units didn’t restart. They were able to get to 8, but when they got to the last 5, the temperature of the water had risen to a level that wouldn’t allow a restart. So they had to power down some computer equipment and go through a more lengthy process to get everything running.

    This sounds bad, but in reality, this is exactly the type of thing I’ve seen in private data centers, who almost never have all the staff they need, or the knowledge necessary, to deal with large-scale failures. While I haven’t seen the chillers, I have seen people trip electrical systems and be unable to restart or reset UPS’s or generators for hours until qualified staff could come in. If you read the incident history, there is a good retroactive of what happened, and then some actions taken to try and prevent this in the future. They increased staff levels but also identified some places where the previous staffing level would have been fine with some equipment and protocol updates.

    I wish more organizations would review incidents and examine them with an eye towards not only what happened and where there were failures, but how to prevent issues in the future. Too often I see people going through this exercise in order to blame someone and “prevent this from ever happening again”, which usually means we fire someone and don’t change anything else. We need psychological safety in reviews of actions to get better.

    As we build more complex systems, or even more complex organizations with lots of teams, people, equipment, procedures, etc., it’s easy to build in lots of points of failure without realizing there will be problems in the future. My goal often these days is to assume I’ll have some inexperienced or less capable staff and design processes and systems to survive issues. To keep things simple, and not get too cute with engineering. I like robust, resilient systems that anyone can operate, not those that require me to ensure my senior superstars are always on call.

    Of course, it often takes the senior superstars to design and test these systems and protocols, which is a good use of their time.

    Many businesses struggle with staffing, in many areas. Technology groups are no different, and we have to learn to work smarter, not assume we will just get more staff and solve our problems.

    Steve Jones

    Listen to the podcast at Libsyn, Spotify, or iTunes.

  • How Many Days Can You Survive?

    I saw a great post from DCAC on disaster plans and using them after a fire in an LA data center. DCAC wasn’t affected, and I wouldn’t expect them to be. Denny, Joey, Monica, John, Kerry, and Meagan are all experts in running systems efficiently and effectively. They think about how to ensure that things keep working in the event of disasters and have delivered presentations all over the world helping you become better at managing your own servers.

    However, Denny brings up a great question in the post. How many days could you survive without IT systems? Days? Weeks? Have you ever been in that situation? I’m sure many of us have experienced a failure of some sort. An application crash, hardware dying, or most commonly, Internet access being cut. All of these are small disasters, which typically are fixed quickly. Hopefully, these aren’t issues you experience every week. If you do, maybe you need to call DCAC or someone else and fix more fundamental issues.

    While it’s unlikely that you lose all your systems, it could happen if you concentrate your resources in one data center, one region, one cloud provider, etc. Having a single point of failure is something we try to avoid in IT, and that is true not just for one application, but for your infrastructure design. Most of us depend on one authentication system, and a failure of our Active Directory could lock everyone out of a system. Those are rare, and hopefully, your administrators have enough redundancy (and backups) to recover from this type of disaster.

    I have experienced a few large failures at one enterprise. We had a few viruses, including the SQL Slammer worm in the early 2000s. Our network was shut down for a couple of days, when almost all systems from email to CRM to ticketing systems weren’t available. Everyone had to use whatever paper systems they could to keep business running. While we likely lost some revenue from these outages, we learned we could survive a few days without our network. We also learned that we needed better virus scanning and education for employees, as well as a few more resources for tech people. Before those events, everyone assumed an outage was bad, but had no idea how bad. I have no idea how much this cost us, but it didn’t appear in a 10-Q, so it must not have been too bad.

    I think there are lots of businesses that could find ways to continue to work if some systems were down. However, there can be costs, sometimes significant. Even if the company doesn’t go out of business, perhaps some people get terminated because of less revenue. That might not be the tech person in the short term, but how would you feel if your DR plan didn’t work (or you didn’t have one) and some co-workers were let go?

    The move to the cloud, and the move to more software-as-a-service systems, might help you better survive local disasters, but if you have too many systems concentrated in one place, it is worth preparing for some contingency. After all, even if this fire were to happen in an Azure or AWS data center, it’s possible that their process to move and restore all the systems from one data center to another could take time. Your systems might not even be their top priority as cloud vendors have some large customers. It probably won’t take months, but I wouldn’t want to bet my job on any cloud vendor getting everything moved in less than a week.

    If you’re not in the cloud, make sure you have a plan. If you don’t know how to do that, call DCAC or another consultant to help you.

    Steve Jones

    Listen to the podcast at Libsyn, Stitcher, Spotify, or iTunes.