Tag: disaster recovery

  • DR Prep Can Miss the Little Things

    The other day we had a blizzard in Colorado and we weren’t quite prepared. My wife and I were away on a trip and watching the weather. We were concerned about flights and about the conditions at the ranch with someone else in charge of horses. We were lucky in that the weather didn’t come as quickly as predicted and we made it home before the snow started.

    Depending on moisture and temperature, we sometimes put blankets on horses, which is a chore that I’m not very good at completing. Fortunately the kids were around and at 8pm, we all dressed in warm clothes and headed out to the barn. The horses depend on us to provide for some of their needs, so we each had jobs to do. One kid secured a temporary holding area to keep horses nearby. One gathered 14 or 15 blankets from storage and laid them out. My wife made grain to facilitate attracting the horses and giving us a way to hold them while blankets went on.

    Me? I had the envious task of double checking electricity and tank water headers. Our employee had thought that one was shorting out and at night, with a flashlight and plug tester, i got to debug 4 water heaters, including one that’s semi-remote from the barn. Testing for shorts involves touching water to see if there is any flowing power (this is very low amperage), which isn’t terribly dangerous, but is a little daunting. We managed to get everything done, and felt like we were prepped for a 4-10″ snowfall, 50mph winds, and a 0-5F temperatures.

    Later I lay in bed, trying to relax with a little TV before calling it a night. While the wind was howling, I heard a pop and the power went out. We have a generator, and I wasn’t too worried, but I did want to ensure it came on. I walked downstairs and across the house to listen for the engine noise. I heard cranking, but the engine didn’t catch. It was then I remembered the tank was low the last time I checked, and with a weekly diagnostic auto run for 5 minutes, I was likely out of propane. Since we needed power to heat water for horses, I had to get up around 11pm and take care of things.

    Fortunately I keep a spare bottle of propane and I went outside in the blowing snow to change it. A true work-at-home-techie, I did this without pants and got things running. Back to bed, though worried a bit about how to get more propane in the morning, just in case of an extended outage. The power came back on in the morning, but I still needed to get more propane just in case.

    Extended disasters sometimes cause problems with our plans because many of us focus on the immediate reactions. Longer term, things like additional fuel, food for humans, even shelter and child care are issues that I’ve had to deal with in teams that no one had planned for. We’ve seen this over and over again in the world, especially when supply chains break down quickly in disaster situations.

    I could have prevented some of the issues and stress if I’d prepped things better. I should have filled propane tanks earlier, knowing that bad weather can come at any time, and likely will at some time. Double checking heaters and power earlier might have made the blanketing process go quicker. There are always things we miss in prep, often ones we don’t consider to be important, but may be difficult to deal with in the moment. Take a few minutes and think about the little things you might dismiss as not important. Imagine how hard some of them might be to deal with under the pressure and stress of a DR event. If you can do a little more prep, some maintenance, or get some work done early, now might be the time to do it.

    Steve Jones

    The Voice of the DBA Podcast

    Listen to the MP3 Audio ( 5.8MB) podcast or subscribe to the feed at iTunes and Libsyn.

  • DR Planning

    Data professionals know that protecting and ensuring our data is safe is one of our primary jobs. Many of us worry that if we were unable to recover a database, our job might be jeopardy, which makes sense. If we lose data, an organization might make the decision that we aren’t trustworthy enough to manage a database system. This leads many of us to plan and prepare for different types of disasters, with procedures and systems in place to ensure that we can quickly get things running after an issue.

    Natural disasters, or other large scale disruptions of business sometimes exceed the plans we’ve put into place. Fortunately they’re rare, but I ran across an article that talks about a few items that we might not consider when making plans. Even if you can fail your database over with an Availability Group running in another location, there are still some points you might want to think about.

    Perhaps the biggest one for me is that testing is not optional. I’ve seen some amazing plans put into place, but never tested because no one wanted to disrupt ongoing business activities. Testing is really critical, perhaps because of the second more important item I see cause issues: failback. This isn’t always as simple as we’d like, even with some of the amazing work that Microsoft has done with SQL Server. We want to ensure that we do know that if something happens to our primary systems and we move them, we can come back. After all, these are the primary systems for a reason.

    While other items such as making DR planning something your organization cares about can matter, I think the viability of this depends on how likely these disasters are. If you’re in a seasonal hurricane path, or some other disaster, you likely need to have better plans than if you’re in a stable region. Not that you can avoid any plans, but most of the time a disaster is a rare event and it’s really IT’s job to ensure that systems are resilient.

    I do think it is worth noting that regulatory compliance isn’t optional. While your auditor might have sympathy if they are likewise affected by the disaster, that sympathy dissipates over time. A disaster next month might not factor into an audit review in 20 months. More laws are on the books and more are coming than most of us have dealt with in the past. Design your system and process to be compliant, even in a disaster.

    Lastly, the people side of business can’t be emphasized enough. While some of us might be expected to shoulder a greater workload during a disaster, keep in mind that we would still need support outside of work, perhaps even a replacement worker to handle some of our work tasks or even help with personal items. The longer our lives are disrupted, the less likely we are to work as hard, or remain focused. Part of any plan should be how can our organization help support those that are struggling themselves. After all, the effects of a localized disaster on business could get compounded if staff stops coming to work.

    Steve Jones

    The Voice of the DBA Podcast

    Listen to the MP3 Audio ( 4.5MB) podcast or subscribe to the feed at iTunes and Libsyn.

  • Your Recovery Time

    Disaster Recovery is both an exciting challenge for many DBAs and also a dreaded event that many hope they never experience. When a true disaster befalls your system, there is a tremendous amount of stress as we work to get things running again so clients can access data. Even the best laid plans will often have a glitch, and while it’s a great technical challenge, most of us would rather keep practicing and gaming possibilities than actually experience a disaster.

    The majority of our system disasters are localized to the actual machine handling our workloads. We often don’t have outside disasters, like hurricanes, with disruption and damage to other parts of our infrastructure and ever our lives. When that happens on top of a down system, it’s a level of stress that can affect our health. I’ve been lucky in that I’ve experienced both kinds of disasters, but never at the same time. I hope you can say the same thing.

    Let’s assume some local disaster defalls your system. Hardware, software, it doesn’t matter, but you end up with a corrupted database of some sort. This week I’m wondering if you have an idea of what the recovery time would be? How long before clients are up and running? Maybe more important, how long before you have system rebuilt?

    If you have some sort of High Availability (HA) plan, then you might be back up for clients quickly, in minutes or even seconds. The disaster really isn’t over, however, since you are now running on less hardware than you planned. Until you can rebuild the downed node, you’re still in disaster mode. If you’re like most of the organizations that have employed me, you’ll also be stressed as your formerly well-designed HA setup is now a single point of failure and you’re scrambling to get the main node rebuilt.

    We often think of a disaster as the time we’re down because of some event. Once clients are being served again, we tend to relax and think the disaster is over. It’s not, because until you replace the affected systems and bring them back online, you’re even more vulnerable than you were previously to another Murphy’s Law incident.

    Steve Jones

    The Voice of the DBA Podcast

    Listen to the MP3 Audio ( 3.0MB) podcast or subscribe to the feed at iTunes and Libsyn.

  • Cloud Backup

    I think that backup and restore are the most critical things for any database professional to master. Whether you’re a professional DBA, a developer setting up a system, or a seasoned DBA, I would argue having your data safe is the primary task. Security comes next, then performance, availability, and a number of other tasks that could go in any order, but if we don’t have a way to restore our data when hardware fails, we’re in trouble. And given enough time, or enough difference pieces and parts, something will fail. If you can’t restore a system when a problem occurs, that’s what Grant would call an RGE.

    Over the years companies have moved to many different technologies to handle backups. Tape was common early in my career, but all disk systems, with de-duplication capabilities have become popular. I really don’t think about anything other than getting a second (or third) copy of data these days, so I can’t speak to any particular way of managing backups. However, for SQL Server, I do want the option to set full, differential, and log backups based on my RPO and RTO requirements.

    There seems to be a new trend for companies that I ran across: they’re moving to the cloud. Here’s a short slideshow of some stats that show cloud use is increasing. This is a survey, so it’s not all companies, but the trends are clear. More data is being backed up, and it’s likely easier (and cheaper) to use cold storage in the cloud. That makes sense since it’s data that you expect you’ll very rarely need to use in a restore.

    There are a couple of other interesting items I saw in the survey. The number of companies that are backing up more than 100TB grew quite a bit, even as the number of companies backing up < 25TB fell. That’s a sign that we’re capturing more data. Whether we need to, or whether legislation like the GDPR will get companies to trim some of that data, remains to be seen.

    Another interesting item is that more companies are testing their DR plans. Fewer never test them, but the frequency is increasing with more companies testing quarterly or monthly. That’s smart as we become more dependent on computer systems. I know some organizations can’t roll back to paper, as we’ve seen in a number of airline IT issues. If you never test your DR plan, I hope that you don’t have an issue when you can ill afford to find another job, becuase you might suffer those consequences. Really, I hope that you actually know this isn’t professional and you start working on ways to test a restore of service.

    More companies are moving to the cloud, and the resistance for security, cost, privacy, etc. reasons is going down. I think many of the concerns that both executives and IT professionals have had in the past are proving to be non-issues. This is especially true as more vendors institute government rated or more secure data centers as a part of their product offering.

    If you don’t like the cloud, that’s fine. If you don’t know anything about it except rumor, guesses, or hearsay, I’d suggest you learn more. The cloud is likely coming into your career, so learn a bit about it. At least enough to give reasons why you don’t want to move your data there. If you do, you might be surprised that the cloud is not that bad a place to be, at least for some workloads.

    Steve Jones

    The Voice of the DBA Podcast

    Listen to the MP3 Audio ( 5.3MB) podcast or subscribe to the feed at iTunes and Libsyn.