Tag: disaster recovery

  • Deployment Failures

    Just Ship
    Shipping is good, but your deployment process needs to be solid.

    Years ago the company I worked for would patch the majority of our servers one Friday night each month. The Microsoft patches for the month, and other software patches, would be bundled up into SMS (Systems Management Server) packages and deployed to thousands of servers. We had an amazing administrator who built these packages, and it was quite an experience to walk into the data center and hear thousands of servers shut down and fans spin down for a moment before rebooting.

    That was the smoothest deployment system for vendor patches, but I worked in another place that deployed changes to a web application (the system that generated all our revenue and paid our salaries) every Wednesday night. We did this for over 18 months, over 70 deployments, pushing out changes on a consistent basis. We only rolled back three times, but we did roll back three times.

    Other jobs have had various levels of success at deploying changes. Many of the companies worked with the ad hoc, patch one machine at a time manually, process. Not very efficient, and probably not even possible at the numbers of systems many companies have today. I wanted to ask you this week how successful your company is.

    How often do you have problems during the deployment of some software change?

    Do you think that you have issues more often than not? Do you roll back when you have issues? I doubt that. In my experience, even broken deployments are often pushed forward, with the expectation that developers, vendors, or admins will fix things over the next few days. I’ve never thought that was a good plan, since we often fine broken features limping along for months or years, but organizational momentum can be hard to slow down.

    Let us know if you think you work inside of a smooth, strong deployment process, or one that’s more fragile and brittle.

    Steve Jones


    The Voice of the DBA Podcasts

    We publish three versions of the podcast each day for you to enjoy.

  • Disaster After Disaster

    Running out of diesel in a disaster is not something you want to happen.
    Running out of diesel in a disaster is not something you want to happen.

    There was a large hurricane in the US a short while back. It was a devastating storm for many people, and my heart goes out to those that suffered or are still suffering. There are lots of lessons to be learned in many areas, but a few surprising ones for those people that run technology infrastructures. A number of data centers were shut down because of physical flooding, but others were shut down after electrical substations failed and they were unable to run their generators.

    When I evaluated data centers a decade ago, I was always shown the number of UPSes on site, and the high capacity diesel generators with large fuel tanks that were available just in case of extended outages. Salespeople would bring out their contracts that showed suppliers would commit to refilling their fuel tanks, providing for every contingency.

    Except a lack of diesel. In Denver we wouldn’t have the need for a staircase bucket brigade, but we might have the need for a roadside chain gang carrying containers in a blizzard. You cannot plan for every contingency because there are many factors out of your control. When a disaster gets large enough, it doesn’t matter what you have contracted for. There will be outside influences, like the lack of elevators or the inability of trucks to physically reach your location.

    One of the Red Gate customers was in New York and almost lost their data center after the storm. They asked our technical team to help them prepare scripts to restore data backups for their clients in the event they had to send the full, diff, and log backups (along with application code) to the customers. It was an last-ditch effort to allow their customers to continue to run their service. Fortunately they never had to use any of the scripts, but it did help the company realize they need to build more options for business continuity in the event of future disasters.

    I hope none of you ever experiences anything like Hurricane Sandy. You can’t completely prepare, but you can practice your recovery skills on a regular basis and be prepared to respond when disaster strikes.

    Steve Jones


    The Voice of the DBA Podcasts

    We publish three versions of the podcast each day for you to enjoy.

  • Grace Under Pressure

    Grace Under Pressure
    How will you react when things go poorly? Will you maintain your composure?

    I once worked at a large, 10,000+ person company. We had a large data center with hundreds of machines, where we one day we lost power. Not power from the electric utility and had our UPSes and generator kick in. We lost power when some maintenance caused all of our UPSes to trip off line and cut power to all the servers.

    I was in the data center, surprised by the sudden quiet. Unfortunately one of our senior executives was also in the data center and proceeded into the raised floor area. As various technicians and sysadmins attempted to restore power and reboot systems, this senior executive watched, commenting, questioning, and often berating the employees. Not a good situation for anyone, least of all the people trying to reconnect high voltage wires together.

    Most of you will never experience a large disaster and need to recover your systems. Even fewer of you will recover from disasters with anyone other than your peers or a direct manager watching you. However you shouldn’t count on being that lucky. Whether the disaster is small or large, your fault or a natural occurrence, I hope that you are able to successfully restore your systems with some professionalism and grace under pressure.

    The key to a strong performance in a stressful situation is the same in technology as it is in sports, music, and almost any other endeavor with an audience. They key is practice.

    Simulate disasters, pretend that refresh of a development system is really a restore after a fire. Think about the various possible scenarios that might require you to recover a system and incorporate practice time into your daily routine.

    Steve Jones


    The Voice of the DBA Podcasts

    We publish three versions of the podcast each day for you to enjoy.

  • SQL in the City Slides

    I’m making the slides available from my SQL in the City 2012 talks. You can download them here

    If you attended, you’ll get an email with links to videos to watch soon.