Tag: career

  • The Remote DBA

    I’ll start this week with a question, which I hope some of you answer in the discussion: would you like to, or do you, work at home the majority of the time?

    I remember when I worked in a company and needed to leave my desk and walk to a room somewhere to get some work done on a server console. Some of those rooms were cold rooms, which necessitated me keeping a jacket at my desk. I still remember going to a large company that had cables run from the data center to a couple specific workstations near the administrators’ cubes. At one point we installed a remote IP device allowing us to get to the console of any server without having to walk downstairs, or even use RDP, which was just becoming to Windows machines.

    That was the end of my visiting servers in person, and since then, the only times I’ve ever really needed to look at a server was when a critical error prevented me from connecting remotely. Even then, at many of the co-location facilities I’ve contracted with, I could call and have an individual go press a power button. These days, with cloud providers and virtual machines, even that is unnecessary.

    Those of us that have worked with SQL Server typically understand that we always make a network connection to work with the server. Even when we’re working on the server console, SSMS, SQLCMD, and more all make a “connection” to the database server. Therefore, is it really necessary that we ever work near a particular system?

    I’ve been working from home as a telecommuter for about 8 years. My wife worked in technology from home for nearly 20 years. More and more people are doing so, in fact, there was a piece on Fast Company recently that noted more people work from home than ever before. Unfortunately, this study shows that most people end up working more hours, adding a few from home to the 40 or more they spend at work.

    That’s not good, but if we can get some work done at home, why not more? I know meetings and face to face time matter, especially in some jobs, but more and more I find that lots of people that need time working alone could do a portion, perhaps a significant portion, of their work away from the office.

    So this week, would you like to do more work from home (or elsewhere)? Do you want more virtual meetings, more communication over email, Slack, Skype, or some other tool? Let me know.

    Steve Jones

    The Voice of the DBA Podcast

    Listen to the MP3 Audio ( 4.0MB) podcast or subscribe to the feed at iTunes and Mevio . feed

  • A New Look

    I was browsing around the Internet recently, collecting links for Database Weekly when I came to Michael J Swart’s blog. I love Michael’s writing and drawing. On a whim, I asked if he’d do an avatar.

    He agreed, and sent me some costs. He’s creating art, and I was happy to pay for a new image. We went back and forth a bit and ended up with this:

    SteveJones-Big

    I like it. It captures me life in sunny, outdoor Colorado and has a nice Hawaiian shirt in it. Plus, it has my colors. I’m a Virginia grad and Broncos fan, and like the blue/orange look.

    Photo Feb 26, 11 32 19 AM (1)

    If you want an avatar, I’m sure you can contact Michael. I know it’s not for everyone, but it was fun for me. It might make a great gift for someone in your life as well.

  • Triple Check Your Restores

    I used to work at a nuclear power plant. Not really in the plant, but as a network administrator in the building next door to the plant. Probably a good thing since I struggled to get to work on time and everyone going into the plant had to go through a metal detector like that at most airports. My tenure might have been shorter if I had been late every day to my desk.

    However, there was one thing that got drilled into me by each person I knew who did work closely with the power generation side of the business. Everything was (at least) triple redundant. Not only did they want a backup for a component, or a system, but they wanted a backup for the backup. There was a limit to the paranoia, but in critical places, where radiation was concerned, we had multiple backups. One notable, and impressive, area was the power for the control rooms and pumps. In addition to batteries for short term power loss, there was a large diesel generator for each of our two reactors, plus a third that could take the load if either of the first two failed. Those were impressive engines, each about the size of a very large moving truck and jump started with a few dozen large canisters of compressed air that could spin the crankshaft in a split second.

    This week there was a report that the database for the US Air Force Automated Case Tracking System had crashed. Apparently the database became corrupted, which happens. However, the surprising part of this story is that the company managing this system reported they didn’t have backups and had lost some data going back to 2004. They are looking to see if there are copies in other places, which I assume might mean exports, old backups or something else, but the reports make this seem like a completely unacceptable situation. I assume this is an RGE event for a few people, perhaps all of the staff working the system.

    I was reminded of my time at the nuclear plant because we had a similar situation. We didn’t lose any data, but we found a backup system hadn’t been working for months. Those days we had a tape drive system that automatically rotated 7 tapes. I think this would last us about 4 or 5 days, so it was a once a week job for one administrator to pull the used tapes and replace them with new ones. We had a system where tapes were used 4 or 5 times before being discarded, and our rotation had a tape being used every 3-4 months. However, the person managing the system rarely restored anything.

    One day we decided to run a test. I think this was just my boss giving us some busy work to keep us occupied but in a useful way. When we went to read a tape, it was blank. Assuming this was just a mix-up, we grabbed one of the tapes from the previous day and tried it.

    Blank.

    At this point, my coworker turned a bit red and started to stress. He was in his 40s, with a family and mortgage. I was in my early 20s and had no responsibility here, but I could appreciate his concern. We frantically loaded tape after tape, even looking at the oldest tapes we’d just received from our off-site provider. None were readable, and most were blank. We nervously reported this to our boss, who had us request a sample of tapes from off-site storage going back over 6 months.

    Eventually we realized that we hadn’t had any backups for about 4-5 months. The tape drive had stopped working properly, hadn’t reported errors, but dutifully kept retrieving files and rotating tapes each week, unable to properly write any data. No databases, no email, no system was being backed up.

    A rush order to our computer supplier had been placed the first day to get us two working tape drives that we manually loaded tapes in each day, and checked them the next morning. Eventually we replaced the drive in our tape leader and instituted random weekly restores to be sure we had working backups. I’m not sure if the plant manager or upper IT management was ever told, but I’m glad we never had to deal with a hard drive crash during that period.

    Backups are something we all need to perform. I note this as the #1 thing a new DBA or sysadmin should perform on systems. However, backups are only good if you can read them and actually restore data. I’ve made it a point to regularly practice restores as a DBA, randomly restoring backups with diffs, logs, or to a point in time. Not only do I test the backup, but I test my skills. I’ve also tried to keep an automated process around that restores all production systems to another server to test both the restore as well as run a DBCC CHECKDB. Corruption can live in databases for a long time. It flows through backups, at least in SQL Server, and this is something to keep in mind.

    I’d suggest that you make sure you ensure that your backup plan is actually working by performing a few restores. Build an automated process, but also run some manual restores periodically. You want to be sure that you can really recover data in the event of an emergency.

    Steve Jones

  • The RGE

    I first heard this little acronym from Grant Fritchey (b | t). He used it when talking about backups and restores, and I like it. However I realize that I’ve never actually noted what it is, so a short blog to do so today.

    An RGE is a Resume Generating Event. This is usually when you make a mistake so egregious that you’ll be packing up your personal effects and exiting the building. If it’s really bad, such as releasing financial or other confidential information, you might be escorted out and someone else packs up your things. I’ve seen it happen, and it will shake you. Don’t do this.

    We talk about forgetting about backups, or writing bad code or some important task we often perform as causing an RGE. In my experience, that doesn’t happen too often. Companies usually have a fairly high tolerance for mistakes.

    However, that tolerance is usually extended only once. Don’t make the same mistake again. I’d also note that some managers can be very short tempered, and a single, large issue might be an RGE in their eyes.

    I don’t usually worry about causing an RGE, but I keep the acronym in mind. Especially when I do something that could affect the core parts of my organization’s business.