Tag: databases

  • A Billion Transactions

    These are some of the sensors generating that half billion transactions/day.
    These are some of the sensors generating that half billion transactions/day.

    How long would it take your systems at work to process a billion transactions? You’d expect some, heavily used and highly visible systems to be involved. The stock market systems process billions of trades a day, but I’m sure most of the systems in single companies, even large companies, deal with fewer transactions on a daily basis. A billion transactions a day is 11,000+ transactions a second, sustained across the entire day. That’s a heavy load, but it might be the level of transactions that more and more of us will see over time as our systems gather more data.

    The Microsoft corporate headquarters in Redmond consists of over 100 buildings on 500 acres. It’s grown over the years from its original 88 acres, which is also the title of a story about Microsoft and the relatively unknown work in automating their infrastructure. Not the computer systems their software developers use, but rather their facilities and physical buildings. It’s a fascinating story that outlines the way in which Microsoft saves millions of dollars in maintenance and repairs by using software.

    Across those buildings, Microsoft collects a huge amount of data, from disparate systems, which is the presented to the facilities personnel. The sensors and systems don’t process a billion transactions day; they process half a billion. Still an amazing amount of data, just from physical buildings and the infrastructure that ensures Microsoft employees have a pleasant place to work every day. Using a combination of SQL Server, Office, and Azure, Microsoft has built a software system that corrects many faults itself within sixty seconds. Those that can’t be fixed remotely often end up generating one of the 30,000 work orders produced for personnel every quarter. The system is forecasted to save 6-10% of the energy that might otherwise be wasted with a less efficient system.

    It’s a great read, and perhaps is a good case study for an application that is well suited for cloud services. There are a few great quotes from the article as well that are particularly pleasing to a data professional. “Give me a little data and I’ll tell you a little,” he (Darrell Smith) says. “Give me a lot of data and I’ll save the world.” That ought to be the model for data analysts. As SQL Server professionals and developers, we should be helping others to do just that.

    Steve Jones


    The Voice of the DBA Podcasts

    We publish three versions of the podcast each day for you to enjoy.

  • Modeling the Earth

    The idea of modeling the Earth is an incredible challenge.
    The idea of modeling the Earth is an incredible challenge.

    Whether you agree with the science of climate change or not, the ability to work on a project like this one from Microsoft Research would be cool. The issue of carbon pollution and the potential impact on our world is huge. If we accept that global climate predictions of problems are true, we may severely impact world economies with the changes that some people have suggested. If we discount the problems and they turn out to be true, we may end up in an even worse position. It’s also entirely possible that climate changes are natural cycles of the planet and we have no need, or possibility, to alter the way the world is evolving.

    No matter what your position, it seems the Microsoft Research isn’t trying to make a stand for either position, but rather attempting to clear the technical hurdles that would allow other groups to compare and debate about their models with regards to any global issue. Their goal isn’t to push users in a direction, but give governments or other organizations tools they can use to either consider future actions, or react to new data or information that can be added to a model. There’s a short interview from CNN with one of the researchers.

    As I browse through the list of projects around the world at the various Microsoft Research facilities, it’s an interesting mix of somewhat practical ideas with pure research into areas that may never become projects. One thing is clear, and that’s much of this work involves large amount of data. Quite a few of the projects themselves deal with data issues, from searching to visualization to analytics. I have to think a few of them will influence or impact SQL Server over the next decade in some way.

    Steve Jones


    The Voice of the DBA Podcasts

    We publish three versions of the podcast each day for you to enjoy.

  • Car Data

    I don't know if I like this layout, but I like more data being available from my car usage.
    I don’t know if I like this layout, but I like more data being available from my car usage.

    I really like cars. In my lifetime I have owned more than my share of vehicles, and I always look forward to renting new makes and models when I travel, just to drive something different. As cars have evolved over the last few decades, there are some things about the changes I love, and some things I dislike. Personally I like the idea of a key fob that enables me to unlock the car and start the engine with a button without pulling the keys out of my pocket. However, as someone that’s lost my share of keys, I’d prefer a real key as a backup mechanism. Unfortunately that doesn’t seem to be an option many manufacturers want to provide. I dread having to replace a $200 “smart key” at some point in order to drive my car.

    Cars have implemented a wide array of technology over time, some of which drivers are not even aware. These days cars gather an impressive amount of data, though probably not quite at the scale of the new Dreamliner. Or maybe they do. According to this article, cars can produce “hundreds of MB/s” in data acquisition. I would guess most of that data is thrown away, but some may not. I was quite impressed with the amount of data Tesla logged during the recent test drive controversy with one of their vehicles.

    All this data, and the potential need to manage it, mine it, and perhaps make it available to other applications, is another sign of just how important our jobs as data professionals may be in the future. More and more of the things we encounter on a daily basis are creating data that we may turn into information for business decisions through creative uses of software. While some parts of our database systems may become easier to use, I suspect there are no shortage of new skills we will need to learn in the future to make sense of our data.

    Mobile technology, whether with cell phones or transportation (planes, trains, and automobiles), will become more prevalent in our lives in the next decade. I suspect this will mean many more opportunities for data professionals. Especially as I’m sure there will be new regulations and requirements that will keep software developers and database professionals busy modifying applications as legislation tries to catch up with the creativity of technologists.

    Steve Jones


    The Voice of the DBA Podcasts

    We publish three versions of the podcast each day for you to enjoy.

  • Zettabytes and Beyond

    Billions and billions of stars don't seem so large anymore.
    Billions and billions of stars don’t seem so large anymore.

    The old Carl Sagan quote about billions and billions of stars in the universe doesn’t seem so large anymore. In fact, a billion of anything, while a large number, seems immense only until we talk about the scale of data. How much data is there in the world? I’m not sure. As I was researching this for a new presentation, I’m not sure I can even conceive of the scale of data creation occurring in the world today, much less how much data we have.

    Think back 30 years ago, as computers were just starting to become household items. The high density floppy disk (not really very floppy) was a 1.44MB disk. At the time, this held what felt like lots of data in terms of text pages. However many songs we listen to today wouldn’t fit on this media. As we’ve progressed through CDs and DVDs to flash drives, we’ve grown the storage capacity of our hardware by unbelievable amounts. My phone has 64GB of storage, which is a level of growth so far removed from the Apollo guidance computer’s 2kb that comparisons don’t do it justice.

    We used to create data storage analogies by listing the number of books that would fit on the device. Today that’s meaningless. The 30,000 ebooks on Bookworm fit in 20GB. That’s a number of books that’s hard to conceive of. My local library branch has about 20,000 books, so I can somewhat grasp that scale, but not really. Two libraries worth of books is a level of words and knowledge that I’m not sure I appreciate. Trying to understand the amount of storage a library like the Vatican needs, is beyond comprehension.

    The grasp of how much digital data we create is even more mind boggling. I saw a talk recently that said 24 hours worth of digital video is being uploaded to YouTube every second. Every second. That’s an impressive statistic, but I’m not sure we can even comprehend what that means. An even more daunting statistic is that all the knowledge recorded from the Gutenberg printing press invention through the next 500 years totaled about 1 exabyte. At current rates, we create an exabyte of digital data in less than a month and that’s only going to increase. If you read some of the analogies in this report from EMC, they’re almost silly, and certainly not something most of us can relate to. I certainly can’t picture 75 billion iPads.

    These days the amount of data we are dealing with is growing faster than ever, and that means it’s a good time to be in the data business. From “Big Data” to data warehousing to the common OLTP databases we manage, there is no shortage of bits and bytes we will get the chance to manage. For pay.

    Steve Jones


    The Voice of the DBA Podcasts

    We publish three versions of the podcast each day for you to enjoy.