Tag: big data

  • More and More Data Growth

    I wrote an editorial a few years ago about a Zettabyte (ZB) of data being created in a year in the world. This was 2012, and those seemed like crazy predictions. At the time I was doing some work with full text search, Filestream and Filetable, and wanted to do a presentation on those topics. I did some research and found some neat facts, which I incorporated into the presentation. At that time, there was an estimate that we’d have 50x that data by 2020. I also wrote I was carrying 3.5GB in my pocket.

    Those values seem crazy now. Crazy small, that is. At a recent event, I was carrying 1.6TB in my pockets. My pockets, not my laptop bag. With 64GB in my phone, an mSata 512GB drive and a 1TB mSata drive. I was carrying a smaller physical amount of electronics than was in my first 10MB hard drive on my person, both in size and mass. What’s funny is that there was only another 1.5TB in my bag, between my laptop and another external 2.5″ drive.

    Between the growth of storage and the Internet of Things (IoT), data growth estimates are rising. At the recent SQL Nexus conference, part of the keynote was given by Dr. Troels Peterson, a physicist with the Neils Bohr Institute, working with the Large HADRON Collider. The work there means big data has a new meaning. Dr. Peterson noted that during his work with the Atlas detector, they can generate 1 PB/s of data.

    Per second.

    That makes the few TB I carry around seem puny. There were two other really interesting items from the keynote. One is that computers cannot process that level of data so hardware sensors must make decisions on what data to capture and store for analysis. The other item is that much of the data will not really be analyzed, and can’t easily be analyzed directly. Instead, algorithms and patterns are used to determine which data is good and should be used. They know there are errors in the data, so the trick is to actually use computers to find the good data.

    The world is a little different for those of us that deal with customers and orders instead or particles and uncertainties, but we are still seeing lots of data growth. Our challenge is to find ways to better manage this growth, while still making data available and useful for our users.

    Steve Jones

    The Voice of the DBA Podcast

    Listen to the MP3 Audio ( 3.4MB) podcast or subscribe to the feed at iTunes and LibSyn.

  • Getting Big Data

    Many of us work with data every day. Often we do so at work, but plenty of us also look at data outside of our jobs. The recent Power BI Report contest shows that our colleagues look at all sorts of different kinds of data. What is more interesting to me is that there is a huge variety of data that people choose to analyze, from sports to goverment to healthcare to airlines.

    I encourage you to get a set of data and work with it. I know that I’ve often just created my own sets when building demos, but there are so many interesting sets out there that I would think many of you could find one that might be of interest to you in building experiments. After all, knowing something about the data is important in deciding how to analyze or visualize the information contained within.

    One of the things that I’ve had people ask about is where can we find data. How can you get a set of data that you might be interested in. What’s amazing is that there are all sorts of sources out there, many of them free. I ran across a list of data sets at Forbes that contains quite a bit that many of you might want to play with. I know I’ve used a few of these in the past and am currently downloading various machine learning sets from UCI with which to experiment.

    Certainly some of these data sets cost money if you want specific information, which makes sense. It takes resources to compile data, and a nominal charge makes sense. However most of these sources will let you download large sets to play with if you choose. You might need some PowerShell, Python, or other scripting skill to get lots of files, but that’s a good excuse to learn another skill.

    Be aware that many of these sets are CSV files, so you’ll need to work on your ETL skills as well to load these into a SQL Server database for fun. Yet another excuse to learn. And if not, you could just take them as is for your very own Power BI dashboard. And to help you along, and provide a resource, I’ve started my own list of sets at SQLServerCentral.

    Steve Jones

    The Voice of the DBA Podcast

    Listen to the MP3 Audio ( 2.9MB) podcast or subscribe to the feed at iTunes and LibSyn.

  • A Good Use for Hadoop

    Netflix is an amazing company. I’ve been a customer for over a decade, starting with the well known red envelopes that moved DVDs through the mail. Since then I’ve continued to subscribe as a streaming company and have enjoyed quite a few of their original programming efforts.

    However I’ve also been fascinated by the way in which Netflix has made use of technology to grow a unique kind of company. They own very little, having used Amazon’s AWS to host their infrastructure and mostly serve content they acquire from contracts with various media companies. They are famous for their Chaos Monkey approach, having machines randomly fail to test the fault tolerance of their entire infrastructure. I read recently about the closing of their last, small data center, so outside of employees’ laptops, they keep all systems in the cloud.

    However I ran across a post on the scale of their data flow, and it’s amazing. Apparently their events are generating 1.6PB a day of data. That’s incredible, and a scale at which very, very few of us will ever work. Personally, I think that’s interesting, but I’d prefer not to be working on a PB of changing data a day, but I might feel differently if I worked at a company like Netflix that has obviously been successful with dealing with data at that scale.

    The post notes that they put this data into Hadoop, where it can then be queried. I’ve often wondered what the domains are for using Hadoop. Most of the people I know that have tried it aren’t really working at large scales. I think that a well built ETL process would allow their data to easily work in a SQL Server data warehouse, and possibly even the Azure SQL Data Warehouse. However at 1.6PM a day, Hadoop seems like a much better choice.

    I’m curious how often that data is queried. Is most of that 1.6PB actually used in reporting, or is much of it lost in the shuffle and ignore? Is there a process to aggregate some of this raw data and then delete the older values? Can 1.6PBx365 actually be useful in analyzing your business? I’m sure some is, but wouldn’t a lot of those events lose their value over time?

    Steve Jones

    The Voice of the DBA Podcast

    Listen to the MP3 Audio ( 3.7MB) podcast or subscribe to the feed at iTunes and LibSyn.

  • An Inconceivable Scale

    One of the talks I give on SQL Server deals with unstructured data. I start out this talk looking at the scales of data we deal with and was been amazed by the research I did about how much data we humans have created. What’s even more interesting is that the growth is outracing predictions made just a few years ago.

    When I started working with computers, we talked about kb of data, thousands of characters. That’s an amount of data we humans can easily comprehend. In fact, we used to talk about floppy disks and the number of average sized books that could be stored in kb, or single digits of MBs. As humans, we can comprehend that scale. Most of us have seen hundreds or thousands of books in a library.

    When we move to GB, things get harder, though at 4GB for a DVD, many of us can conceive what multiple GBs can mean. However terabytes? Can we conceive the scale of data? Sure. A TB is about 40 Blu-ray disks. While we might not appreciate how much data that is, we can picture it.

    A PB? That’s 41,000 Blu-ray disks. I can’t even conceive of what that looks like, much less imagine the billion MB sized pictures.  That’s a scale that has no reference. However as humans, we will create multiple exabytes of data this year. As individuals working with data, few of us will work with EB in our organizations, but some of us will. I read recently that Paypal processes 1.1PB of data regularly. Regularly processes, not just has in cold storage.

    We have zettabytes and yottabytes, but who could possibly conceive of what those mean? There’s not frame of reference I can imagine, though that may change. I expect that we will become used to PB at some point, just like a TB is no big deal right now. In fact, I really think I’ll see a TB on my phone sometime before the end of this decade.

    Steve Jones

    The Voice of the DBA Podcast

    Listen to the MP3 Audio ( 2.8MB) podcast or subscribe to the feed at iTunes and LibSyn. feed