Tag: data

  • The Exciting World of Data

    I was honored to have the chance to give the keynote at SQL Saturday #520 in Cambridge this year. This was a quick keynote, and it was fast. I didn’t record it, but people seemed to enjoy it, and I decided to share some of my thoughts on the Exciting World of Data, the title of the talk.

    We love data. At least, I do. It’s a way of learning more about the world around us, describing it, modeling it, understanding it, even enhancing it. And the world of data is changing. Size is increasing. We’ve moved from bits to bytes, to kilobytes to megabytes to gigabytes to terabytes, just in our hands. We can’t even really conceive of what this amount of storage means in a physical sense. Our large systems have grown to petabytes, perhaps exabytes and zettabytes one day and eventually to yottabytes and beyond.

    We have to transfer data quicker as well. My first modem was 300 baud Hayes Smartmodem. This was at university, where I could watch the text crawl across the screen. From here I had a few upgrades and standarized on 28.8k for quite some time before moving briefly to 56k speeds. When I started SQLSeverCentral, I had an ISDN line in my house. I’ve configured T1 lines at work, and upgraded to faster DSL and Cable routers at both home and work. I never worked in the OC space, but some of you may have transferred data across OC-256 or even OC-768 lines.

    Our mobile data moved from SMS to GPRS to Edge, where with a Sidekick, where I could actually type real messages on a keyboard. When I got to 3G, I thought was all the speed I’d need for a long time. I think I spent 2 or 3 years with an iPhone 3GS. However, like many of you, I upgraded to 4G and LTE, which are amazing speeds, faster than many of the early networks I had at work. We’re testing 5G and 6G and maybe we’ll keep going to subspace radio? Who knows.

    Here on earth, we move more data in our systems. Some of you may have worked with tape storage. My first PC had a tape drive. So I was quite pleased to get a floppy disk drive, first 5.25″ and then 3.5″. I thought we’d have those forever, but I’ve migrated to hard disks to solid state disks to 3D drives. I think 3D SSD technology is going to fundamentally change the world, with latencies that will require our software to be very, very efficient.

    Our interfaces have improved, to allow us to move more data, quicker. From SMD to ESDI to ATA to IDE to SATA to SCSI to Wide SCSI to Fast SCISI to Fast Wide SCSI to Ultra SCSI to Ultra Wide SCSI. SCSI 2 to SCSI 3 to Fibrechannel, infiniband and beyond. USB to Firewire 400 to USB 2 to Firewire 800 to USB 3, 3.1, eSata, Thunderbolt, Thunderbolt 2, Thunderbolt 3, and what’s next? Who knows?

    Our computers used to be the room, but we moved to minis, with the computer in the room. Then we got desktops and portable luggable machines, moving to laptops that we can carry one handed to handhelds computers in our pockets. We even went to tiny devices that we found were too small. So we’ve gone the other way with smartphones and phablets and iPads and tablets. Soon the small things will be larger and the world  around us will become enhanced with virtual reality and Hololens. Maybe.

    Our world is using all this technology to monitor, mark, chip, tag, record, watch, measure, and gather data. We get to work with that data. We get to gather, store, manage, index, backup, transfer, clean, and care for all that data. We need to work with it. We’ve got to move it with text files, CSVs, Excel, Word, PDF, MP3, MP4 and more.

    We send data over TCP, FTP, SMB, AirDrop, VPN, Web services, REST, jQuery, and more. We share data with files, messages, texts, clicks, likes, tweets, pings, drops, shares, snaps, hangouts, and once in awhile, we communicate with phones.

    What do we do with all that data? Why, we can do anything. We have PowerPivot, Power query, Power View, Power map, and Power BI. It seems Microsoft really believes data has power.

    We have the chance and potential to build amazing visualizations. We can analyze our business progress, producing tables, charts, graphs, animations, and of course, reports. We can map our own activities and events, tracking how we interact with the world, experience it, perhaps even using the data to relive, remember, or reinvent the world around us.

    But, we have so many things to learn in order to reach our potential in working with data. Fortunately, we’ve got all sorts of resources to help us, no shortage of books, articles, blogs, podcasts, tutorials, classes, and most impotantly, friends. I hope you take advantage of the resources to learn more.

    And you can start today.

    Steve Jones

    The Voice of the DBA Podcast

    Listen to the MP3 Audio ( 6.8MB) podcast or subscribe to the feed at iTunes and Libsyn

  • Opening Up Data

    Tim O’Reilly has been an advocate of open data access and standards for some time, especially from governments. He’s pushed for more interoperability and certainly more accessability from all sorts of groups. He gave an interview earlier this year to LinuxVoice where he talked about a variety of things, but data was foremost on his mind.

    There are some good thoughts, but I was pleased to see him looking for more software to adapt how it works with data rather than asking data to match the application. An interesting thought he had was in the area of control systems. Does every device or sensor need a separate application and way of interacting or should we have some guiding design principles that let similar applications work in similar ways with different data? That almost sounds like good data modeling and normalization principles in action, backing a data driven application.

    I also liked his acknowledgment of the fact that so much of our data isn’t very portable. Between social networks and proprietary storage, it becomes hard to move data around. The pattern of downloading data, perhaps editing, perhaps not, and then uploading elsewhere works great with ETL tools, but it’s cumbersome for many users and applications to deal with. Building ways for us to interact with disparate data, allowing for queries to remote sources, sometimes transforming and copying data, all of this needs to be easier to implement and integrate inside software.

    In some ways, I think the 3.0 model of our Internet interaction will take place around data. I think SSIS will continue to be one of the most valuable tools in SQL Server (along with lots of demand for work), but it still needs improvement and enhancement to catch up to other ETL tools. I really hope Microsoft believes this and continues to invest in the tool.

    I also think that the data professionals that really stand out in the next decade will be those that learn to make the choices about when to use R, JSON, XML, HADOOP, or whatever non-RDBMS tool to meet a need. But also when not to use these tools. The better data professionals will make good decisions about when to query data, and when to move it to another system.

    It’s an exciting time to work with data as the opportunities and rewards continue to expand and grow. I look forward to what the future will bring us.

    Steve Jones

    The Voice of the DBA Podcast

    Listen to the MP3 Audio ( 3.2MB) podcast or subscribe to the feed at iTunes and LibSyn.

  • Data Lakes

    I heard a new phrase this week: the data lake. It comes from a Radar Podcast episode where Edd Dumbill talks about data as an asset that should be exposed everywhere in an organization. There’s a blog post on the subject as well.

    The idea is somewhat centered on Hadoop, but it could apply elsewhere. Data comes into an organization and then tends to seek and get stored with other data in a large lake. Applications are just a way of accessing the data in the lake, but all the data really lives in a large Hadoop lake of information. In some sense, this isn’t far away from the “single view of the truth” that I’ve seen plenty of organizations attempt. In a relational world this means all data moves from OLTP systems to a large data warehouse, and is then moved to smaller data marts (really warehouse subsets) and is accessed from there.

    It’s a good idea in theory, but in the practice of trying to move data around with any velocity between users, with all the copying, cleaning, transforming, and more going on doesn’t work. OLTP systems are needed because there are transactional actions that must be completed quickly and accurately. Moving this data to other systems becomes harder as data volumes and the sheer number of clients (whether users or other systems) increase. The idea that we can keep all of our processes working quickly enough that users won’t get frustrated is likely a dream. The more a value exists in a set of data, the more users will access it. The more accesses, the slower it often becomes, which starts a cycle of smaller subsets of data and applications that subsist on those small data puddles.

    Excel is probably the most common example of a data puddle that exists in your organization. A set of data, perhaps out of date, but useful enough to make decisions based on. Infinitely flexible and convenient enough that updates, changes, and more often spawn more and more puddles where the information never gets transferred back to the large lake of a database, whether that’s an RDBMS, Hadoop clusters, or something else.

    I think the idea of a large data lake is great, but in a practical sense, much of an organization’s data will never live in the lake. If it does, it will most likely be data that’s been superceeded by information in a puddle somewhere on an employee’s laptop, tablet, or personal cloud.

    Steve Jones

  • Deleting Data

    Like Tim Berners-Lee, I find the “right to be forgotten” law in Europe to be dangerous. The tremendous growth of data in our world means that searches become increasingly important in order for us to find data. If data is removed from searches, then for all practical purposes, it might cease to exist.

    I think that is disturbing. As a data professional, I try to ensure that the quality and integrity of data is maintained. I really to try to avoid ever deleting data (preferring to archive or hide it) in case it’s ever needed later. Far, far too often I’ve had someone ask me to remove something, only to have them ask for it’s return a short time later.

    I know that this law isn’t removing the data. The original sources will still contain the data, but search engines won’t return it. On one hand that means that many of us won’t ever realize that the data exists. In some sense, this will return us to the very limited, analog search engines of the past, where we depending on indexes compiled on microfiche to find information stored in libraries. On the other hand, perhaps this will give rise to data investigators that will develop their own methods and archives that enable them to offer services where they can find the “forgotten” data.

    I can’t decide if I think this law is a huge step back, or a correction against some of the overzealous data creation that occurs in the spur of the moment. Certainly data about past events is valuable and important, but the way much of the world uses a search engine, clicking on a link or two and accepting the information as valid, can be misleading. Time will tell if this is a good idea or not, and I’m curious to see how well the law performs.

    Steve Jones

     

    The Voice of the DBA Podcast

    Listen to the MP3 Audio ( 2.0MB) podcast or subscribe to the feed at iTunes and LibSyn.