Category: Editorial

  • What Will 2017 Bring?

    What do you think will happen in the database world in 2017?

    That’s a question I want to ask you today, the last work day of 2016. When most of us come back to work next week, a new year starts, though it won’t really mean much to most of us. We’ll continue on with the projects we’ve been working on, managing the same systems and dealing with similar issues to those we face today. Budgets may reset, which could be a good thing if you can find a way to divert some of that money for your own training or pet project use. In general, next week will just be a continuation of the work many of us have been doing.

    If I look forward and try to imagine where 2017 will take us, I envision focus in a few areas, and perhaps a few things that won’t change. As much as I find the progress our industry has made in the last ten years amazing, I also think that year to year we tend to make small changes. It’s rare that a huge advance in computing drives us forward. Usually we can see the technology emerge, gain momentum, and then grow very quickly. That happened with SSDs. The first models were exciting, and expensive, but also prone to failure and burnout. Across a few years, quality improved, prices dropped, and all of a sudden most new machines now use SSDs. In fact, it seems most people working with databases wouldn’t consider purchasing hardware without at least some SSD storage.

    There’s a lot of media attention being paid to Artificial Intelligence and Machine Learning these days, and I think these will grow more rapidly in 2017. As we get more tools that make it easier to build applications that incorporate AI and ML, I envision pressure on many developers to begin incorporating these features into applications. Those that are able to do this effectively will likely help their organizations gain a significant advantage over competitors. The maintenance and evolution of these systems is harder than expected. While our tooling will lower the bar to get started, my suspicion is that most of the applications that get built will not provide the expected results and will get abandoned.

    That brings me to the second area that I think will grow in 2017. Data Science, and the idea of somehow analyzing all the Big Data (from 2015) to gain amazing new insights will be an even bigger focus in 2017. The media attention and hype have so many managers thinking they need more data science. The interest means higher paychecks, which have everyone that passed Introduction to Statistics in university claiming some data science skills to get a larger paycheck. This will snowball in 2017, as more and more people on both sides press the issue.

    The reality is that data science and more complex analysis is hard. It’s much harder than most people realize, requiring lots of knowledge and patience to experiment with data. I’m not sure that most people are willing to make the investment in time and resources that it takes to become a good data scientist. Of course, if it’s like many of the other career paths in technology, even average skills can end up with a successful career.

    Security will continue to be a problem, especially as more and more hackers continue to probe organizations and systems for vulnerabilities and weaknesses. Some do this for profit, some (maybe most) for basic vandalism, but we’ll continue to find SQL Injection, phishing, and other attack vectors to be problems. As we connect more and more systems together, and as there is pressure to build more distributed systems, whether with cloud services or business partners, we will have more and more weak points. I wish I could say that 2017 would be the year that our aging systems will see pressure to improve their security and patch more rapidly, I suspect that far, far too many large organizations will continue to tolerate poorly written systems (from a security standpoint) and allow their developers to deploy code that doesn’t remotely adhere to best coding practices. Perhaps one day…

    In our SQL Server world, I think we’re going to have an interesting year. All signs point to another SQL Server version, one that runs on both Windows and Linux, so we’ll start to have the challenges and excitement of moving databases across platforms. Add in the ability to run inside containers, and I think we’ll see some crazy growth in SQL Server deployments that will stretch the ability of administrators to keep track of how many instances are actually running. I suspect that we’ll see lots of data loss from small instances that are deployed by developers quickly and easily, without a backup plan. I wouldn’t be surprised to see a whole new set of issues when transaction logs grow to fill disks from instances running in containers that never back them up. Maybe we’ll get the simple recovery model as the default for SQL Server 2017? I can hope.

    All in all, I think 2017 is just another new year, but it is the time for you to take stock of your life, decide what you like, what you don’t, and where you want to go. Perhaps you should take another step on your epic life quest, maybe you want to improve your skills, maybe you want to find a new job, or perhaps you have something outside of work that will be important this year. No matter where you are in life, the end of one year and beginning of a new one is a good time to reflect, reminisce, and dream about the future.

    Steve Jones

    The Voice of the DBA Podcast

    Listen to the MP3 Audio ( 7.5MB) podcast or subscribe to the feed at iTunes and Libsyn.

  • Looking Back at 2016

    We’re coming to the end of a crazy year. 2016 has seemed to be one of the craziest of my life with world events like Brexit and the US election as well as an astounding number of data breaches. More people who impacted my life passed in 2016 than in any other year I can remember, and I traveled far, far too much this year. Quite a change from the beginning of the year when the Denver Broncos won SuperBowl 50. 2016 has also been a very interesting year in the data world.

    Certainly the release of SQL Server 2016 was exciting for many of us. For the first time since 2012, or really since 2008, I thought this was a true, major release of the platform. I was surprised and pleased by the amount of features added and improvements made to this version. I very much liked to see the inclusion of a number of security features. While some of these need some maturity and work, they do bring us some additional capabilities that I think start to help us implement better data protection for our database systems.

    We’ve also seen a few things I’ve written about for years coming true in the SQL Server world. We have a Linux version in CTP status, due to be released next year. Whether adding a Linux edition is a good idea or not remains to be seen, but I am glad that Microsoft is making an attempt to port SQL Server to other host platforms. With SQL Server 2016 SP1, we also have a common programming surface, allowing almost all of the T-SQL features that were previously only in Enterprise edition to be used in other editions. This means we’re closer to paying for SQL Server based on the scale of data we process. I think this is a good move that makes sense for Microsoft and customers. While some might lament the hardware limits on Standard Edition, I think they are fine. I just wish it wasn’t sure a big jump to move from Standard to Enterprise, or there were an option in between the two.

    The cloud has grown tremendously in 2016, in many areas, but certainly for data. While AWS and Azure grew in size, they also lowered prices for users. It’s not clear how much of this usage is just for database work, but I certainly think that more and more organizations are looking at moving a portion of their data to the cloud. When you can store data cheaply and scale your query computing up and down, this starts to look like a viable option for some workloads. I don’t know that I think most RDBMSes used for on-premise applications make sense in the cloud, but some do, especially when your customer base is distributed and your workload has predictable spikes.

    I think the idea of cloud databases for analytics and warehousing makes more sense. Those are the workloads that require larger hardware for peak workload levels and become expensive for local systems. Getting your data to the cloud is a challenge, but I suspect that data movement, gateways, and other innovative ETL (or ELT) solutions are coming. The Azure SQL Data Warehouse is a very interesting product to me, as is the Azure Data Lake, and I look forward to seeing how people start to use these solutions in the future. Certainly the cloud is going to continue to play an interesting role for data professionals in the future.

    This was an interesting year of hardware for me. The DevOps movement has said that we should treat servers like cattle, not pets. I’ve started to try and do this with hardware as well. I got a new laptop (VAIO Z Canvas) in 2016, and after setting up my old one with Chocolatey, I did the same with the new laptop, becoming productive with my new machine in a few hours. It helps to have various distributed data services like Evernote, Dropbox, and remote Git Repos, but I suspect many of you have similar services inside of your organization. When I rebuilt my desktop this year, I was using it about an hour after I rebooted the new hardware thanks to Chocolatey. This really make me rethink of my individual machines as cattle. Provided I have some good way to remotely keep various data accessible. That brings me to another interesting issue.

    Data breaches were an issue in 2016, as in years past. They seem to be occurring on a regular basis and increasing in size, though it’s hard to determine whether or not the impact to individuals is greater. The Yahoo breaches were incredible in size, with over 1 billion accounts affected. However, other security issues are just as worrisome. The DDOS attack on DYN shut down a number of sites, but what if a more subtle attack managed to change DNS entries. I’d worry that as more of our data is accessible through public networks, the compromise of credentials could lead to more data loss issues.

    However, one of the most common issues with computer security has been shown to be out of date software with vulnerabilities where patches have been available. To me, this means we could thwart a significant number of breaches by keeping software up to date with patches. This brings to mind plenty of other issues, but ensuring our platforms are patched seems to be important. Perhaps ensuring we can patch our application software quickly if there are issues from patching platforms is one way to improve security.

    Perhaps one of the items that I think dramatically changed in 2016 was the growth of data analysis. Whether we look at the call for data scientists, R analysis of data, visualizations with tools like Power BI, big data or something else, it seems like these were important topics in 2016. This might be the first year where I think that business intelligence truly took a leap forward with tools designed to make analysis easier, and lower the bar. This is interesting, and perhaps profitable, for the average data professional to use in their jobs, but all these tools and options for analyzing data don’t necessarily mean that we will better analyze data. I suspect that many of the initiatives started by individuals and organizations will be abandoned in the short term because the experiments aren’t well designed, and the effort to cleanse and prepare the data for some predictive analytics is greater than what most companies want to invest in this project. However, there are more tools and ways for people to begin their journey to better understanding statistical methods and use them for analysis. The data professional of the future (say 10 years from now) will have a better understanding of data analysis techniques as they begin their careers. Just as we expect most data professionals to have some understanding of relational databases and SQL today.

    It’s been a long year, and I, for one, am glad to see it ending. Some great memories and trips, but too much travel for me. I only took 16 work trips (and 5 personal ones), but three of those were over two weeks and I spent about 80 nights in hotels. Hopefully I can substantially lower those numbers in 2017.

    Steve Jones

    The Voice of the DBA Podcast

    Listen to the MP3 Audio ( 11.1MB) podcast or subscribe to the feed at iTunes and Libsyn

  • Syntactic Sugar

    There are all sort of features and enhancements that Microsoft can make to the SQL Server platform. If you look around Connect, you’ll see suggestions for improvements, such adding common checks, as well as additions like adding virtual tables. I’m sure many of you would like to see simple things, like Regex added to the T-SQL language. In fact, you might find that some of these small changes, which we can code around or build, should just be added. After all, if Dynamic Data Masking can be added, shouldn’t some other simple features be included, such as helping us solve the “string or binary data truncated” error?

    These handy, useful, simple changes to the platform are often called syntactic sugar. These are changes that are simple, many of us could easily code them, but they make development easier. Or even administration in the case of the SQL Server platform. These are not necessarily expanding the power or capability of the platform, but they can make working on the system more enjoyable.

    Should Microsoft create more syntactic sugar for SQL Server? Certainly they do at times, but perhaps not as much as many of us would like. The thing to keep in mind is that making changes to the SQL Server platform can be very difficult and time consuming. Adding small features that might be helpful, while enticing, can slow the evolution and development of more core product features. Or they can cause more problems than we might expect in other parts of the platform. Would you rather have something like DDM that makes obfuscating some data easier, or a more robust replication engine that recovers from more problems? Do you want stronger security features like Always Encrypted, or more robust Always On features, or is it more important to get Regex added? I think we might have differing opinions here.

    These can be really hard questions, and certainly I think our feedback can help influence Microsoft. After all, if there is a nagging issue that is constantly causing issues, then maybe it’s worth a syntactic sugar improvement, even if this takes resources away from some other area. The one thing I hear from Microsoft over and over is that specific business cases and issues are more important than complaints. Express the issues you have with the platform in terms of workarounds, developer time lost, or specific performance issues rather than just a “I don’t like this” or something “doesn’t work as expected.”

    Ultimately Microsoft is a business, and they do look to add new features regularly to the platform to increase sales. I get that, and I try to temper my requests and complaints. I like to see them focus a portion of resources on core improvements to systems that work, and while I think this does happen, I’d like to see a bit more improvement in existing features. Certainly our backup and log reading systems have improved over time. SSMS is decoupled and being updated regularly. There is more work to be done, and if lots of us provide specific feedback, I’m sure we will see even more core improvement over time.

    Steve Jones

    The Voice of the DBA Podcast

    Listen to the MP3 Audio ( 3.8MB) podcast or subscribe to the feed at iTunes and Libsyn.

  • Data Preservation

    This editorial was originally published on Dec 3, 2013. It is being re-run as Steve is on holiday.

    Maintaining data across time isn’t something many of us think about. We work with data in the here and now, and in the database world, we typically only need to recover or restore data from a short window. Like most of you, I would usually plan on recovering data that’s only a few days old. Being forced to restore a database from two weeks on any of my systems would make me cringe. It certainly would be embarrassing for me personally if it were my fault I couldn’t restore to a point in time that was more recent than that.

    In planning to recover our systems, we typically know the versions of software we have to recover from, and we can easily re-download copies of SQL Server or the patches we need. Most of us are dealing with SQL Server 2000 or later, which is good since those are the only versions still documented on MSDN. If you need SQL Server 7.0 or SQL Server v6.5 documentation, I hope you have copies.  The same goes for the media. You can still download SQL Server v6.5, and SP5, but if you needed SP3, it isn’t easily available. I ran into that situation about 10 years ago, and we had to make a special request through our TAP manager to get someone in Redmond to dig up a copy.

    In some ways it might not be important to worry about long term storage. Most of us will end up transferring our data to newer systems (and formats) over time. As we upgrade SQL Server, our databases move along to newer formats, or we abandon them because they are no longer needed. That’s fine for some data, but not all.

    Long term archival and storage is a challenge, as you can see in this short look at how old films are maintained. It just touches the edges of what’s being done, and doesn’t address costs. Plenty of old films have been lost forever, and perhaps that doesn’t matter, but it does concern me. I have thousands, maybe tens of thousands of digital images. While I love the ease with which I can share them with family, and make extra copies, I am worried that perhaps the lack of a physical copy means my great-grandchildren will struggle to find evidence of my generation if there is a catastrophe or storage formats change.

    This is one area of our industry in which we have a lot of maturing to do, and I hope that we can come up with some new ideas for maintaining our data for the long term, across not months or years, but decades or centuries.

    Steve Jones