Author: way0utwest

  • The Uninteresting But Necessary Work

    There has been a rising tide of data legislation in the last few years that asks for organizations, especially private companies, to better protect their data. The GDPR is one of the most well known, taking effect with regards to enforcement earlier this year, and I’ve been doing quite a bit of work in relation to this law. There are plenty of other laws, such as California’s CCPA, Australia’s NBD, Japan’s APPI, and more that we ought to be aware of as data professionals. These laws affect personal data about people in a variety of ways, and they can affect how we process and use portions of the data we store.

    It’s not as simple as it might sound to change our data handling practices. In fact, it might not be that easy for many of us to do this now unless we’ve actually done something in advance: we need to have classified our data. We need to understand the impact of the various columns in our tables, the exports of flat files or reports, and even the development processes that make copies of our production databases.

    I’ll be honest, classification work is mind-numbingly boring and uninteresting. This almost feels like busy work to me, especially once we get past the obvious tax IDs and birthdays of people. When we start examining other data, the task feels like it ought to be delegated to junior staff, but many of them lack the experience to make the decisions. What can be more frustrating is that most of them lack the status to get others in the organization to respond to questions, which means the task ultimately falls on more senior people. This also means they often do it once and then forget it, leading to out of date information.

    We don’t like doing classification, but we need to do it. Without having some mechanism that allows us to determine if data can be moved or used in another system/database/report/etc., we end up just ping-ponging around. We assume all data is sensitive and try to lock it all down. That leads to complaints, as well as staff working to circumvent the rules until they appear meaningless. At this point we might give up on controlling data and just trust people. That leads to audit problems, potential data loss from security incidents, and plenty of embarrassment about why we didn’t implement some simple controls.

    Then the cycle starts again.

    Classifying data is simple in some ways, but not easy to ensure the data is available, up to date, and easy to find for any size team. I’ve seen simple solutions that rely on spreadsheets. I’ve seen complex software packages that are expensive and cumbersome to implement with other applications. Microsoft has started to help with a few changes in SSMS, but this doesn’t seem like a long term solution, though SQL Server 2019 might help. Redgate has spent some time on this as well, thinking about the issue and we have an early access program now.  All of these are partial solutions that might work for some organizations, but not all.

    Ultimately, this is something like security, that we ought to be building into our systems from day one. Every proof-of-concept or prototype ought to be classifying data from the beginning. We won’t be perfect, and won’t get every label correct, but if we’re always thinking about the data, we can always correct our label and more tightly or loosely decide to handle data. I’d also like to think that if we conservatively label the data early, we’re unlikely to get into positions where we are mishandling data in a way that makes it more likely that we accidentally lose data.

    Steve Jones

    The Voice of the DBA Podcast

    Listen to the MP3 Audio ( 6.0MB) podcast or subscribe to the feed at iTunes and Libsyn.

  • No More Mysterious Truncation

    If you read the Microsoft White Paper on SQL Server 2019, there’s a gem buried on page 17. It mentions a trace flag, which some of you might appreciate.

    Here’s a little repro:

     2018-09-25 15_00_20-SQLQuery1.sql - Plato_SQL2019.Sandbox (PLATO_Steve (65))_ - Microsoft SQL Server

    As you can see, the plee I made has actually been heard. I don’t know if they listened to me, or if the collective complaints over the year grew to the point that this got fixed.

    In any case, thank you, Microsoft. This is a very nice enhancement.

  • Practical Refactoring

    Today I hosted a webinar with Gene Kim (@RealGeneKim) and we had a fantastic discussion. I was slightly star struck since I’ve been reading his work and quoting him for years in talks about Database DevOps. It feels like I got to work with someone really famous, and I’m hoping I didn’t appear too nervous on the webinar.

    In any case, we had a great discussion, and I think you can still register to watch the recording. If not, we should have this on our Redgate YouTube page soon. We discussed the State of DevOps report, and specifically how the findings relate to databases. It was a good discussion, but when we talked testing, we both had some links to ways that we could build better software.

    In my case, I referenced this talk, Practical Refactoring, which I think is great. A bit is the technical approach, but mostly I find the philosophy and freedom that comes with having tests in place to be invaluable.

    This is based on a real project that these consultants worked on. The code is mocked, so don’t get caught up in the actual methods and structure, but think about how you could apply the ideas to your own work. How can you make the code better in a few minutes.

  • SQL Server v.Next is Coming in 2019

    Yesterday the 2018 Ignite conference kicked off with a number of announcements from Microsoft. You can re-watch some talks and the keynote, catching up on Windows, CosmosDB, and more. The big announcement concerning most of us is about the next version of SQL Server. This has been called v.Next in NDA briefings, as we’ve learned of different pieces of work, but now there’s a name: SQL Server 2019.

    The preview version, CTP 2.0, was released yesterday for public download. This is the first public release of the next version, and actually the first version I’ll be installing. To date I haven’t had time to even try to work with any previews. The keynote covers a bit of the product, but to see What’s New, check out Books Online. There are some database engine changes, such as UTF-8 support, better index rebuilds, improvements in Always Encrypted, Java programmability extensions and more. Big Data Clusters come as well, with Spark and better HDFS support. I don’t know much about it, but lot of friends that work in analytics are excited about Spark support.

    There are also some enhancements for SQL Server on Linux, which start to bring the two platforms closer together. Replication has been added, as well as DTC support, which have both been blockers for some users. AG support in containers is really interesting, though I’m not positive that this is that helpful. Machine learning services and OpenLDAP support are worthwhile additions as well.

    I am most excited about the secure enclaves for Always Encrypted. These will finally allow AE to be a more useful technology, and I’ll be updating my security session with this information. Maybe we’ll actually start to see AE deployed in more situations where high security is required as most of the operations we’ve needed, such as LIKE and range evaluations, haven’t been possible. I’m excited about the possibility of better security, though as most of us know, the weakest links are still the human and the client computer.

    There are plenty more enhancements, including more database scoped configuration items, better query processing for some opertions, more synchronous AG replicas, and maybe better, auto redirection of AG clients without a listener. There aren’t any new enhancements to the data classification options, though I’m hoping that will change before RTM. One last note, SQL Operations Studio has been renamed to Azure Data Studio. I’m not a big fan of naming changes, and I don’t like either of these, but I am curious if any you think this is the way forward for our toolset.

    All in all, this looks like a nice evolution for SQL Server, but not a major release with lots of new features. Perhaps my BI colleagues that use Spark or Java programers will disagree, but I don’t see anything that would be worth upgrading for the SQLServerCentral servers. Even in most of my jobs, other than getting the Always Encrypted enhancements, I don’t know many of these features would provide enough of an ROI. If you feel differently, let me know. There is some other coverage at Brent Ozar, MSSQLTips, and SQL Performance.

    One last item, if you’re looking to get started with any of the new Microsoft technologies, there are a few options. Certainly we will cover some items here, but we tend to focus on the data platform. Microsoft has announced Microsoft Learn, with content that covers Azure, PowerApps, and more. This might be a good resource for your learning plan this year. Pick a lunch or two a week, a weekend morning, or some other time and try to slowly improve your skills.

    Steve Jones

    The Voice of the DBA Podcast

    Listen to the MP3 Audio ( 6.0MB) podcast or subscribe to the feed at iTunes and Libsyn.