Category: Editorial

  • A Large GDPR Victim

    During the last two years, Redgate has been preparing for the GDPR to take effect in the European Union. As a company based in the UK, we recognized that there were both challenges and opportunities for our business. We needed to ensure we were compliant with the regulations, which would likely require us to change processes and educate our employees. At the same time, our customers would face similar challenges and there was an opportunity to help them achieve their own compliance with software tools.

    The GDPR enforcement began last May, though fairly slowly with few fines and decisions being handed down. Across the EU, it seems to have been a quiet period with few companies told they were non compliant. Most organizations likely think all their preparation has been worth the effort and likely believe that they are prepared for any complaints from customers or investigations from regulatory authorities.That confidence may have been shaken in the last week as Google was assessed a fine over $50 million for violations. In particular, the EU regulators in France found that Google had not obtained the consent needed for using certain data in personalizing ads. They also decreed that Google had not clearly presented information about how users data would be handled and stored, as well as creating a difficult process to opt out.

    This fine isn’t much for the tech giant, but it’s just the start and will likely force Google to change the way they handle data. It may also have implications for other tech companies of all sizes. Google is appealing the decision, and this will be an interesting case to follow for data professionals since we may need to ensure that we can comply with the final ruling. Many of us view the data in our organizations as belonging to our employer, with free reign in how we handle, process, and store it. That may change quickly if the ruling is upheld.

    Much of the decisions about how companies will deal with data is made by others, but data professionals often need to ensure that we do comply with whatever rules our organizations decide to use. This means a number of practices that we must consider. At a high level, we need to know what data is affected by the GDPR, or any other privacy regulation. This requires that organizations have a data catalog that allows them to track which data is sensitive and must be handled carefully. Few organizations have a comprehensive data catalog already, so this will be an area in which to focus resources during 2019.

    Once we are aware of where our sensitive data is stored, we must take precautions to protect this data throughout our organization. Most companies have implemented security in their production environments, but their data handling practices in test and development areas are often not the same. The GDPR calls for anonymization, randomized data, encryption, and other protections, which data professionals will need to implement in a consistent manner throughout their IT infrastructure.

    Finally, accidents and malicious attacks will take place. This means that every organization really needs a process to detect data loss and a plan for disclosing the issues to customers. Auditing of activity, forensic analysis, and communication plans need to be developed, practiced, and distributed to the employees that may be involved in security incidents.

    There may be other preparations needed, and the larger the company, the more work that will be required. Tools are critical to ensuring this process can be completed in a timely manner, both to save time in implementing processes and also to show regulators that actions are underway to better protect data. Fine levels aren’t mandated, and the more effort put into achieving compliance, the less likely that regulators will assess a fine equivalent to 4% of your annual revenue.

    There will be plenty of other GDPR fines in the future, and it is worth following this case with Google to see how stringently the regulations will be enforced. The world of data handling practices is changing and all organizations need to get used to better disclosure of practices, tooling for customers, and protection of the data assets they hold.

    Steve Jones

    The Voice of the DBA Podcast

    Listen to the MP3 Audio ( 5.5MB) podcast or subscribe to the feed at iTunes and Libsyn.

  • The Devil is in the Monitoring Details

    Monitoring a database server is something that many of us know is important, but we often take the process for granted. Whether we’ve purchased a tool, like SQL Monitor, or we’ve built our own system, we often set up a watcher for our systems and rarely view the details unless something goes wrong. I’m not sure that’s the wrong approach as part of the reason monitoring is set up is to allow data capture in the background and remove one more task from our daily workload.

    Monitoring isn’t necessarily simple, however, and while I still debate the best way to do this in many organizations, I realize that monitoring isn’t necessarily something I want to build in house. There is enough work to just work with the data that other systems might output that I really want some other software in place that is built to perform monitoring for specific technologies. In reading about the complexity for the Stack Overflow monitoring systems, I realize that this can become very complex for a “set it up and let it run in the background” configuration.

    The team at Stack Overflow built their own system for monitoring various systems, including SQL Server, but I think part of the mission of Stack Overflow was to build a system from scratch, which isn’t the job for most of us. Plenty of us have other tasks to deal with as a part of our job, and software development for monitoring or alerting or some other administrative task isn’t one of those jobs. I know I wouldn’t want to stop and think about data management and gathering, and more as a software process. If I’m a DBA, I want to just get the data and use it to ensure systems are running well.

    Monitoring can be a way for us to proactively look for developing issues and mitigate them before clients know there is a problem. It’s important that a system is in place and handling data. It’s even more important that there is some alerting application in place as well to ensure that when something does start to go wrong, the DBAs are alerted early enough to prevent widespread problems.

    If you read about all the thought and details of the Stack Overflow system, you quickly realize that there is a lot to consider when setting up the monitoring for your systems. I’d encourage you to think about what is important and ensure that you’ve got some way to gather and analyze that data. When something goes wrong, and something will go wrong, you’ll appreciate the time spent on the details of the monitoring system.

    Steve Jones

    The Voice of the DBA Podcast

    Listen to the MP3 Audio ( 3.2MB) podcast or subscribe to the feed at iTunes and Libsyn.

  • Your Recovery Time

    Disaster Recovery is both an exciting challenge for many DBAs and also a dreaded event that many hope they never experience. When a true disaster befalls your system, there is a tremendous amount of stress as we work to get things running again so clients can access data. Even the best laid plans will often have a glitch, and while it’s a great technical challenge, most of us would rather keep practicing and gaming possibilities than actually experience a disaster.

    The majority of our system disasters are localized to the actual machine handling our workloads. We often don’t have outside disasters, like hurricanes, with disruption and damage to other parts of our infrastructure and ever our lives. When that happens on top of a down system, it’s a level of stress that can affect our health. I’ve been lucky in that I’ve experienced both kinds of disasters, but never at the same time. I hope you can say the same thing.

    Let’s assume some local disaster defalls your system. Hardware, software, it doesn’t matter, but you end up with a corrupted database of some sort. This week I’m wondering if you have an idea of what the recovery time would be? How long before clients are up and running? Maybe more important, how long before you have system rebuilt?

    If you have some sort of High Availability (HA) plan, then you might be back up for clients quickly, in minutes or even seconds. The disaster really isn’t over, however, since you are now running on less hardware than you planned. Until you can rebuild the downed node, you’re still in disaster mode. If you’re like most of the organizations that have employed me, you’ll also be stressed as your formerly well-designed HA setup is now a single point of failure and you’re scrambling to get the main node rebuilt.

    We often think of a disaster as the time we’re down because of some event. Once clients are being served again, we tend to relax and think the disaster is over. It’s not, because until you replace the affected systems and bring them back online, you’re even more vulnerable than you were previously to another Murphy’s Law incident.

    Steve Jones

    The Voice of the DBA Podcast

    Listen to the MP3 Audio ( 3.0MB) podcast or subscribe to the feed at iTunes and Libsyn.

  • Faster Cloud Warehouses

    I think the cloud is a perfect place for a data warehouse. In many organizations, I’ve found that a data warehouse system is often the largest SQL Server database, both in size of database and also in terms of resources allocated. These systems often handle many complex queries for business users and are allocated a large number of CPUs as well as lots of RAM. Even then, many ad hoc BI tools or lots of “what-if” queries can bring the system to its knees, often causing lots of stress for database administrators during the periods of time when the system is in heavy use.

    Fortunately, many of these systems aren’t in use all the time. Often these are systems used by financial departments to “close the books” at month, quarter, or year end. It’s at these times when lots of resources are needed. Outside of these times, the data warehouse might be one of the least used systems, which makes it a perfect choice for a cloud, scale on demand, environment. Scale up when needed, down when not, limit your costs to the resources you need, when you need them.

    Microsoft has increased the capabilities of the Azure SQL Data Warehouse quite a few times across the last few years. I was thrilled to see ASDW separate out storage from compute, allowing customers to scale up the query, or compute, nodes independently of the storage used. Changes last year improved the ability to move data around between compute nodes as well as increased the number of concurrent queries.

    These improvements are perfect for data warehouses, and if you are looking to build a more responsive SQL Server based warehouse, you ought to take a look at ASDW. More and more customers are finding it valuable and cost effective in their businesses. What seemed to once be a niche idea has grown into a business that quite a few customers are using, with the demand growing.

    What’s more, the move to Big Data Clusters in SQL Server 2019 seems to have adapting some of this technology to the regular SQL Server product many of us use. These will separate out the storage from compute, something that should help many of us scale our systems to meet the demand of our workloads. I haven’t tried a big data cluster yet, but I’m looking forward to seeing how well one works and if it truly scales SQL Server further than I would have dreamed.

    Steve Jones

    The Voice of the DBA Podcast

    Listen to the MP3 Audio ( 3.6MB) podcast or subscribe to the feed at iTunes and Libsyn.