Tag: data privacy

  • The Data Industrial Complex

    At a keynote this week (video, transcript), Tim Cook, CEO of Apple, outlined some of the issues that we face in this world where technology is being used to gather, hoard, sell, and use tremendous amounts of data. He used the term data industrial complex, a take off of the military industrial complex that has industry working closely with government in a way that might make some of us uncomfortable. Or may align interests that benefit the few over the many.
    I know that Apple has been on both sides of this debate in recent years. If you get all your data, you might be surprised. There have been concerns about tracking, about the storage of messages, even after phones are wiped, and other issues. At the same time, Apple has provided tools and protection that have stymied law enforcement and upset governments with their encryption.
    This isn’t to defend or laud Apple in any way, but rather to note that privacy issues around data are complex and of a concern. I like the general topics that Mr. Cook outlines, and I do think that we have lots of work to do in these areas, both as private organizations and with public laws and frameworks that outline rights and responsibilities, while still allowing innovation and creative use of data.
    That is a hard balance to strike, and I really don’t quite know how I want my data used and protected. At times I want to prevent the use of data for purposes I haven’t agreed to, but at other times I appreciate the creative and helpful use of information about my life. I’m not even consistent about the ways in which I treat my own data at times, which reminds me of the complexity in dealing with sensitive data about individuals.
    The GDPR, which took effect this year, and similar laws (CCPA, AUS Privacy Act, etc.), are a good move forward, in my opinion. They might go too far in some ways, or be too lax in others, but we need to start the discussion and examine the effects of some concrete rules in the real world. My hope is that we have a regular and constant debate on how we should treat data in an increasingly connected world that gathers larger volumes and more types of data than ever before in our history.
  • We Need Data Privacy Consistency

    For most of the last year, I’ve had quite a bit of my time devoted to the GDPR and related topics. My company is affected, as it’s based in the UK. Not only must we comply, but we know many other companies must as well. As a result, some of our product focus was aimed at helping companies solve their data privacy issues, especially with regards to data.

    That continues to be a good idea as the GDPR isn’t the only regulation out there affecting organizations’ data handling practices. There are other laws around the world, but the US is a big market, one of the biggest we have, and we are seeing increased need in the US for the same types of data privacy and protection solutions mandated by the GDPR.

    California recently passed their own data protection legislation, and it’s leading the way in the US. Tim Ford wrote a short piece on how this affects his company. He notes that as a consumer, he’s glad to see stricter data handling practices being required. However, as a business owner, he’s concerned and I think there is some basis to be worried.

    There are other laws that might pass soon in the US. New York has a bill, Colorado has signed a weaker, but still new, law. Other states are considering items, but the US Congress has yet to really move forward on any legislation, which might lead us to have multiple data handling practices that are required. That would be a nightmare, much more difficult than the hiring and tax practices of different states.

    I couldn’t imagine having to work with different processes, and certainly wouldn’t want to have more restrictive laws being passed in the future that might cause us to change practices multiple times. I can only hope that the US gets a common law for all our states, and that the practices are in line with what the GDPR requires. Other countries have used that as a basis for their laws, and I can only hope the US does the same.

    Steve Jones

  • What Data is Really Needed?

    The GDPR took effect in May of this year, at least with regards to enforcement. A few days after the May 25 date, a German court ruled against ICANN, the company that registers domain names on the Internet and manages the global WHOIS database. The case revolves around the information collected when you register a domain. ICANN wants multiple contacts, which they’ve required for decades. However, a company in Germany that is a partner, argued that the additional technical and administrative contacts were not required for fulfilling the business that both ICANN and EPAG (the German registrar) are engaged in. ICANN Is appealing the ruling, citing the need for clarification of what this means with regard to the law.

    This is interesting to me, because a) it concerns data, and b) there is an interesting argument here to be made about what data is needed for a business purpose. I could see this being argued successfully either way, and not just in court. As a domain holder, does the registrar really need multiple different sets of personal information from me? Arguably, this is a convenience for them, one that is based on tradition. However, one could argue the other way.

    It is a little scary that a court, with no expertise in some industry (Internet domain registration, in this case), will decide if there is an actual business need. After all, can a lawyer or judge really understand what data a business needs in their daily activities?

    Maybe, maybe not, but I do think this forces businesses to actually stop and think about what data they collect, have a justification, and document that. That’s a good thing, because often I find business people just asking to collect data without any idea what they’ll do with the information. I also find technical people collecting data, not maliciously, but often to anticipate what might be asked of a system, or because they want to avoid rework and just decide to collect everything they can.

    Data is precious, and while I don’t want to put many limits on what data businesses can collect, I also don’t want to them be able to collect anything, not disclose what they’ve collected, and not secure it properly. Having some limits, or at least forcing them to consider the risk of holding old, useless data, is likely a good thing for all of us.

    Steve Jones

    The Voice of the DBA Podcast

    Listen to the MP3 Audio ( 3.7MB) podcast or subscribe to the feed at iTunes and Libsyn.

  • Anonymisation Confusion

    The GDPR starts getting enforced in a few weeks. It’s been law for a couple years, but the authorities have given companies time to comply. I know various entities are frantically working towards compliance as I keep getting updates to Terms of Service. My company is among them, and we are diligently ensuring we can prove that we aren’t violating any rules. That’s good because I’m sure fines will reduce any bonus we might earn this year.

    As I’ve been reading over the law and talking with customers, I’ve learned quite a bit. Redgate Software builds products to help with compliance and we’re updating guidance and information on how to work with data. As I’ve helped to update information and explain concepts I keep running into the term “pseudonymisation”. If you listen to the podcast, you’ll probably hear me struggle to pronounce it, but more importantly, I was initially confused about what this actually meant.

    The Data Protection site from Ireland has a great description of how this differs from anonymisation. You can read through the document, but anonymisation means that the data can’t be some how reverse engineered to find the original data. In terms of privacy, an anonymised set of my data wouldn’t allow anyone to determine the data is about me.

    If the data is pseudonumised, data about me would be replaces with a token, but there might be other means of discovering a data set is about me. As an example, in an eCommerce system, you might have an order key in a dev data set that’s copied directly from production. However, my name would be replaced with something else, like Bob Smith. It’s not apparent that it’s my data, but the protection is limited. If the data were anonymised, the order key would be replaced as well to prevent reverse engineering.

    Many of us have gotten used to being lax with dev and test data, often just restoring from production. It’s handy, convenient, and allows you to find problems in production or verify changes using known values. The downside of this is that we have poor data security. There are no shortage of data breaches from dev and test systems. Certainly plenty of data has been lost from developer laptops as well. Even if you had your laptop encrypted, there’s no real excuse for using real data in less secure environments.

    We need to learn to use anonymised data, and become comfortable with the idea of working on secure data sets. That also means we need the skills to ensure we build good, useful, valid datasets with production-like characteristics.

    Steve Jones

    The Voice of the DBA Podcast

    Listen to the MP3 Audio ( 4.2MB) podcast or subscribe to the feed at iTunes and Libsyn.