Tag: data privacy

  • A Data Controversy

    Quite a bit has changed since this article about airlines and the US government.  Since very few people are flying, or even can fly, perhaps this disagreement is moot, but I bet it comes up again. Now, separate from the idea of the actual disagreement here, there is an interesting discussion about the data involved here. In short, the US government wants airlines to collect data about passengers to help track the COVID-19 virus. Airline executives say they can’t easily get this data, other than on paper, without spending a few months on development.

    Certainly having a way to gather additional information in a digital form can require some development work. There are all sorts of software decisions to be made about when, where, and how users might input information. We have mobile devices, kiosks, laptops, and more, all of which might require separate interfaces for software changes. There is also the testing, validation, and verification we want to ensure the software works well and doesn’t introduce instability.

    In today’s world, with growing legislation, there is also a question of privacy. These requests may or may not conflict with other laws that airlines are bound by. There is likely to be more conflict here as the world changes and laws are slow to change and converge in some type of consistency. Rapidly changing requirements, as have been pushed during the COVID-19 pandemic, can potentially put us technical people in a difficult position. We have to balance the urgency of meeting requirements with the potential liability of violating privacy. I’d hope we could find some balance there, especially in a crisis.

    We do need to be flexible and ready to adapt to changing requirements. If regulations change, our organizations ought to be able to prioritize these changes and rapidly deploy them. In today’s world, where many high performing DevOps companies can get new software out in hours or days, governments may expect large companies to be prepared to follow suit. In that case, especially where new data is needed, having a software development process that includes the database is critical.

    Steve Jones

    Listen to the podcast at Libsyn, Stitcher or iTunes.

  • Automatic Refaction of PII

    Data privacy and protection has become a hot topic. Across the last few years, as I’ve worked with customers that deal with the GDPR or other laws, they’ve become more concerned and careful with their use of data. I find less and less resistance from developers about using sensitive production data in development environments, but still too much.

    Using data that’s stored in databases or other text files is one thing. What about data in less structured forms? I’ve dealt with a few customers over the years that recorded customer interactions for various purposes. Call center or financial organizations commonly do this, and sometimes deal with sensitive information. I know I’ve given my date of birth, credit card number, bank account number, and other data to representatives at various times.

    Those calls are often recorded, and the IT staffs have often had to ensure extra security is applied to these files. Not everyone has access to listen for various reasons, but certainly when there is sensitive information inside the audio (or video), this data needs the same protection we’d apply to data in other forms. Providing protection, or redacting the information, isn’t an easy task.

    I saw recently that the Amazon Transcribe service will now redact some PII information automatically. This service can be configured to automatically remove the information in text, which is fantastic. This is a great way to start to use technology in a safe way to ensure that we have less data  leakage when we re-use data. Certainly people might look at this data to better train reps, but it’s also likely someone wold look through transcripts to determine why customers are calling in and use that information to better design applications. In either case, there isn’t any need to expose PII data to them.

    This doesn’t protect against the data inside the audio, but perhaps companies can delete and remove those recordings sooner with transcripts available and more quickly reduce their potential attack surfaces. We’ll always have some liability, but reducing that and not unnecessarily creating issues is part of what we want to do when protecting data.

  • The Changing Nature of Data

    Are addresses sensitive or private information? It’s a good question to ask since many of us have address data in our databases. I asked this recently at a SQL in the City event and the room was split. I come down on the side of “no”, for addresses in and of themselves. After all, the domain of addresses is known. It’s public information in most every country.

    A few people pointed out that while the address isn’t private data, when it’s linked to a particular person, it is private. It’s not the address, but the linkage. To me this should give data modelers pause when trying to set up a schema, whether set in an RDBMS or a schema on read in some other type of data store. Separating the user from the address, and having a link that doesn’t necessarily disclose private information can reduce the surface area of sensitive data in your system.

    A second question: have you ever worried about your name being on a door or mailbox? I know some people in larger cities have, but that might be a minority. As I’ve visited friends, a name is often valuable to see on a mailbox, especially in my rural area where houses aren’t very visible from a road. That might change, or need to change. An article in the Washington Post notes that in Vienna names are being replaced with numbers. The linkage to an actual person is being removed in response to a complaint. It this overkill? I don’t know, but it is worth thinking about.

    Google Street View and similar services might be affected. The service blurs faces, but it might need to start blurring addresses or even houses. I’m not sure I think that the images are problematic from a privacy perspective, but I also know that the ability to harvest data remotely and create linkages occurs at a scale and with a creativity that I would never have imagined.

    Could a set of thieves search for people posting a vacation notice, image search for a house and then start correlating those images with Google Street View to find addresses? Sure, though arguably a search of public records for ownership might be easier. Many people rent, so maybe this is a bigger issue than I think? I’m not sure, and really, trying to determine how criminals might use data hurts my head.

    I do try not to be too paranoid, but I do get concerned about data privacy. The stories of abuse I hear in the world are truly stunning. The creativity of criminals is scary. I don’t know where to draw the lines, but I do think that we should neither be cavalier with data nor paranoid. There’s a balance to be found, but one that needs debate and deep thought, not casual dismissal or overreaching concern. I hope as a society that we move in the direction of careful consideration as we derive some framework for both the protection and use of personal data.

    Steve Jones

    Listen to the podcast at Libsyn, Stitcher or iTunes.

  • Fines for Data Access

    Most of us know about the principle of least privilege. After I wrote about network segmentation recently, I’d hope that most of us know that limiting access to production data from all workstations might also be a good idea. Many of us also know that unusual patterns of access might indicate an issue. I wonder how many of us have a system in place to look for unusual access, especially this is something that might help us prevent, or at least detect, potential hacking activities.

    I wonder if we’ll get to the point where we need to do more and not only implement better auditing of system, but also data access. Will we need to actually monitor what data is accessed and ensure there is a valid need to do so? That was probably needed in this case, where various employees accessed a woman’s DMV data for who knows what purposes. This is a creepy story, and I’d hope that it’s the rare man that actually does this in any company, for any kind of data.

    As a general recommendation: please don’t spy on someone that you want to date. It’s very much an invasion of privacy and no way to develop a relationship with anyone.

    There were logs of access in this case, which isn’t the case in some states, but there should be more logging of access. The default “black box” trace in SQL Server doesn’t give us much information, though if you have a monitoring system in place, you likely do get more data. The logs of who and what are important, and they are something I’d like to see built into SQL Server by default, an easy way to log activity, with some archival/management of the data.

    Too many of us don’t log anything, and maybe it’s not important for some applications, but in the cases where sensitive data is stored, I think logging ought to be required, and I’m hoping GDPR 2.0 or other laws mandate this, perhaps with some retention of a year. I know I’d certainly like to see this built into more applications.

    The way we handle personal information has been poor for most of the digital age, and I would like to see that changed in a number of ways. I think requiring compliance against a standard might be the best way, though without specifying an implementation. This way we accept the responsibility of safeguarding data that we already should feel.

    Steve Jones

    Listen to the podcast at Libsyn