Tag: data analysis

  • The Data Request

    Spotify implemented a way to get your data from your profile. This looks at the request and initial data download.

    The Request

    When you go there, you see a “Manage Your Data” section.

    2019-07-08 16_47_45-Privacy Settings - Spotify

    Below this is a “Download your data” section, with three steps. The first is a request to verify your identity. Clicking this will send an email to your registration email. I like the idea of verifying your identity somehow rather than just letting you get the data. Minor security hurdle, but still one.

    After clicking on the email, step two lights up and shows you that it can take up to 30 days to get data. Similar to a GDPR request, I guess this sets some expectations if they get overloaded, but for now it’s annoying. We’ll see when the data comes.

    Step 3 is the download link, that you can resend.

    Overall, a good process to get what might be private data. It will be interesting to see what they keep on me and what songs and artists appear, both in popularity and unpopularity.

    Getting the Link

    About nine (9) days after sending my request, I woke up and found this in my email:

    2019-07-16 09_02_42-Your Spotify personal data is ready to download - sjones@dkranch.net - DkRanch M

    I clicked the link and went to the same manage data page shown in the first section, but before I could get the download, I saw this. Nice to see they are verifying access to get data.

    2019-07-16 09_02_58-Privacy Settings - Spotify

    When I verified my password, a file downloaded immediately. As expected, it’s zipped.

    2019-07-16 09_40_19-my_spotify_data.zip

    What’s in here? A number of JSON files, of various sizes. There are some items that are less interesting, like the plan and payments.

    2019-07-16 09_40_30-001-way0utwest

    If I open the playlist.json file, I see this:

    2019-07-16 12_52_44-C__Users_Steve_OneDrive_SQL_SpotifyData_Playlist.json - Sublime Text

    One cool thing here, is that this first playlist is one I created yesterday, the day before I received this link. Meaning, I think the notification was someone processing my request, not the time to get the data. The data was assembled with a process, that used data as of yesterday. I could be wrong, but I have a list of songs in a playlist.

    Going on to StreamingHistory, which is where I thought I might see interesting data, I see this:

    2019-07-16 12_56_09-C__Users_Steve_OneDrive_SQL_SpotifyData_StreamingHistory.json - Sublime Text

    This looks like I only have a bit of history, not a lot. The first row is April 2019, the last row is July 2019, when I got the download. Disappointing, as I’d assume I’ll have a much different set of history this year than last year. I’ll have to load this into a database and try to decode it. That’s for another post.

    It is nice to get some of this data, and I wonder what it looks like if I download it again in a week or two. Especially after another trip or two.

    If I lose the data, I can go back and get the file again, at least for 30 days.

    2019-07-16 09_45_55-Privacy Settings - Spotify

    I’m glad that companies are starting to give us some data, at least for us data people. Certainly most people might just like a visualization of analysis of data, but allowing users to get their data means that *anyone*, or more *anyplace* can do this analysis for you. Opportunities that your data gatherer might not create.

    If you use Spotify, I’d be curious what you think of your data. Feel free to write your post and leave me a comment on how you loaded and examined this data. I’ll do some follow-ups on this later.

  • We Don’t Have Perfect Information

    I was discussing the PASS Summit with someone and they were wondering about building their schedule. Actually, they wanted to pick sessions, but see the choices in a calendar format, but the schedule wasn’t out. My suggestion was to just build the schedule and then sort out conflicts later.
    A few people have mentioned over the years that they want to build a schedule and be ready for the event to maximize their experience and be efficient. I think that’s a common, normal, technical person thing to do. We’re Type-A, we like knowing and having a set schedule.
    The problem is that we don’t have perfect information. Even if the descriptions and abstracts included perfect information about the agendas, what is covered, and to what depth, including demos, we’d still not necessarily assimilate and recognize all that data. We’d think a session on database design covered fourth normal form, even when the text said third normal, or we’d expect that an SSIS data load talk included something on CSVs when the presenter described the talk as being with flag text files.
    We’re human, and that means we have flaws in how we deal with the world. This includes the ways in which we model and analyze data. While we can make mistakes in our analysis, we often may simplify our view of a problem to the point where our analysis is inherently flawed.
    I try to remember this when I write reports from systems that others will use. I won’t have every piece of information that might affect a system, but I try to ensure I have the most important, or significant, data. At least, the data I (and the users) feel is significant. The important thing to remember is that out data is always incomplete, and it’s entirely possible that we have missed a valuable piece of data.
    When that happens, we have to adapt and adjust our systems, just like our conference schedule. We’ll learn more across time and we can use that information to change our system. I know that my view of a conference like the PASS Summit today, or even a week before the event, will be different than what I know, and how I feel, at the event. I should have a plan, but be willing to flex as circumstances change.
    Steve Jones
  • Super Nerds

    I enjoy the game of baseball, playing as a kid and then for a decade as a 40 year old adult. I gave it up a few years ago, worried about the wear and tear of sporadic play. I’m happy now coaching kids and participating in individual sports for exercise, but I still love the game.

    A few years back I read the book, Moneyball. The book is about the use of metrics and analysis in addition to the human scouting evaluation in the Oakland As organization. There was a movie made as well, and both are worth consuming. Since that time,  other baseball teams have adopted some of the ideas, and I’ve been hearing that both basketballNFL football, and other sports in the US are using more metrics. There’s even a Revisionist History episode that talks about football/soccer and the way that makes the most sense to improve your team.

    Recently I ran across an article with a great headline: Jayson Werth rails against ‘super nerds’ that are ‘killing the game’. In the article, a player rants about the way in which data and statistics are being used to make decisions. Instead of allowing players to just play, analytics become a part of the decision to play a certain way. For example, bunting and stealing have dramatically declined, mainly do to the analytics that show these are lower percentage actions compared to other choices.

    It’s interesting to hear players rant against the user of statistics. I completely get the annoyance at losing control of your choices, and certainly appreciate that the game becomes less exciting at times. However, I also know that players, and even coaches, may make emotional decisions, or base decisions on poor information. Most of us humans can’t remember all the tendencies and likelihoods. In a modern world where skill levels have dramatically increased in many ways, there are often better ways to build a strategy.

    Data is valuable, and certainly can help in sports. It isn’t the end-all be-all, and it can be misleading when applied to individual humans. At the professional level, where more data is available, I think it makes sense to use data more as a significant part of your decisions, though not the only factors. For me, as a coach of younger kids where I have relatively little data, I still use the eye-test for most things, relying on data to double check my thoughts. The world moves fast, and it can be easy to forget how individual players have performed across an event, especially when I’m trying to manage the game as well. Data helps me remember how the day is going.

    Steve Jones

    The Voice of the DBA Podcast

    Listen to the MP3 Audio ( 5.1MB) podcast or subscribe to the feed at iTunes and Libsyn.

  • Finding the Right Data

    For a couple years, Big Data was heavily hyped, and Hadoop became incredibly popular. In fact, so popular that Microsoft build HDInsight and Polybase to allow us to take advantage of these technologies and integrate them into our own systems. While the year has seen less hype on “big data” specifically, more and more of us are dealing with large amounts of data every day. There isn’t a good definition of Big Data I’ve seen, but whatever you thought it meant five years ago has surely changed to mean larger volumes today.

    One of the important things that many organizations are learning is that they don’t necessarily need more bits and bytes of all their data. They’re increasingly learning that they need more of the right data, which is the data that is useful to them. Often this is the data that lets them make decisions that improve their revenue, profits, efficiency, etc. As we move to GDPR this spring, it might also me more auditing data that prevents problems or satisfies regulators.

    I ran across an interesting article that talks about companies needing the right data, which can often mean unstructured data outside of their traditional OLTP databases when dealing with customers. The article focuses on NLP (natural language processing) and social media data, but it could just as easily mean audio/video data from customer calls (or emails) or even sensor data from systems that are managed and track the ability of customers to use your product or service.

    As the world of computing advances, many of us know that we need to find new and better ways to provide value to our employers. This might be with managing and gathering or tracking a wider variety of data, perhaps meaning that some of us need to keep some of those tweets or posts inside our systems. It could be that we need to provide new ways of analyzing data, maybe with some sort of ML (Machine Learning) or AI (Artificial Intelligence) processing. Perhaps it’s that we need to collate and collect detailed auditing information we can produce on demand to ensure our organization complies with legal requirements.

    Perhaps there’s some other way that our work as data profressionals will change, but I’m sure it will continue to change and evolve across the next decade. I’m also sure this means lots of new and different data opportunities for us, if we’re willing to grab them with the right data.

    Steve Jones

    The Voice of the DBA Podcast

    Listen to the MP3 Audio ( 3.9MB) podcast or subscribe to the feed at iTunes and Libsyn.