Tag: data

  • Being Responsible for Data

    For much of my career, I’ve run SQL Server Central. A large part of the popularity of the site is from the forums, where people can pose questions about their struggles with SQL Server and get answers from the community. There are also some off-topic forums, where people discuss various things outside of databases. In here, we have discussions a about life, sports, and more. While we do expect people to maintain an air of professionalism and respect others, we don’t try to moderate content.

    That’s how much of the Internet has worked, with various sites allowing users to post content, but not having any responsibility for what has been posted. The liability for that lies with the person doing the posting, which creates a thorny issue when users post anonymously. Setting that aside, I’ve been a proponent of this, not believing that Facebook or LinkedIn, or SQL Server Central ought to be liable for what users write and post. I do thinks users bear that responsibility.

    However, in the US, there is a Supreme Court case that may change our view, and that of many others. This case deals not with the data itself, but rather the algorithms that might display or recommend some of that data to others. That’s an interesting approach to the case law that has shielded many tech companies from their users’ poor behavior. Essentially the plaintiffs argue that Google and Twitter bear responsibility for their algorithms, which in this case aided terrorist recruitment. Meaning that the code they wrote to analyze data, essentially the queries that promoted content to users, were harmful.

    There are four possibilities listed in the article for what could happen, and I find them fascinating from a data analysis standpoint. Essentially a ruling against tech companies could shape how many of these companies process data in the future. While we might like to ensure these companies do not promote harmful content, think about this from the data analysis view? Do you want these companies to moderating how they provide results? Would this mean that we need to more carefully craft our search terms? In the context of tremendous floods of information, we often depend on Google, Bing, or some search algorithm to distinguish among the various meanings of words to bring back results relevant to us. At the same time, we might wish that everyone got the same results from the same search terms.

    Separate from the results, what about related results, or suggested items that might be related. I find the quality of these can vary for me, but often there is something “sponsored” or “I might like” that is helpful to me. Or just interesting. The infinite scrolling that many people live, getting similar recommendations is a double edged sword. It can increase learning, pleasure, etc. It can also send someone down a rabbit hole of anger and reinforcement of negative emotions. I think this also is one way that the content of the Internet creates division and disagreement among many.

    While I think users are responsible for their words, I also think that the way that these companies recommend and showcase content likely bears some responsibility. At the same time, I can’t imagine how you regulate this, and I do not want to see a constant battle of lawsuits over how we interpret rules. The sex, drugs, and rock and roll issues of the past, where we tried to legislate morality, didn’t work well. I don’t want to see that again.

    There isn’t a good answer here for me, and of the four possibilities, I fall somewhere between two and three. Some changes to section 230 (the legal writing) but not heavy changes or an abandonment of the way this has been interpreted. What do you think? Should we start to hold companies responsible for how they present content? I don’t know I worry for SQL Server Central, but it might change other sites. For us, we just show things from the last 24 hours. It’s not much of an algorithm, but it is one that likely isn’t going to get us sued.

    Steve Jones

    Listen to the podcast at Libsyn, Stitcher, Spotify, or iTunes.

  • Looking Back at 2022

    This is the last workday of 2022. Next week starts a new year, and as I’ve often done, I wanted to look back at the year. This time I decided to look back month by month, at some of the headlines and memorable data-related topics. I’m tackling things month-by-month.

    In January there was a set of “tech experts” who shared their thoughts on the best database management systems. This one is worth a read for the humor involved. I wouldn’t really consider many of these to be DBMSes. I thought about including this one as an April 1 joke, but it was a real story. I get asked for my opinion at times by writers researching a topic they don’t understand. I hope I don’t come across like a few of these people.

    February was another sad data breach story. In this case, from the state of Washington where many tech people live and have startups. It wasn’t clear initially what happened, but later articles noted this was from a stolen device. To me, this was a great reminder why dev machines (and databases) should NOT have PII. Mask/obfuscate/anonymize that data please. Or at least delete my name from your dev systems.

    March had another humorous story (to me): Oracle is going to lure people away from AWS and SAP with their new offering. I could believe the latter, but not the former. Oracle hasn’t ever been good about pricing and it seems more people are leaving Oracle than coming to it.

    April is the month of April Fools, but this isn’t a joke. Another data breach, again from a dev system. This one from Fox News, which included information of not only employees, but guests and celebrities. The claim that this was a dev system and not production doesn’t matter if the data is real. Please people, keep prod PII out of dev.

    May had a funny post from Hacker News. Someone put their whole life in a database. The comment that caught my eye on HackerNews: “Men will literally devote hundreds of hours to building a bespoke database tracking every moment of their lives instead of going to therapy.”

    I worked a lot on my weight and diet in 2022. June had me finding a public database to help me choose better food. A public database on processed foods. Great idea, but everyone has an agenda. I hope this has some crowdsourcing and reasoning and isn’t just one person’s opinion.

    In July, another data breach. This time in China with information for 1 billion people. Wow. I dislike large databases for this reason. It’s also a good reminder why you ought to remove information from your databases over time, at least the PII part. Again, delete my name, if nothing else.

    There are so many types of database platforms. Have you heard of a vector database? Apparently, it’s for managing vector embeddings, whatever those are. An August article on the strange growth of the database market.

    September started this crazy AI art craze. There was a call to remove living artists from the database of works that an AI uses. Makes sense to me. I think artists deserve support and while I like AI doing new things, maybe wait until the artist isn’t producing work.

    October showed a good reason why we need ongoing patches or open-sourcing of code for retired systems. There was a 22-year-old vulnerability reported in SQLite.

    In November, what other news could there be than Lego Steve? It was a tiring week.

    December is just ending, but I’m ending on a reason why databases without auditing are a problem. Men behaving badly in this one. Gathering public information (or even semi-public) at scale can be problematic, and the information gets abused. Better controls, but also more auditing and triggering of some actions to prevent this (and other) sort of abuse.

    Let me know if you remember these events, or perhaps if there’s a favorite memory of the 2022 data world that you wish I’d included.

    Steve Jones

    Listen to the podcast at Libsyn, Stitcher, Spotify, or iTunes.

  • My Stats from 2022

    I wrote last week about my travel, 23 trips in 2022. However, I’ve been gathering some other stats about my life and what I do, so I wanted to share a few of them here.

    Music

    I mostly listen to music on Spotify these days. One thing I loved are the Spotify wrapped updates at the end of the year. For 2022, I had these stats:

    Writing and Speaking

    It’s my job. A few stats:

    • Editorials written – 151
    • Coping Tips – 257
    • Speaking sessions delivered: 34 talks
    • Events: 19 events
    • Blog views – 74,000+
    • Visitors: 55k+

    Workouts

    Taking care of my health is important to me. I’ve always exercised regularly, and at one point I ran every day of the year for a few years. I’ve relaxed a bit, and settled into more of a routine. I miss a few days and I don’t stress about it, but I make my best effort to work by body a bit most days.

    Stats from 2022. I might have missed some, but these are the gross totals. This adds up to more than 365 since I do two things some days. A lot of weight days have cycline or walking or something as a warm-up

    • Yoga: 89
    • Cycling (indoor and outdoor): 84
    • Weights 63
    • Walking 22
    • Elliptical: 22
    • Swimming: 14
    • Cardio classes: 2
    • Rowing: 2

    Total workout days: 232

    The total days is low for me, 232, which is 63% of the year. My aim is 75% of the days, but I had ankle surgery this year, which threw me off for a month, and I spent way too much time on the road with travel and missed working out some of those days.

    Driving

    I don’t know exactly how much I drove, but the Tesla give me some stats and I’ve been tracking a few. For the Tesla,  show about 17k miles this year, which is mostly me. For all the times my wife might drive (or the kids), likely I put miles on the X5 or Suburban (or a few in the Ram 3500), so this is not a bad set of mileage for me.

    I also got to drive in two countries and a few states. Rough stats:

    • Miles driven 17,000
    • Countries: 2 (US and Portugal)
    • States: 5 (Colorado, New York, California, Florida, Nevada)
    • Top driving locations from the Tesla: Lifetime Fitness (135), Safeway (98), and a local gas station (89). Yes, I have a diet soda addiction.

    Reading

    I spend a lot of time reading books. They are an escape for me from life, and a way to improve myself. In 2022, I finished

    • 101 books
    • 35000+ pages
    • Avg. screen time on Kindle app for Dec: 10 hours/week
  • A Big Year for Travel

    I do tend to travel a good amount as my kids have gotten older. The pandemic slowed things for a year, but only then. Someone remarked on this year being a lot of trips, and it has been. Not the most ever, but the most fun ever.

    I decided to look at the travel numbers for the last few years, and see how this compares. Here are my trip numbers, for both business and personal, as best as I can tell.

    • 2013- 18
    • 2014 -17
    • 2015 – 17
    • 2016 – 16
    • 2017 – 15
    • 2018 – 20
    • 2019 – 23
    • 2020 – 1 (Las Vegas for my wife’s birthday, as the pandemic was starting)
    • 2021 – 11
    • 2022 – 23

    Looking at this, 2022 ties the biggest year, but it didn’t feel as bad as 2013-2105.I don’t have exact numbers, but I know I was closer to 30 trips those years and felt more burned out.

    I think because we have some fantastic vacations (Bruge, Hawaii, Venice, Lisbon) plus 3 volleyball trips and 3 trips to see my daughter in college.

    Hopefully 2023 will turn out to be just as big and exciting with more places that we’ll visit and enjoy.