Category: Editorial

  • Words vs Data

    I would guess that most of you reading this are very comfortable looking at data for insights and answers. You might even prefer to provide a result set instead of a picture or chart to a user when they are asking for help with data analysis. However, do you add any words to your analysis to help? Any descriptions, summaries, or conclusions that could be drawn from the data or the picture?

    I ran across a blog asking about the right ratio of words to data. The post uses the childhood story of Goldilocks and the Three Bears. Many of you might know the story and have drawn your own conclusions of what the story shows or means. If you read this post, you will find a very different interpretation. While some of you may not think that’s a valid interpretation, it’s possible that some thought that when they first heard the story.

    The point of the post is that we can provide data and pictures, but others might interpret things differently. Each of us has our own point of view, our own experiences, and our mood. That last one might lead us to focus on a piece of data or a part of the picture that the author didn’t intend for us to focus on, or didn’t think was relevant. Without any sort of guidance on the narration from the author, we don’t know how closely our interpretation matches theirs.

    Many of us have certainly seen others spin data, especially aggregates and statistics, to suit a narrative. However, the idea of providing some narrative isn’t to hide or mislead, but rather give context to what you see in the report. As the blog notes, don’t leave their interpretation to chance. Give them a “well-crafted, objectively reasonable narrative that is supported by your data.”

    Or, if you don’t have one, let them know that and ask them to send you one back showing what they see or what they expect.

    Steve Jones

    Listen to the podcast at Libsyn, Spotify, or iTunes.

    Note, podcasts are only available for a limited time online.

  • Data Debt

    I had never heard of data debt until I saw this article on the topic. In reading it, I couldn’t help thinking that most everyone has data debt, it creates inefficiencies, and it’s unlikely we’ll get rid of it. And by the way, it’s too late to get this under control. I somewhat dismissed the article when I saw this: “addressing data debt in its early stages is crucial to ensure that it does not become an overwhelming barrier to progress.” I know it’s a barrier, as I assume most of you also know, but it’s also not stopping us. We keep building more apps, databases, and systems, and accruing more data debt. Somehow, most organizations keep running.

    The description of debt might help here. How many of you have inconsistent data standards, where you might define a data element differently in different databases? Maybe you have duplicated data that is slow to update (think ETL/warehouses), maybe you have different ways of tracking a completed sale in different systems. Maybe you even store dates in different formats (int, string, or something weirder). How many of you lack some documentation on what the columns in your databases mean? Maybe I should ask the reverse, where the few of you who have complete data dictionaries can raise your hands.

    For most of my career I’ve heard a couple of terms that I’ve never really seen implemented. There’s the famous “single version of the truth” for a system, which seems to break down whenever we add a reporting or warehousing system. Even inside a single database, often an OLTP one, it’s hard to get a truth because values are changing so fast. The other term is MDM (master data management), which promises to ensure that every element is tracked and tagged the same way. No misspelled customer names or outdated addresses. There have been no shortage of products I’ve seen to help people tackle this problem, but ultimately I think the amount of data debt is too high. When we realize we need MDM, we’ll never pay down that debt, mostly because too many developers have too many habits and legacy ways of capturing data that will never get integrated into any MDM dictionary.

    The article seems like a great academic set of principles. Make sure you label all your data. Put governance in place, with good access controls. Train workers, establish accountability to properly manage data. Invest in scalable architectures. How many of you can add scale to your system easily? It’s always taken me jumping through a variety of hoops to do that. The cloud makes it easy.

    For a month. Then when the bill comes, you’ll be scaling back down.

    Really, the chaos of the real world, where organizations are not one thing, but a large number of people and groups, each with their own goals and processes, just trying to get enough done to keep the organization moving forward is where we live. There’s no real time to deal with data debt.

    Except if you’re the ETL person. We mostly pay you to move data around and clean it as best you can. At least then the problem remains hidden from the report readers, who trust you’ve actually done the T portion of ETL correctly.

    Steve Jones

    Listen to the podcast at Libsyn, Spotify, or iTunes.

    Note, podcasts are only available for a limited time online.

  • Creating vs. Maintaining

    If your job as a developer or DBA has been like mine, it’s a constant stream of requests to change something, often without enough information and short deadlines that create a bit of stress. There’s always more work to be done, and while it might be a great job, you’re often trying to finish something quickly enough to get to the next thing.

    In this mode, how often do you think about creating (or modifying) the thing you’re working on for today vs maintaining it for tomorrow. In other words, do you consider how easily your work can be understood, is documented, is designed to allow for flexibility, and can be enhanced without many (any?) side effects, or anything else.

    In other words, is it maintainable?

    We often build things to solve a problem and we can be very creative. We design a solution that works well, solves a problem, and may work very efficiently. However, I know I often haven’t thought about future maintenance. I haven’t considered how difficult it may be to understand the context around what I put together by anyone else. Sometimes I’ve built things that things that were very clever, but weren’t easily understood by my team or able to easily adapt to new requirements when those arise.

    Often I’ve been looking at the problem from only a “it’s this tree that’s important” perspective, and forgetting that my particular thing is part of a wider system (a forest). I hadn’t been considering future maintenance, which often led someone in the future (often a future-me) to tear it down and rebuild something new.

    That’s technical debt.

    When some structure isn’t maintainable, it’s debt. It’s a burden for the team in the future. Designing things quickly and building them within a deadline, while making them maintainable requires some knowledge, experience, and also discipline to work with the patterns, and avoiding the anti-patterns, that make code difficult. The same thing applies to managing systems. Custom jobs on every server and separate configurations make life hard. At the same time, a one-size-for-everyone approach also isn’t maintainable. We need a balance of well-written solutions that solve our problems, but are easily maintainable.

    Part of becoming a better engineer or admin is learning how to build maintainable things that others will continue to use for a long time, not because they have to, but because they want to because the code works well.

    Steve Jones

    Listen to the podcast at Libsyn, Spotify, or iTunes.

    Note, podcasts are only available for a limited time online.

  • The Era of Cloned Humans

    AI-technologies are evolving at an alarming rate. The ability of LLMs to produce drafts, review work, even write some code continues to improve to the point where junior level workers are in danger of having less opportunity than in the past.

    Perhaps even more alarming is the ability of AI technologies to mimic what they find in the real world, which can include voices. As with any technology created to make the world better, criminals will find a way to use it for nefarious purposes. This article notes that the FBI has warned families to have a secret word or phrase, as criminals are using AI to clone the voice of a family member asking for financial help.

    Can you imagine getting a call from your spouse, parents, or children, saying that they’re in trouble and need some money right away? With AI tech, the voice could even respond to your questions, mimic anxiety or distress, cry, or who knows what.

    Last year Microsoft announced a text to speech technology that can closely simulate a person’s voice with just 3 seconds of audio. I’m sure that capability has improved in a year and soon we may not be able to trust the voices we hear. We certainly can’t trust the photographs we see, something that we enjoy when we see artists photoshop images for fun, but when we look at an image  in the news, we want it to be real, but we can’t trust that an organization hasn’t manipulated a photo. When anyone could simulate another’s voice, things can go wrong quickly.

    I assume the ability to fake video in real time is coming soon. With enough hardware and some imagery, I would guess AI models will be able to hold a Facetime-type call as a human, fooling most people that might not know the person extremely well. At some point, since all our images and video are digital anyway, I assume that without some security measure, the tech could likely fool almost everyone.

    Even as I write this I doubt how well the quality is, but I’m sure that it will continue to improve to the point where we might be loathe to trust remote interactions. This has the potential to be a security nightmare for some people, and I worry about the scams and losses criminals will inflict upon the unsuspecting.

    This is one type of technology whose negatives will far outweigh the positives.

    Steve Jones

    Listen to the podcast at Libsyn, Spotify, or iTunes.

    Note, podcasts are only available for a limited time online.