Category: Editorial

  • Send Metrics Not Logs

    This is part of a series on observability, a concept taking hold in modern software engineering.

    One of the interesting things I saw in an engineering presentation on Observability from Chik-Fil-A was that they are sometimes bandwidth-constrained at remote sites. In an early version of their platform, they sent logs back to HQ, and their logs used all the available bandwidth, so they were unable to process credit card transactions.

    While most of us don’t deal with lots of remote offices sending data back to a central data warehouse, we do often work in distributed environments, and we may send data to/from a cloud or even employees’ remote offices. Or maybe we send a lot of data between components. Bandwidth is very good in many parts of the world, but it isn’t infinite.

    In the presentation, they talked about a tool, called Vector, that can work with lots of data, slice/dice/aggregate/sample/etc. the data, and then send the results to a sink location. This works like many other ETL tools that have a source and sink, along with various transforms that operate on the data.

    It’s an interesting philosophy to try and send back metrics that might be useful to developers or Operations staff in understanding the performance of their system. By only sending metrics, the load on downstream systems is reduced. This also allows us to store less data and read metrics sooner rather than storing all the data and processing it each time someone needs a metric.

    The flip side of this is that taking this approach means that the consumers of the metrics need to ensure they are getting useful and actionable information. Determining what is needed will be like any development project, something built, iterated, re-tested, and repeated. This might even be an ongoing part of building software as new features and logging are added to your software or system.

    In general, I prefer to have more data over less, but the volumes of logging and instrumentation data have grown dramatically. Some systems are producing more log data than actual data on a daily basis. Like audit data, we likely need to reduce and limit the amount of data stored long-term. However, we want to keep the important data that we find useful.

    I am looking forward to trying out Vector and seeing what’s possible. Having good CLI-based tools that can work with data is becoming more important all the time, especially as more of us move to DevOps flows, coding our systems operation in text, storing it in version control, and deploying on demand.

    If you’ve used Vector, let us know what you think, and if you prefer another tool, share why today.

    Steve Jones

    Listen to the podcast at Libsyn, Spotify, or iTunes.

    Note, podcasts are only available for a limited time online.

  • Responding to a Disaster

    Ryan had a planes, trains, and automobiles situation a few weeks ago. On the same day, I didn’t. A delay for a few hours, but an easy trip home. Lots of other people didn’t have a smooth trip, and I had a few friends who spent an extra night somewhere or had flights canceled and decided not to take a trip. The Crowdstrike outage hit the entire world, causing problems everywhere. Just before that happened, Azure had an outage in the central region.

    If you were affected by these, you have my sympathies. If your IT job included responding to these, you get all the virtual hugs from me (and Brent). I know what it’s like when someone calls you to handle an outage, and I know it can turn your life upside down. Hopefully, it hasn’t been too stressful to you.

    In Brent’s post, he talks about reviewing your DR plans and presenting your findings to your boss. Good advice, and worth reviewing yearly as most of our environments will experience some amount of change. If you follow this advice, make sure you present this in a logical, dispassionate way. I see plenty of technology professionals who get upset when DR isn’t given enough priority. That doesn’t help the situation. Channel your inner Spock and present your results, allowing someone else to decide what to do and accepting their priorities. You can advocate for change, but keep in mind that there is often no shortage of work and limited time. If someone else decides you can accept the risk of poor plans, document that and move on.

    Plenty of us will experience a disaster at some point. I’ve had more than a few occur at various positions, and I know many colleagues who have gone through a wide variety of failures. Even for companies that have extensive DR plans, there will be challenges when a disaster occurs. We’re certainly not at the point where any AIs can handle reading our DR document and implementing fixes. Instead, those of us responding will have a guide, maybe a great guide, for how to respond, but we will need to adapt our plans to the current situation. Having the mindset that our plans might not be perfect (as Mike says), will help you deal with any challenges that arise.

    Aside from the technical situation, there are also mental challenges in a disaster. You will feel pressure from your employer, but also pressure at home. Your partner might not understand the extra hours. If you miss commitments made to your family, they will be disappointed, and I hope you are as well. You might have expected your schedule to include fun events, and you likely didn’t expect to get less sleep. Your diet might suffer in times of crisis and long hours.

    All of these situations can be managed, but it helps to think about them when you’re calm and make plans. Warn friends and family. Think about how to limit unhealthy foods or situations, set expectations with yourself that help you manage stress, and most of all, treat yourself with kindness. Many of us want to help and support our co-workers but ensure that you (and them) get breaks. Imagine you’re in a Crowdstrike situation that lasts for multiple days. Talk about how that will look, even with just a few people and you’ll be better prepared when a disaster event occurs.

    Hopefully, it won’t be a large-scale event, but better to prepare for that and experience a minor situation than the reverse.

    Steve Jones

    Listen to the podcast at Libsyn, Spotify, or iTunes.

    Note, podcasts are only available for a limited time online.

  • A Quick Turnaround

    I visited a customer last week and attended SQL Saturday Baton Rouge 2024. Both were fun events, and an enjoyable week, though I was away from my wife for 4.5 days, which she didn’t love. Today, I had a quick turnaround, heading to Wisconsin Dells for THAT! Conference, which I attended last year and enjoyed. I didn’t submit this to the event, but got asked to go as part of my job. I accepted, and I’m gone from home for a week between these two trips.

    I do travel a lot, but these trips got me thinking about how many of us might handle the unexpected demands from our companies. In this case, I had planned on SQL Saturday Baton Rouge, but not a customer visit. I got asked a month ago to add this in, which was fine, but two weeks ago the trip got extended by another day. Just before I got the update from that call, I agreed to go to THAT! for a quick presentation.

    As a developer, I’ve been given last-minute work. It might have been a bug that was discovered or a new request that was high priority. In those cases, I’ve often had to work more hours for a short period, usually long days or even weekends. I’ve lost personal time because of the hours, but also the stress, as my mind is elsewhere. As an Operations person, I sometimes get stuck at work overnight or for long parts of many days.

    In most of my jobs, I’ve been a strong performer and had a good relationship with management. I’ve almost always been able to negotiate comp time for the extra hours spent. In some cases, I’ve been paid for them, but usually, there is the equivalent of shorter days for a period or even missing some days of work. Often this is off-the-books as most HR systems don’t cope well with this.

    For those of you who get unexpected demands, how do you handle things? Do you get something back from the company? Time, days, compensation, maybe even a thank you and a gift? It’s not something that I think is universal, but I’d like to think it’s common.

    And if you’re at THAT, say hello to me this week. I don’t know when I’ll get some shorter days, but likely they’ll come in the next month as I try to catch up on ranch work.

    Steve Jones

    Listen to the podcast at Libsyn, Spotify, or iTunes.

    Note, podcasts are only available for a limited time online.

  • Thinking About Technology

    Technology has dramatically changed the world over time. The advent of cars dramatically changed the US, as people could go places and meet others in a way that was difficult and slow before. The telephone let us communicate with people all over the world at a pace that was previously impossible. Computer technology has furthered this at a truly amazing pace, especially since the adoption of mobile devices by so many people. The flexibility in how we can integrate computer technology into our lives has been incredible.

    However, each technology change brings about plenty of negatives and potential problems as well. I ran across a piece from L. M. Sacasas that has some questions we might ask about any technology, including the software we build. The start of the piece is that most of us don’t think about how our work might be misused, which can lead us to dismiss security risks or moral misuse risks. We often don’t consider the malicious ways people view applications.

    The piece is interesting to read, but it ends with several questions that we might ask ourselves as we build something. I think a lot of these questions might not apply to our work with databases or corporate technology, but some do. I think many of them might apply if we think about the tools we use, especially AI.

    Most things many of us build are re-hashes of something else. We might smooth the flow of work with better UX in an application. We might rewrite code to more efficiently use resources. We might implement features or reports in response to a business request, but we often are lightly evolving our software not changing it. We do, however, build things that others might misuse, either accidentally or maliciously.

    The list of questions is interesting, but I also think we need to consider that others might use your system differently, not just from a UX perspective. Consider how they might exploit your software to achieve other aims. We may not be security experts, but others are experts and there are plenty of tools available to scan code and identify vulnerabilities. Use these tools with the knowledge that just because you wouldn’t use the software in a particular way doesn’t mean others won’t.

    Steve Jones

    Listen to the podcast at Libsyn, Spotify, or iTunes.

    Note, podcasts are only available for a limited time online.