Category: Editorial

  • The Vast Expansions of Hardware

    At the Small Data conference recently, one of the talks looked at hardware advances. It was interesting to see a data perspective on hardware changes, as many of us only worry about the results of hardware: can I get my data quickly? In or out, most of us are more often worried about performance than specs. However, today I thought it might be fun to look at a few changes and numbers to get an idea of how our hardware has changed, in the march towards dealing with more and more data. Big data anyone?

    In thinking about disks, I saw a chart that looked at the changes from HDD (hard disk drives) to SDD (solid state drives) to NVMe (Nonvolatile Memory Express). These show read speeds going through the list from 80MB/S to 200MB/s to 5000+MB/s. That’s a dramatic change, and not one only in high-end arrays. There are off-the-shelf drives you can put in a desktop that read this fast. If you think about some of the early IBM drives, which read at 8800b/s. Growth in disk speed, inside the timeline of our careers, has grown by a few orders of magnitude in read speed.

    Write speed hasn’t grown as much but capacity has. My early career work used HDDs with a 100MB capacity. These days we can get TB range storage on all of these mediums, with many laptops having 0.5TB or more on them. Desktops often have plenty more. My current workstation at home has 3.5TB of storage. Contrast that to the early IBM drive linked above, which had 5MB. These days people regularly demo hundreds of TB, or even 1PB queries from a database.

    Many of us just expect the network to work well. In fact, I assume many of us won’t complain to network people since they are never at fault for performance issues. I started my career with Arcnet connections between machines. Those ran at 2.5Mb/s. We were moving those and 4Mbps Token ring to Ethernet at 10Mbps with Thicknet, Thinnet, and eventually RJ-45 connections. When we got 100Mpbs bridges, I thought we were cutting edge for our SQL Server Central servers. If we look back 20 years, 1Gbps was more the standard then, but today we see growth up into the 800Gbps with Infiniband. While I don’t know many data centers doing that, there are plenty running in the 50Gbps range.

    If we think about CPUs, I started my career on a 386 machine running at 25MHz. I helped upgrade some 286 machines, but most of our servers were 486 class machines at 25 MHz. I still remember being excited about the early Pentium processors for a large system. There were many Pentium variants and later families of processors, but back in the 2000s, almost all machines were single-core. The first multi-core chips were released and slowly became more common over time. These days, many new laptops have multiple cores, including the new on I got, which has 12 cores. If you want, you can purchase an AMD Epyc 9004 processor with 96 cores. That’s on one chip. Since most servers can take more than one CPU, you can have hundreds of cores running if you want. If you want to get really crazy. the Nvidia Blackwell has thousands of cores for their GPU-based AI calculations.

    Memory has likewise grown, though it seems most servers are much less than a TB of RAM, which is a much lower growth over time than storage and networking. Maybe because of those two changes, memory has had less of a reason to grow into common multi-TB-sized capacities in our systems. In fact, for you reading this, what are the common memory sizes you have in servers? I see many VMs and other machines set up with somewhere between 128GB and 1TB for memory, even as their data sizes have grown much, much larger. However, there are plenty that don’t have anything near 128GB.

    That was one of the interesting things I realized about the Small Data conference, and one reason the event was created. Most of our data sets, especially usable sets, and most of our queries can run on a laptop if not a mobile device. The focus on big data seems overblown, especially as most of our companies don’t have anything approaching 100TB, much less 1PB. If you need it, there is hardware out there for you, but some of the amazing advances made over time are lost on me as the common, average capabilities out there on the majority of systems could handle the majority of my needs.

    With some well-written queries.

    Steve Jones

    Listen to the podcast at Libsyn, Spotify, or iTunes.

    Note, podcasts are only available for a limited time online.

  • The Modern Algorithm of Chance

    These days algorithms rule much of the world. From how supply chains are managed to how vehicles run their engines to the media that many of us watch on the various streaming services. I assume that most of you know that algorithms drive what you see on social media, on YouTube, and even the search results you get, and what you see might be different than what I see. There is a constant search for a perfect, or at least, very targeted way of getting you what you want.

    Or at least what the algorithm thinks you want. However, is that the best way for algorithms to be designed? It is for the companies that want to profit from your attention, but is this intense personalization better for us?

    There is an interesting article on music discovery, focusing on Spotify, since they are one of the largest streaming services. The article talks about the algorithm and how it tries to match selections to our tastes, basically a complex data analysis of our choices along with metadata that’s been created around data that’s hard to classify. There are attributes assigned to songs, but are these the attributes that make sense? That’s a topic for another day. The result of this is that Spotify tends to recommend more of what we already listen to, which has also driven artists to change how they produce songs since the algorithm matters.

    This seems like a similar challenge to what I’ve seen with the written word. A long time ago many of us consumed the words (with less choice) in newspapers, books, and other physical media. However, we often ran into random things that were different because of our physical paths in life. We might encounter books in a shop or library and be attracted to a cover for some random reason. We might pick up an unexpected work lying adjacent to one in which we were interested and discover something new.

    The way we look at books, or anything, changes when we browse and randomly wander the world. These days, we have less of that, with algorithms in electronic systems that guide us further on a path we’re walking, not allowing for chance encounters, or even wildly different thoughts because we stumbled on something. Even in our social media, this doesn’t often happen. I’d hope that we might encounter a recommendation from another we wouldn’t otherwise see, but the promotion of certain feeds and the glut of viral re-sharing often ensures that we don’t see many random things. Instead, most of us see the same thing that many others do.

    Those of us who have studied computer science know random things are hard to create in computer systems. Building algorithms that embrace randomness isn’t something many of us focus on, instead trying for matches that reinforce or duplicate something our clients already want/use/see/etc. That has helped create many businesses in the digital world, but I’m not sure that those businesses are always good for the world.

    I don’t have a good solution for random chance, other than talking with others, especially those who live different lives from you, and embracing the way they view the world. Hopefully that leads to a book, movie, or other chance encounter that you might not otherwise have.

    Steve Jones

    Listen to the podcast at Libsyn, Spotify, or iTunes.

    Note, podcasts are only available for a limited time online.

  • The Role of Databases in the Era of AI

    I’m hosting a webinar tomorrow with this same title: The Role of Databases in the Era of AI. Click the link to register and you’ll get some other perspectives from Microsoft and Rie Merritt.

    However, I think this is an interesting topic and decided to try and synthesize some thoughts into an editorial today, partially to prep for tomorrow and partly because I’m fascinated by AI and how this technology will be used in the future.

    The title says the role of databases, not data professionals. You might worry an AI is going to take your job as a DBA or developer, or you might think there is no way an AI can do your job. I tend to think the latter, but only if you are above average in your role and you add value by understanding your employer’s business. In those cases, the AI will help you (as a co-pilot, not a pilot) and allow you to get more work done or work done faster. You choose. If you churn out average, or below-average work, or cut/paste from Stack Overflow or SQL Server Central or anywhere on the Internet, then yes, you should worry.

    Databases store lots of information, and extracting that out is hard. I see no shortage of poor data models, no shortage of overloaded data in fields, de-normalized structures, repeated information, and more. Humans jump through lots of hoops to build reports or screens or other interfaces to present to humans looking for answers. We may load join data in Excel with values in a database or vice-versa. I’m sure many of you have plenty of stories on how you get data to move between some data store and a text format. I’m sure you also have no shortage of frustrations from your efforts.

    AIs will get good at this. At the Small Data 2024 conference, I saw many people working at using AI without a semantic layer, which I think is possible, but will likely fail. We store data in too many crazy ways, and companies will need to make it easy for customers to create a semantic layer that describes what data is stored in each place. They’ll also get the AIs to help not only with this but with creating a way to simulate Master Data Management without requiring every application to use Redgate Software, Inc. as a name. We need to ensure Redgate, Red-gate, Redgate Software, and RG stored in different fields can all joined as if they were the same value. Which they are.

    Fuzzy matching is the domain where AIs can shine, as the models can do this quicker than humans, without getting annoyed and with fewer mistakes. AIs can adapt with our feedback as we find ways to train the models better and overload the AI prompts with semantics that help translate the (extremely) poor data models in our databases, data lakes, spreadsheets, and even PDF documents. Companies that require a semantic layer can ease the process of building one with AI assistance so that customers can quickly start to query their wide array of data sources.

    The best use I’ve seen for AIs is as an easy-to-use, context-aware, powerful search engine. When we learn how to tune these for specific sets of data, such as all the datastores and spreadsheets in a company, we’ll start to see some amazing gains in information analysis. I don’t know that humans will analyze any better than they do today, but the process of getting the information to analyze will be easier. I think AIs will also help in the analysis phase, but that’s going to require more co-work between humans and AIs to improve the quality of analysis.

    There are other things, but I see databases as incredible stores of information that AIs will make easy to access. I’m also positive AIs will be used to more easily update information in databases and assist in easily moving data from one format to another or one location to another.

    Tune into the webinar tomorrow and see what Microsoft thinks and ask any questions you have.

    Steve Jones

    Listen to the podcast at Libsyn, Spotify, or iTunes.

    Note, podcasts are only available for a limited time online.

  • Serverless Gets Faster

    When the Azure SQL Database serverless option was introduced, I was a bit disappointed that I couldn’t get the database to pause any sooner than 1 hour. That meant I needed to ensure clients didn’t access the system for an hour, but also, that I burned an hour of compute after the last access.

    Recently I saw an announcement that this time frame has come down to 15 minutes. While this might seem like a very simple change from a technical standpoint (just alter a timer option), I’m sure there was more work needed. I’m also sure there was a lot of debate on the sales/marketing side to decide if this would lose a lot of revenue.

    I’m sure this costs Azure some compute revenue in the short term, but it might also create opportunities from customers who consider using this in new situations since it can shut down quickly. I certainly think this makes the use of an Azure SQL database for QA/staging type work more attractive. This might also get more people to take a look at serverless and realize the auto-scale benefits are pretty cool.

    My request would be to drop this down to 5 minutes and increase the range of auto-scale as well. Maybe allow me to go from 2-16 vCores if needed with corresponding memory jumps. I don’t know I need this by the minute, but I would like to have things shut down fairly quickly if we stop a workload and aren’t using the system.

    I’d also like a better retry on startup other than trapping an error on the client and re-sending my request to connect. It’s just embarrassing that we still have that happening for a cloud PaaS service.

    Steve Jones