Tag: software development

  • The Challenges of Splitting a Table

    I ran across a discussion on Reddit about splitting a table. In this case, the original post had to do with a vertical partition of data, which is a technique that can help you better manage data in your database. However, I haven’t often seen this technique employed in the real world.

    I wonder how many of you have considered a vertical partition when you are modeling data. Often we may not think about this early in the lifecycle of an entity, but as it grows, you might think about reducing the amount of data you often query in some way, and a vertical partition can help.

    Is there some criteria that you might use in deciding this? Or how you can evaluate if there is a need? I once worked on a system with a very hot table, lots of queries, lots of updates against this table from our online system. In response to some requests, the developers wanted to add some columns to the table. This was important, and we needed to capture the data.

    These were valid columns, but they were large in terms of data size, and not every one would always be used. This was before the option of sparse columns, so that wasn’t an option. I had no interest in an EAV table, despite the fact that it might have worked well at this scale. Instead, this was a situation where I thought a vertical partition would work. In fact, I thought a few of the other columns in this entity could be moved as well, as they were rarely queried and contained significant data.

    We split the table, and performance actually improved for the main table, as it had less data. Just like an index, we had more rows on every page and less IO for range queries, and even key lookups for data that wasn’t already indexed.

    There are lots of good techniques in database development for dealing with the challenges we face in data modelling and with performance. I’d urge you to learn about some of them and understand when they can be useful. I would also practice implementing them, making changes to existing tables, and learning how you can deploy them if the need arises.

    Steve Jones

    Listen to the podcast at Libsyn, Stitcher or iTunes.

  • The Challenges of Resetting Databases

    I was working on a demo recently where we had a database in version control and a development database. This was a team environment, with a few of us making changes and syncing them across our dev systems using git. We had some advanced technology with our dev environments in containers, Flywaydb, and GitHub. Once we had our scenarios working, someone wanted to reset our git repo and capture a new database image.

    However, when we reset the repo back, we had some issues with the database. In this case, there were changes in the database that didn’t exist in the repo, giving us a mismatch. Not a big problem, but cleaning things out to get the db to match the repo, without putting those changes into the repo, was a challenge.

    A developer I was working with got a little frustrated, because when working in C#, there is no state. If we reset the repo and sync our local copy, we have everything ready to go. However, a database repo isn’t the same because there is often a database that exists separately.

    This is the main challenge when working with databases, relational or otherwise, in a development environment. Experiments, bug fixes, even testing data changes persist over time. Resetting data to repeat tests, or even automating tests, can be hard.

    This is one reason I think containerization and subsetting of production datasets will become very important over time as we try to ensure we can react to business requirements and keep our teams coordinated. These technologies ensure we always have a known starting point for our databases. At least at any particular moment. We certainly need to update this foundation as we deploy changes to databases.

    Hopefully Microsoft, more vendors, and us as developers help advance these technologies, as well as help all developers build more skills. I’m grateful to Andrew Pruski, Anthony Nocentino, and others for the information they share about containers and databases. Hopefully we see more people engaging in these areas over time.

    Steve Jones

    Listen to the podcast at Libsyn, Stitcher or iTunes.

  • Microservices

    Recently I wrote about software trends from 2019 and how applicable they might be today. One of those was microservices, which I don’t see implemented that often with relational databases. I do see developers wanting to use them, but I don’t see them often with a RDBMS backend.

    I did see a piece on microservices, which mirrors that view. Lots of talk about the architecture, but not a lot of implementation. O’Reilly did a survey in Feb 2020, just before the world changed, and the found that only 10% reported complete success, but 54% felt mostly successful, and 92% some success. That’s better than I thought.

    The respondents are self-selecting, so it’s not surprising that there might be some success. It’s not surprising that a third of these people are moving over half of their legacy systems to this architecture. That doesn’t mean they’ll be successful in the move, or that they’ll continue.

    They also note that containers are important, and I think this might be more the nature that containers lend themselves to microservices. The companies I’ve know that heavily use containers tend to use microservices, almost as if the two technologies reinforce each other.

    I think culture plays an important part in any success, and if I were in an organization, I’d focus on this. Perhaps just like DevOps, with a small team proving some success and using that as evidence to start convincing others to move. However, I’d also need to have a good argument about why we should rewrite applications, or even bet heavily on a new architecture. I’ve yet to see a good reason why microservices work well with an RDBMS, but I’m open to someone convincing me.

    Steve Jones

    Listen to the podcast at Libsyn, Stitcher or iTunes.

  • Looking Back at Software Development Trends

    In some ways, the world of software development hasn’t changed much. The same sorts of skills and techniques I saw people using on COBOL programs hitting DB/2 and C++ over Oracle 6 are used today in React and C# against Azure SQL database. On the other hand, it does seem that we are more mature in how we work together and the flexibility with which we design systems.

    I saw the results from a survey from 2019 that Atlassian ran for software developers. This was a look at what modern trends might exist, though over a year later, perhaps the world is dramatically changed again. Let me look at each of the four trends they point out.

    I’ve been hearing about microservices for years, but have found relatively few customers using them. I don’t hear a lot about them in the RDBMS space, and I think this is because the idea of separating out each entity (or small set) in a database and having data access only through a front end component doesn’t make sense. There’s power in using an RDBMS to enforce data integrity rules and allow aggregations. I’ve also seen some companies looking for miniservices, not micro ones.

    Manual testing is still very prevalent for customers. While there are lots of unit tests, including some against a db, they don’t always extend through CI to more complex scenarios. Quite a few customers still have bottlenecks where humans look at the application in a larger sense. I think the high cost of tools that run more complex tests is a part of the problem. I’m not sure how we improve this, though I do hope to see more unit or functional tests for db code.

    Feature flags are extremely complex, though I do see them more as a defensive measure, where features are released, but then turned off if there are issues. This prevents a rollback from the app and db standpoint. I also don’t see a lot of use of dark deploys for database features, perhaps because until the app is working, we aren’t sure the data model is correct. Feature flag cleanup is certainly an issue for some clients.

    The last trend is one I rarely see implemented. Looking at customer outcomes, and not just immediate sales, is something that few people seem to do. Perhaps because developers are evaluated on the work they complete, not whether it’s in use. Managers are evaluated based on getting developers to do work, or on the sales that are produced, but it seems that few organizations try to measure the customer impact. We’ve started doing that at Redgate, and I’m interested to see how this evolves.

    Keeping developers motivated, excited about their jobs, and productive with creative solutions is tough. The trends listed from the article seem to me more aspirational for most of the organizations with which I deal. Most clients I know see developers as interchangeable parts, similar to factory workers. I think if they invested a little more in the well-being of developers, both from their mental focus and the growth of their skills, they might find a lot more benefits accruing from their software.

    Steve Jones

    Listen to the podcast at Libsyn, Stitcher or iTunes.