Collecting Data is Hard

·

Data is the lifeblood of much of the world today. Not necessarily big data, and certainly not perfect data, and definitely not just digital data. Organizations, individuals, governments, really everyone out there are making decisions based on data. You might think it’s going to rain, so you cut the grass today, or maybe defer adding fertilizer. Your organization sees demand for a product increase, so it orders more and produces more. Government is always using data to make decisions about resource allocation. We might not think governments make great decisions, but they do use data and data matters.

Recently I was reading a science fiction book (I, Starship) about the future, where a person’s brain (Henry) becomes uploaded to manage a starship. This ship will travel light years away for 80 years and they need the human crew asleep in hibernation to survive the journey. The interesting thing, to me, was a part in the book where there is a discussion of why Henry was uploaded and why AIs aren’t advanced enough to run the starship. There’s this quote: “The first generations (of LLMs) performed well, but as time went on, we entered a situation where more and more of the data available to train them on was itself machine-generated. So, instead of mimicking high-quality human output, the outputs got more garbled.”

I worry about this as the current models are sucking up so much data to learn, but so little of the new data is being generated by humans. We already see plenty of AI-slop on the Internet, with fewer and fewer articles, blogs, etc. being human generated. I’m sad because people don’t share as much of their own thoughts, knowledge, etc. This is especially true in light of the AI companies taking individuals’ work for training without compensation. Indeed, I worry that many places will go the route of Stack Overflow, where they essentially fail.

I’m worried about that here at SQL Server Central, as I see less questions being asked by humans and less discussion about the nuances of database challenges.

However, there are AIs out there also polluting the world. This was a piece from last year that more AIs are taking surveys and polls. Reddit is seeing questions being asked, and I’m sure there are AIs answering them. How long before the amount of AI generated traffic dwarfs human generated traffic? I mean new data, not consumption. I expect plenty of humans are going the way of the people in WALL-E and just consuming data. They’ll continue to watch untold numbers of reels, shorts, Tik-Toks, etc.

Are we going to see less “real” data and more generated data? I already have seen no shortage of issues from customers trying to use synthetic data for testing. It doesn’t match the real world well, but if they stop getting real data from customers and more from other bots, maybe it won’t matter. Of course, I’m not sure how well their systems will perform in the real world.

GIGO is a real issue, and I expect a lot of companies will learn this as the volume of AI-generated data increases.

Steve Jones

Listen to the podcast at Libsyn, Spotify, or iTunes.

Note, podcasts are only available for a limited time online.

Comments

Leave a comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.