Tag: sql server

  • T-SQL Tuesday #26 – Second Changes with Date/Time

    TSQL2sDay150x150

    I missed the very first T-SQL Tuesday, so when this month’s topic of second chances came up, I decided to write that one.

    If you are unsure of what T-SQL Tuesday is, follow the link to this month’s topic to get the rules and description and then write a blog post.

    Date/Time Challenges

    I picked an easy one, but one that I continue to see asked in the forums by people new to SQL Server. I suspect we’ll see less questions over time as more people take advantage of the new DATE and TIME datatypes in SQL Server 2008 and later, but maybe not. Lots of people are still sure that they need to keep those items together.

    In any case, have you ever seen sales data like this:

    OrderID     OrderDate               CustomerID  OrderAmount
    ———– ———————– ———– ————
    1           1982-05-19 06:31:48.950 1           579040.5070
    2           1994-11-27 17:14:41.790 2           348808.5860
    3           1972-11-08 17:40:01.170 3           758992.3650
    4           1972-05-31 01:19:05.530 4           779853.1990
    5           1994-12-22 10:40:57.410 5           666173.8040
    6           1974-04-03 01:42:29.490 6           134218.2330
    7           1976-06-22 15:21:18.910 7           322938.6950
    8           1953-08-05 23:00:34.620 8           14169.7580
    9           1971-08-16 22:33:28.970 9           586057.3820
    10          2002-03-28 13:08:00.420 10          632785.0760

    Here’s some DDL, you can create your own data, but here are a few rows:

    CREATE TABLE SalesOrders
    ( OrderID INT IDENTITY(1,1) , OrderDate DATETIME , CustomerID INT , OrderAmount NUMERIC(10, 4) ) INSERT SalesOrders (OrderDate, CustomerID, OrderAmount) VALUES ( '1982-05-19 06:31:48.950', 1, 579040.5070), ( '1994-11-27 17:14:41.790', 2, 348808.5860), ( '1972-11-08 17:40:01.170', 3, 758992.3650), ( '1972-05-31 01:19:05.530', 4, 779853.1990), ( '1994-12-22 10:40:57.410', 5, 666173.8040) 

    If I want to get all the sales in May of 1972, I get query them all like this:

    SELECT OrderDate
    , OrderAmount
     FROM SalesOrders
     WHERE OrderDate > '1972/5/1' AND OrderDate <= '1972/5/31' 

    I get 4 rows back. That’s something people often write when they get input from a user. A user has a start and end edit box, they enter “1972/5/1’” in the start box (or use a calendar picker) and then enter “1972/5/31” in the other. I’m using ISO dates to make this clear, though in the US it would normally display like “5/1/1972” and in the UK as “1/5/1972”.

    However, that isn’t quite correct. If I run this query:

    SELECT OrderDate
    , OrderAmount
     FROM SalesOrders
     WHERE MONTH(OrderDate) = 5
      AND YEAR(OrderDate) = 1972

    I actually get 6 rows. The data rows for May 1972 are:

    OrderDate               OrderAmount

    ———————– —————————-

    1972-05-31 01:19:05.530 779853.1990

    1972-05-13 09:26:52.590 676848.9700

    1972-05-28 21:05:28.840 923425.0510

    1972-05-07 10:59:09.930 266079.6480

    1972-05-01 04:21:01.250 464241.6480

    1972-05-31 01:19:05.530 779853.1990

    What’s happening?

    If you look at the OrderDates for May 31, you see two values that have a time of 1:19:05am. Those are excluded from the query, which has an end date of “1972/5/31”. Why? This query:

    SELECT CAST( '1972/5/31' AS DATETIME) 

    shows why. It returns:

    1972-05-31 00:00:00.000

    That’s midnight between the 30th and 31st, which is before 1:19:05am. When a datetime value is converted in SQL Server, without a time component, it defaults to the beginning of the day. That works great for the start date, not so good for the end date.

    Fixing this

    There are two real fixes here. Well, maybe more. You can query on the month and year, but those functions can disrupt indexes, so I don’t recommend them. The two main fixes are:

    • add a time component
    • add a day

    The first fix is the addition of the last time of the day to your query.

    SELECT OrderDate
    , OrderAmount
     FROM SalesOrders
     WHERE OrderDate > '1972/5/1' AND OrderDate <= '1972/5/31 23:59:59.997PM' 

    In some type of code, it looks more like this:

    DECLARE @end DATETIME SELECT @end = '1972/5/31' SELECT @end = @end + '23:59:59.997' SELECT OrderDate
    , OrderAmount
     FROM SalesOrders
     WHERE OrderDate > '1972/5/1' AND OrderDate <= @end

    Assume the first select is actually coming from the user.

    The second fix is to add a day, and actually query like this, which is what I’d recommend:

     SELECT OrderDate
    , OrderAmount
     FROM SalesOrders
     WHERE OrderDate > '1972/5/1' AND OrderDate < '1972/6/1' 

    In this case, instead of querying for May 31, we move to June 1 (the next day) and then change the <= to a < only. This gets all orders occurring up until the end of May 31, but before June 1.

    Easy fixes, but so often there’s code that doesn’t allow for the time component. Take a minute and check your reports, and be sure you aren’t underreporting any data to your clients or customers.

    Also be sure to check out the new datatypes in SQL Server 2008 and later:

  • Referencing Remote Data

    Maybe if I searched more, I'd use synonyms

    One of the features added to SQL Server a few years ago were the ability to create synonyms and use those to reference other objects. The ability to create synonyms is something that I had wanted for years in SQL Server, but when they were released, I found them to be a tool that I rarely reached for. Whether it was because this was something I rarely needed to accomplish, or because my habits were too ingrained, I’m not sure, but I have only created synonyms for testing purposes and not for use in any production databases.

    When a developer needs to reference data in another database (or on another server), they have a variety of ways in which they can do this. Some people prefer a three or four part naming convention, others use a view local to the database, and still others might use synonyms. While they all work, from a maintenance standpoint, I think a view or synonym provide a nice layer of abstraction while minimizing the potential maintenance headaches of future changes.

    If you find synonyms more useful than local views, I would be interested in knowing why. They seem to almost operate in the same way to me, but for some reason I find views to be easier to track and manage. Perhaps it’s just an ingrained habit from years of making do with views, or maybe it’s my habit of browsing for objects, instead of using a tool like SQL Search.

    Whichever method you use, I do urge you to always consider a layer of abstraction. That other database you reference might be located on the same instance today, but in the future it might grow and require it’s own instance. If you have three part naming buried in all of your stored procedure or application code, it might not be as simple as a global search and replace to make changes. If it’s not, then you are wasting development time down the road by not having a layer of abstraction implemented at the beginning.

    Steve Jones

    Steve Jones


    The Voice of the DBA Podcasts

  • Capacity planning for new hardware

    I get asked this question a lot: When getting a new application and database, what kind of hardware do I need?

    The answer is easy: it depends.

    It depends on

    • the load you will put on the SQL Server instance.
    • on the SLA and performance numbers you need.
    • the uptime you need to maintain (RPO/RTO factor in here)

    There are other factors, but essentially the server needs enough hardware to handle the workload in the time you need it to handle things.

    It doesn’t depend on

    • the number of users
    • the number of databases
    • the size of the databases*

    A slight asterisk on the last one. The size of the databases matter for the space you need to buy, but they don’t necessarily affect the RAM or CPU you need. The number of users and databases can contribute to load, but those numbers of a vacuum don’t affect the hardware. I have a test machine with dozens of databases, and it generates no load. Why not? No workload, or not much of one.

    The same thing applies to users. If the users do a lot of work, making changes, querying large data sets, it can be a loaded database. However I had a database one time that had 5,000 clients. It was updated by agent software on desktops in our company every hour. However each update was a singleton update to a specific row, and it was comfortably hosted on a 2CPU, 2GBRAM Standard SQL Server 2000 instance.

    How do I plan?

    Photo Jan 03, 8 50 26 AM

    This is a tough question in many ways. I have searched all over, and asked questions, and even written a SQL University week on it (overview, disk, other). There are two main scenarios here, but they both get similar treatment.

    You have to test your workload against hardware and see what happens. It can be a simulated one, or if you have an existing instance being upgraded, take a real workload.

    There are numerous replay tools or testing tools, but the bottom line is you have to simulate the way you use SQL Server on actual hardware. You can do some extrapolation, but don’t expect it to be a linear change.

    For example, if I have a dual core CPU of xx type, with 1GB of RAM and a 2 drive R1 array, I can’t assume I’ll get twice as much performance with a quad core of the same type with 2GB of RAM and 2 R1 drives. It almost becomes an art to examine the memory usage, the IOPS you are generating, and the percentage of CPU. As you scale up, the usage of those three values might changes and shift. More RAM can reduce IOPS and CPU, especially in reads, but not necessarily. If you have a 1TB table you’re constantly summarizing, you might not get much better performance going from 1 -> 2GB.

    Ultimately it’s a bit of a guess for most of us. Fortunately most of our workloads can be handled by hardware since most databases are relatively small. Test, make some guesses, and go a little big, especially on RAM and CPU. Those are the hardest to add later, from my experience. It seems people expect disks to fill up and we need more space, but it’s often harder to approve CPU/memory upgrades, especially if it requires newer motherboards.

    I wish I had a better answer, and there are numerous articles that don’t seem to me to do a better job, but read as much as you can, and try to learn how others view their systems. Then make your own guesses about what is best for your system.

  • Read-only Data

    We keep gathering, storing, and managing more and more data. Many of our systems could use an archiving plan to migrate older data to another database or system where it can be accessed, but it won’t impact the performance of queries against our current data. If you don’t have any type of archive plan, you might consider building one for any future tables you design.

    Do you manage read only data differently?

    Once you migrate data to a new set of tables, typically you would consider that older data to be read only, and potentially mark it as such in it’s own storage location, perhaps even adding more indexes than you have on the current data. And if the data is static, then it doesn’t change from week to week, and you can reduce the amount of backups that you create from this data.

    However you can’t eliminate backups. There is still the possibility that you might have a disaster situation and need to recover the data. If your last backup of the read only database or file group is 18 months old, will you be able to find it? That can be quite a challenge, and I wanted to get some opinions this Friday about how to handle this situation. The poll this week is:

    How often do you back up read only data?

    I am also curious how you manage tracking these infrequent backups and recovering them if you must perform a restore. I’ll admit that I don’t have a great solution other than scheduling some regular backup interval, something like once a quarter and documenting the location someplace that would be accessible in a disaster.

    Steve Jones


    The Voice of the DBA Podcasts