Showing posts with label capacity metrics. Show all posts
Showing posts with label capacity metrics. Show all posts

Friday, 1 July 2016

Linux Server – Disk Utilization (12 of 17) Capacity Management, Telling the Story

On Wednesday I said today I would share with you an example of a report on disk utilization of a Linux server. 

The report is illustrated below and the reason I chose to share this report is that it is an instance based report, displaying the top 5 disks and their utilization on this system.


You have the ability to pick out our top 5 or our bottom 5 to display to your audience because we don’t want too much ‘noise’ on our chart.


We want to keep things clear and concise, don’t flood reports with meaningless data and keep it relevant to our audience.

On Monday I'll be looking at how you can show correlation on reports, in the meantime don't forget to register for our next free webinar 'Capacity Planning & Forecasting using Analytic Modeling' 
http://www.metron-athene.com/services/webinars/index.html

Charles Johnson
Principal Consultant

Monday, 27 June 2016

Dashboard (10 of 17) Capacity Management, Telling the Story


Dashboard – Overview Scorecard

In the following example of a dashboard we can immediately see that we have a green, 2 reds and some greys. Based on the green, amber and red status we can immediately see that we have an issue with a couple of these categories, memory and I/O.

 
Is this enough information? Who is viewing this information and does it tell them enough? If management were looking at this information they would be worried as they can see red in the status. It does scare senior management when they see a red status, mainly due to the fact that they do not have the time or inclination to see what is behind the issue. They would immediately be on the phone to their capacity management team asking why there are issues and it then causes more pressure further down the tree.
It may be that this particular issue is not an immediate problem, maybe one of the thresholds was breached during a certain time period and needs investigation.
Dashboard – Overview Scorecard Detail
We can drill down and find out some further information on the issue in this case.
In the report below there is still some red showing so it is going to have to be investigated fully and we would need to drill down even further to find out what applications are involved here.
In the further drill down report below we can see that we have some paging activity on Unix that has breached the threshold.
 
These red, amber and green scorecards have to be based on thresholds.
Where the grey is shown this simply means that there is no threshold data attached to that.
We need to get in to the details to understand what the root cause of the issue is and understand whether the issue is serious or not.
On Wednesday I'll be taking a closer look at Unix reports. In the meantime why not take a few minutes and complete our online Capacity Management Maturity Survey to find out where you fall on the Maturity Scale and receive a 20 page report for free.
Charles Johnson
Principal Consultant

Wednesday, 8 June 2016

What is the Capacity Management Story? Capacity Management, Telling the Story (2 of 17)

So what is the Capacity Management Story?


•        What is happening in the current environment – this is typically called the baseline. When analyzing data in our systems we want to identify a ‘normal behaviour’ period which shows the demands an application is making on resources in usual circumstances.


•        Concise information – We need to state the facts and not over complicate matters, presenting clear and concise information.

•        Display forecasts – How do we present this story to our audience? In terms of capacity management we could be using forecasting methods such as:


•        Trends

•        Models

We need to describe to our audience so that they can understand easily the point/ message that we are trying to get across to them. We must also ensure that we are getting the right information to the right people in the right way.

•        Gather further information – Do we need to actually gather any additional information or do we have enough information? You may need to supplement resource data with some business data, or perhaps speak to the Service Delivery Managers to get the SLA information.

When we are forecasting it is important to have as much information as possible from the business because we need a full understanding of what we are forecasting to.

On Friday I’ll be talking about what we’re  trying to display.

Charles Johnson
Principal Consultant


Monday, 6 June 2016

Capacity Management,Telling the Story (1 of 17)

In Capacity Management we need to produce reports and make presentations, sometimes to our technical colleagues and sometimes to more senior people. 

In this blog series I’m going to be discussing the best ways to present technical information, your capacity management story, to all levels of audience.

To begin with what is a Story?
It is either:

a : an account of incidents or events, either fact or fictional

b : a statement regarding the facts pertinent to a situation in question

Data is nothing more than 1’S and 0’s. It’s what we do with it that makes it powerful.

We can collect or capture as much data as we want to from our systems, from our business level, from our service delivery level and store it but it’s what we do with this data that is the important thing.

Data used in a meaningful way is very powerful.

My series will look at ways in which you can tell your capacity management story to your audience in a meaningful way.

On Wednesday I'll begin with 'What is the Capacity Management story'?
In the meantime sign up and come along to one of our Capacity Management workshops.
http://www.metron-athene.com/services/online-workshops/index.html

Charles Johnson
Principal Consultant

Tuesday, 31 May 2016

The perils of the wrong type of aggregation


Looking at data and picking out the “story” it tells is often as much an art as a science when it comes to Capacity Management.  A gentle disbelieving of anything you are told also often makes for a better and quicker outcome than taking everything on face value.  Here’s a recent anecdote.

A customer raised an issue with Metron that after a software upgrade a database re-index job was taking a long time and that the application using the database had been working really slowly.  A hasty conference call / screen-sharing session was set up with us, the customer, his boss, and a SQL Server database administrator.  The conversation started with words to the effect of “this started after the upgrade, what’s going on?” - so we looked and we talked for a little while, then it came out that the database re-index job failed because it ran out of disk space.  The next comment “has the database got bigger because of the upgrade, then?”  That’s not our experience, but you never know….so we looked at a graph of the database size over time with a nice simple trend line over the top – the customer had already had this to hand. It looked a little like this:


With the disk size confirmed as 600 GB what was going on? This clearly shows housekeeping of the database as it grows, is shrunk down, grows again.  Even the trend line appears to be going down slightly.  The upgrade occurred in mid-April, so there was clearly no obvious jump up in database size at that time. 

The clue was the x-axis of the graph.  The dates were just the beginning of each month.  The chart seemed to have a nice neat shape to it – too neat, perhaps?
Looking further into the chart, the data was aggregated from the original 15 minute intervals up to the average for an entire day.  So what happens when we plot a chart of some of the later data points, showing each interval instead of the aggregated ones?

When did the re-index job start?…at 07:00 on May 1st.

That’s what polished off the remaining disk space.  The DBA killed the job and manually shrank the database at about 12:00.  During the time the disk had become full, and the application using this database detected this shortage of disk space and shut itself down to avoid losing data.  Only when it was restarted, which was some time after the space had been freed up, did it carry on.

Looking back at previous weekends the shape of the graph was the same each time - gently rising disk usage with a sharp, usually short increase during the time the re-index was taking place, dropping back to previous levels afterwards.

The previous weeks had survived, just, in those cases, so no-one had noticed how close the limit the disk space had become.  Now some more disk space has been added to cater for these weekly “spikes”, and the customer has a better handle on the growth rate of the database and the effect of the necessary but heavy weekly database maintenance.

Having the ability to aggregate large quantities of data into a simplified overview is useful for some things, but you do need to consider the “story” you are trying to tell with the resulting numbers.  Instead of an average of a large set of numbers that lowered the effective number, perhaps the better aggregation would have been something like “the peak value per day”, or “the aggregated hour containing the peak value”.  A straight average in this case hid the issue from sight, even with a trend line to try and predict how things were moving.

Metron’s athene® makes visualizing data simple, quick and intuitive and helps to keep your systems running by giving you that “over the horizon” view of what’s coming up, helping you run IT systems with no capacity surprises and having time to think about the best solution for the way ahead.
http://www.metron-athene.com/products/athene/index.html
Nick Varley
Chief Services Officer