Showing posts with label modeling. Show all posts
Showing posts with label modeling. Show all posts

Thursday, 14 July 2016

VMware, Virtual Center Headroom (17 of 17) Capacity Management, Telling the Story

Today I’ll show you one final report on VMware, which looks at headroom available in the Virtual Center.

In the example below we’re showing CPU usage. The average CPU usage is illustrated by the green bars, the light blue represents the amount of CPU available across this particular host and the dark blue line is the total CPU power available.
 
VMware – Virtual Center Headroom
 
 
 
We have aggregated all the hosts up within the cluster to see this information.
We can see from the green area at the bottom how much headroom we have to the blue line at the top, although actually in this case we will be comparing it to the turquoise area as this is the amount of CPU available for the VM’s.
This is due to the headroom taken by VMkernel which has to be taken in to consideration and explains the difference between the dark blue line and the turquoise area.
 
Summary

To summarize my blog series, when reporting:

•        Stick to the facts

•        Elevator talk

•        Show as much information as needs to be shown

•        Display the information appropriate for the audience

•        Talk the language for the audience

….Tell the Story
Hope you've enjoyed the series, if you have any questions feel free to ask. If you're interested in VMware Capacity Management don't forget to book on to our workshop http://www.metron-athene.com/services/online-workshops/index.html#vmwarevsphere
Charles Johnson
Principal Consultant


Friday, 8 July 2016

Model – Linux server change & disk change (15 of 17) Capacity Management, Telling the Story

Following on from Wednesday's blog today I'll show the model for change in our hardware.

In the top left hand corner we are showing that once we reach the ‘pain’ point and then make a hardware upgrade the CPU utilization drops back to within acceptable boundaries for the period going forward.




In the bottom left hand corner you can see from the primary results analysis that the upgrade would mean that the distribution of work is more evenly spread now.

The model in the top right hand corner has bought up an issue on device utilization with another disk so we would have to factor in an I/O change and see what the results of that would be and so on.

In the bottom right hand corner we can see that the service level has been fine for a couple of periods and then it is in trouble again, caused by the I/O issue.

Whilst this hardware upgrade would satisfy our CPU bottleneck it would not rectify the issue with I/O, so we would also need to upgrade our disks.

When forecasting modeling helps you to make recommendations on changes that will be required and when they will need to be implemented.

On Monday I'll take a look at some examples of VMware reports.

In the meantime why not register for our next webinar Capacity Planning and Forecasting using Analytic Modeling http://www.metron-athene.com/services/webinars/index.html

Charles Johnson
Principal Consultant

Wednesday, 6 July 2016

Modeling Scenario (14 of 17) Capacity Management, Telling the Story

I have talked about bringing your KPI’s, resource and business data in to a CMIS and about using that data to produce reports in a clear, concise and understandable way.

Let’s now take a look at some analytical modeling examples, based on forecasts which were given to us by the business.

Below is an example of an Oracle box, we have been told by the business that we are going to grow at a steady rate of 10% per month for the next 12 months. We can model to see what the impact of that business growth will be on our Oracle system.

In the top left hand corner is our projected CPU utilization and on the far left of that graph is our baseline. You can see that over a few months we begin to go through our alarms and our thresholds pretty quickly.

Model – oracleq000 10% growth – server change



In the bottom left hand corner we can see where bottlenecks will be reached indicated by the large red bars which indicate CPU queuing.

On the top right graph we can see our projected device utilization for our busiest disk and we can see that within 4 to 5 months it is also breaching our alarms and thresholds.

Collectively these models are telling us that we are going to run in to problem with CPU and I/O.

In the bottom right hand graph is our projected relative service level for this application. In this example we started the baseline off at 1 second, this is key.

By normalizing the baseline at 1 second it is very easy for your audience to see the effect that these changes are likely to have. In this case, once we’ve added the extra workload we can see that we go from 1 second to 1.5 seconds (a 50% increase) and then jumped from 1 second to almost 5 seconds. From 1 to 5 seconds is a huge increase and one that your audience can immediately grasp and understand the impact of.

We would next want to show the model for change in our hardware and I'll be looking at this on Friday.

In the meantime why not join our Community and get access to a wealth of Capacity Management Resources http://www.metron-athene.com/_resources/

Charles Johnson
Principal Consultant

Friday, 1 July 2016

Linux Server – Disk Utilization (12 of 17) Capacity Management, Telling the Story

On Wednesday I said today I would share with you an example of a report on disk utilization of a Linux server. 

The report is illustrated below and the reason I chose to share this report is that it is an instance based report, displaying the top 5 disks and their utilization on this system.


You have the ability to pick out our top 5 or our bottom 5 to display to your audience because we don’t want too much ‘noise’ on our chart.


We want to keep things clear and concise, don’t flood reports with meaningless data and keep it relevant to our audience.

On Monday I'll be looking at how you can show correlation on reports, in the meantime don't forget to register for our next free webinar 'Capacity Planning & Forecasting using Analytic Modeling' 
http://www.metron-athene.com/services/webinars/index.html

Charles Johnson
Principal Consultant

Monday, 20 June 2016

Display of different presentation types (7 of 17) Capacity Management, Telling the Story

As discussed today I'll be looking at the types of presentations that you can use.

Below is a selection of them:


Humans like visual representation so using these charts in the right way and gauging which are right to represent the information to your audience is crucial.

Dashboard - More aligned to presenting real time information. The key thing to remember is that any dashboard you use should auto update.


Analysis - Presents the drill down of a problem. This is the root cause analysis, where we know there is a problem and we want to drill down and show what is causing the issue. Where’s the bottleneck? Was there a change?

Advice - Provide some automatic advice, automatic interpretation of the data that you are reporting on.

Virtualization - Report on virtualization data, make it easy to understand what  is happening in your virtual environment.

Business - We have discussed about bringing in business data to the CMIS. Why do we want to do that? We can look for correlations by measuring component data against business metrics and show these in our business reports.

Trending - We can show ‘what-if’ trending which can give you a ‘time to live’ value.

Modeling - More accurate prediction reports to show are modeling reports. You can show things like future system response times or identify where any future bottlenecks are likely to occur.

Breakdown - Shows further analysis on the data in an easy to understand way.

Don't forget the key is to always remember to tailor your presentations to suit your audience.

On Wednesday I'll be dealing with what we will see at the management level. 

There's a great webinar coming soon 'Capacity Planning & Forecasting using analytical modeling', don't miss it!
http://www.metron-athene.com/services/webinars/index.html

Charles Johnson
Principal Consultant



Friday, 10 June 2016

What are we attempting to display? ( 3 of 17) Capacity Management, Telling the Story

When we’re telling our capacity management story what are we trying to display?

•         Display current state of the environment – this is our baseline, a period or periods which reflect our normal pattern of behavior. You need an understanding of the business to class a period as ‘normal’ behavior.

•         Display possible anomalies – this involves looking at periods and determining whether the peaks in a set period of time are an anomaly or whether they are just normal user busy periods of the day, explaining peaks in the data. Has there been some kind of change made that you are aware of? You may need to perform some root cause analysis to get to the bottom of peaks where you don’t already have an understanding of what is causing them.

•         Display forecasts

•     Trends – are very useful. You can trend to a limit, you can trend to a threshold etc but if you wish to get a better degree of accuracy to your forecasting you may wish to model.

•         Models – analytical modeling can provide you with a better degree of accuracy when  forecasting, especially when it comes to things like utilizations where you have to take in to consideration the utilization curve.

Are there different stories for different audiences? I’ll be dealing with this on Monday.

In the meantime register for our community and get access to free capacity management white papers, on-demand capacity management webinars and more

 
Charles Johnson
Principal Consultant 

Monday, 17 August 2015

Why do we model in the UK, monitor in Japan, and manage in the USA?

Many times we look at Capacity Management from our own perspective and the culture in which we live.

I'll be presenting at a webinar which looks at the current and future practice of capacity management and considers why it might be different in different countries.

I have experience in the EU, USA and Japan and will concentrate on these, with particular emphasis on Japan as it is possibly the least known of the three.

My webinar will cover the following:
  • Introduction
    • Capacity Management
    • ITSM, ‘do more with less’
  • Objectives of ITSM
    • Find a better way for IT Service Management
    • Assess impact of capacity management
  • Capacity Management review
    • Comparison between USA, EU/UK and Japan
    • What are the differences and why?
  • How to regain customer trust?
    • Case study
  • Who is doing it and why/why not?
  • Summary & Conclusion
It's taking place this Wednesday August 19, 2015 (8am PST, 9am MST, 10am CST, 11am EST, 4pm UK, 5pm CEST) - register for your place now
http://www.metron-athene.com/services/webinars/index.html

Charles Johnson
Principal Consultant

Monday, 27 July 2015

Forecasting, when modeling is not the only choice - Forecasting Techniques (3 of 5)

In my opinion there are four main techniques that are used in forecasting: analytical modeling, simulation modeling, trending and headroom charts.

Let’s take a look at each of these:
Analytical Modeling                                        

•        The ability to predict environment changes based upon the arrival of work
·      CPU
·      IO
•        Based upon
·      Baseline timeframe (seasonal peaks and troughs, highest transaction times etc)
·      Calibration of model – Does it match real life?
·      What-If changes to model
•        Spend virtual $’s, helping you to determine what is the best cost benefit that you can get out of it
•        Low cost modeling technique, as you’re able to take the data and apply it going forward without making costly mistakes.
The graph below illustrates how this can be considered, always taking in to account how the ‘end user’ will be affected.
Simulation Modeling
There are many organizations such as Wall Street companies, who feel the need to run simulation models but it does take longer to set up and can be very costly.
•        Replication of environment
•        Has to stay close to production environment
•        Usage of tools to replicate workload into simulated environment
•        Longer lead time to setup
•        Provides granular detail on environment changes
•        Can be costly
Trending
With trending you are usually looking at one metric for each chart, this doesn’t prohibit you from having several charts looking at different metrics for the same scenario.
•        Provide prediction based upon an interval of time
•        Quick to produce
•        Usually looking at one metric
•        Can show steps in growth changes
•        Confidence is based on length of interval and amount of historical data
The more data that you have to feed in to your trend chart then the more confidence you can have in the end result.
On Wednesday I'll be taking a more in-depth look at Trend types.
Charles Johnson
Principal Consultant

Monday, 20 April 2015

Automatic reporting and alerting(10 of 10)

What do we actually need? What do we actually want to report on?  How often do we want to report on it? 

If we are applying threshold based alerting to our reports, we need to ensure that the correct values are set.  These values may be the utilizations or response times stated within an SLA or based on maximum resource usage.  Failure to set the correct values may lead to incorrect alerts being produced, leading to unnecessary investigations, stress and panic.  By including availability and response time information within your capacity reports, you improve both the accuracy and increase confidence in your forecasts whilst providing potential SLA breach information in advance. 

When creating forecast models, whether trend or analytical models or both, it is important to make sure that the inputs into your model are as accurate as possible so we can make these predictions to avoid any costly or unnecessary performance issues or SLA breaches.

So let's ensure we have a Brighter Outlook.  It is crucial that we get the information at all levels as described and store this information typically in a centralised database, so that we have it readily available.  Production of adhoc reports on infrastructure usage and current performance of our applications along with implementing automated reports specifically based on our SLA thresholds, enables us to produce early warning alerts on potential breaches and take appropriate action as necessary.  

Guest and Host consolidation.  If you have over provisioned systems, look at the usage of your virtual machines against the configured resources to see if there is scope to consolidate your guests onto smaller numbers of ESX hosts.   You may also be able to have multiple applications running within the same VM rather than have many VMs running a single application. 

Plan ahead.  By producing trend reports, producing analytical models and predicting what impact is likely to happen to your infrastructure and application performance running within it.  Then make the necessary recommendations on upgrades or configuration changes that prevent you encountering any SLA breaches and associated impacts on services.  All of this information should be included within a Service Capacity Plan, allowing us to make decisions on whether we need to upgrade, whether we need to standardise our hardware and what the associated costs are likely to be, so budgets can be accurately planned.

It can also help us decide on whether a more powerful and expensive server is actually required when maybe a less expensive, slightly less powerful server, will do just as good a job.   Creating analytical models gives you the information you require to make those decisions.

Regular consultation and information sharing with Application teams and other Service Delivery teams will assist you in making the best decisions going forward and allows you to explain why you have made the stated recommendations.

I'll leave you with a quick look at monetary savings on Capital Expenditure (CAPEX) and Operational Expenditure (OPEX)

CAPEX

All the way through this series I have mentioned being able to possibly reduce the numbers of servers required to host all of your virtual machines.  This enables you to possibly make savings on the number of licenses you require and on the actual hardware costs.  As you start to reduce your physical estate and consolidate ESX hosts, you start to look at the possibility of reducing the size of your datacenter.

OPEX 

Make savings by reducing the amount spent on maintenance and support as you reduce the numbers of servers required in your infrastructure hosting your applications and services.  By performing application sizing, we can assist in accurately provisioning resource requirements and help eliminate any potential overspend by over provisioning.  Further to this, we can actually reduce the physical server count leading to a reduction in the size of a datacenter. 

The savings from this approach such as Power Usage -  servers & cooling / lighting, reduction in emissions through reduced power consumption but also through  a reduction in components as we consolidate servers and finally through usage, by optimally sizing and provisioning.

In my series I‘ve covered what Cloud computing is and how it is underpinned by Virtualisation, the benefits it can provide, things we should be aware of and how putting in place effective Capacity Management can save you time and money. 

If you'd like further information on Capacity Management there are a selection of papers available to download http://www.metron-athene.com/_downloads/index.html and don't forget to register for my webinar 'Understanding VMware Capacity' http://www.metron-athene.com/services/training/webinars/index.html

Jamie Baker
Principal Consultant