Showing posts with label VMware ready time. Show all posts
Showing posts with label VMware ready time. Show all posts

Monday, 10 July 2017

Understanding VMware Capacity - Ready Time (4 of 10)

Imagine you are driving a car, and you are stationary. There could be several reasons for this. You may be waiting to pick up someone, you may have stopped to take a phone call, or it might be that you have stopped at a red light. The 1st two of these (pick up, phone), you decided to stop the car to perform a task. But in the 3rd case, the red light is stopping you doing something you want to do. You spend the whole time at the red light ready to move away as soon as you get a green light. That time you spend waiting at a red light  is ready time. When a VM wants to use the processor, but is stopped from doing so it accumulates ready time. This has a direct impact on the performance of the VM.

Ready Time can be accumulated even if there are spare CPU MHz available. For any processing to happen all the vCPUs assigned to the VM must be running at the same time. This means if you have a 4 vCPU VM, all 4 vCPUs need available cores or hyperthreads to run. So the fewer vCPUs a VM has, the more likely it is to be able to get onto the processors. You can reduce contention by having as few vCPUs as possible in each VM. And if you monitor CPU Threads, vCPUs and Ready Time for the whole Cluster, then you’ll be able to see if there is a correlation between increasing vCPU numbers and Ready Time.


Here is a chart showing data collected for a VM. In each hour the VM is doing ~500 seconds of processing. The VM has 4 vCPUs. 
Despite just doing 500 seconds of processing, the ready time accumulated is between 
~1200 and ~1500 seconds. So anything being processed spends 3 times as long waiting to be processed, as it does actually being processed. i.e. 1 second of processing could take 4 seconds to complete.


Now lets look at a VM on the same host, doing the same processing on the same day. Again we can see ~500 seconds of processing in each hour interval. But this time we only have 2vCPUs. The ready time is about ~150 seconds. i.e. 1 second of processing takes 1.3 seconds.
By reducing the number of vCPUs in the first VM, we could improve transaction times to somewhere between a quarter and a third of their current time.

Here’s a short video to show the effect of what is happening inside the host to schedule the physical CPUs/cores to the vCPUs of the VMs. Clearly most hosts have more than 4 consecutive threads that can be processed. But let’s keep this simple to follow.




1.      VMs that are “ready” are moved onto the Threads.

2.      There is not enough space for all the vCPUs in all the VMs. So some are left behind. (CPU Utilization = 75%, capacity used = 100%)

3.      If a single vCPU VM finishes processing, the spare Threads can now be used to process a 2 vCPU vm. (CPU Utilization = 100%)

4.      A 4 vCPU VM needs to process.

5.      Even if the 2 single vCPU VMs finish processing, the 4 vCPU VM cannot use the CPU available.

6.      And while it’s accumulating Ready Time, other single vCPU VMs are able to take advantage of the available Threads

7.      Even if we end up in a situation where only a single vCPU is being used, the 4 vCPU VM cannot do any processing. (CPU utilization = 25%)



As mentioned when we discussed time slicing, improvements have been made in the area  of co-scheduling with each release of VMware. Amongst other things the time between individual CPUs being scheduled onto the physical CPUs has increased, allowing for greater flexibility in scheduling VMs with large number of vCPUs.
Acceptable performance is seen from larger VMs.

Along with Ready Time, there is also a Co-Stop metric. Ready Time can be accumulated against any VM. Co-Stop is specific to VMs with 2 or more vCPUs and relates to the time “stopped” due to Co-Scheduling contention. E.g. One or more vCPUs has been allocated a physical CPU, but we are stopped waiting on other vCPUs to be scheduled.

Imagine the bottom of a “ready” VM displayed, sliding across to a thread and the top sliding across as other VMs move off the Threads. So the VM is no longer rigid it’s more of an elastic band.
Phil Bell
Consultant


Friday, 7 July 2017

Understanding VMware Capacity - 5 key VMware metrics (3 of 10)

There are plenty of metrics provided by VMware. Some of these are familiar to anybody who’s ever monitored a computer (CPU%, IOPS etc.). So for a short list of 5 I’ve selected 5 metrics that are important and available in VMware environments.

To some extent I’ve cheated to make a list of 5, because for some I’m looking to get the same metric from different levels in the environment. Most of these I will cover in greater detail in the sections following this.

CPU MHz
Why not %? Well because we can’t compare the CPU% Used of a VM with the % of the Host or the Cluster. We want to judge the amount of processing power an individual VM requires. That way when we want to move it to another cluster, or are considering the size of the existing cluster, we have a comparable value to work with. I fully admit it’s not perfect. The processing you can do with 1MHz today is greater than that of 10 years ago. But it’s still better than %. 

Ready Time
It’s a measure of CPU contention. The bigger the number the less happy your users are. But it doesn’t always mean you are short of CPU power.

Active Memory
The amount of memory the VM has accessed in the last 20 seconds. If there isn’t space for all the active memory in the cluster, then performance will be sub optimal.

Ballooned Memory
Memory ballooning works by “robbing Peter to pay Paul”. When memory is in short supply and/or over committed, pages of RAM in a VM OS can be filled by a “balloon”. These pages are then not actually available to the OS, and the space freed up and will be used by other VMs that need the RAM resources. As memory demand goes down, VMware deflates the VMs balloon and the RAM is again available to the VM. Again this is a sign of contention for resource and introduces additional overhead to the VMware hypervisor.

Host Disk Latency
Among the metrics that are provided for Hosts are the Disk Latency metrics. Disk IO performance is still the greatest performance challenge faced by most organizations I meet.

In my next blog on Monday I'll take a closer look at Ready Time.

Phil Bell
Consultant