IT capacity management in ITIL 4 is handled by the capacity and performance management practice, whose purpose is to ensure that services achieve the agreed and expected levels of performance and satisfy current and future demand in a cost-effective way. It is one of the 17 service management practices among ITIL 4’s 34, which also include 14 general management and 3 technical management practices.
How much cloud and infrastructure capacity is wasted?
Flexera’s 2026 State of the Cloud report puts wasted cloud spend at 29%, the first increase in five years. In the same survey, 85% of organizations name managing cloud spend as a top challenge.
Overprovisioning used to surface as a procurement problem at hardware refresh time; with cloud billing, it shows up on an itemized invoice every month.
Under-provisioning is the opposite failure, and it is expensive too. The Uptime Institute’s 2026 Annual Outage Analysis reports that 57% of respondents said their most recent major outage cost more than $100,000, and for the second consecutive year 1 in 5 reported costs exceeding $1 million. Uptime does not break those costs down by cause, so the figures show what an outage costs, and capacity saturation is one of several causes that can lead to one.
Capacity failure now shows up as an unexpected invoice as well as an outage, and ITIL 4 gives one practice responsibility for both. Monitoring tools usually catch the outage, while the overspend often goes unnoticed until the bill arrives. Among the 34 ITIL 4 practices, capacity and performance management is the one that covers both risks.
What is the difference between capacity management and performance management?
Capacity management checks whether there is enough of a resource to meet demand; performance management checks whether the service responds fast enough.
ITIL 4 pairs them inside a single practice so that neither question gets answered alone, because adding capacity is both the most common way to buy performance and the most common way to waste money. The purpose statement above carries both halves deliberately, and cost-effectiveness constrains how you satisfy either.
In ITIL v3 the two were separate, and Capacity Management was a standalone process inside the Service Design lifecycle stage, with three sub-processes covering business, service and component capacity, plus Capacity Management Reporting. ITIL 4 renamed the discipline capacity and performance management and recast it as a service management practice. The analysis work stayed recognizably the same, and the emphasis moved from sizing infrastructure to sustaining an agreed outcome.
A common error about this practice is that ITIL 4 merged capacity with availability, when it actually merged capacity with performance. Availability management remains a separate ITIL 4 practice that asks whether the service is usable when it is supposed to be.
A service can be entirely available and unusably slow.
What are the three levels of capacity management?
The three levels are business, service and component capacity.
The model comes from ITIL v3, where each level was a formal sub-process; in ITIL 4 the three levels remain useful as perspectives. They are still the quickest way to see which parts of the practice an organization actually runs.
The three levels of capacity management
| Level | What it asks | Typical owner | What it consumes |
|---|---|---|---|
| Business capacity | What will the business do next, and how much load will that create? | Service owner with business stakeholders | Business plans, seasonality, project pipeline |
| Service capacity | Can this service hold its agreed targets at that load, end to end? | Service owner | Service model, transaction volumes, response times |
| Component capacity | Is this server, database, link or instance family running out of headroom? | Platform, database and network teams | Utilization telemetry, thresholds, vendor limits |
In many organizations the component level is well covered and the business level barely exists. An infrastructure team can then say a disk is 80% full but cannot say which business service is closest to breaching its target. Planning stalls when someone has to translate transaction growth into cores and gigabytes, because that translation needs a maintained map from each business service down to its components, which Matrix42 builds from discovery and CMDB data.
Is capacity management still relevant with cloud autoscaling?
Yes. With autoscaling, the decision moves from how much capacity to buy toward where the limits sit and what they cost when they are reached.
A practitioner analysis published by Proscaler makes the operational case; it is informed opinion and does not include measured data. Autoscaling is bounded by spin-up delay, by cascading infrastructure limits such as Classless Inter-Domain Routing (CIDR) blocks, shared gateways and persistent storage, and by application design. A scaling policy covers none of these limits, so they tend to surface the first time a scale-out event stalls on something nobody modeled.
Physical limits are also bringing the practice back into focus. IDC put artificial intelligence (AI) infrastructure spending at $89.7 billion in the first quarter of 2026, up 33.1% year over year, and raised its full-year 2026 forecast to $497 billion. The Uptime Institute’s 2026 global data center survey, covering more than 800 owners and operators, reports a growing number of operators at peak rack densities of 30 kW or higher, which makes power, cooling and floor space hard constraints again.
What is the difference between capacity management and demand management?
Demand management shapes and influences demand; capacity and performance management sizes for it.
Demand management changes the shape of the load curve through scheduling, policy, throttling or pricing, while capacity and performance management treats the curve as given and makes sure the estate can carry it.
Capacity and performance management and its adjacent practices
| Practice | The question it answers | What it produces | How it relates |
|---|---|---|---|
| Capacity and performance management | Is there enough, and is it fast enough? | Capacity plan, sizing decisions, performance analysis | Consumes demand patterns and utilization data |
| Demand management | What drives the load, and can it be shaped? | Demand patterns, user profiles, influencing measures | Supplies the input this practice sizes against |
| Availability management | Is the service usable when it is needed? | Availability targets, designs, recovery measures | Shares telemetry with this practice |
| Monitoring and event management | What is the current state, and what changed? | Events, thresholds, alerts, utilization data | Supplies the data this practice analyzes |
Most disagreements arise at the handoff between these practices. Monitoring and event management owns whether a threshold fires; capacity and performance management owns whether the threshold was set at the right number. IT operations management (ITOM) tools blur the two when their default thresholds are never reviewed. In a subscription estate, reshaping a peak through demand management is usually cheaper than buying capacity for it.
What is a capacity plan in ITIL?
A capacity plan records current utilization, forecast demand, and the resource decisions that follow from the gap between them, with dates and costs attached.
In practice it holds usage per service, expected growth, the performance targets being defended, planned increases or reductions, and the assumptions behind every forecast in it.
Capacity plans usually go stale because the service model behind them is out of date, and extra effort does not fix that. Without a maintained map of which components carry which service, the plan documents server utilization and says little about the capacity of services. Service configuration management is the practice that maintains that map, and the configuration management database (CMDB) is the tool that stores it.
The plan also depends on finance, since a capacity plan that names no budget line is a wish list with graphs. Service financial management turns forecast headroom into approved spend, or into a documented decision to accept the risk of not buying it. It is also what makes the cost-effectiveness half of the purpose measurable. Matrix42 holds configuration and asset records in one platform, so a forecast can name the component, the service it carries and the license that limits it.
What are the KPIs for capacity and performance management?
The most widely used key performance indicators (KPIs) for capacity and performance management are seven metrics published by IT Process Maps: Incidents due to Capacity Shortages, Exactness of Capacity Forecast, Capacity Adjustments, Unplanned Capacity Adjustments, Resolution Time of Capacity Shortage, Capacity Reserves and Percentage of Capacity Monitoring. ITIL 4 has no official public KPI list.
The AXELOS practice guide for this practice sits behind PeopleCert MyAxelos membership, so every KPI set on the open web is a secondary-source proposal. The seven-metric list above was last edited on 17 June 2019 and follows the ITIL v3 (2011) process model. It is a useful starting point, but it is not authoritative, and reports should say so when a steering committee asks where the numbers came from.
The seven metrics fall into two groups, depending on whether they warn before a service degrades or count what has already gone wrong.
- Leading indicators Exactness of Capacity Forecast, Capacity Reserves, Percentage of Capacity Monitoring and planned Capacity Adjustments warn before a service degrades.
- Lagging indicators Incidents due to Capacity Shortages, Unplanned Capacity Adjustments and Resolution Time of Capacity Shortage count problems that have already happened.
Reporting packs often lean on the lagging three because the ticket system produces them with no extra work. The leading four require a forecast that somebody is willing to be wrong about in writing. They are harder to produce, and they turn measurement, reporting and performance analytics into the regular review of a published forecast.
What is the difference between capacity management and FinOps?
FinOps and capacity management optimize the same resources against different objectives: cost for FinOps, delivered performance for capacity management.
FinOps, the discipline of cloud financial operations, owns the unit economics of cloud spend: what a workload costs, who pays for it, and whether the commitment model fits the usage. Capacity and performance management owns whether the service still delivers the performance it promised. With 63% of organizations running established FinOps teams, according to Flexera’s 2026 report, the two functions increasingly work on the same estate. On the cost side, Matrix42 Cloud Cost Management brings billing data from Microsoft Azure, Amazon Web Services (AWS) and Google Cloud Platform (GCP) into one view of cloud contracts, costs and usage.
The usual point of conflict is right-sizing. A FinOps analyst looking at a fleet sitting far below its provisioned ceiling sees an obvious saving. A capacity practitioner looking at the same fleet sees the headroom that absorbs quarter-end. Both are correct about their own objective, which is why neither function can settle the argument on its own.
Service level management decides when cheaper and slower is acceptable, because it holds the agreed target in the service level agreement (SLA). If the smaller instance still meets the committed response time, the saving comes at no service cost. If it does not, the commitment has to be renegotiated before the instance is downsized; otherwise the first sign of the problem is usually an SLA breach.
Does ITIL 5 change capacity and performance management?
PeopleCert has not published any change to capacity and performance management for ITIL (Version 5).
PeopleCert released ITIL Foundation (Version 5) on 12 February 2026, positioned as AI-native by design, with emphasis on digital product and service management. PeopleCert’s own published pages say nothing about practice counts, practice categories, or this practice specifically.
Claims about what Version 5 does to the practice structure, including the content-split percentages some training providers publish, are unconfirmed. The underlying work is unlikely to change: forecasting demand, defending an agreed performance target and paying only for the headroom that is needed. The AI-native framing may even raise the stakes on the capacity side, because AI workloads are among the hardest loads an estate has to size for.
Key takeaways
- Capacity and performance form one practice ITIL 4 merged capacity with performance, and availability management remains a separate practice. Checking whether there is enough capacity without checking whether the service is fast enough leaves estates overprovisioned and still slow.
- The business level is the one most often missing Business, service and component capacity remain useful perspectives. Where only the component level is covered, a utilization report cannot name the service that is about to breach its target.
- The capacity plan depends on the service model A forecast is only as accurate as the service model beneath it, so thin configuration and asset data limits IT capacity management whatever the forecasting method.
- Every public KPI list is a secondary source The authoritative practice guide sits behind membership, and the widely copied seven-metric list dates from 2019 and follows ITIL v3. Attribute it when you use it, and report leading and lagging indicators separately.
- Autoscaling still has limits Spin-up delay, shared infrastructure limits and application design still bound elasticity, and rising rack densities have made physical capacity a live constraint again.
See which services your busiest components actually carry
Many capacity failures are predictable, because the forecast existed and nobody signed off on the spend; the cost then arrives later as an outage or an invoice.
Utilization data becomes actionable once it points at a service. Matrix42 Discovery and Dependency Mapping (DDM) identifies applications, devices, services and configuration items (CIs) and their dependencies, on-premises or in the cloud. It builds multi-layer baseline service maps across the application, network, virtualization and storage layers, and keeps the CMDB underneath them current. Matrix42 FireScope Service Performance Manager (SPM) then monitors end-to-end business services against SLA-demanded performance levels, with business service level predictive analytics.
Matrix42 does not provide application performance monitoring (APM) or workload modeling. Third-party monitoring telemetry reaches the platform through the FireScope connector and a library of 400+ integrations, and the dependency-aware service map links that telemetry to the services it affects.
Connect utilization data to your services
See how discovery, the CMDB and asset records share one data model in Matrix42 IT service management.
Explore the future of ITSM→FAQs
ITIL 4 names no mandatory job title. Accountability usually sits with a capacity manager or a service owner who signs off the capacity plan, while platform, database, and network teams supply the forecasts for their own components. In smaller organizations the role is absorbed by infrastructure leads or a FinOps analyst, with service level management owning the targets being defended.
Capacity planning in ITIL starts from the demand patterns of the business. Practitioners profile the business activity that drives load, translate it into service workload, then into component requirements such as compute, storage, memory, and bandwidth. Forecasts are tested against agreed performance targets and against budget, and the resulting plan is reviewed on a fixed cycle so that aging assumptions stay visible.
The practice draws on infrastructure and application monitoring platforms, application performance tracing, cloud provider utilization and cost consoles, and the service management toolset that holds the service model. Matrix42 covers the configuration, asset and cloud cost side, linking observed utilization to the services and licenses behind it. No single tool spans all three levels, so the data usually has to be correlated deliberately.
ITIL practice guides define practice success factors as the things a practice must do well to fulfill its purpose. For this practice they split in two: meeting agreed capacity and performance requirements for services that already exist, and making sure future demand can be met. The authoritative wording sits in the AXELOS practice guide behind MyAxelos membership, so versions reproduced on the open web are paraphrase.
Monitoring and event management observes state and raises an event when a threshold is crossed. Capacity and performance management consumes that data to judge whether the threshold is set correctly and what the service will need next quarter. Monitoring is a continuous detection capability; capacity and performance management is an analysis and planning discipline that depends on it.
No. Service design was a lifecycle stage in ITIL v3, and ITIL 4 replaced the lifecycle with the service value system and the service value chain. Capacity and performance management is a service management practice that feeds several value chain activities, most visibly plan, design and transition, and deliver and support. Sources still filing it under service design are describing ITIL v3.
No. ITIL 4 merged capacity with performance, and availability management remains its own practice that answers a different question about the same service. The confusion comes from how closely the two work together, since both draw on the same monitoring data and are often staffed by the same engineers.
Yes. Elastic infrastructure changes what the practice controls. For Software as a Service (SaaS), the unit of capacity becomes subscriptions, application programming interface (API) quotas and tenant limits instead of processor and disk capacity. For Infrastructure as a Service (IaaS), it becomes instance families, reservations and autoscaling boundaries. The typical failure changes too, from a saturated server to a throttled API or an unplanned jump in the renewal quote.
Service level management agrees what performance is owed to the customer. Capacity and performance management is the practice that has to make the committed number achievable and keep it achievable as load grows. Targets negotiated without a forecast behind them tend to be met in the first quarter and missed in the fourth, when demand catches up with the original sizing.
It needs the service model showing which components carry which service, which the configuration management database supplies. Without that mapping, utilization data only describes servers, so a saturated component cannot be traced to the business activity that will suffer. Asset records add the commercial limits: licenses, contractual ceilings, and refresh dates that constrain how much capacity can actually be added. Matrix42 keeps those asset records alongside the configuration data the service model is built from.
Related articles
ITIL 4 practiceWhat is ITIL 4 Availability Management? A comprehensive practice guideRead the article →
ITIL 4 practiceWhat is ITIL 4 Monitoring and Event Management? Categories, AIOps and MTTD guideRead the article →
ITIL 4 practiceWhat is ITIL 4 Service Level Management? SLAs, metrics and practice guideRead the article →
ITIL 4 practiceWhat is ITIL 4 Service Financial Management? Budgeting, charging and practice guideRead the article →
ITIL 4 practiceWhat is ITIL 4 Measurement and Reporting? Metrics, KPIs and practice guideRead the article →
ITIL 4 overviewITIL 4 practices for IT Service ManagementRead the article →
Sources
- 1 AXELOS / PeopleCert, “Capacity and performance management: ITIL 4 Practice Guide,” 2020.axelos.com/resource-hub/practice/capacity-and-performance-management-itil-4-practice-guide
- 2 ITSM.tools, Sophie Danby, “ITIL 4 Management Practices Explained: Full List and Purposes,” 2026.itsm.tools/34-itil-4-management-practices
- 3 IT Process Wiki / IT Process Maps, “Capacity Management,” 2023.wiki.en.it-processmaps.com/index.php/Capacity_Management
- 4 IT Process Wiki / IT Process Maps, “KPIs Capacity Management,” 2019.wiki.en.it-processmaps.com/index.php/KPIs_Capacity_Management
- 5 Flexera, “Flexera Finds Cloud Value is Rising While AI Waste Grows: 2026 State of the Cloud Report,” 2026.flexera.com/about-us/press-center/flexera-finds-cloud-value-is-rising-while-ai-waste-grows
- 6 Uptime Institute, “Annual Outage Analysis Report 2026,” 2026.uptimeinstitute.com/about-ui/press-releases/uptime-announces-annual-outage-analysis-report-2026
- 7 Uptime Institute, “16th Annual 2026 Global Data Center Survey,” 2026.uptimeinstitute.com/about-ui/press-releases/16th-annual-2026-global-data-center-survey-deployment-of-high-density-racks-rising-fast-operators-face-continued-recruiting-and-retention-pressures
- 8 IDC, “AI Infrastructure Spending Holds Near $90 Billion in Q1 2026; 2026 Forecast Raised to $497 Billion,” 2026.idc.com/resource-center/blog/ai-infrastructure-spending-holds-near-90-billion-in-q1-2026-as-arm-overtakes-x86-in-accelerated-servers-2026-forecast-raised-to-497-billion
- 9 Proscaler, “Why capacity management is more relevant than ever in the cloud era,” n.d.proscaler.de/post/why-capacity-management-is-more-relevant-than-ever-in-the-cloud-era
- 10 PeopleCert, “New ITIL Explained for Certified Professionals: ITIL (Version 5),” 2026.peoplecert.org/news-and-announcements/itil-version-5-explained
- 11 PeopleCert, “ITIL Foundation (Version 5),” 2026.peoplecert.org/browse-certifications/it-governance-and-service-management/ITIL-1/itil-5-foundation-version-50-4154