Assessing Network Management Effectiveness

Now that we have reviewed the factors and management features that contribute to management's business impact, how can we actually assess how well management is working? What is needed is a set of metrics by which the effectiveness of management can be measured. An understanding of such metrics, along with an understanding of the different ways in which management technology contributes to those metrics, can provide invaluable guidance for determining priorities in the development and deployment of management technology.

Ultimately, what counts is the bottom line—hence, metrics used for network management should include the same type of metrics that would be used for any other type of business. We therefore start by taking a look at some metrics that capture business impact. As we shall see, those metrics alone cannot capture the entire picture; they need to be augmented by other metrics that capture how effective various factors are that contribute to the bottom line. That aspect is discussed later in the chapter.

In effect, there is a management value chain. This chain starts with individual management features, such as threshold-crossing alerts, programmability, and sort and search functions. As discussed previously, those features contribute to various factors in management effectiveness, such as user productivity, operational robustness, and scalability of operations. Those factors in turn influence the bottom line, as Figure 12-3 shows.

Figure 12-3 Management Value Chain and Associated Metrics

Metrics

/ \

\ /

Scale

\ i

/ |TCAs| n 1 Programmability 1

User Productivity

Input validation

Event correlation

N'

Usage Complexity

! "1

Interface Consistency

impact

Integrity

-i |Rollback|

Sort & Search

Management features

Factors in management \effectiveness /

Sort & Search

Management features

Factors in

Increased Network Value

Addl. Revenue Opportunity

TCO Reduction management \effectiveness /

Management Metrics to Track Business Impact

As mentioned earlier, the business impact of network management consists of its contribution to cost savings, revenues, and improved network services themselves, specifically regarding reliability and availability. We start with metrics that can be applied to assessing the cost side of the equation. The metrics that are presented here do not constitute a complete list. Instead, they are intended to provide a sense of the areas on which management has an impact and how this impact can possibly be captured.

For one, there is the total cost of operations. This includes the cost of the operations staff and the amortization cost of the operations support infrastructure. The more effective the management, the lower the total cost.

Of course, total numbers can be hard to compare because the total cost does not take into account the complexity and scale of the network being managed. Normalized metrics that put the cost into perspective regarding what is actually being achieved are often more meaningful. For example, total cost of operations can be normalized over the following:

■ Number of devices (operational cost per device). Because devices can be vastly different in nature and scale, this metric is not always meaningful unless dealing with very homogeneous environments.

■ Number of ports (operational cost per supported access port). This provides a more level playing field and allows operational cost to be compared across different types of equipment. For example, this measure should reflect better economies of scale if large-scale pieces of equipment are used. Such equipment will likely involve a higher operational cost per device; however, taking into account the additional capacity paints a more appropriate (and favorable) picture.

■ Number of service instances—for example, operational cost per voice extension.

■ Number of end users—for example, the operational cost per employee in an enterprise.

Other metrics track the contribution of individual cost factors. An obvious set of metrics concerns operator productivity, measuring the number of devices, ports, service instances, or end users that a single network operator can support. A variation of this is the amount of service revenue supported per operator. Another example concerns the number of truck rolls—for example, truck rolls per service activation or truck rolls per year of maintaining an already activated service. Truck rolls refer to the need to send personnel to a customer (in many cases, driving in a truck) to activate and maintain a communications service. Not only are truck rolls very expensive, but they also inconvenience customers and negatively impact customer satisfaction. A third example concerns the number of customer calls to a help desk that are required to resolve a case. Clearly, the number of calls should be as low as possible because each call costs money and inconveniences customers.

Sometimes the ratio between different cost factors can provide useful information as well. A telling metric in that regard concerns the ratio of operations support infrastructure amortization cost to personnel cost. A high ratio could be an indication that there is a high buildout of operations support infrastructure, leading to potentially diminished returns on further investment in that area. A low ratio, on the other hand, could be an indication that investment is too low and that operators could be better leveraged by beefing up infrastructure, particularly if the normalized cost of operations is relatively high.

There are also metrics concerning network management's impact on revenue. The problem with many of those metrics is that although management is a contributing factor, how much revenue can really be attributed to management is not always clear. In many cases, there are also many other factors at work. Nevertheless, there are metrics that are still useful, such as the following:

■ Customer loss rate from dissatisfaction with the service level. (This is similar to the popular business metric of customer retention rates, but it focuses on aspects that might be attributable to poor management of the service.)

■ Mean time to fill customer orders—time from service order to service activation. The shorter the time, the better because the revenue potential of a service is more fully exploited.

■ Average size of customer order backlog. Again, this number can be translated directly into customer dollars. The smaller the average backlog size, the better.

■ Customer defections from slow order fulfillment (services cancelled and, hence, revenue lost from customers canceling services before they were activated).

■ Total penalties for service level agreement violations, reflecting total lost revenue.

■ Ratio between penalties and (theoretical) service revenue, reflecting the portion of revenue that was lost from violated SLAs.

Finally, reliability and availability of the network and services need to be assessed. The classical metrics here are the following:

■ Availability—The percentage of time during which a service is functioning properly. Two aspects of availability should be distinguished: availability as a whole and availability that factors scheduled maintenance into the equation. The second aspect is much more relevant than the first because it captures only unavailability that comes as a negative surprise.

■ MTBF—Mean time between failures. A classical measure of reliability, this gives an indication of how often a service or device becomes unavailable, regardless of the duration of the failure. Imagine a case in which a service is overall highly available, but occasional very brief glitches occur that require users to redial a phone call. Clearly, this is much less acceptable than if there is just one outage, even if that outage lasts slightly longer. Figure 12-4 depicts this situation.

■ MTTR—Mean time to repair. Another classical measure, this indicates how long it takes for services that are impacted by failures to be restored.

Figure 12-4 Availability and MTBF

V >

XJ

FallIré

y/ ,

->-

(a) Low availability, high MTBF

(a) Low availability, high MTBF

Mf J

(b) Higher availability, low MTBF

Session completed Session aborted due to glitch

Of course, for all these metrics, the question needs to be asked: "Availability/MTBF/MTTR of what?" The metrics can be applied to individual devices in a network and to the service as a whole. However, it is quite clear what counts from the perspective of the customer: that customer's service. It is important to realize that availability of a service and of a device in the network that carries the service are not the same. The good news is that service availability can be much higher than availability of the network or network devices. For example, when a device fails, services might automatically be switched over to other devices in the network. Nevertheless, it is also important to keep an eye on these measures at the device level because poor reliability and availability of network equipment could ultimately affect services as well. In addition, it leads to negative impact on operational cost because failures must be dealt with.

As with many other metrics, it often makes sense to further differentiate these metrics and distinguish between different types of devices, different types of services, different geographical regions, and different domains of responsibility between organizational subgroups. This can yield data for interesting comparative evaluations. For example, if one group in the organization consistently scores higher on certain management metrics, it might be possible to learn important lessons from that group and apply them to other parts of the organization.

Finally, in the eyes of the customer, maintenance is another factor that impacts availability. Needing to account for maintenance operations is an inconvenience, although, of course, it is by far preferable to unexpected service outages. One meaningful metric with regard to the effectiveness of maintenance is the ratio between unplanned outages and scheduled (maintenance) intervals—for services and for managed devices. Clearly, this ratio should be as close to 0 as possible. Another metric is mean time to maintenance. Although the ratio between unplanned outages and scheduled maintenance intervals might be improved simply by increasing the number of maintenance intervals, doing so would come at the expense of this metric. In addition to being a factor in availability, how often maintenance needs to be performed is a cost factor.

Management Metrics to Track Contribution to Management Effectiveness

Metrics that can be used to assess management business impact are one way to assess management effectiveness. However, those metrics do not tell the entire story. Many factors impact management effectiveness, each of which is the result of a combination of features. For example, some features of management applications make network managers more productive. Other features improve operational robustness by making it harder for network managers to introduce inconsistent configurations, the operational equivalent of shooting themselves in the feet. An important question, therefore, concerns whether it is possible to assess the individual contribution of each of those features to management effectiveness. Although the metrics that have so far been discussed allow assessment of the impact that management features collectively have, they provide no indication of their individual effectiveness. In some cases, it is possible to do better. This is the topic of the following discussion.

Metrics for Complexity of Operational Tasks

Perhaps the most important aspect that contributes to management effectiveness concerns the complexity of operational tasks that network managers are faced with. Reducing complexity is important for a number of reasons: It reduces operational cost because operators generally have fewer steps to perform for a given task and those steps may be simpler, requiring less training. It allows the completion of more of the same task in the same time, thereby increasing operational throughput. It reduces the opportunity to introduce errors, which might impact availability and demand complicated recovery. The management metrics that we have encountered so far allow us to see the impact of complexity on business. However, they do not allow us to assess complexity directly.

Complexity of operational tasks can be assessed in a number of ways, many of which were first articulated in a paper by Brown, Keller, and Hellerstein that was published at the IM 2005 conference (the reference is listed in Appendix B, "Further Reading"). Three different categories of complexity are distinguished:

■ Execution complexity consists of two aspects:

— The number of steps that the task involves.

— The number of context switches involved in a task. Roughly, this reflects how many different targets (for example, different pieces of networking equipment, or different types of device interfaces) the steps are directed at.

■ Parameter complexity concerns the complexity of the parameters that are required as part of each step. This complexity is determined by several aspects, including the following:

— The number of different parameters that are required for all steps, in total.

— The average number of parameters that are required in each step. (Some parameters are required in more than one step.)

— The ease with which a network manager can obtain those parameter values. For each parameter, a corresponding score is assigned. A high score is applied when a parameter represents an obscure value that requires a good deal of experience to choose; a low score is applied for parameters that are "obvious" or that require only a simple lookup. The scores of all parameters are aggregated.

■ Memory complexity concerns the amount of memory that is required in the mind of a network manager to perform the steps. It takes into account the number of parameters that must be remembered, the length of time they must be retained in memory, and how many intervening items were stored in memory between uses of a remembered parameter. Memory complexity can be expressed by the largest depth of a stack that would be required to hold the required items.

To obtain management metrics from these complexity measures, a number of benchmark tasks should be selected for which complexity is assessed according to the mentioned criteria. For example, tasks such as the following are good candidates:

■ Adding a device to the network in a rudimentary configuration.

■ Adding a device to the network in an advanced configuration, such as a redundant failover configuration.

■ Configuring all interfaces of the device with the same set of parameters. For example, this could be configuring all DS0 ports on a line card for voice service with a certain echo-cancellation setting.

■ Configuring a loopback on an interface.

■ Configuring an IPSec tunnel (for a reference network topology).

■ Adding a subscriber to a service.

Clearly, it is possible for different management applications to score quite differently when their operational complexity is evaluated against those benchmarks. For example, consider the operational complexity that is involved in adding a subscriber to a service. An operations support environment that includes an application that offers cookie-cutter templates to facilitate this task will achieve a much better score than an operations support environment in which every aspect of the service has to be configured individually (because more steps would need to be performed and more parameters would have to be remembered). Of course, we expected this outcome, but it is nice to have it backed up by quantifiable data.

Related to metrics that are intended to capture operational complexity for network managers are metrics that attempt to capture operational complexity for management applications. Instead of the number and complexity of individual operational steps, key in this case are the number and complexity of communication exchanges that need to take place. As before, it is advisable to select a number of benchmark tasks that are representative for tasks that management applications face. Here are some examples:

■ Keeping an application synchronized with the actual device configuration, with mean time between configuration changes 48 hours and maximum permissible time lag to reflect a changed configuration 15 minutes

■ Collecting a set of performance counters periodically, every 15 minutes over an interval of 24 hours

■ Running a diagnostic test pattern every 15 minutes over the period of 24 hours, and being notified when an abnormal result occurs

Subsequently, the complexity of exchanges between managed network and management applications is assessed.

■ The number of communications exchanges that are required.

■ The number of dependencies between those exchanges, measured by number of exchanges that need to be serialized as one exchange relies on a piece of information provided by another. This also indicates the complexity of rainy-day scenarios in which things go wrong and need to be recovered from.

Again, different capabilities of systems in the agent role yield vastly differing results. For example, those results show that to keep a management application synchronized with a managed device, a much lower number of communication exchanges is required when reliable configuration change events are supported, compared to cases in which periodic synch-up with devices is required.

Metrics for Scale

Another aspect that is important to management effectiveness concerns scale. Scale is an area in which the need for metrics is widely recognized and often many different metrics are available.

Not all metrics that are offered to support claims of great scale are meaningful, even though they might sound reasonable and impressive at first. A good example is the "number of managed objects supported" by a management application. Several terms in that metric are quite fuzzy: What exactly is meant by a "managed object"? Is it an object in the object-oriented sense that abstracts a managed resource (such as a line card) on a device? If so, how complex is the model? How many managed objects are used to model a device? Or does every attribute count as its own managed object? What about data types? Can they be complex, such as structs or arrays, or are they simple only, such as strings or integers? And even if the meaning of the term "managed object" is clarified, what is exactly meant by "supported"? Does it simply mean that an application will not crash when populated with so many objects? If it indeed means that a network manager will still be able to do actual work, with what performance?

However, there are many excellent metrics for scale that are more appropriate. Here is a small sample:

■ Time to provision 1000 instances of a particular type of service.

■ Provisioning throughput, which is basically the inverse of the previous metric: Number of service instances that can be provisioned per time unit (minute, hour).

■ Time to synchronize a management application with a network of a given size.

■ Number of events that can be sustained per second without dropping events.

■ Number of managed devices of a given type (or a given mix) that are supported, accompanied by a meaningful performance measure—for example, while being able to perform a complete network audit within 60 minutes. (Note the difference to the less meaningful metric of the "number of managed objects supported.")

Other Metrics

Many other aspects contribute to management effectiveness. For example, one aspect concerns the effectiveness of events that are reported from the network and of events that management applications bring to the network manager's attention. It is important that the reported event information be rich enough to convey what is going on, while focused and relevant enough not to distract from the real problems. One metric that is useful in this context is the number of observed alarms per root cause—effectively, a "management signal to management noise" ratio. We would like to see this number go down—ideally, to 1, in which case there is no noise and each alarm is indeed indicative of a distinct problem. In some cases, it might make sense to exclude events from the ratio that are explicitly used to report the impact of a root cause and that refer to other events that point to the root cause. An example is service alarms that indicate which and how many customers are impacted by a service outage, which is caused by a failure that is indicated by a separate alarm.

Another aspect concerns operational robustness and the error-proneness of management interfaces. An interesting metric here is the percentage of network outages from operational error. This is a number that every network provider wants to see trend down over time.

There are other aspects still, such as the consistency between management interfaces of different managed systems, or between interfaces of management applications. As mentioned previously, the discussed metrics are by no means intended as a comprehensive list, but as an illustration of how the contribution of different factors to management effectiveness can be captured in a quantifiable manner. We leave this discussion at this point and turn to the topic of how to determine which metrics should be applied in a particular situation.

Developing Your Own Management Benchmark

By now, it should be clear that no single metric can be used to assess management effectiveness or the business contribution that management provides. Instead, many metrics should be looked at collectively to obtain the overall picture. (Of course, it is possible to combine different metrics into one aggregate using some formula.) In the end, the goal is to provide a set of criteria that can be used to assess the effectiveness and business contribution of management and management technology under different criteria.

Because there are so many metrics, which ones should you use in your particular case? The answer is: It depends. And: Use common sense. Obviously, when assessing the effectiveness of fulfilling telephony service orders, a metric to assess provisioning throughput will be more meaningful than a metric to measure the capacity of processing alarms. To identify meaningful metrics, it helps to think about the properties that are most important in the particular context. For example, is the main concern cost or improved availability? Depending on the answer, focus on metrics that provide information on one versus the other, and combine them into a management benchmark, custom-tailored to your specific situation. Also, be creative—do not hesitate to invent your own metrics as needed. Usually, a good question to ask is, "What use cases need to be assessed?"

Use-case analysis is a software-engineering technique that is used to identify system requirements. In a nutshell, use-case analysis works as follows: First, a set of "actors" is identified. Actors are entities that derive value from the system being analyzed—network operators, for example. Then a set of basic scenarios called use cases is identified. For each use case, the individual steps and interactions that need to occur between the actors and the system are specified, along with preconditions and post-conditions that are supposed to hold before and after a use case is executed, and an explanation of exceptions that can occur.

Nothing prevents use cases from being used not only to derive requirements, but also to capture the factors that best characterize how effective the use case is executed. The result is a set of metrics that is custom-tailored to a particular context. One way in which metrics can be derived is to apply the complexity measures that were discussed earlier to the operational task that corresponds to the use case. For example, it can be meaningful to compare different systems with regard to the number of steps that are required for a given use case, along with the complexity of information that needs to be exchanged at each step.

Assessing and Tracking the State of Management

Now that we know what types of metrics can be applied, what values should these metrics indicate? What constitutes a "good" or "acceptable" value versus one that is "bad" and "needs improvement"? The general answer here again is, it depends.

What is perhaps most significant about those metrics is not their actual value at any one moment in time, but the fact that they can serve as a basis for quantifiable comparisons.

Those comparisons can be conducted between alternatives. For example, metrics can be used to compare management effectiveness between different service providers or IT departments, providing insight into their relative strengths and weaknesses. It might not be really all that interesting to know that it takes a service provider organization 8 hours to provision 1000 instances of a particular type of service, until you hear that it takes the competition only 15 minutes ("What are they doing that we don't?") or 15 days ("Wow, we're way ahead of them!"). Other metrics can be used to compare the effectiveness of different management applications, and of different device types and devices from different vendors that play a similar role in the network.

Comparisons can also be conducted over time, allowing assessment of progress. Again, the absolute value of a metric is often of lesser significance, as long as its trend points in the right direction. If your ratio of customer-reported incidents to incidents overall stands this month at 20 percent, this might be good news if it was 21 percent last month and 30 percent a year ago because it means you have gotten better at identifying problems yourself before they impact customers. However, it is bad news if it used to hover around 15 percent.

Another aspect to consider is that no single metric tells the whole story, but many metrics combine to provide an overall picture. It is easy to get lost in the amount of data that can be collected. What can be extremely helpful in such situations is a way to graphically visualize the data. No one common established technique exists for this. Figures 12-5 and 12-6 depict one way in which this can be accomplished. Figure 12-5 represents a management metrics coordinate system with multiple axes: one axis for every metric. Axes should be arranged so that axes of metrics that relate to similar factors are adjacent. For example, metrics that relate to operator productivity are in one quadrant, metrics related to availability as impacted by management are in a second quadrant, metrics that relate to the integration tax in are a third quadrant, and so on. Then the value for each metric is assessed. The corresponding coordinates are connected and the enclosed area is shaded. The result is a visual profile, as depicted in Figure 12-6. To compare different alternatives, their respective profiles can be superimposed to identify areas where over- and underlap occur, exposing their respective strengths and weaknesses. As management effectiveness improves over time, the values of the coordinates increase and the overall enclosed area grows.

Figure 12-5 Visualizing Management Effectiveness: A Metrics Coordinate System Operational Complexity-Related Scale-Related

W, »/\ (B ra

Event throughput (events/sec) ^

Service availability

i

Reliability/availability-related Revenue-related

Reliability/availability-related Revenue-related

Of course, it is also possible to compile a single value across a set of metrics according to some formula. For example, each metric might be assigned a weight factor that it is multiplied by, before being added up. Although this approach is easy to use and convey, one drawback is that some information is lost in the course and the resulting picture is less differentiated.

Figure 12-6 A Management Effectiveness Profile

Figure 12-6 A Management Effectiveness Profile

Using Metrics to Direct Management Investment

Metrics such as those discussed in this chapter can be of tremendous help when trying to determine where to direct management investment and to assess whether a particular management investment might even be worth the while. A general approach to take is as follows:

■ In the beginning, establish the objectives for further investment—for example, is the most important goal to reduce operational cost, or to increase network availability, or to accelerate rollout of a particular service?

■ With the objectives in place, identify metrics that can be used to assess the degree to which those objectives are met.

■ With the metrics identified, determine different alternatives that will have an impact on those metrics. We saw many examples of this in the subsection on factors that determine management effectiveness. Assess and try to quantify the expected impact of those factors in terms of the identified metrics.

■ Finally, select the alternative that yields the most attractive return on investment (ROI)—that is, the highest benefit in relation to the cost. Additional investment should be made as long as the expected ROI is deemed favorable.

ROI is generally determined by dividing the benefit expected from an investment by the investment amount. A related measure calculates the time until the benefit will have paid for the investment. In our case, the benefit might not always be expressed in terms of dollars; however, the same general idea can be applied.

0 0

Post a comment