What Other ISPs Are Doing

Here are just a few examples of what ISPs from all over the Internet are using to manage their network. Randy Bush ([email protected]) asked major ISPs in the United States on the NANOG mailing list what they used for traffic analysis. His summary follows. Notice especially the number of UNIX script-based tools. Readers can find more information by looking at the NANOG mailing list archives at www.nanog.org.

We do SNMP polling every 15 minutes at SESQUINET on every line over which we have administrative control and over every peering point. We produce a daily report on errors and usage. We are getting ready to switch to Vulture or NetScarf (or some combo) to give us more interactive information.

We perform measurement of certain basic network parameters, such as usage (bandwidth used/total bandwidth) and line error rates on all of our noncustomer links. We perform CPU usage, memory usage, and environmental monitoring of all our routers. We also perform the line usage and error rate on all customer lines. We monitor all of our customers' routers unless they say otherwise and notify them of any problems. Finally, we monitor select points throughout the Internet (root name servers and so on) on a four-times-an-hour basis using pings. We accomplish this monitoring using the following items: an in-house built package that uses SNMP, traceroute, and ping to provide graphs and tabular statistical information. We use Cabletron's Spectrum for a quick network overview.

We do SNMP MIB-II stuff, plus the cflowd stuff and something we call mxd (which measures round-trip times, packet loss and potential reason, and so on, from a whole bunch of different points in our network to a bunch of other points in our network. We use it to create delay matrices, packet loss reports, and other reports. There are some other things, but these are the biggies.

The mxd thing was originally just sort of a toy for neat reports, but in the last year, it has become a critical tool for measuring delay variance for one of our VPDN customers that does realtime video stuff (and is, to some extent, helping us figure out where we have delay jitter and why; however, it's also raising more questions).

Because most of my professional career has been in the enterprise world, I can offer you what we used to measure availability to our mail servers, Web servers, DNS servers, and so on at one of my previous employers.

We employed several application tests, along with network performance tests. Our primary link was via UUnet, a burstable T1. We purchased an ISDN account from another local provider who wasn't directly connected to UUnet. Probably a good example of a joe-average-user out there.

Every five minutes, we measured round-trip response times to each of the servers and gateway router (via ping) and recorded it. We also had application tests, such as DNS lookups on our servers, timing sendmail test mails to a /dev/null account and time to retrieve the whole home page.

This wasn't meant to be a really great performance-monitoring system; it was actually meant to 1. check how our availability looked from a "joe user" perspective on the Net (granted, reachability/availability wasn't perfect because it was only one point in the net), and 2. look at response time trends/application trends to see if our hardware/software was cutting it.

We use a traffic flow monitoring system from Kaspia Systems (www.kaspia.com). The Kaspia product collects all sorts of data from router ports and RMON probes, stores the data, and performs various trend analysis. We collect traffic flow, router CPU usage, and router memory information plus various errors. There is a data-reduction process that runs once a day and a very nifty Web interface. The product is not cheap, but the system definitely fills a void here.

Maybe I should organize a talk on what we are doing with it for an upcoming NANOG. As an old instrumentation engineer, I think the basis of our use of the tool is pretty solid. Plus, I actually developed a means for calibration of the accuracy of the flow data. I have not had time yet to work out a validation for the trends, but I'll get to it one of these decades.

Also, the Kaspia people will give you a 30-day trial on their product at no charge.

For nonintrusive stuff, we keep a log of all interface status changes on our routers, and we pull five-minute byte-counts inbound and outbound on each interface, which we graph against port speed. Watching the graphs for any sort of clipping of peaks gives a pretty good indication of problems, and watching for shifts of traffic between ports on parallel paths does likewise.

As for intrusive testing, we do a three-packet min-length ping to the LAN-side port of each of our customers' routers once each five minutes, and we follow that up with additional attempts if those three are lost. We log latency, and if we have to follow up with a burst, we log loss rate from the burst. Pinging through to the LAN port obviously lets us know when CPE routers konk out; occasionally we see hung routers that still have operational WAN ports talking to us. Likewise, simply watching VC-state isn't a reliable enough indicator of the status of the remote router. Plus, it tells you if the customer has kicked the Ethernet transceiver off their equipment, for instance. It wouldn't matter to you probably, but our demarc is all the way out at the WAN port because we own and operate our customers' CPE.

I think a bit about what more we could be doing; flows analysis and whatnot. . .. It's nice to think about, and eventually we'll get around to it, but programmer time is relatively precious and other things have higher priority because the current system works and tends to tell us most of what we seem to need to know to provide decent service.

We place quite a bit of emphasis on network stats. Currently we have about three years of stats online, and we are working on converting our in-house engine to an rdbms so we can more easily perform trend analysis. Besides Kaspia, other commercial packages include trendsnmp (www.desktalk.com) and concord's packages (www.concord.com). Our in-house stuff is located at http://netop.cc.buffalo.edu/, if you are curious about what we do.

We are Neanderthals right now—we use a hacked rcisco to feed data to nocol. We watch bandwidth (separately as well) on key links—and also watch input errors and interface transitions (for nocol)—all done with Perl and expect-like routines, parsing "sho int's" every few minutes.

Emergency stuff goes through nocol; bandwidth summaries are mailed to interested parties overnight.

We have running here now the MRTG package that generates some fancy graphics, but, in my opinion, these graphics are useless, and looking in detail to some of the reports, they are not accurate. Several of our clients request the raw data, but this package only maintains raw data just to generate the graphs.

In the past we used to have a kind of ASCII report (Vikas wrote some of the scripts and programs) generated from information obtained using the old SNMP tool set developed by nysernet, but I guess that nobody maintained the config files and I believe that the SNMP library routines used aren't working.

We have been using the MRTG package, which is basically a special SNMP agent that queries the routers for stats and then does some nice graphing of the data on the Web.

SNMP queries with a heavily modified version of MRTG from the nice guy in Germany. It works very nicely. We have recently installed NetScarf 2.0, and are contemplating merging NetScarf 3.0 with the MRTG front end.

I'm researching whether I can rewrite Steve Corbato's fastpoll program using the fastsnmp library from the NetScarf people. I think this will allow fastpoll to scale better. I've successfully written a quick C program that uses the library to collect the required data for a router—now I've just got to make it so that we can manage it easily (in other words, autogenerated config files from our databases).

My goal is to be able to collect 1- to 2-minute period data on all links that are greater than 10 Mbps— 15 minutes of data for everything else. The two-minute collection period will allow the bandwidth to scale up to 280 Mbps before experiencing two counter rollovers within a polling interval. Hopefully that will hold us over until the interface counters are available as Counter64 objects with SNMPv2 (if that ever happens).

What fastpoll collects now is ifInOctets, ifOutOctets, ifInUcastPkts, ifOutUcastPkts, ifInErrors, and ifOutDiscards. Rather than storing the raw counters, it calculates the rate by dividing the delta by the period. Getting the accurate period is actually the hard part—I am having SNMP send me the uptime of the router in each query and using that to calculate the interval between polls and to detect counter resets due to reboots. The other trick to handle is the fact that, although IOS Software updates the SNMP counters for process-switched packets as they are routed, it looks like the counter for SSE switched packets on C70X0 routers get updated only once every 10 seconds.

Continue reading here: Committed Access Rate CAR

Was this article helpful?

0 0

Readers' Questions

  • Marcel
    Is line with other isps reducing?
    1 year ago
  • Yes, other ISPs are also reducing the cost of their services. This trend is likely to continue as new technologies such as fiber optics and 5G become more widely available and make it cheaper for ISPs to provide services. Additionally, many governments are now introducing legislation to regulate the pricing of internet services, which can also contribute to cost reductions.