Voice Quality Issues

Overall voice quality is a function of many factors, including delay, jitter, packet loss, and echo. This section discusses these factors and ways to minimize them.

Packet Delays

Packet delay can cause voice quality degradation. When designing networks that transport voice, you must understand and account for the network's delay components. Correctly accounting for all potential delays ensures that overall network performance is acceptable.

The generally accepted limit for good quality voice connection delay is 150 milliseconds (ms) one-way. As delays increase, the communication between two people falls out of synch (for example, they speak at the same time or both wait for the other to speak); this condition is called talker overlap. The ITU describes network delay for voice applications in recommendation G.114; as shown in Table 8-3, this recommendation defines three bands of one-way delay.

Table 8-3 ITU G.114 Recommended Delays for One-way Voice Traffic

Readers' Questions

  • manuela foerster
    Which of the following contributes to your voice quality?
    2 months ago
    1. Breathing technique
    2. Vocal range
    3. Diction and enunciation
    4. Vocal resonance
    5. Articulation
    6. Pitch and intonation
    7. Volume control
    • medardo
      What is the maximum recommended oneway delay for voice traffic?
      2 months ago
    • The maximum recommended one-way delay for voice traffic is 150 milliseconds (ms).

      Delay

      Effect on Voice Quality

      0 to 150 ms

      Acceptable for most user applications.

      151 to 400 ms

      Acceptable provided that the organization is aware of the transmission time and its impact on the transmission quality of user applications. Note that this is the expected range for a satellite link.

      Longer than 401 ms

      Unacceptable for general network planning purposes; however, this limit is exceeded in some exceptional cases.

      Voice packets are delayed if the network is congested because of poor network quality, underpowered equipment, congested traffic, or insufficient bandwidth. Delay can be classified into two types: fixed network delay and variable network delay.

      Fixed Network Delays

      Fixed network delays result from delays in network devices and contribute directly to the overall connection delay. As shown in Figure 8-26, fixed delays have three components: propagation delay, serialization delay, and processing delay.

      Propagation Delay

      KEY POINT

      Propagation delay is the length of time it takes a signal to travel the distance between the sending and receiving endpoints.

      This form of delay, which is limited by the speed of light, can be ignored for most designs because it is relatively small compared to other types of delay. A popular estimate of 10 microseconds/mile or 6 microseconds/kilometer is used for estimating propagation delay.

      NOTE Propagation delay has a noticeable impact on the overall delay only on satellite links.

      Serialization Delay

      KEY POINT

      Serialization delay is the delay encountered when sending a voice or data frame onto the network interface. It is the result of placing bits on the circuit and is directly related to the circuit (link) speed.

      The higher the circuit speed, the less time it takes to place the bits on the circuit and the less the serialization delay. Serialization delay is a constant function of link speed and packet size. It is calculated by the following formula:

      (packet length)/(bit rate)

      A large serialization delay occurs with slow links or large packets. Serialization delay is always predictable; for example, when using a 64-kbps link and 80-byte frame, the delay is exactly 10 ms

      NOTE The previous example is calculated as follows:

      ■ 64 kbps = 64,000 bits/sec * 1 byte/8 bits = 8000 bytes/sec = 8000 bytes/1000 ms = 8 bytes/ms

      ■ Serialization delay = (packet length)/(bit rate) = (80 bytes)/(8 bytes/ms) = 10 ms

      NOTE Serialization delay is a factor only for slow-speed links up to 1 Mbps.

      Processing Delay

      KEY POINT

      Processing delay is the time the DSP takes to process a block of PCM samples.

      Processing delays include the following:

      ■ Coding, compression, decompression, and decoding delays: These delays depend on the algorithm used for these functions, which can be performed in either hardware or software. Using specialized hardware such as a DSP dramatically improves the quality and reduces the delay associated with different voice compression schemes.

      ■ Packetization delay: This delay results from the process of holding the digital voice samples until enough are collected to fill the packet or cell payload. In some compression schemes, the voice gateway sends partial packets to reduce excessive packetization delay.

      Variable Network Delays

      Variable network delay is more unpredictable and difficult to calculate than fixed network delay. As shown in Figure 8-27 and described in the following sections, the following three factors contribute to variable network delay: queuing delay, variable packet sizes, and dejitter buffers.

      Figure 8-27 Variable Delays Can Be Unpredictable

      Queuing Delay Queuing Delay Queuing Delay H H

      Figure 8-27 Variable Delays Can Be Unpredictable

      Queuing Delay Queuing Delay Queuing Delay H H

      Queuing Delay and Variable Packet Sizes

      KEY POINT

      Congested output queues on network interfaces are the most common sources of variable delay.

      Queuing delay occurs when a voice packet is waiting on the outgoing interface for others to be serviced first. This waiting time is statistically based on the arrival of traffic; the more inputs, the more likely that contention is encountered for the interface. Queuing delay is also based on the size of the packet currently being serviced; larger packets take longer to transmit than do smaller packets. Therefore, a queue that combines large and small packets experiences varying lengths of delay.

      Because voice should have absolute priority in the voice gateway queue, a voice frame should wait only for either a data frame that is already being sent or for other voice frames ahead of it. For example, assume that a 1500-byte data packet is queued before the voice packet. The voice packet must wait until the entire data packet is transmitted, which produces a delay in the voice path. If the link is slow (for example, 64 or 128 kbps), the queuing delay might be more than 200 ms and result in an unacceptable voice delay.

      Link fragmentation and interleaving (LFI) is a solution for queuing delay situations. With LFI, the voice gateway fragments large packets into smaller equal-sized frames and interleaves them with small voice packets. Therefore, a voice packet does not have to wait until the entire large data packet is sent. LFI reduces and ensures a more predictable voice delay. Configuring LFI

      fragmentation on a link results in a fixed delay (for example, 10 ms); however, be sure to set the fragment size so that only data packets, not voice packets, become fragmented. Figure 8-28 illustrates the LFI concept.

      Figure 8-28 LFI Ensures That Smaller Packets Do Not Get Stuck Behind Larger Packets

      WAN Interface

      Output Queue

      Voice Packet (small)

      File Transfer Packet (Large)

      Without LFI, Voice Packet Will Have to Wait Until Complete File Transfer Packet Has Been Sent

      WAN Interface

      Last Fragment

      Voice Packet

      Fragment 3

      Fragment 2

      File Transfer Fragment 1

      With LFI, Voice Packet Will Be Interleaved with Fragments of File Transfer Packet, Resulting In:

      Last Fragment

      WAN Interface

      Fragment 3

      Voice Packet

      Fragment 2

      File Transfer Fragment 1

      Dejitter Buffers

      Because network congestion can occur at any point in a network, interface queues can be filled instantaneously, potentially leading to a difference in delay times between packets from the same voice stream.

      KEY POINT

      The variable delay between packets is called jitter. The next section describes jitter.

      Dejitter buffers are used at the receiving end to smooth delay variability and allow time for decoding and decompression.

      NOTE The dejitter buffer is also referred to as the playout delay buffer.

      On the first talk spurt, dejitter buffers help provide smooth playback of voice traffic. Setting these buffers too low causes overflows and data loss, whereas setting them too high causes excessive delay.

      Dejitter buffers reduce or eliminate delay variation by converting it to a fixed delay. However, dejitter buffers always add delay; the amount depends on the variance of the delay.

      Dejitter buffers work most efficiently when packets arrive with almost uniform delay. Various QoS congestion avoidance mechanisms exist to manage delay and avoid network congestion; if there is no variance in delay, dejitter buffers can be disabled, reducing the constant delay.

      KEY POINT

      When using dejitter buffers, delay is always added to the total delay budget. Therefore, keep the dejitter buffers as small as possible to keep the delay to a minimum.

      Jitter

      KEY POINT

      Jitter is a variation in the delay of received packets.

      At the sending side, the originating voice gateway sends packets in a continuous stream, spaced evenly. Because of network congestion, improper queuing, or configuration errors, this steady stream can become lumpy; in other words, as shown in Figure 8-29, the delay between each packet can vary instead of remaining constant. This can be annoying to listeners.

      Figure 8-29 Jitter Is the Variation in the Delay of Received Voice Packets

      Steady Stream of Packets

      Time

      Same Packet Stream After Congestion or Improper Queuing

      When a voice gateway receives a VoIP audio stream, it must compensate for the jitter it encounters. The mechanism that handles this function is the dejitter buffer (as mentioned previously in the "Dejitter Buffers" section), which must buffer the packets and then play them out in a steady stream to the DSPs, which convert them back to an analog audio stream.

      Packet Loss

      Packet loss causes voice clipping and skips. Packet loss can occur because of congested links, improper network QoS configuration, poor packet buffer management on the routers, routing problems, and other issues in both the WAN and LAN. If queues become saturated, VoIP packets might be dropped, resulting in effects such as clicks or lost words. Losses occur if the packets are received out of range of the dejitter buffer, in which case the packets are discarded.

      The industry-standard codec algorithms used in the Cisco DSP can use interpolation to correct for up to 30 ms of lost voice. The Cisco VoIP technology uses 20-ms samples of voice payload per VoIP packet. Therefore, only a single packet can be lost during any given time for the codec correction algorithms to be effective.

      KEY POINT

      For packet losses as small as one packet, the DSP interpolates the conversation with what it thinks the audio should be, and the packet loss is not audible.

      Echo

      In a voice telephone call, an echo occurs when callers hear their own words repeated.

      KEY An echo is the audible leak of the caller's voice into the receive path (the return path). POINT

      Echo is a function of delay and magnitude. The echo problem grows with the delay (the later the echo is heard) and the loudness (higher amplitude). When timed properly, an echo can be reassuring to the speaker. But if the echo exceeds approximately 25 milliseconds, it can be distracting and cause breaks in the conversation.

      KEY POINT

      Perceived echo most likely indicates a problem at the other end of the call. For example, if a person in Toronto hears an echo when talking to a person in Vancouver, the problem is likely to be at the Vancouver end.

      The following voice network elements can affect echo:

      ■ Hybrid transformers: A typical telephone is a two-wire device, whereas trunk connections are four-wire; a hybrid transformer is used to interface between these connections. Hybrid transformers are often prime culprits for signal leakage between analog transmit and receive paths, causing echo. Echo is usually caused by a mismatch in impedance from the four-wire network switch conversion to the two-wire local loop or an impedance mismatch in a PBX.

      ■ Telephones: An analog telephone terminal itself presents a load to the PBX. This load should be matched to the output impedance of the source device (the FXS port). Some (typically inexpensive) telephones are not matched to the FXS port's output impedance and are sources of echo. Headsets are particularly notorious for poor echo performance.

      When digital telephones are used, the point of digital-to-analog conversion occurs inside the telephone. Extending the digital transmission segments closer to the actual telephone decreases the potential for echo.

      NOTE The belief that adding voice gateways (routers) to a voice network creates echo is a common misconception. Digital segments of the network do not cause leaks; so, technically, voice gateways cannot be the source of echo. However, adding routers does add delay, which can make a previously imperceptible echo perceptible.

      An echo canceller, shown in Figure 8-30, can be placed in the network to improve the quality of telephone conversation. An echo canceller is a component of a voice gateway; it reduces the level of echo leaking from the receive path into the transmit path.

      Figure 8-30 Echo Cancellers Reduce the Echo Level a

      Echo Canceller Block Diagram

      Central Office

      Central Office

      Adaptive

      t

      Filter

      Echo cancellers are built into low-bit-rate codecs and operate on each DSP. By design, echo cancellers are limited by the total amount of time they wait for the reflected speech to be received. This is known as an echo trail or echo cancellation time and is usually between 16 and 32 milliseconds.

      To understand how an echo canceller works, assume that a person in Toronto is talking to a person in Vancouver. When the speech of the person in Toronto hits an impedance mismatch or other echo-causing environment, it bounces back to that person, who can hear the echo several milliseconds after speaking.

      Recall that the problem is at the other end of the call (called the tail circuit); in this example, the tail circuit is in Vancouver. To remove the echo from the line, the router in Toronto must keep an inverse image of the Toronto person's speech for a certain amount of time. This is called inverse speech. The echo canceller in the router listens for sound coming from the person in Vancouver and subtracts the inverse speech of the person in Toronto to remove any echo.

      The ITU-T defines an irritation zone of echo loudness and echo delay. A short echo (around 15 ms) does not have to be suppressed, whereas longer echo delays require strong echo suppression. Therefore, all networks that produce one-way time delays greater than 16 ms require echo cancellation. It is important to configure the appropriate echo cancellation time. If the echo cancellation time is set too low, callers still hear echo during the phone call. If the configured echo cancellation time is set too high, it takes longer for the echo canceller to converge and eliminate the echo.

      Attenuating the signal below the noise level can also eliminate echo.

      Continue reading here: Voice Coding and Compression

      Was this article helpful?

      +1 0