Voice Coding and Compression

Voice communication over IP relies on voice that is coded and encapsulated into IP packets. This section provides an overview of the various codecs used in voice networks.

NOTE The term codec can have the following two meanings:

■ A coder-decoder: An integrated circuit device that typically uses PCM to transform analog signals into a digital bit stream and digital signals back into analog signals.

■ A software algorithm: Used to compress and decompress speech or audio signals in VoIP, Frame Relay, and ATM.

Coding and Compression Algorithms

KEY POINT

A codec is a device or software that encodes (and decodes) a signal into digital data stream.

Each codec provides a certain quality of speech. Advances in technology have greatly improved the quality of compressed voice and have resulted in a variety of coding and compression algorithms:

■ PCM: The toll quality voice expected from the PSTN. PCM runs at 64 kbps and provides no compression, and therefore no opportunity for bandwidth savings.

■ Adaptive Differential Pulse Code Modulation (ADPCM): Provides three different levels of compression. Some fidelity is lost as compression increases. Depending on the traffic mix, cost savings generally run at 25 percent for 32-kbps ADPCM, 30 percent for 24-kbps ADPCM, and 35 percent for 16-kbps ADPCM.

■ Low-Delay Code Excited Linear Prediction Compression (LD-CELP): This algorithm models the human voice. Depending on the traffic mix, cost savings can be up to 35 percent for 16-kbps LD-CELP.

■ Conjugate Structure Algebraic Code Excited Linear Prediction Compression (CS-ACELP): Provides eight times the bandwidth savings over PCM. CS-ACELP is a more recently developed algorithm modeled after the human voice and delivers quality that is comparable to LD-CELP and 32-kbps ADPCM. Cost savings are approximately 40 percent for 8-kbps CS-ACELP.

■ Code Excited Linear Prediction Compression (CELP): Provides huge bandwidth savings over PCM. Cost savings can be up to 50 percent for 5.3-kbps CELP.

The following section details voice coding standards based on these algorithms.

Voice Coding Standards (Codecs)

The ITU has defined a series of standards for voice coding and compression:

■ G.711: Uses the 64-kbps PCM voice coding technique. G.711-encoded voice is already in the correct format for digital voice delivery in the PSTN or through PBXs. Most Cisco implementations use G.711 on LAN links because of its high quality, approaching toll quality.

■ G.726/G.727: G.726 uses the ADPCM coding at 40, 32, 24, and 16 kbps. ADPCM voice can be interchanged between packet voice and public telephone or PBX networks if the latter has ADPCM capability. G.727 is a specialized version of G.726; it includes the same bandwidths.

■ G.728: Uses the LD-CELP voice compression, which requires only 16 kbps of bandwidth. LD-CELP voice coding must be transcoded to a PCM-based coding before delivering to the PSTN.

■ G.729: Uses the CS-ACELP compression, which enables voice to be coded into 8-kbps streams. This standard has various forms, all of which provide speech quality similar to that of 32-kbps ADPCM.

For example, in G.729a, the basic algorithm was optimized to reduce the computation requirements. In G.729b, voice activity detection (VAD) and comfort noise generation were added. G.729ab provides an optimized version of G.729b requiring less computation.

■ G.723.1: Uses a dual-rate coder for compressing speech at very low bit rates. Two bit rates are associated with this standard: 5.3 kbps using algebraic code-excited linear prediction (ACELP) and 6.3 kbps using Multipulse Maximum Likelihood Quantization (MPMLQ).

Sound Quality

Each codec provides a certain quality of speech. The perceived quality of transmitted speech depends on a listener's subjective response.

The mean opinion score (MOS) is a common benchmark used to specify the quality of sound produced by specific codecs. To determine the MOS, a wide range of listeners judge the quality of a voice sample corresponding to a particular codec on a scale of 1 (bad) to 5 (excellent). The scores are averaged to provide the MOS for that sample. Table 8-4 shows the relationship between codecs and MOS scores; notice that MOS decreases with increased codec complexity.

Table 8-4 Voice Coding and Compression Results

Algorithm

ITU Standard

Data Rate1

MOS Score

PCM

G.711

64 kbps

4.1

ADPCM

G.726/G.727

16/24/32/40 kbps

3.85 or less

LD-CELP

G.728

16 kbps

3.61

CS-ACELP

G.729

8 kbps

3.92

ACELP/MPMLQ

G.723.1

6.3/5.3 kbps

3.9/3.65

1 Data rates shown are for digitized speech only. In addition to this coded digital stream payload, RTP, UDP, IP, and Layer 2 headers are needed.

1 Data rates shown are for digitized speech only. In addition to this coded digital stream payload, RTP, UDP, IP, and Layer 2 headers are needed.

KEY G.729 is the recommended voice codec for most WAN networks (that do not do multiple POINT encodings) because of its relatively low bandwidth requirements and high MOS.

The Perceptual Speech Quality Measurement (PSQM) is a newer, more objective measurement that is overtaking MOS scores as the industry quality measurement of choice for coding algorithms. PSQM is specified in ITU standard P.861. PSQM provides a rating on a scale of 0 to 6.5, where 0 is best and 6.5 is worst. PSQM is implemented in test equipment and monitoring systems. It compares the transmitted speech to the original input to produce a PSQM score for a test voice call over a particular packet network. Some PSQM test equipment converts the 0-to-6.5 scale to a 0-to-5 scale to correlate to MOS.

Codec Complexity, DSPs, and Voice Calls

A DSP is a hardware component that converts information from telephony-based protocols to packet-based protocols (such as IP).

KEY POINT

A codec is a technology for compressing and decompressing data; it is implemented in DSPs. Some codec compression techniques require more processing power than others.

KEY POINT

Codec complexity is divided into low, medium, and high complexity. The difference between the complexities of the codecs is the CPU utilization necessary to process the codec algorithm and the number of voice channels that a single DSP can support.

The number of calls supported depends on the DSP and the complexity of the codec used. For example, as illustrated in Table 8-5, the Cisco High-Density Packet Voice/Fax DSP Module (AS54-PVDM2-64) for Cisco voice gateways provides high-density voice connectivity supporting 24 to 64 channels (calls), depending on codec compression complexity.

Table 8-5 Code Complexity and Calls per DSP on the AS54-PVDM2-64 Voice/Fax DSP Module

Low Complexity (Maximum 64 Calls)

Medium Complexity (Maximum 32 Calls)

High Complexity (Maximum 24 Calls)

G.711 a-law

G.729a

G.723.1: 5.3/6.3 kbps

G.711 Mu-law

G.729ab

G.723.1a: 5.3/6.3 kbps

Fax Passthrough

G.726: 16/24/32 kbps

G.728

Modem Passthrough

T.38 fax relay

Modem relay

Clear-channel codec

Cisco Fax Relay

Adaptive multirate narrow band: 4.75, 5.15, 5.9, 6.7, 7.4, 7.95, 10.2, and 12.2 kbps, and silence insertion descriptor

Continue reading here: Bandwidth Considerations

Was this article helpful?

0 0