RTP Header
As stated in RFC 3550, the RTP header has a 12-octet mandatory part followed by an optional header extension. The header has the format illustrated in Figure 4-2.
Figure 4-2 RTP Header
Sequence Number
Time Stamp
Synchronization Source (SSRC) Identifier
Contributing Source (CSRC) Identifier(s)
Payload Header (Optional Depending on the Codec) Payload
The following sections describe the octets in the RTP header shown in Figure 4-2.
First Octet in the Header
The fields in this first octet of the RTP header are described as follows:
■ Version (V): 2 bits—This field identifies the version of RTP. The Version field is set to a value of 2 in most RTP implementations to denote the RTP profile defined in RFC 3551.
■ Padding (P): 1 bit—If the padding bit is set, the packet contains one or more additional padding octets at the end, which are not part of the payload. The last octet of the padding contains a count of how many padding octets should be ignored, including itself. Some encryption algorithms with fixed block sizes might need padding to carry several RTP packets in a lower-layer protocol data unit.
■ Extension (X): 1 bit—If the extension bit is set, exactly one header extension must follow the fixed header.
■ Contributing Source (CSRC) count (CC): 4 bits—The CSRC count contains the number of CSRC identifiers that follow the fixed header. CSRC is explained in much more detail later in this chapter in the section "Contributing Source Identifiers."
■ Marker (M): 1 bit—The interpretation of the marker is defined by the RTP profile in use. The M bit is intended to allow significant events such as frame boundaries to be marked in the packet stream. The M bit is helpful in video streams because it allows the endpoint to know that it has received the last packet of the frame so that it may display the full image. Without the M bit, the receiver would need to wait for one additional packet to detect a change to a new frame number.
Payload Type
RFC 3550 defines payload type as a 7-bit field that identifies the codec type and sample rate of media carried in the packet. When the endpoint or conference server receives an RTP packet, it uses the payload type to determine how to interpret the payload. The numeric value of the payload type may be predefined (called static payload types in the range of 0 to 96) or can be dynamically assigned during the capability negotiation between the conference server and the endpoint. There is one important distinction: For the static payload types, the clock rate is specified in the payload format. When using SIP signaling with dynamic payload types, the clock rate should be defined in the appropriate attribute line of the Session Description Protocol (SDP) offer. For example, G.711|>Law uses a static payload type of 0, and the clock rate is defined in RFC 3551. H.264 is a dynamic payload type, and the clock rate is 90 kHz, which is specified in the SDP as follows:
m = video rtp port number RTP/AVP 97 a = rtpmap:97 H.264/90000
Sequence Number
The sequence number is a two-octet field that identifies the order in which RTP packets were transmitted. The sequence number allows the receiver to detect packets that were dropped on the network and allows the receiver to handle out-of-order packets. The sender increments the sequence number by 1 for each RTP packet it sends. As defined in RFC 3550, the endpoint or conference server should choose the initial value of the sequence number at random, rather than starting from 0, to prevent known-value encryption attacks.
Time Stamp
The time stamp is a 32-bit integer that increments at the media-dependent rate. As stated in RFC 3550, the time stamp reflects the sampling instance of the first octet of the media data in the RTP packet. As with sequence numbers, senders should choose a random value for the time stamp of the first packet, rather than starting at 0. The time stamp will also wrap around to 0 if it exceeds its maximum 32-bit value. The sender must transmit packets according to the real-time rate of the media, which means that if the sender issues packets with a fixed number of media samples, the delay between RTP packet transmissions should also be fixed. Table 4-1 shows audio sampling rates and their packet sizes.
|
Sampling Rate |
Packet Size in RTP Time-Stamp Units |
|
Audio 10 milliseconds (ms) G.711 at 8000 Hz |
80 |
|
Audio 20 ms G.711 at 8000 Hz |
160 |
|
Audio 30 ms G.711 at 8000 Hz |
240 |
|
Video 30 frames per second at 90,000 Hz |
3000 (1130 * 90,000) |
|
Video 25 frames per second at 90,000 Hz |
3600 (1/25 * 90,000) |
In the case of MPEG bitstreams, which transmit frames out of order, the sender may transmit the RTP packets with out-of-order time stamps, but the sequence numbers will still increase. The receiver must reconstruct the data and play out the media accordingly based on the RTP time stamps. Also, note that a frame of video bitstream may be fragmented across multiple packets, which means that each packet will have the same RTP time stamp, but the sequence numbers will increase.
RTP packetization for audio codecs uses an RTP time-stamp clock that is the same as the sample clock, which means that the sampling clock increases by 1 for each sample. As a result, RTP time stamps for audio are essentially sample indexes. For example, an endpoint uses an audio codec with an 8000-Hz sample rate and an H.261 video codec. Because the sample rate of the audio stream is 8000 Hz, the RTP time stamp uses a sample clock of 8000 samples/second. If an audio stream packet has a size of 20 ms, the number of samples in that packet is 160, and therefore, the size of the packet is 160 RTP time-stamp units. H.261 uses an RTP sample clock of 90 kHz, which means a 29.97 FPS. An H.261 video stream will have a duration between frames of 33.37 ms, and the RTP time-stamp duration will equal 33.37 ms * 90,000 samples/second = 3003 RTP time-stamp units. The sender must assign RTP time stamps based on the absolute position in the source stream, which means that the RTP time-stamp sequence must account for packets not sent because of silence suppression at the sender.
Synchronization Source Identifier
The SSRC is a 32-bit field that serves as a unique identifier for an instance of an RTP stream. The originator of the RTP connection should choose this value at random. No two RTP streams within the same RTP session can have the same SSRC value. If the endpoint or the conference server changes the source IP address, the RTP packet stream must change to use a new SSRC value.
Contributing Source (CSRC) Identifiers
In a conference session, each endpoint transmits audio and video RTP packets to the audio mixer. The audio mixer then picks the top three or four speakers, mixes them, and sends the resulting output stream back to the endpoints. The output RTP packets should include the CSRC field, which is a list of the SSRC values of all participants selected for the mix. The audio mixer sets the CC bit in the first octet of the header to indicate the presence of a CSRC list. Many conferencing systems do not include this CSRC list because the endpoints are not conference-aware.
Payload Header
The RTP packetization method for the media is defined by a payload format definition, which is unique to each codec. A payload format might define a payload header, which resides in each RTP payload. The primary purpose of this header is to convey the state of the encoder to the destination. If the network drops packets, the receiver may use this state information to continue decoding the bitstream after the dropped packets. For instance, Figure 4-3 shows the format of the H.263 RTP packet.
Figure 4-3 H.263 RTP Packet
RTP Header
H.263 Payload Header H.263 Bitstream
The payload header identifies (among other things) the group of blocks (GOB), slice, or macroblock (MB) index for data at the start of the packet. It also indicates whether this packet is part of an I-frame.
Payload
The payload is the actual media data sent and received between endpoints and the conference server. The payload may contain multiple audio frames, which means that the decoder may need to parse the bitstream to determine whether the packet contains more than one frame. RTP packets generally do not contain more than one frame of video; instead, single frames of video typically fragment across multiple RTP packets.
The following shows an example of the RTP header data structure with one CSRC identifier:
Continue reading here: RTP Header Extensions
Was this article helpful?