(19)
(11) EP 0 912 063 B1

(12) EUROPEAN PATENT SPECIFICATION

(45) Mention of the grant of the patent:
13.02.2002 Bulletin 2002/07

(21) Application number: 98119908.6

(22) Date of filing: 21.10.1998
(51) International Patent Classification (IPC)7H04N 7/24, H04N 7/26, H04N 7/36, H04N 7/50

(54)

A method for computational graceful degradation in an audiovisual compression system

Verfahren zur rechnerisch graziösen Degradierung in einem audio-visuellen Kompressionssystem

Procédé pour dégradation gracieuse de calcul dans un système de compression audiovisuel


(84) Designated Contracting States:
DE FR GB

(30) Priority: 24.10.1997 SG 9703861

(43) Date of publication of application:
28.04.1999 Bulletin 1999/17

(60) Divisional application:
01118323.3 / 1154650
01128012.0

(73) Proprietor: MATSUSHITA ELECTRIC INDUSTRIAL CO., LTD.
Kadoma-shi, Osaka 571-8501 (JP)

(72) Inventor:
  • Tan, Thiow Keng
    08-506 Singapore 470601 (SG)

(74) Representative: Eisenführ, Speiser & Partner 
Martinistrasse 24
28195 Bremen
28195 Bremen (DE)


(56) References cited: : 
EP-A- 0 650 298
EP-A- 0 677 961
WO-A-99/12126
EP-A- 0 676 899
EP-A- 0 739 141
US-A- 5 402 146
   
       
    Note: Within nine months from the publication of the mention of the grant of the European patent, any person may give notice to the European Patent Office of opposition to the European patent granted. Notice of opposition shall be filed in a written reasoned statement. It shall not be deemed to have been filed until the opposition fee has been paid. (Art. 99(1) European Patent Convention).


    Description

    BACKGROUND OF THE INVENTION


    1. Field of the Invention



    [0001] The present invention relates to a method for computational graceful degradation in an audiovisual compression system. This invention is useful in a multimedia encoding and decoding environment where the computational demands for decoding a bitstream is not well defined. It is also useful in cases where channel capacity is limited and some form of quality of service guarantee is required. It is also useful for inter working between two video services of different resolutions.

    2. Description of the Related Art



    [0002] It is common in the case of software decoding to employ some form of graceful degradation when the system resources is not sufficient to fully decode all of the video bitstream. These degradation ranges from partial decoding of the picture elements to dropping of complete pictures. This is easy to implement in the case of a single video stream.

    [0003] In the proposed new ISO/IEC SC29/WG11 standard of MPEG-4, it is possible to send multiple Audiovisual, AV, objects. Therefore, the total complexity requirements no longer depend on one single stream but on multiple streams.

    [0004] In compression systems such as MPEG-1, MPEG-2 and MPEG-4, a high degree of temporal redundancy is removed by employing motion compensation. It is intuitive to see that successive pictures in a video sequence will contain very similar information. Only regions of the picture that are moving will change from picture to picture. Furthermore, these regions usually move as a unit with uniform motion. Motion compensation is a technique where the encoder and the decoder keep the reconstructed picture as a reference for the prediction of the current picture being encoded or decoded. The encoder mimics the decoder by implementing a local decoder loop. Thus, keeping the reconstructed picture synchronized between the encoder and decoder.

    [0005] The encoder performs a search for a block in the reconstructed picture that gives the closest match to the current block that is being encoded. It then computes the prediction difference between the motion compensated block and the current block being encoded. Since the motion compensated block is available in the encoder and the decoder, the encoder only needs to send the location of this block and the prediction difference to the decoder. The location of the block is commonly referred to as the motion vector. The prediction difference is commonly referred to as the motion compensated prediction error. These information requires less bits to send that the current block itself.

    [0006] In intra-picture coding, spatial redundancy may be removed in a similar way. The transform coefficients of the block can be predicted from the transform prediction of its neighboring blocks that have already being decoded.

    [0007] There are two major problems to be solved in this invention. The first is how to indicate the decoding complexity requirements of the current AV object. In the case where there are multiple AV objects, the systems decoder must decide how much resource should be given to a particular object and which object should have priority over another. In other words, how to model the complexity requirements of the system. A point to be noted here is that the complexity requirements of the decoder is dependent on the implementation of the decoder. An operation that is complex for one implementation may be simple for another implementation. Therefore, some form of implementation independent complexity measure is required.

    [0008] The second problem is how to reduce complexity requirements in the decoder. This deals with the method of reducing the complexity requirements of the decoding process while retaining as much of the information as possible. One biggest problem in graceful degradation is the problem of drift caused by errors in the motion compensation. When graceful degradation is employed the reconstructed picture is incomplete or noisy. These errors are propagated from picture to picture resulting in larger and larger errors. This noise propagation is referred to as drift.

    [0009] EP O 650 298 A1 discloses a method and an apparatus for coding or decoding a time-varying image. Therein the image signals are divided into a first image part which is an inner image portion and a second image part which is the image portion outside the first image part. The first and second image parts are further divided into given independent division units each formed by a plurality of pixels. When the coded information is about given division units belonging to the second image part, an identification code is added to the header of each division unit.

    SUMMARY OF THE INVENTION



    [0010] In order to solve the problems the following steps are taken in the present invention.

    [0011] The AV object encoder encodes the AV object in a manner that would allow different amounts of graceful degradation to be employed in the AV object decoder. Parameters relating to the computational complexity requirements of the AV objects are transmitted in the systems encoder. Implementation independent complexity measure is achieved by sending parameters that gives an indication of the operations that are required.

    [0012] At the systems decoder, estimates of the complexity required are made based on these parameters as well as the implementation methods being employed. The resource scheduler then allocates the appropriate amount of resources to the decoding of the different AV objects. In the AV object decoder, computational graceful degradation is employed when the resources are not sufficient to decode the AV object completely.

    [0013] According to the invention the above mentioned problems are solved by a method for encoding a visual object as claimed in claim 1.

    BRIEF DESCRIPTION OF THE DRAWINGS



    [0014] 

    Figure 1 is an overall block diagram of the present invention;

    Figure 2 shows a block diagram of encoder and decoder;

    Figure 3 illustrates the embodiment of the sub-region and the motion vector restriction;

    Figure 4 illustrates the embodiment for the pan-scan vectors and the sub-region dimensions;

    Figure 5 illustrates the second embodiment for the padding method of the motion compensated prediction at the sub-region boundary; and,

    Figure 6 illustrates the block diagram for the Complexity Estimator.


    DESCRIPTION OF THE PREFERRED EMBODIMENTS



    [0015] Figure 1 shows an overall system block diagram of the present invention. Encoder unit 110 encodes the video sequence to allow computational graceful degradation techniques. The output of encoder 110 is a coded representation of the video sequence that is applied to an encoding buffer 120. At the same time the video sequence and the coded representation are also applied to a complexity parameter encoder 130 where the parameters associated with the operation that are required for decoding is computed and encoded. These information together with the output of the encoding buffer 120 are passed to a System Encoder and Multiplexer unit 140 where a system-multiplexed stream is formed. The system-multiplexed stream is transmitted through a transmission media 150.

    [0016] A Demultiplexer and System Decoder unit 160 receives- the system-multiplexed stream, where the bitstream is demultiplexed into its respective elementary streams. The video elementary stream is passed to a Decoding Buffer 170, and complexity parameters are passed to a Scheduler and Complexity Estimator unit 180. From the Decoding Buffer 170, the video elementary stream is passed to a Decoder unit 190. The decoder 190 waits for the commands coming from the Scheduler unit 180 before decoding.

    [0017] The Complexity Estimator 180 gives the amount of decoder computational graceful degradation that is to be employed. Computational graceful degradation is achieved in the decoder by decoding only a sub-region of the complete picture that is deemed to contain the more important information. The encoder will have to prevent the encoder and decoder from drifting apart under these conditions. After decoding, the decoder unit 190 also feedback information to the Scheduler and Complexity Estimator 180 so that the information may be used to estimate the complexity of the next picture.

    [0018] The following is the embodiment of the various units illustrated in the above invention shown in Figure 1.

    [0019] Figure 2 is a block diagram of the encoder and decoder according to the present embodiment. The input picture to the encoder 110 is segmented into blocks for processing. Temporal redundancy is removed from the picture by subtracting the motion compensated picture of the previous picture from the current picture. The prediction difference is then transformed into the DCT domain in a DCT unit 111. The resulting DCT coefficients are then quantized in a Quantization unit 112. The quantized coefficients are then entropy coded in a Variable Length Coding (VLC) unit 113 to form the compressed output bitstream. The encoder 110 also has a local decoder loop comprising of an Inverse Quantization unit 114, an Inverse DCT unit 115, a Frame Storage 116, and a Motion Compensation unit 117. The local decoder loops mimics the decoder operations by inverse quantizing the coefficients and transforming it back into the spatial domain in the Inverse Quantization unit 114 and Inverse DCT unit 115. The output is then added to the output of the Motion Compensated unit 117 to form the reconstructed picture. This picture is stored in the Frame Storage 116 for motion compensation of the next picture.

    [0020] In this embodiment the encoder units of Motion Estimation unit 118 and Motion Compensation unit 117 are changed so that computational graceful degradation may be performed in conjunction with the motion compensation without causing drift.

    [0021] Figure 3 illustrates the present invention, according to which the picture is divided into two parts 220 and 210. The first part 220 is a sub-region that must be decoded in the decoder regardless of whether computational graceful degradation is employed or not. The second part 210 is the region outside of the sub-region, which may be discarded by the decoder when computational graceful degradation is employed.

    [0022] Figure 3 also show two blocks that are used for motion compensation. When motion compensation is performed on a block 250 that resides in the sub-region 220, the motion compensated prediction block must also come from within the sub-region 220 of the reference picture. In other words the motion vector 260 pointing out of the region is not allowed. This is referred to restricted motion vector. On the other hand, when a block 230 resides outside the sub-region 220, the motion compensated prediction block can come from anywhere in the reference picture. This is the same as where there is no sub-region.

    [0023] Figure 4 shows a method how to indicate the sub-region 220 within each picture. In order to specify the rectangular sub-region 220 for each picture the following parameters must be specified for each picture and be encoded in the picture header of the compress bitstream. In Figure 4, a picture 310 and the sub-region 220 is illustrated. The horizontal offset 330 of the left edge of sub-region 220 from the left edge of the picture, and the vertical offset 340 of the top edge of the sub-region 220 from the top edge of the picture are shown. These two parameters, referred to as the pan scan vectors, are used to indicate the location of the sub-region. The width 350 and the height 360 of the sub-region 220 are the second set of parameters that are required to specify the dimensions of the sub-region 220.

    [0024] In a second embodiment of this invention, the motion vector for a block in the sub-region need not be restricted. It is allowed to point out of the sub-region of the reference picture. However padding is required. This is illustrated in Figure 5 in which the picture 310 and the sub-region 220 are shown. The motion compensated prediction 430 is shown straddling the boundary of the sub-region 220. A portion 431 of the block residing outside of the sub-region 220 is not used for prediction and is padded by repeating the value of the pixel found at the edge of the sub-region 220. A portion 432 of the block residing in the sub-region 220 is used without any padding. A similar padding method is used for the rows and columns for blocks located at the vertical edge and horizontal edge, respectively.

    [0025] Like the first embodiment, the method according to the second embodiment would also enable computational graceful degradation method to discard the portion of the picture outside the sub-region 220 without causing the encoder and decoder to drift apart.

    [0026] Apart from motion compensation that may cause drift in inter blocks, intra blocks at the top and left boundary of the sub-region 220 are also restricted from using any blocks outside of the sub-region 220 for prediction. This is because in the computational graceful degraded decoder, these blocks would not be decoded and thus the prediction cannot be duplicated. This precludes the commonly used DC and AC coefficient prediction from being employed in the encoder.

    [0027] Figure 2 also illustrates a block diagram of a decoder 190. The embodiment of the decoder 190 employing computational graceful degradation is described here. The compressed bitstream is received from the transmission and is passed to a Variable Length Decoder unit 191 where the bitstream is decoded according to the syntax and entropy method used. The decoded information is then passed to the Computational Graceful Degradation Selector 192 where the decoded information belonging to the sub-region 220 is retained and the decoded information outside of the sub-region 220 is discarded. The retained information is then passed to an Inverse Quantization unit 193 where the DCT coefficients are recovered. The recovered coefficients are then passed to an Inverse DCT unit 194 where the coefficients are transformed back to the spatial domain. The motion compensated prediction is then added to form the reconstructed picture. The reconstructed picture is stored in a Frame Storage 195 where it is used for the prediction of the next picture. A Motion compensation unit 196 performs the motion compensation according to the same method employed in the encoder 110.

    [0028] In the first embodiment of the encoder where the motion vector is restricted, no additional modification is required in the decoder. In the second embodiment of the encoder where the motion vector is not restricted, the motion compensation method with padding described above in connection with Fig. 5 is used in the decoder. Finally, intra blocks at the top and left boundary of the sub-region 220 are also restricted from using any blocks outside of the sub-region 200 for prediction. This precludes the commonly used DC and AC coefficient prediction from being employed.

    [0029] In this embodiment the Complexity Parameter Encoder consist of a counting unit that counts the number of block decoding operations that are required. The block decoding operations are not basic arithmetic operations but rather a collection of operations that are performed on a block. A block decoding operation can be a block inverse quantization operation, a block inverse DCT operation, a block memory access or some other collection of operations that perform some decoding task on the block by block basis. The Complexity Parameter Encoder counts the number of blocks that require each set of operations and indicate these in the parameters. The reason block decoding operations are used instead of simple arithmetic operations is because different implementations may implement different operations more efficiently than others.

    [0030] There is also a difference in decoder architecture and different amounts of hardware and software solutions that makes the use of raw processing power and memory access measures unreliable to indicate the complexity requirements. However, if the operations required are indicated by parameters that counts the number of block decoding operations necessary, the decoder can estimate the complexity. This is because the decoder knows the amount of operations required for each of the block decoding operations in its own implementation.

    [0031] In the embodiment of the System Encoder and Multiplexer, the elementary bitstream are packetized and multiplexed for transmission. The information associated with the complexity parameters is also multiplexed into the bitstream. This information is inserted into the header of the packets. Decoders that do not require such information may simply skip over this information. Decoders that require such information can decode this information and interpret them to estimate the complexity requirements.

    [0032] In this embodiment the encoder inserts the information in the form of a descriptor in the header of the packet. The descriptor contains an ID to indicate the type of descriptor it is followed by the total number of bytes contained in the descriptor. The rest of the descriptor contains the parameter for each of the block decoding operations. Optionally the descriptor may also carry some user defined parameters that are not defined earlier.

    [0033] In the Scheduler and Complexity Estimator 180 in Figure 1, the time it takes for decoding all the audiovisual objects is computed based on the parameters found in the descriptor as well as the feedback information from the decoder.

    [0034] An embodiment of the Complexity Estimator 180 is shown in Figure 6. The block decoding operation parameters 181a, 181b and 181c are passed into the complexity estimator 183 after being pre-multiplied with weightings 182a, 182b and 182c, respectively. The complexity estimator 183 then estimates the complexity of the picture to be decoder and passes the estimated complexity 184 to the decoder 190. After decoding the picture the decoder 190 returns the actual complexity 185 of the picture. An error 186 in the complexity estimation is obtained by taking a difference between the estimated complexity 184 and the actual complexity 185 of the picture. The error 186 is then passed into the feedback gain unit 187 where the corrections 188a, 188b and 188c to the weightings are found. The weights are then modified by these corrections and the process of estimating the complexity of the next picture continues.

    [0035] The effect of this invention is that the need for implementations that can handle the worst case is no longer necessary. Using the indications of computational complexities and the computational graceful degradation methods simpler decoders can be implemented. The decoder would have the capabilities to decode most of the sequences, but if it encounters some more demanding sequences, it can degrade the quality and resolution of the decoder output in order to decode the bitstream.

    [0036] This invention is also useful for inter working of services that have different resolutions and/or different formats. The sub-region can be decoder by the decoder of lower resolutions where as the decoder of higher resolutions can decode the full picture. One example is the inter working between 16:9 and 4:3 aspect ratio decoders.


    Claims

    1. A method for encoding a visual object, comprising:

    encoding the visual object to obtain compressed coded data, wherein the compressed coded data is obtained through at least one encoding operation performed on a block by block basis;

    generating a descriptor indicating at least a numeric number of required decoding operations; and

    multiplexing the descriptor with the compressed coded data.


     
    2. A method of encoding a visual object according to claim 1, wherein the decoding operations include at least one of:

    a block entropy decoding operation;

    a block motion compensation operation;

    a block inverse quantization operation;

    a block transforming operation; and

    a block addition operation.


     
    3. A method of encoding a visual object according to claim 1, wherein the descriptor further comprises:

    an identification number identifying a descriptor type;

    a length field indicating a size of the descriptor; and

    a plurality of block decoding parameters.


     


    Ansprüche

    1. Verfahren zur Kodierung eines visuellen Objekts, aufweisend:

    Kodierung des visuellen Objekts, um komprimierte kodierte Daten zu erhalten, wobei die komprimierten kodierten Daten durch wenigstens eine Kodierungsoperation, die auf einer blockweisen Basis durchgeführt werden, erhalten werden;

    Erzeugung eines Descriptors, der wenigstens eine numerische Anzahl von erforderlichen Dekodierungsoperationen anzeigt; und

    Multiplexen des Descriptors mit den komprimierten kodierten Daten.


     
    2. Verfahren zur Kodierung eines visuellen Objekts gemäß Anspruch 1,
    wobei die Dekodierungsoperationen wenigstens eine der folgenden Operationen einschließen:

    eine Block-Entropie-Dekodierungsoperation;

    eine Block-Bewegungs-Kompensationsoperation;

    eine Block-Inverse-Quantisierungsoperation;

    eine Block-Transformierungsoperation; und

    eine Block-Additionsoperation.


     
    3. Verfahren zur Kodierung eines visuellen Objekts gemäß Anspruch 1,
    wobei der Descriptor femer aufweist:

    eine Identifikationsnummer, die einen Descriptortyp anzeigt;

    ein Längenfeld, das eine Größe des Descriptors anzeigt; und

    eine Vielzahl von Blockdekodierungsparametern.


     


    Revendications

    1. Procédé de codage d'un objet visuel, comprenant :

    le codage de l'objet visuel de manière à obtenir des données codées compressées, dans lequel on obtient les données compressées codées par au moins une opération de codage exécutée sur une base bloc par bloc ;

    la génération d'un descripteur indiquant au moins un nombre numérique d'opérations de décodage nécessaires ; et

    le multiplexage du descripteur avec les données compressées codées.


     
    2. Procédé de codage d'un objet visuel selon la revendication 1, dans lequel les opérations de codage comprennent au moins l'une des opérations ci-dessous :

    une opération de décodage d'entropie de blocs ;

    une opération de compensation de mouvement de blocs ;

    une opération de quantification inverse de blocs ;

    une opération de transformation de blocs ;

    une opération d'addition de blocs.


     
    3. Procédé de codage d'un objet visuel selon la revendication 1, dans lequel le descripteur comprend en outre :

    un numéro d'identification identifiant le type de descripteur ;

    un champ de longueur indiquant la taille du descripteur ; et

    un ensemble de paramètres de décodage de blocs.


     




    Drawing