FIELD OF THE INVENTION
[0001] The present invention relates to audio coding.
BACKGROUND OF THE INVENTION
[0003] A basic parametric stereo coder may use inter-channel level differences (ILD) as
a cue needed for generating the stereo signal from the mono down-mix audio signal.
More sophisticated coders may also use the inter-channel coherence (ICC), which may
represent a degree of similarity between the audio channel signals, i.e. audio channels.
Furthermore, when coding binaural stereo signals e.g. for 3D audio or headphone based
surround rendering, also an inter-channel phase difference (IPD) may play a role to
reproduce phase/delay differences between the channels.
[0004] The synthesis of ICC cues may be relevant for most audio and music contents to re-generate
ambience, stereo reverb, source width, and other perceptions related to spatial impression
as described in
J. Blauert, Spatial Hearing: The Psychophysics of Human Sound Localization, The MIT
Press, Cambridge, Massachusetts, USA, 1997. Coherence synthesis may be implemented by using de-correlators in frequency domain
as described in
E. Schuijers, W. Oomen, B. den Brinker, and J. Breebaart, "Advances in parametric
coding for high-quality audio," in Preprint 114th Conv. Aud. Eng. Soc., Mar. 2003. However, the known synthesis approaches for synthesizing multi-channel audio signals
may suffer from an increased complexity. Furthermore, the use of ICC parameters, e.g.
in addition to other parameters, such as inter-channel level differences (ICLDs) and
inter-channel phase differences (ICPDs), may increase a bitrate overhead.
[0005] US 2005/180579 A1 discloses a scheme for stereo and multi-channel synthesis of inter-channel correlation
(ICC) (normalized cross-correlation) cues for parametric stereo and multi-channel
coding. The scheme synthesizes ICC cues such that they approximate those of the original.
For that purpose, diffuse audio channels are generated and mixed with the transmitted
combined (e.g., sum) signal(s). The diffuse audio channels are preferably generated
using relatively long filters with exponentially decaying Gaussian impulse responses.
Such impulse responses generate diffuse sound similar to late reverberation. An alternative
implementation for reduced computational complexity is proposed, where inter-channel
level difference (ICLD), inter-channel time difference (ICTD), and ICC synthesis are
all carried out in the domain of a single short-time Fourier transform (STFT), including
the filtering for diffuse sound generation.
[0006] WO 2003/219130 A1 discloses a combination device that includes: a detection unit that detects active
coded bitstreams that are effective coded bitstreams from a plurality of coded bitstreams
within a predetermined time period; a first combining unit that combines, from a plurality
of downmix sub-streams included in the coded bitstreams, only downmix sub-streams
included in the active coded bitstreams so as to generate a combined downmix sub-stream;
and a second combining unit that combines, from a plurality of parameter sub-streams
included in the coded bitstreams, only parameter sub-streams included in the active
coded bitstreams so as to generate a combined parameter sub-stream.
[0007] US 2003/219130 A1 discloses an auditory scene synthesized from a mono audio signal by modifying, for
each critical band, an auditory scene parameter (e.g., an inter-aural level difference
(ILD) and/or an inter-aural time difference (ITD)) for each sub-band within the critical
band, where the modification is based on an average estimated coherence for the critical
band. The coherence-based modification produces auditory scenes having objects whose
widths more accurately match the widths of the objects in the original input auditory
scene.
SUMMARY OF THE INVENTION
[0008] A goal to be achieved by the present invention is to reduce complexity of a parametric
coding scheme. This goal is achieved by the features of the independent claims. Further
embodiments are apparent from the description, the drawings and from the dependent
claims.
[0009] The invention is based on the finding that combining parametric encoding parameters
such as ICC parameters may reduce bit rate required for representing the parameters
and thus may reduce complexity of the resulting parametric encoding scheme. The combined
encoding parameters may be applied e.g. only to a certain frequency region in order
to improve an audio quality for e.g. speech whereby the complexity and the memory
requirements may further be reduced.
[0010] The invention is described in the independent claim 1 and 6. Further embodiments
are defined in the dependent claims 2-5
[0011] According to a first implementation form, the first and second encoding parameter
may be an inter-channel phase difference.
[0012] According to a second implementation form, the first and second encoding parameter
may be an inter-channel coherence.
[0013] According to a third implementation form, the first and second encoding parameter
may be an inter-channel intensity difference.
[0014] According to a fourth implementation form, the first and second encoding parameter
may be an inter-channel level difference.
[0015] According to a firth implementation form, the parameter generator is configured to
generate the first encoding parameter and the second encoding parameter upon a basis
of a multiplication of values of the first transformed audio signal and of the second
transformed audio signal.
[0016] According to a sixth implementation form, the parameter combiner is configured to
determine a weighted average of the first encoding parameter and the second encoding
parameter using powers of the a first transformed audio signal and the second transformed
signal at the certain frequency as weights to obtain the combined encoding parameter.
[0017] According to a seventh implementation form, the parameter combiner is configured
to determine a weighted average of the first encoding parameter and the second encoding
parameter using a frequency-dependent weight to obtain the combined encoding parameter.
[0018] According to an eighth implementation form, the parameter generator is configured
to generate a plurality of encoding parameters from the first transformed audio signal
and from the second transformed audio signal at a plurality of frequencies, and wherein
the parameter combiner is configured to combine the plurality of the encoding parameters
to obtain the combined encoding parameter.
[0019] According to a ninth implementation form, the parametric encoder further comprises
a signal combiner for combining the first transformed audio signal and the second
transformed audio signal to obtain a down-mix signal.
[0020] According to a tenth implementation form, the parametric encoder further comprises
an inverse transformer for inversely transforming a combination of the first transformed
audio signal and the second transformed audio signal to obtain a down-mix audio signal.
[0021] According to a second aspect the invention relates to a method for parametrically
encoding a multi-channel audio signal having a first audio signal and a second audio
signal, the method having transforming the first audio signal into frequency domain
to obtain a first transformed audio signal, and transforming the second audio signal
into frequency domain to obtain a second transformed audio signal, generating a first
encoding parameter from the first transformed audio signal and from the second transformed
audio signal at a first frequency, and generating a second encoding parameter from
the first transformed audio signal and from the second transformed audio signal at
a second frequency, and combining the first encoding parameter and the second encoding
parameter to obtain a combined encoding parameter.
[0022] Further method steps of implementation forms according to the second aspect are directly
derivable from the functionality of the parametric encoder according to the first
aspect.
BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Further embodiments of the invention will be described with reference to the following
drawings, in which:
Fig. 1 shows a block diagram of a parametric encoder according to an implementation
form;
Fig. 2 shows a block diagram of a parametric decoder according to an implementation
form;
Fig. 3 shows a diagram of a method for parametrically encoding according to an implementation
form; and
Fig. 4 shows a diagram of a method for parametrically decoding according to an implementation
form.
DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] Fig. 1 shows a diagram of a parametric encoder for encoding a multi-channel audio
signal having a first audio signal, x1, and a second audio signal, x2, according to
an implementation form. The parametric encoder comprises a transformer 101 for transforming
the first audio signal into frequency domain to obtain a first transformed audio signal,
and for transforming the second audio signal into frequency domain to obtain a second
transformed audio signal. The transformer 101 may comprise a first transformer 103
for transforming the first audio signal, and a second transformer 105 for transforming
the second audio signal. The transformer 101 and/or the transformers 103, 105 may
be Fourier transformers, by way of example. The first and second transformed audio
signals are provided to a parameter generator 107 for generating a first encoding
parameter from the first transformed signal and from the second transformed audio
signal at a first frequency, e.g. at the i-th frequency or in the i-th band. The i-th
band or "band i" (see also in Fig. 1) refer to a frequency band i at or in which the
parameter generator 107 generates the respective encoding parameter from the first
and second transformed signal, and is also referred to as parameter band i. The parameter
generator 107 is further configured to generate a second encoding parameter from the
first and second transformed audio signal at a second frequency or in a second band.
The first and the second encoding parameters are provided to the parameter combiner
109 which combines the first encoding parameter and the second encoding parameter
to obtain a combined encoding parameter according to a principle described herein.
However, the parameter combiner 109 may separately obtain encoding parameters for
different parameter bands.
[0025] With reference to Fig. 1 and to ICC parameters forming an embodiment of encoding
parameters, the e.g. stereo input audio channels x1 and x2 are converted to a plurality
of sub-bands or parameter bands. In all or in a subset of the parameter bands the
corresponding ICC parameters may be estimated. One or more ICC parameter combining
processes, e.g. one of the processes according to the equations (1)-(4), may be applied
to the ICC parameters of all or subsets of parameter bands, to compute the combined
ICC parameters. At least one combined ICC parameter may be put into a bit stream 111
or transmitted to an audio decoder which is not depicted in Fig. 1.
[0026] The following embodiments are exemplarily described with respect to ICC forming an
embodiment of an encoding parameter. It is, however, to be understood, that the encoding
parameter may be any encoding parameter or of any encoding parameter type used for
parametric encoding, e.g. inter-channel phase difference or inter-channel intensity
difference or an inter-channel level difference or the like, and that the encoder
may be adapted to produce combined encoding parameters according one, some or all
of the aforementioned encoding parameter types and to include combined encoding parameters
of different types in the bitstream 111 as side information.
[0027] The parametric encoder of Fig, 1 may form a parametric stereo encoder which estimates
in parameter bands perceptual spatial cue parameters, such as ICLD, ICPD, and/or ICC.
If the parameter band index is i, then the estimated parameters in that band are denoted
ICLD(i), ICPD(i), and ICC(i). The left and right signal power in a parameter band
are denoted P1(i) and P2(i), respectively.
[0028] In this regard, one or more combined ICC parameters may be computed, e.g. as an average

where I is the set of indices of parameter bands of which ICC are used to compute
the combined ICC parameter and NI is the number of indices in the set I.
[0029] Another way of computing a combined ICC parameter is to use a weighted average, i.e.

wherein P1(i) denotes a signal power of the first audio channel signal in the i-th
band, and wherein P2(i) denotes a signal power of the second audio channel signal
in the i-th band.
[0030] Additionally, different parameter bands (frequencies) can be weighted differently,
when computing the combined ICC:

where
gi is a weight given to frequency (parameter band)
i.
[0031] Another example is to use an average considering not power but frequency weighting:
Additionally, different parameter bands (frequencies) can be weighted differently,
when computing the combined ICC:

[0032] A single full-band ICC performs surprisingly well. In this case, the combined ICC
is computed using all parameter bands, i.e. I contains all parameter band indices.
[0033] According to some implementation forms, the speech quality may be improved by only
using ICC in a limited frequency range. When only parameter bands between 500 Hz and
1.5 kHz are used for generating ICC parameters, less artifacts occur. In this case,
the combined ICC is computed using only parameter bands between 500 Hz and 1.5 kHz,
i.e. I contains only those indices.
[0034] According to some implementation forms, the parametric encoder shown in Fig. 1 may
estimate one or more combined ICC parameters by:
- (a) Combining the ICC parameters from a plurality of parameter bands to a combined
ICC parameter,
- (b) Putting combined ICC parameters into a bit stream, and
- (c) Outputting the bit stream.
[0035] Fig. 2 shows a block diagram of a parametric encoder for decoding a down-mix audio
signal according to an implementation form. The down-mix signal may be provided by
the parametric encoder as shown e.g. in Fig. 1. The parametric decoder comprises a
transformer 201 for transforming the down-mix audio signal, as, to obtain a transformed
down-mix audio signal having a certain frequency, e.g. an i-th frequency of a plurality
of frequencies, or correspondingly a certain band, e.g. an i-th band of a plurality
of bands. The parametric decoder further comprises a provider 203 for providing a
frequency-specific encoding parameter associated with the certain frequency. The frequency-specific
encoding parameter may be derived from the combined encoding parameter. However, the
frequency-specific encoding parameter may correspond to the combined encoding parameter.
The parametric decoder further comprises an audio synthesizer 205, e.g. a stereo synthesizer,
for synthesizing a first audio signal and a second audio signal at the certain frequency
or in the certain band from the transformed down-mix audio signal provided by the
transformer 201 using the frequency-specific encoding parameter as provided by the
provider 203.
[0036] According to some implementation forms, the transformer 201 may be a Fourier transformer,
wherein the audio synthesizer may synthesize the first and the second audio signal
in frequency domain. Thus, the output signal provided by the synthesizer 205 may correspond
to the first and second audio channel signal. However, according to some implementation
forms, the parametric encoder may further comprise an inverse transformer 207 for
inversely transforming the first and second audio channel signal in time domain in
order to obtain a first and second audio channel signal, x1 and x2, in time domain.
[0037] The parametric decoder shown in Fig. 1 uses for all parameter bands, or a subset
J thereof, the combined ICC parameter or a modified version thereof. One can use the
same subset of parameter bands at the decoder as at the encoder,
i.e. J=I, or a different subset.
[0038] With respect to Fig. 2, the parametric decoder may receive the down-mix signal s
and the stereo parameters, i.e. encoding parameters, amongst which at least one combined
ICC parameter may be received. For at least one parameter band an ICC parameter derived
from combined ICC parameters is used. To some bands no ICC synthesis may be applied.
[0039] According to an implementation form, for the "Combined ICC to Band ICC Conversion"
when one combined ICC is used, combined ICC is applied to all parameter bands. Or,
if the combined ICC was estimated only for a subset of bands, J, then the decoder
may apply the combined ICC to all bands, to the same subset, or to another subset,
e.g. a subset of the same subset.
[0040] If two combined ICC are used, representing two distinct frequency regions of the
audio signal. The decoder may apply the combined ICCs to parameter bands corresponding
to the frequency regions from which the combined ICCs were estimated.
[0041] Fig. 3 shows a diagram of a method for parametrically encoding a multi-channel audio
signal having the first and the second audio signal as mentioned above. The method
comprises transforming 301 the first and second audio signal into frequency domain
to obtain a first and second transformed audio signal, generating 303 a first encoding
parameter from the first and second transformed audio signal at a first frequency,
and a second encoding parameter from the first and second transformed audio signal
at a second frequency, and combining 305 the first and second encoding parameter to
obtain a combined encoding parameter. By way of example, the method depicted in Fig.
3 may be performed by the parametric encoder as shown in Fig. 1.
[0042] Fig. 4 shows a block diagram of a method for parametrically decoding a down-mix audio
signal upon a basis of a combined encoding parameter. The down-mix audio signal may
represent a combination, e.g. a superposition, of a first and second audio signal.
The combined encoding parameter may have features as described above.
[0043] The method comprises transforming 401 the down-mix audio signal to obtain a transformed
down-mix audio signal having a certain frequency, providing 403 a frequency-specific
encoding parameter associated with the certain frequency upon the basis of the combined
encoding parameter, according to the principle described herein, and synthesizing
405 the first and second audio signal at the certain frequency from the transformed
down-mix audio signal and from the frequency-specific encoding parameter.
[0044] According to some implementation forms, the method depicted in Fig. 4 may be performed
by the parametric decoder as shown in Fig. 2.
[0045] According to some implementation forms, the parametric decoder shown in Fig. 2 may
be a parametric stereo decoder adapted for
- (a) Receiving one or more combined ICC parameters, and
- (b) Using for at least one parameter band an ICC parameter related to the received
combined ICC parameters.
1. Parametric encoder for encoding a multi-channel audio signal having a first audio
signal and a second audio signal, the parametric encoder having:
a transformer (101) for transforming the first audio signal into frequency domain
to obtain a first transformed audio signal, and for transforming the second audio
signal into frequency domain to obtain a second transformed audio signal;
a parameter generator (107) for generating a first encoding parameter, X(i), from
the first transformed audio signal and from the second transformed audio signal at
a first frequency band i, and for generating a second encoding parameter, X(j), from
the first transformed audio signal and from the second transformed audio signal at
a second frequency band j; and
a parameter combiner (109) for combining the first encoding parameter and the second
encoding parameter to obtain a combined encoding parameter, X, according to the formula

wherein parameter I denotes a set of indices of frequency bands, parameter gi is a weight given to frequency band i, parameter P1(i) denotes a signal power of the first audio signal in the i-th frequency band, parameter
P2(i) denotes a signal power of the second audio signal in the i-th frequency band,
and wherein the first encoding parameter, X(i), and the second encoding parameter,
X(j), is an inter-channel phase difference or an inter-channel coherence or an inter-channel
intensity difference or an inter-channel level difference.
2. The parametric encoder of claim 1, wherein the parameter generator (107) is configured
to generate the first encoding parameter and the second encoding parameter upon a
basis of a multiplication of values of the first transformed audio signal and of the
second transformed audio signal.
3. The parametric encoder of any of claims 1 to 2, wherein the parameter generator (107)
is configured to generate a plurality of encoding parameters from the first transformed
audio signal and from the second transformed audio signal at a plurality of frequency
bands and wherein the parameter combiner is configured to combine the plurality of
the encoding parameters to obtain the combined encoding parameter.
4. The parametric encoder of any of claims 1 to 3, being further configured to combine
the first transformed audio signal and the second transformed audio signal to obtain
a down-mix signal.
5. The parametric encoder of any of claims 1 to 4, further comprising an inverse transformer
for inversely transforming a combination of the first transformed audio signal and
the second transformed audio signal to obtain a down-mix audio signal.
6. Method for parametrically encoding a multi-channel audio signal having a first audio
signal and a second audio signal, wherein the method is configured to operate a parametric
encoder according to the preceding claims.
1. Parametrischer Codierer zum Codieren eines Mehrkanal-Audiosignals, das ein erstes
Audiosignal und ein zweites Audiosignal aufweist, wobei der parametrische Codierer
Folgendes aufweist:
eine Transformationseinrichtung (101) zum Transformieren des ersten Audiosignals in
den Frequenzbereich, um ein erstes transformiertes Audiosignal zu erhalten, und zum
Transformieren des zweiten Audiosignals in den Frequenzbereich, um ein zweites transformiertes
Audiosignal zu erhalten;
einen Parametergenerator (107) zum Erzeugen eines ersten Codierungsparameters, X(i),
aus dem ersten transformierten Audiosignal und aus dem zweiten transformierten Audiosignal
in einem ersten Frequenzband i und zum Erzeugen eines zweiten Codierungsparameters,
X(j), aus dem ersten transformierten Audiosignal und aus dem zweiten transformierten
Audiosignal in einem zweiten Frequenzband j; und
einen Parameterkombinierer (109) zum Kombinieren des ersten Codierungsparameters und
des zweiten Codierungsparameters, um einen kombinierten Codierungsparameter, X, zu
erhalten, gemäß der folgenden Formel

wobei der Parameter I eine Menge von Indizes der Frequenzbänder bezeichnet, der Parameter
gi ein dem Frequenzband i gegebenes Gewicht ist, der Parameter P1(i) eine Signalleistung des ersten Audiosignals in dem i-ten Frequenzband bezeichnet
und der Parameter P2(i) eine Signalleistung des zweiten Audiosignals in dem i-ten Frequenzband bezeichnet,
und
wobei der erste Codierungsparameter, X(i), und der zweite Codierungsparameter, X(j),
ein Zwischenkanal-Phasenunterschied oder eine Zwischenkanal-Kohärenz oder ein Zwischenkanal-Intensitätsunterschied
oder ein Zwischenkanal-Pegelunterschied sind.
2. Parametrischer Codierer nach Anspruch 1, wobei der Parametergenerator (107) konfiguriert
ist, den ersten Codierungsparameter und den zweiten Codierungsparameter auf einer
Grundlage einer Multiplikation von Werten des ersten transformierten Audiosignals
und des zweiten transformierten Audiosignals zu erzeugen.
3. Parametrischer Codierer nach einem der Ansprüche 1 bis 2, wobei der Parametergenerator
(107) konfiguriert ist, aus dem ersten transformierten Audiosignal und aus dem zweiten
transformierten Audiosignal in mehreren Frequenzbändern mehrere Codierungsparameter
zu erzeugen, und wobei der Parameterkombinierer konfiguriert ist, die mehreren Codierungsparameter
zu kombinieren, um den kombinierten Codierungsparameter zu erhalten.
4. Parametrischer Codierer nach einem der Ansprüche 1 bis 3, der ferner konfiguriert
ist, das erste transformierte Audiosignal und das zweite transformierte Audiosignal
zu kombinieren, um ein Abwärtsmischsignal zu erhalten.
5. Parametrischer Codierer nach einem der Ansprüche 1 bis 4, der ferner eine Einrichtung
für die inverse Transformation zum inversen Transformieren einer Kombination aus dem
ersten transformierten Audiosignal und dem zweiten transformierten Audiosignal, um
ein Abwärtsmisch-Audiosignal zu erhalten, umfasst.
6. Verfahren zum parametrischen Codieren eines Mehrkanal-Audiosignals, das ein erstes
Audiosignal und ein zweites Audiosignal aufweist, wobei das Verfahren konfiguriert
ist, einen parametrischen Codierer nach den vorhergehenden Ansprüchen zu betreiben.
1. Codeur paramétrique destiné à coder un signal audio multi-canaux ayant un premier
signal audio et un second signal audio, le codeur paramétrique comportant :
un transformateur (101) pour transformer le premier signal audio dans un domaine de
fréquence pour obtenir un premier signal audio transformé, et pour transformer le
second signal audio dans un domaine de fréquence pour obtenir un second signal audio
transformé ;
un générateur de paramètres (107) pour générer un premier paramètre de codage, X(i),
à partir du premier signal audio transformé et à partir du second signal audio transformé
au niveau d'une première bande de fréquences i, et pour générer un second paramètre
de codage, X(j), à partir du premier signal audio transformé et à partir du second
signal audio transformé au niveau d'une seconde bande de fréquences j ; et
un combineur de paramètres (109) pour combiner le premier paramètre de codage et le
second paramètre de codage afin d'obtenir un paramètre de codage combiné, X, selon
la formule

dans lequel le paramètre I désigne un ensemble d'indices de bandes de fréquences,
le paramètre gi est une pondération donnée à la bande de fréquences i, le paramètre Pl(i) désigne une puissance de signal du premier signal audio dans la ième bande de fréquences, le paramètre P2(i) désigne une puissance de signal du second signal audio dans la ième bande de fréquences,
et dans lequel le premier paramètre de codage, X(i), et le second paramètre de codage,
X(j), est une différence de phase entre canaux ou une cohérence entre canaux ou une
différence d'intensité entre canaux ou une différence de niveau entre canaux.
2. Codeur paramétrique selon la revendication 1, dans lequel le générateur de paramètres
(107) est configuré pour générer le premier paramètre de codage et le second paramètre
de codage en fonction d'une multiplication de valeurs du premier signal audio transformé
et du second signal audio transformé.
3. Codeur paramétrique selon l'une quelconque des revendications 1 à 2, dans lequel le
générateur de paramètres (107) est configuré pour générer une pluralité de paramètres
de codage à partir du premier signal audio transformé et à partir du second signal
audio transformé au niveau d'une pluralité de bandes de fréquences et dans lequel
le combineur de paramètres est configuré pour combiner la pluralité des paramètres
de codage afin d'obtenir le paramètre de codage combiné.
4. Codeur paramétrique selon l'une quelconque des revendications 1 à 3, étant en outre
configuré pour combiner le premier signal audio transformé et le second signal audio
transformé pour obtenir un signal mixé abaissé.
5. Codeur paramétrique selon l'une quelconque des revendications 1 à 4, comprenant en
outre un transformateur inverse pour transformer de façon inverse une combinaison
du premier signal audio transformé et du second signal audio transformé pour obtenir
un signal audio mixé abaissé.
6. Procédé de codage paramétrique d'un signal audio multi-canaux ayant un premier signal
audio et un second signal audio, le procédé étant configuré pour exploiter un codeur
paramétrique selon l'une quelconque des revendications précédentes.