FIELD OF THE INVENTION
[0001] The invention relates to generation and/or processing of an encoded audio signal
for a multichannel audio signals and in particular, but not exclusively, to generation
and/or encoding of stereo signals.
BACKGROUND OF THE INVENTION
[0002] Spatial audio applications have become numerous and widespread and increasingly form
at least part of many audiovisual experiences. Indeed, new and improved spatial experiences
and applications are continuously being developed which results in increased demands
for audio processing and rendering.
[0003] For example, in recent years, Virtual Reality (VR) and Augmented Reality (AR) have
received increasing interest, and a number of implementations and applications are
reaching the consumer market. Indeed, equipment is being developed for both rendering
the experience as well as for capturing or recording suitable data for such applications.
For example, relatively low-cost equipment is being developed for allowing gaming
consoles to provide a full VR experience. It is expected that this trend will continue
and indeed will increase in speed with the market for VR and AR reaching a substantial
size within a short time scale. In the audio domain, a prominent field explores the
reproduction and synthesis of realistic and natural spatial audio. The ideal aim is
to produce natural audio sources such that the user cannot recognize the difference
between a synthetic or an original one.
[0004] A lot of research and development effort has focused on providing efficient and high-quality
audio encoding and audio decoding for spatial audio. A frequently used spatial audio
representation is multichannel audio representations, including stereo representation,
and efficient encoding of such multichannel audio based on downmixing multichannel
audio signals to downmix channels with fewer channels have been developed. One of
the main advances in low bit-rate audio coding has been the use of parametric multichannel
coding where a downmix audio signal is generated together with parametric data that
can be used to upmix the downmix audio signal to recreate the multichannel audio signal.
[0005] In particular, instead of traditional mid-side or intensity coding, parametric multichannel
audio coding uses a downmix of a multichannel input signal to a lower number of channels
(e.g. two to one) and multichannel image (stereo) parameters are extracted. Then the
downmix audio signal is encoded using a more traditional audio coder (e.g. a mono
audio encoder). The data of the downmix is combined with the encoded multichannel
parameter data to generate a suitable audio bitstream. This bitstream is then transmitted
to the decoder, where the process is inverted. First the downmix audio signal is decoded,
after which the multichannel audio signal is reconstructed, guided by the encoded
multichannel image spatial parameters.
[0006] An example of stereo coding is described in
E. Schuijers, W. Oomen, B. den Brinker, J. Breebaart, "Advances in Parametric Coding
for High-Quality Audio", 114th AES Convention, Amsterdam, The Netherlands, 2003, Preprint
5852. In the described approach, the downmixed mono signal is parametrized by exploiting
the natural separation of the signal into three components (objects): transients,
sinusoids, and noise. In
E. Schuijers, J. Breebaart, H. Purnhagen, J. Engdegård, "Low Complexity Parametric
Stereo Coding", 116th AES, Berlin, Germany, 2004, Preprint 6073 more details are provided describing how parametric stereo was realized with a low
(decoder) complexity when combining it with Spectral Band Replication (SBR).
[0007] However, although Parametric Stereo (PS) and similar multichannel downmix based encoding/
decoding approaches were a leap forward from traditional stereo and multichannel coding,
the approach is not optimal in all scenarios. In particular, known approaches tend
to introduce some distortion, changes, artefacts etc. that may introduce differences
between the (original) multichannel audio signal input to the encoder and the multichannel
audio signal recreated at the decoder. Typically, the audio quality may be degraded
and imperfect recreation of the multichannel occurs. Further, the data rate may still
be higher than desired and/or the complexity/ resource usage of the involved processing
may be higher than preferred.
[0008] Hence, an improved approach would be advantageous. In particular, an approach allowing
increased flexibility, improved adaptability, an improved performance, improved audio
quality, improved audio quality to data rate trade-off, reduced complexity and/or
resource usage, facilitated implementation, and/or an improved spatial audio experience
would be advantageous.
SUMMARY OF THE INVENTION
[0009] Accordingly, the invention seeks to preferably mitigate, alleviate or eliminate one
or more of the above mentioned disadvantages singly or in any combination.
[0010] According to an aspect of the invention there is provided an apparatus for generating
an encoded audio signal for a multichannel audio signal, the apparatus comprising:
a receiver arranged to receive the multichannel audio signal; a downmixer arranged
to generate a downmix audio signal by downmixing the multichannel audio signal; an
encoder arranged to encode the downmix audio signal to generate encoded downmix audio
data; a trained artificial neural network arranged to determine first scale factors
for (sets, blocks, contiguous groups of) time frequency intervals of the multichannel
audio signal, the trained artificial neural network having input nodes for receiving
samples of the multichannel audio signal and/or the downmix audio signal, and output
nodes providing the first scale factors; a signal generator arranged to generate a
first modified multichannel audio signal by applying the first scale factors to samples
of the time frequency intervals; a parameter circuit arranged to generate a first
set of spatial parameters for the downmix audio signal (from the modified multichannel
audio signal), the first set of spatial parameters representing relative properties
for channels of the first modified multi-channel audio signal; and an output circuit
arranged to generate the encoded audio signal to comprise the encoded audio data and
first set of spatial parameters.
[0011] The approach may provide an improved audio experience in many embodiments. For many
signals and scenarios, the approach may provide improved representation of a multichannel
audio signal allowing improved generation/ reconstruction of the multichannel audio
signal with an improved perceived audio quality. The approach may provide a particularly
advantageous arrangement which may in many embodiments and scenarios allow a facilitated
and/or improved possibility of utilizing artificial neural networks in audio processing,
including typically audio encoding and/or decoding. The approach may allow an advantageous
employment of artificial neural network(s) in generating a multichannel audio signal
from a downmix audio signal.
[0012] The approach may provide an efficient implementation and may in many embodiments
allow a reduced complexity and/or resource usage.
[0013] The approach may in many cases allow improved representation of particular signal
components with high perceptual impact, such as e.g. transients. The approach may
in many cases in particular mitigate or avoid temporal smearing and/or frequency distortion
thereby allowing improved quality of the resulting multichannel audio signal.
[0014] The set of spatial parameters may comprise parameter (values) relating properties
of the downmix audio signal to properties of the multichannel audio signal. The set
of spatial parameters may comprise data being indicative of relative properties between
channels of the modified multichannel audio signal. The set of spatial parameters
may comprise data being indicative of differences in properties between channels of
the modified multichannel audio signal. The set of spatial parameters may comprise
data being perceptually relevant for the synthesis of the multichannel audio signal.
The properties may for example be differences in phase and/or intensity and/or timing
and/or correlation. The set of spatial parameters may comprise data including at least
one of interchannel intensity differences (IID), interchannel timing differences (ITD),
interchannel correlations (ICC) and/or interchannel phase differences (IPD) for channels
of the modified multichannel audio signal. In particular, the set of spatial parameters
may include one or more of an ICC, IPD, IID parameter as known from Parametric Stereo
encoding/decoding.
[0015] The artificial neural network(s) may be a trained artificial neural network(s) trained
by training data including training multichannel audio signals; the training employing
a cost function dependent on a (spectro-temporal) envelope/shape/ difference between
the multichannel audio signal and a multichannel audio signal synthesized from the
encoded audio signal. The cost function may be dependent on a correlation between
the input multichannel audio signal and the synthesized multichannel audio signal,
and specifically may provide a decreasing cost function for a decreasing correlation.
The cost function may be dependent on a difference between statistical properties
of the synthesized multichannel audio signal and the original multichannel audio signal,
and specifically may provide a decreasing cost function for a decreasing correlation.
Depending on the preferences and the decorrelation calculation/measurements/metric,
it may in some embodiments be appropriate for the cost function to be dependent on
a correlation between the synthesized multichannel audio signal and the original multichannel
audio signal, and specifically to provide an increasing cost function for a decreasing
correlation.
[0016] The scale factors may be provided for individual time-frequency intervals/ tiles.
In some cases, one scale factor may be provided for each time-frequency interval of
the multichannel audio signal.
[0017] The multichannel audio signal may specifically be a stereo signal. The downmix audio
signal may specifically be a mono downmix audio signal.
[0018] In some embodiments, one or more of the trained artificial neural network(s) may
be a neural network trained using a perceptually weighted loss function. The loss
function may be perceptually weighted by applying a higher weight to time frequency
intervals representing a perceptually significant signal component, such as a transient.
In some embodiments, the perceptual weighted loss function may be generated by applying
a higher weighting to (time) frames/segments comprising a perceptually significant
signal component/event, such as a transient or loud sound, relative to frames/segments
that do not include perceptually significant signal component/events.
[0019] According to an optional feature of the invention, the apparatus further comprises
a scale factor circuit arranged to determine second scale factors for the time frequency
intervals of the multichannel audio signals; and the signal generator is arranged
to generate a second modified multichannel audio signal by applying the second scale
factors to samples of the time frequency intervals; the parameter circuit is arranged
to generate a second set of spatial parameters for the downmix audio signal from the
second modified multichannel audio signal, the second set of spatial parameters representing
relative properties for channels of the second modified multi-channel audio signal;
and the output circuit is arranged to generate the encoded audio signal to comprise
the second set of spatial parameters.
[0020] This may provide a particularly efficient and high-performance operation in many
scenarios and may typically result in substantially improved audio quality of a multichannel
audio signal generated from the encoded audio signal.
[0021] According to an optional feature of the invention, the scale factor circuit is arranged
to determine the second set of scale factors from the first set of scale factors.
[0022] This may provide a particularly advantageous and high performance operation and/or
implementation. The second set of scale factors may be determined so that the first
and second set of scale factors have a constant combined value (w
2=1-w
1).
[0023] According to an optional feature of the invention, the scale factor circuit comprises
a further trained artificial neural network arranged to determine the second scale
factors, the further trained artificial neural network having input nodes for receiving
samples of the multichannel audio signal and/or the downmix audio signal, and output
nodes providing the second scale factors.
[0024] This may provide a particularly advantageous and high performance operation and/or
implementation.
[0025] According to an optional feature of the invention, a data rate for the first set
of spatial parameters is different than a data rate for the second set of spatial
parameters.
[0026] This may provide a particularly advantageous and high performance operation and/or
implementation.
[0027] According to an optional feature of the invention, the first scale factors are complex
values.
[0028] This may provide a particularly advantageous and high performance operation and/or
implementation.
[0029] According to an optional feature of the invention, the apparatus comprises an additional
trained artificial neural network arranged to generate feature set values from the
multichannel audio signal, the additional trained artificial neural network having
input nodes for receiving samples of the multichannel audio signal, and the output
circuit is arranged to include the feature set values in the encoded audio signal.
[0030] This may provide a particularly advantageous and high performance operation and/or
implementation.
[0031] According to an aspect of the invention, there is provided an apparatus for generating
a multichannel audio signal from an encoded audio signal, the apparatus comprising:
a receiver arranged to receive the encoded audio signal comprising encoded downmix
audio data for a downmix of a multi-channel audio signal and a first set of spatial
parameters, the first set of spatial parameters representing relative properties for
channels of a first modified multi-channel audio signal resulting from applying first
scale factors to samples of time frequency intervals of the multichannel audio signal;
a decoder arranged to decode the encoded audio data to generate the downmix audio
signal; a decorrelator arranged to decorrelate the downmix audio signal to generate
a decorrelated signal; a trained artificial neural network arranged to determine the
first scale factors for time frequency intervals of the multichannel audio signal,
the trained artificial neural network having input nodes for receiving samples of
the downmix audio signal and output nodes providing the first scale factors; an upmixer
arranged to generate the multichannel audio signal by upmixing the downmix audio signal
and the decorrelated signal in dependence on the first set of spatial parameters and
the first scale factors.
[0032] According to an optional feature of the invention, the upmixer is arranged to generate
each channel of the multichannel audio signal as a linear combination of the downmix
audio signal and the decorrelated signals where weights of the linear combination
are dependent on the first set of scale factors and the first set of spatial parameters.
[0033] This may provide a particularly advantageous and high performance operation and/or
implementation.
[0034] According to an optional feature of the invention, the encoded audio signal comprises
a second set of spatial parameters, the second set of spatial parameters representing
relative properties for channels of a second modified multi-channel audio signal resulting
from applying second scale factors to samples of time frequency intervals of the multichannel
audio signal; the apparatus comprises a circuit arranged to generate the second scale
factors; and the upmixer is arranged to determine a first upmixed multichannel audio
signal from the downmix audio signal and the decorrelated signal using the first set
of spatial parameters and the first scale factors, to determine a second upmixed multichannel
audio signal from the downmix audio signal and the decorrelated signal using the second
set of spatial parameters and the second scale factors, and to include a combination
of the first upmixed multichannel audio signal and the second upmixed multichannel
audio signal in the multichannel audio.
[0035] This may provide a particularly advantageous and high performance operation and/or
implementation.
[0036] According to an optional feature of the invention, the encoded audio signal comprises
feature set values for a trained network and the trained artificial neural network
has input nodes for receiving the feature set values.
[0037] This may provide a particularly advantageous and high performance operation and/or
implementation.
[0038] According to an optional feature of the invention, the trained artificial neural
network of an encoder apparatus is different from the trained artificial neural network
of a decoder apparatus.
[0039] This may provide a particularly advantageous and high performance operation and/or
implementation.
[0040] According to an optional feature of the invention, one or more of the trained artificial
neural networks is trained using training data comprising a number of training multichannel
audio signals and a cost function dependent on a perceptual difference measure for
the multichannel audio signal received by the encoder apparatus and the multichannel
audio signal generated by the decoder apparatus.
[0041] This may provide a particularly advantageous and high performance operation and/or
implementation.
[0042] According to an aspect of the invention, there is provided a method of operation
for an audio apparatus for generating an encoded audio signal for a multichannel audio
signal, the method comprising: receiving the multichannel audio signal; generating
a downmix audio signal by downmixing the multichannel audio signal; encoding the downmix
audio signal to generate encoded downmix audio data; a trained artificial neural network
determining first scale factors for time frequency intervals of the multichannel audio
signal, the trained artificial neural network having input nodes for receiving samples
of the multichannel audio signal and/or the downmix audio signal, and output nodes
providing the first scale factors; generating a first modified multichannel audio
signal by applying the first scale factors to samples of the time frequency intervals;
generating a first set of spatial parameters for the downmix audio signal, the first
set of spatial parameters representing relative properties for channels of the first
modified multi-channel audio signal; and generating an encoded audio signal to comprise
the encoded audio data and first set of spatial parameters.
[0043] According to an aspect of the invention, there is provided a method of operation
for an audio apparatus, the method comprising: receiving the encoded audio signal
comprising encoded downmix audio data for a downmix of a multi-channel audio signal
and a first set of spatial parameters, the first set of spatial parameters representing
relative properties for channels of a first modified multi-channel audio signal resulting
from applying first scale factors to samples of time frequency intervals of the multichannel
audio signal; decoding the encoded audio data to generate the downmix audio signal;
decorrelating the downmix audio signal to generate a decorrelated signal; a trained
artificial neural network determining the first scale factors for time frequency intervals
of the multichannel audio signal, the trained artificial neural network having input
nodes for receiving samples of the downmix audio signal and output nodes providing
the first scale factors; and generating the multichannel audio signal by upmixing
the downmix audio signal and the decorrelated signal in dependence on the first set
of spatial parameters and the first scale factors.
[0044] These and other aspects, features and advantages of the invention will be apparent
from and elucidated with reference to the embodiment(s) described hereinafter.
BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Embodiments of the invention will be described, by way of example only, with reference
to the drawings, in which
FIG. 1 illustrates some elements of an example of an audio apparatus in accordance
with some embodiments of the invention;
FIG. 2 illustrates some elements of an example of an audio apparatus in accordance
with some embodiments of the invention;
FIG. 3 illustrates some elements of an example of a time to frequency converter for
an audio apparatus in accordance with some embodiments of the invention;
FIG. 4 illustrates some elements of an example of an audio apparatus in accordance
with some embodiments of the invention;
FIG. 5 illustrates an example of a stereo signal including a transient;
FIG. 6 illustrates an example of a neuron for an artificial neural network;
FIG. 7 illustrates an example of a structure of an artificial neural network;
FIG. 8 illustrates an example of a structure of a diffusion model approach;
FIG. 9 illustrates an example of a structure of an artificial neural network;
FIG. 10 illustrates an example of a structure of an artificial neural network;
FIG. 11 illustrates an example of a structure of an artificial neural network arrangement;
FIG. 12 illustrates an example of elements of an arrangement for training an artificial
neural network;
FIG. 13 illustrates an example of an inter-channel level difference; and
FIG. 14 illustrates some elements of a possible arrangement of a processor for implementing
elements of an audio apparatus in accordance with some embodiments of the invention.
DETAILED DESCRIPTION OF SOME EMBODIMENTS OF THE INVENTION
[0046] FIG. 1 illustrates some elements of an audio apparatus arranged to generate an encoded
audio signal representing a multichannel audio signal and FIG. 2 illustrates elements
of an audio apparatus arranged to generate a multichannel audio signal from an encoded
audio signal such as that generated by the audio apparatus of FIG. 1. An audio distribution/communication
system may accordingly comprise the audio encoder apparatus of FIG. 1 generating an
encoded audio signal for a multichannel audio signal with the encoded audio signal
being transmitted to an audio decoder apparatus of FIG. 2 which may recreate a local
replica of the multichannel audio signal from the received encoded audio signal.
[0047] The audio encoder apparatus comprises a receiver 101 which is arranged to receive
a multichannel audio signal that is to be encoded and represented by the encoded audio
signal. The multichannel audio signal may typically be a stereo signal, and the following
description will focus on this specific example.
[0048] The processing by the audio encoder apparatus is typically performed predominantly
in the frequency/ subband domain and the following description will focus on such
embodiments. However, in some embodiments, some or all of the processing may be performed
in the time domain using time domain processing and indeed in many embodiments the
processing may include both processing of a time domain and frequency domain representations
of the signals.
[0049] In many embodiments, the multichannel audio signal may be received as a time domain
signal and a frequency representation may be generated and used in the processing.
[0050] In some embodiments, the frequency domain multichannel audio signal may be received
from a remote source but in many other embodiments a remote source may generate a
time domain representation of the audio signal and a local frequency transformer may
be arranged to generate a (time-)frequency representation of the multichannel audio
signal.
[0051] In particular, in the example of FIG. 1, the audio apparatus comprises a filter bank
103 which is arranged to generate a frequency (subband) representation of a received
time domain multichannel audio signal. Typically, the audio apparatus may comprise
a filter bank 103 that is applied to each channel signal of the multichannel audio
signal such that this is divided into frequency subbands.
[0052] The filter bank may be a Quadrature Mirror Filter (QMF) bank or may e.g. be implemented
by a Fast Fourier Transform (FFT), but it will be appreciated that many other filter
banks and approaches for dividing an audio signal into a plurality of subband signals
are known and may be used. The filter-bank may specifically be a complex-exponential
modulated pseudo QMF bank, resulting in e.g. 32 or 64 complex-valued sub-band signals.
[0053] The processing is furthermore typically performed in time segments or time slots.
In most embodiments, the audio signal is divided into time intervals/segments with
a conversion to the frequency/subband domain by applying e.g. an FFT or QMF filtering
to the samples of each signal. For example, each channel of the downmix audio signal
may be divided into time segments of e.g. 2048, 1024, or 512 samples. These signals
may then be processed to generate samples for e.g. 64, 32 or 16 subbands. Thus, a
set of samples may be determined for each subband of the downmix audio signal.
[0054] It should be noted that the number of time domain samples is not directly coupled
to the number of subbands. Typically, for a so-called critically sampled filterbank
of N bands, every N input samples will lead to N sub-band samples (one for every sub-band).
An oversampled filterbank will produce more output samples. E.g. for every N input
samples, it would generate k*N output samples, i.e., k consecutive samples for every
band.
[0055] In some embodiments, the subbands are generated to have the same bandwidth but in
other embodiments subbands are generated to have different bandwidths, e.g. reflecting
the sensitivity of human hearing to different frequencies.
[0056] FIG. 3 illustrates an example of an approach for a filterbank 103 where different
bandwidths are generated by a hybrid filter bank approach.
[0057] In the example, this is realized by a combination of a complex-exponential modulated
pseudo QMF bank 301 and a small filter bank 303 for the lower frequency bands to realize
a higher frequency resolution, as desired for binaural perception of the human auditory
system. The result is a hybrid filterbank with logarithmic filter band center-frequency
spacings that follow that of human perception similar to equivalent rectangular bandwidths
(ERBs). In order to compensate for the delay of the filtering by the small filter
bank 303, a delay 305 is introduced for higher frequency subbands.
[0058] In the specific example, a time-domain signal
x[
n] is fed through a downsampled complex-exponential modulated QMF bank with
K bands. Each frame of 64 time domain samples
x[
n] results in one slot of QMF samples
X[
k, m] with
k = (0, ...
, K - 1) at slot m. The lower slots are then filtered by additional complex-modulated
filterbanks splitting the lower bands further. The higher slots are delayed ensuring
that the filtered signals of the lower bands are in sync with the higher bands as
the filtering introduces a delay. This finally results in a structure where for every
64 time-domain samples
x[
n], one slot m of hybrid QMF samples
Y[
l, m] is produced with
l = (0, ...,
L - 1) at slot
m, e.g. with a total number of hybrid bands M = 77.
[0059] The receiver 101 is fed to a downmixer 105 which proceeds to generate a downmix audio
signal and typically a mono downmix audio signal. In many embodiments and scenarios,
the downmixer 105 may generate a mono downmix audio signal from a received stereo
signal, e.g. simply by summing the channels of the stereo signal in accordance with
a standard Parametric Stereo (PS) downmix approach.
[0060] The downmixer 105 is coupled to an audio encoder 107 which is arranged to encode
the downmix audio signal to generate encoded downmix audio data for the downmix audio
signal. In many embodiments, a mono audio encoding may be used to encode a mono downmix
audio signal. It will be appreciated that any suitable audio encoding approach may
be used including in particular any suitable encoding Standard or algorithm for encoding
an audio signal to generate audio data representing the audio signal. In particular,
the mono downmix audio signal may be encoded using a standard encoding.
[0061] The audio encoder 107 is coupled to an output circuit 109 which is arranged to generate
the encoded audio signal. The encoded audio signal is generated to include the encoded
downmix audio data, and it will be appreciated that the encoded audio signal may be
generated in accordance with any suitable structure, format etc. In many embodiments,
the output circuit 109 may be arranged to generate the encoded audio signal to follow
a Standard data format of audio signals.
[0062] The encoder audio apparatus of FIG. 1 is further arranged to generate and include
spatial (upmix) parameters in the encoded audio signal. However, rather than use a
conventional approach, such as e.g. known from PS, where the spatial parameters are
generated from the multichannel audio signal and are generated to reflect properties
of the multichannel audio signal, the encoding audio apparatus of FIG. 1 is arranged
to perform specific processing to generate signal components of the multichannel audio
signal and with the spatial parameters being generated to reflect these.
[0063] In particular, the audio encoder apparatus includes a first (encoder) trained artificial
neural network 111 which determines scale factors, also referred to as the first (encoder)
scale factors, for time frequency tiles of the multichannel audio signal. In many
embodiments, the trained artificial neural network 111 may generate scale factors
for individual time frequency intervals/times as generated by the filter banks 103.
The first trained artificial neural network 111 is arranged to receive samples of
the generated mono downmix audio signal and/or the multichannel audio signal and to
generate the scale factors based on these values. The first trained artificial neural
network 111 has input nodes that receive these samples and output nodes that provide
scale factors for the time frequency tiles.
[0064] The first trained artificial neural network 111 may be arranged to generate scale
factors that provide a filtering/windowing function which is applied to the multichannel
audio signal to estimate/emphasize perceptually relevant parts of the channel signals
of the multichannel audio signal. The scale factors may form a window function that
may be a masking function between that weighs the relevancy of time-frequency tiles
across each subband of (a frame of) the multichannel audio signal. In many cases,
the scale factors may be determined as values within a given interval such as between
0 and 1. The first trained artificial neural network 111 may be trained using a perceptual-based
loss function which assigns a relevancy weighting to training samples that include
perceptually relevant signal components to help guide the scale factor/window estimation
process to emphasize such components after training is completed. The first trained
artificial neural network 111 may thus be trained to generate scale factors that filter/mask/extract/estimate
perceptually relevant parts of the multichannel audio signal.
[0065] It will be appreciated that, as described in more detail later, the first trained
artificial neural network 111, may be implemented in many different ways and that
any suitable approach or neural network implementation may be used. Similarly, as
will be described later, different approaches for training the trained artificial
neural network may be used in different applications and embodiments.
[0066] The first trained artificial neural network 111 is coupled to a signal generator
113 which is also coupled to the receiver 101. The signal generator 113 receives the
multichannel audio signal and the first scale factors and is arranged to generate
a first modified multichannel audio signal by applying the first scale factors to
samples of the time frequency intervals of the multichannel audio signal. Specifically,
in many embodiments, each frequency sample of the multichannel audio signal may be
multiplied by the scale factor determined by the first trained artificial neural network
111 for the time frequency tile to which the frequency sample belongs. The signal
generator 113 thus generates a modified multichannel audio signal that has been generated
based on the scale factors determined by the first trained artificial neural network
111.
[0067] As will be described in more detail, a second modified multichannel audio signal
may often be generated by the signal generator 113, and indeed the second modified
multichannel audio signal may often be generated as a residual signal such that the
first and second modified multichannel audio signal may form a decomposition of the
original multichannel audio signal. In some embodiments, more than two modified multichannel
audio signals may be generated and in particular the signal generator 113 may perform
a signal decomposition of the multichannel audio signal into more than two (modified)
multichannel audio signal components.
[0068] The modified multichannel audio signal(s) are fed to a parameter estimation circuit
115 which is coupled to the signal generator 113 and to the output circuit 109. The
parameter circuit 115 is arranged to determine spatial parameters for the modified
multichannel audio signal(s) where the spatial parameters are indicative of relative
properties between the channels of the modified multichannel audio signal(s). Specifically,
the parameter circuit 115 may determine a first set of spatial parameters for the
first modified multichannel audio signal where the first set of spatial parameters
represent relative properties of the first modified multichannel audio signal.
[0069] Typically, the upmix/spatial parameters may be indicative of time differences, phase
differences, level/intensity differences and/or a measure of similarity, such as correlation,
between channels of the modified multichannel audio signal(s). Typically, the spatial
parameters are provided on a per time and per frequency basis (time frequency tiles).
For example, new parameters may periodically be provided for a set of subbands. Parameters
may specifically include Inter-channel phase difference (IPD), Overall phase difference
(OPD), Inter-channel correlation (ICC), Inter-channel Intensity/Level difference parameters
as known from Parametric Stereo encoding (as well as from higher channel encodings).
[0070] For stereo applications, the spatial parameters may specifically be spatial parameters
as used for encoding a stereo signal using a Parametric Stereo (PS) encoding of the
stereo signal, such as for example by a mono downmix audio signal given as:

where the parameter
c is chosen such that the power of the stereo signal is preserved in the downmix, the
power being defined using the 2-norm:

and thus e.g.:

[0072] It should be noted that in some embodiments, the input multichannel audio signal
may be weighted by windows (time direction) and grouping/kernels (frequency direction).
In some embodiments, this may be affected by the scale factors determined by the first
trained artificial neural network 111 and the parameter estimation and/or downmixing
may also take these into account.
[0073] The spatial parameters are typically provided for specific time frequency tiles,
and thus specifically each parameter value is generated/provided for a given frequency
subband and for a given time segment.
[0074] The parameter circuit 115 may determine spatial parameters that are indicative of
relative properties of the channels (channel signals) of the stereo signal.
[0075] The spatial parameters may comprise sets of spatial parameters, each set of spatial
parameters comprising at least one of: a level difference parameter indicative of
a level difference between channels of the multichannel audio signal; a correlation
parameter indicative of a coherence between channels of the multichannel audio signal;
a timing difference parameter indicative of a timing difference between channels of
the multichannel audio signal, and a phase difference parameter indicative of a phase
difference between channels of the multichannel audio signal.
[0076] The output circuit 109 is arranged to include the generated spatial parameters in
the encoded audio signal and thus the audio encoder apparatus is arranged to generate
a representation of a multichannel audio signal by a (typically mono) downmix signal
and associated spatial data which however is not providing data on relative properties
of the channels of the multichannel audio signal but rather is indicative of relative
properties of only a particular component/part of the channels of the multichannel
audio signal with this component/part being determined/identified by a trained artificial
neural network.
[0077] The audio decoder apparatus of FIG. 2 comprises a receiver 201 which is arranged
to receive an encoded audio signal and specifically it may receive the encoded audio
signal from the audio encoder apparatus of FIG. 1. Thus, the receiver 201 may receive
an encoded audio signal that comprises encoded downmix audio data for a downmix of
a multi-channel audio signal and a first set of spatial parameters where the first
set of spatial parameters represent relative properties for channels for the first
modified multi-channel audio signal that is generated from applying first scale factors
to samples of time frequency intervals of the multichannel audio signal. The first
set of spatial parameters may be spatial parameters for a multichannel audio signal
component of the multichannel audio signal where the multichannel audio signal component
is related to the multichannel audio signal by a set of scale factors for time frequency
intervals of the multichannel audio signal.
[0078] The audio decoder apparatus comprises a decoder 203 arranged to decode the encoded
downmix audio data to generate the downmix audio signal
xm(
n). The downmix audio signal is fed to a decorrelator 205 which is arranged to decorrelate
the downmix audio signal to generate a decorrelated signal
xd(
n) that is perceptually similar to, but uncorrelated with
xm(
n). It will be appreciated that many different algorithms are known for decorrelating
an audio signal and that any suitable approach may be used by the audio decoder apparatus
of FIG. 2.
[0079] The downmix audio signal, the decorrelated signal, and the received spatial parameters
are fed to an upmixer (207) which is arranged to generate (a local copy of) the multichannel
audio signal by upmixing the downmix audio signal and the decorrelated signal.
[0080] The audio decoder apparatus further comprises a trained artificial neural network,
henceforth also referred to as the first decoder trained artificial neural network
209 which receives the downmix audio signal. The first decoder trained artificial
neural network 209 is arranged and trained to determine scale factors, henceforth
referred to as the first decoder scale factors, for time frequency intervals of the
downmix audio signal. The first decoder trained artificial neural network 209 has
input nodes for receiving samples of the downmix audio signal and output nodes providing
the first scale factors. In many embodiments, the first decoder trained artificial
neural network 209 may correspond directly to the first encoder trained artificial
neural network 111 and indeed in many embodiments both of these trained networks may
receive only the downmix audio signal for generating the scale factors and the first
decoder trained artificial neural network 209 may typically generate scale factors
that are identical or very close to those generated by the first encoder trained artificial
neural network 111.
[0081] In many embodiments, the first encoder trained artificial neural network 111 at the
audio encoder apparatus is arranged to generate the first encoder scale factors based
(only) on the downmix audio signal which is also available at the audio decoder apparatus.
Accordingly, by having the first decoder trained artificial neural network 209 determining
the first decoder scale factors based on the locally available downmix audio signal,
a similar processing can be performed by the trained artificial neural networks resulting
in identical or similar scale factors being determined.
[0082] The upmixer 207 may specifically be arranged to generate a first multichannel audio
signal by upmixing (at least) the downmix audio signal and the decorrelated signal
based on the received first set of spatial parameters. The audio decoder apparatus
comprises the first decoder trained artificial neural network 209 which may generate
scale factors corresponding to those generated at the audio source device 101. Indeed,
in many embodiments, the trained networks may be identical and both operate on the
downmix audio signal and thus will generate (substantially) identical scale factors.
The audio decoder apparatus can accordingly apply the scale factors resulting in it
being possible to generate (a local replica) of the modified multichannel audio signal.
The audio decoder apparatus may then generate an output multichannel audio signal
directly corresponding to the modified multichannel audio signal, or may typically
combine the recreated modified multichannel audio signal with other signal components,
such as typically (a local replica of) a second modified signal representing a residual
signal. Thus, in many embodiments, the output multichannel audio signal may be generated
as a combination of a plurality of generated multichannel audio signals of which the
local replica of the first modified multichannel audio signal is one.
[0083] In the approach, a multichannel audio signal may accordingly be represented by a
downmix signal with spatial parameters that are representing relative properties for
a signal component extracted from the multichannel audio signal rather than from the
multichannel audio signal as a whole. The spatial parameters specifically reflects
properties of a signal component of the multichannel audio signal that has been extracted/determined
using a trained artificial neural network as will be described in more detail later.
Further, at the decoder side, the signal component may be recreated from the mono
downmix audio signal and a complementary artificial neural network may allow the scale
factors to be determined allowing for the output signal to be generated while appropriately
scaling/weighting the signal component.
[0084] In many embodiments, the audio encoder apparatus may include a scale factor circuit
117 which is arranged to further determine a second set of spatial parameters which
describe properties of a second signal component. Specifically, the signal generator
113 may be arranged to generate a second modified multichannel audio signal by applying
the second scale factors to samples of the time frequency intervals. In many embodiments,
the second scale factors may be determined such that the second modified multichannel
audio signal is a residual signal for the first modified multichannel audio signal.
The second scale factors may for example be determined such that the sum of the first
and second scale factors are constant (unity) for all time frequency tiles. For example,
in many embodiments, second scale factors may be generated as 1-w1 where w1 is the
first scale factor for the same time-frequency tile.
[0085] The parameter circuit 115 may accordingly also be fed the second modified multichannel
audio signal, which specifically may be a residual signal, and may proceed to determine
a second set of spatial parameters that represent/describe relative properties for
channels of the second modified multi-channel audio signal. The same approaches as
used for the first modified multichannel audio signal may be used and e.g. the second
set of spatial parameters may include IID, ICC, and IPD values. The second set of
spatial parameters are fed to the output circuit 109 which includes it in the encoded
audio signal.
[0086] At the audio decoder apparatus side, the second set of spatial parameters are also
fed to the upmixer 207 which is arranged to upmix the received mono downmix audio
signal and the decorrelated signal based on the second set of spatial parameters to
generate a second upmixed multichannel audio signal. The audio decoder apparatus may
further proceed to generate second scale factors to correspond to the second decoder
scale factors. For example, in the case where the second modified multichannel audio
signal is a residual signal generated by the second scale factors and the first scale
factors summing to a fixed, constant value, the second decoder scale factors may be
generated by subtracting the first decoder scale factors determined by the first encoder
trained artificial neural network 111 from the fixed, constant value.
[0087] The first and second upmixed multichannel audio signals may then be combined based
on the first and second scale factors. For example, a linear combination of the frequency
domain samples of the first and second upmixed multichannel audio signal with the
weights of the first upmixed multichannel audio signal corresponding to the first
scale factors and the weights of the second upmixed multichannel audio signal corresponding
to the second scale factors.
[0088] Thus, in some embodiments, trained artificial neural networks may be used to isolate
and extract signal components for which individual spatial parameters are determined
and communicated with the audio decoder apparatus performing differentiated upmixing
based on the different spatial parameters sets. In particular, in some embodiments,
the audio encoder apparatus may perform a signal decomposition operation on the multichannel
audio signal and generate individual and separate spatial parameters for the different
signal components. The different components may be separately upmixed at the decoder
side with the resulting upmixed signals being combined. The approach has been found
to provide a particularly advantageous operation in many embodiments, applications,
and scenarios.
[0089] As a specific example, the multichannel audio signal may be a stereo signal and the
audio decoder apparatus may employ the structure of FIG. 4 where two complementary
modified stereo signals may be generated as:

where b and m are indexes representing time frequency times (frequency time band
b and time interval/frame m),
WNN (
b, m) represent the scale factor for time frequency tile b,m, and
WNN (
b,
m) = 1
- WNN(
b, m)
.
[0090] In some cases, a windowing function
WT(
m) across frames may also be applied, such as e.g. a triangular windowing function:

[0091] The parameter circuit 115 may then proceed to generate a set of spatial parameters
for each of the modified stereo signals, and specifically the spatial parameters

, and those based on the complementary window
WNN, ICC(
b),
IID(
b),
IPD(
b), may be determined. The spatial parameters are then included in the encoded audio
signal.
[0092] The bit stream then includes the set of windowed and complementary PS parameters

and {
ICC(
b'),
IID(
b'),
IPD(
b')}, where b' is used to emphasize different groupings of the complementary PS parameters.
[0093] Because the windowing function/scale factors
WNN are determined from the downmix signal
Xm, the window functions/scale factors themselves do not have to be included in the encoded
audio signal as they can independently be estimated at the audio decoder apparatus.
[0094] At the audio decoder apparatus, the downmix signal is decoded by the mono decoder
203 and is by the first decoder trained artificial neural network 209 used to (re)estimate
the set of scale factors/ windows,
WNN. The estimated windows/scale factors, together with the mono downmix subband signal
Xm(
k, m) and the decorrelated signal
Xd(
k, m) generated by the decorrelator 205, are fed to the upmixer 207. The upmixer 207
may then perform upmixing to generate local replicas of the two modified stereo signals
Y(
b, m) and
Ŷ(
b, m)
. These may then be combined into a single stereo output signal with the weighting
of each being set to match the scale factors.
[0096] The upmixer 207 may then upmix each set of signals using the received spatial parameters
for the individual signal component to produce the reconstructed left and right stereo
signal
L(
k, m) and
R(
k, m)
. This is achieved using upmix matrices according to:

and

where the entries

and
Hij(
k) =
f(
ICC(
b),
IID(
b),
IPD(b)), where
f(·) is a function that translates the PS parameters into the corresponding upmix matrix
entries per subband and is known in the prior art.
[0097] Finally, the two components may be combined and specifically summed to produce the
reconstructed left and right stereo signals:

[0098] A synthesis filterbank may then be used to reconstruct left and right time-domain
signals, e.g. using an overlap-add process.
[0099] In contrast to e.g. traditional PS encoding, the current approach may provide a highly
advantageous operation in many embodiments and may in particular in many embodiments
provide an improved perceived audio quality.
[0100] In particular, in comparison to conventional approaches where the spatial parameters
are typically generated to represent signal properties over an entire frame, the current
approach may dynamically adapt to focus the parameters on particular perceptually
relevant signal components.
[0101] Specifically, current PS encoding has a low update rate for spatial parameters in
order to keep the bit rate low (typically one parameter set per frame) which results
in the parameters reflecting average properties over a frame. Indeed, if e.g. some
static windowing is introduced, this will tend to statically emphasize the middle
of the frame thus averaging out the parameter estimation over frequency bands and
biasing the estimate towards the most dominant component, which is usually the stationary
background in this case.
[0102] For example, consider the depiction of a foreground transient component amid background
noise as exemplified in FIG. 5 where a single transient T is present in the right
channel R and for only a short time interval of the frame. FIG. 5 provides a depiction
of left-panned transient spectrogram amid background applause. The checkered background
noise is used to emphasize the random nature of the time-frequency distribution. The
asterisk represents the binaural parameter estimation point. Transients can typically
display wideband characteristics relative to background noise which has a fall-off
given by room absorption/reflectivity. Here, ideally for low frequencies where the
power of the background is larger than the transient, the PS parameters should reflect
the background signal's properties, while for the foreground component, above a certain
frequency, its parameters should dominate. However, given a static symmetric window
centered at the asterisk (boundary between current and previous frame), the transients
PS parameters will be biased to the weak background (not shown here).
[0103] This results in the loss of specific binaural cues related to signal components such
as transients which have a shorter time-span relative to the frame size. The described
approach may employ a neural network-based approach that may be coupled to a relevancy
weighting applied in the loss function during training, producing data-driven subband
scale factors that can be used to generate a signal component that highlight perceptually
relevant components and in particular with frequency selective emphasis on perceptually
relevant components.
[0104] In some embodiments and scenarios where a plurality of sets of spatial parameters
are included for different signal components, the data rate/size for the different
sets may be set to be different. For example, the data rate/size for spatial parameters
for the first modified multichannel audio signal determined by the scale factors from
the first trained artificial neural network 111 may be substantially higher than for
the spatial parameters for the residual signal. For example, the quantization (whether
in frequency or of the actual value) may be different for different signal components
and thus the encoding of the generated parameter values may be different. In many
embodiments, the average number of bits allocated/used per parameter value may be
different for (at least two) different sets of spatial parameters.
[0105] Thus, a different trade-off between accuracy and data rate may be implemented for
the different sets of spatial parameters. This may allow improved audio quality versus
data rate in many scenarios.
[0106] For example, for the PS example described above, the PS parameters for the signal
component based on the complementary adaptive window
WNN may have a lower resolution and correspond to a single triplet of PS parameters (
ICC, IID, IPD) than for the signal component based on the adaptive window
WNN. This may typically result in a scenario where the data rate of the generated encoded
audio signal is only negligibly higher than the data rate for a standard PS coder
while allowing a better audio quality for the reproduced stereo signal.
[0107] In some embodiments, the scale factors may be generated as scale values and thus
may simply be scalar gain values for frequency domain values of the multichannel audio
signal(s). However, in most embodiments, the trained artificial neural networks are
arranged to generate complex valued scale factors and the application to (typically
complex) frequency domain values of the multichannel audio signal(s) are performed
in the complex domain (specifically by performing complex multiplications). Thus,
in most embodiments, the trained artificial neural networks are arranged to not only
generate values that include an amplitude compensation but also to include a phase
property and modification.
[0108] As described above, in some embodiments, the audio decoder apparatus may in some
embodiments be arranged to generate the second set of scale factors for determining
the second modified multichannel audio signal from the first set of scale factors.
For example, for each frequency band, the audio decoder apparatus may determine a
scale factor for the second set of scale factors as a function of a scale factor from
the first set of scale factors. In particular, in many embodiments, the second set
of scale factors may as previously described be determined such that a sum of the
first and second scale factors for a given frequency band add up to a fixed value.
The fixed value may specifically be unity and thus

where c
1,k,m and c
2,k,m are the scale factors of respectively the first and second set of scale factors for
frequency band k and time slot m. It will be appreciated that the audio decoder apparatus
may proceed to perform the same operation to generate second scale factors and that
these may be used in the upmixing/combination to generate the output multichannel
audio signal as previously described.
[0109] In some embodiments, the audio encoder apparatus may be arranged to determine second
scale factors for a second multichannel audio signal which represents a second perceptually
relevant signal component that is not (necessarily) a residual signal.
[0110] In particular, in some embodiments, the audio encoder apparatus may include a second
trained artificial neural network 117 which may generate second scale factors from
the downmix audio signal and/or from the multichannel audio signal. The second trained
artificial neural network 117 has input nodes for receiving samples of the multichannel
audio signal and/or the downmix audio signal, and output nodes providing the second
scale factors. The second trained artificial neural network 117 may typically be trained
to extract a second modified multichannel audio signal that comprises perceptually
relevant information that is not included in the first modified multichannel audio
signal. For example, the second modified multichannel audio signal may represent a
second transient, a dominant point source signal, etc.
[0111] The parameter circuit 115 may then proceed to determine a second set of spatial parameters
for the second modified multichannel audio signal and the output circuit 109 may include
the second set of spatial parameters in the encoded audio signal. Thus, the generated
encoded audio signal includes spatial parameters for upmixing two (or more) intermediate
multichannel audio signals that may represent different perceptually relevant properties/components
of the original multichannel audio signal.
[0112] In many embodiments, the audio decoder apparatus may in addition to generating scale
factors for two or more multichannel signal components also proceed to generate a
residual signal comprising the signal components that are not included in the two
or more multichannel signal components.
[0113] Correspondingly, the audio decoder apparatus may include additional trained artificial
neural networks that determine the set of second scale factors. A corresponding intermediate
multichannel audio signal may be generated by an upmixing based on the second spatial
parameters and the resulting multichannel audio signal may be included in the combination
to generate the output multichannel audio signal with the weights being dependent
on the generated second scale factors (or equivalently the weighting may be included
in the upmix operation).
[0114] In many embodiments, the trained networks of the audio encoder apparatus and the
audio decoder apparatus are substantially identical and may in particular be identical
in terms of design, structure, dimensions etc. Further, in many embodiments, the trained
artificial neural networks may be trained identically, and indeed by be initialized
with values that result from the same training process. Indeed, in some embodiments,
the encoded audio signal may include configuration data for the trained artificial
neural network(s) of the audio decoder apparatus and the audio decoder apparatus may
be arranged to initialize the trained artificial neural network with the received
data values.
[0115] However, in other embodiments, the trained artificial neural networks of the audio
encoder apparatus and the audio decoder apparatus may be different. In particular,
in some embodiments, a trained artificial neural network of the audio decoder apparatus
may have a different structure, size, dimension, or configuration than the corresponding
trained artificial neural network of the audio encoder apparatus. For example, the
audio decoder apparatus may in many embodiments employ a trained artificial neural
network that is substantially less complex and resource demanding than for the audio
encoder apparatus thereby allowing implementation in smaller and less capable devices,
such as e.g. portable devices.
[0116] In such cases, the trained artificial neural networks may for example be trained
by a joint training process simultaneously including both the audio encoder apparatus
and the audio decoder apparatus and where a loss function is e.g. based on comparing
the original multichannel audio signal to the replicated multichannel audio signal
and with both weights/parameters of the trained artificial neural network of the audio
encoder apparatus and the audio decoder apparatus being updated.
[0117] The approach may thus utilize a trained artificial neural network at both the audio
encoder apparatus and the audio decoder apparatus. In some embodiments, the same neural
network architecture and weights are used at both encoder and decoder, with the neural
network weights being shared by the encoder and decoder during training.
[0118] In another embodiment of the invention, the same or different neural network architectures
may be used at encoder and decoder with different weights being trained during training.
In such a case, additional loss terms may be included in the total loss including
the difference between the estimated windows of the encoder neural network and decoder
neural network,

[0119] In some cases, a non-neural network-based approach (e.g., based on signal processing
and segmentation methods) could be used at the encoder side to estimate the subband
windows
WNN and the neural network at the decoder side may learn to adapt its weights to produce
an identical set of windows at the decoder.
[0120] In some embodiments the neural network at the decoder side may not produce the same
sets of subband windows
WNN as at the encoder side, due to reconstruction errors introduced by the mono downmix
audio decoder.
[0121] In yet other embodiments, as an alternative to the approach where the encoder and
decoder effectively both have the same artificial neural network to estimate the scale
factors, it may be desirable to trade-off some of the complexity of the decoder by
transmitting intermediate latent space variables (feature values) e.g. representing
the scale factors.
[0122] Thus, in some embodiments, the audio encoder apparatus may further include a feature
set trained artificial neural network 119 which is arranged to generate feature values
that are fed to the output circuit 109 which includes the feature set in the encoded
audio signal. Thus, in some embodiments, the audio encoder apparatus is arranged to
generate the encoded audio signal to include feature set values for the trained artificial
neural network of the audio decoder apparatus.
[0123] Correspondingly, the trained artificial neural network 209 of the audio decoder apparatus
may have input nodes for receiving the feature set values and thus the generation
of the scale factors may be dependent on dedicated data generated by a trained artificial
neural network of the audio encoder apparatus. This may provide a highly efficient
approach for the audio encoder apparatus to direct and assist the operation of the
audio decoder apparatus leading to a typically improved determination of scale factors
and thus an improved overall audio quality.
[0124] In practice, the feature values are typically also quantized at the encoder. The
quantization may be achieved using a Vector Quantized Variational Auto-Encoder (VQ-VAE)
neural network to learn a discrete set of codebook values during training. The feature
set values are thus quantized at the encoder using a VQ-VAE. These discrete codewords
are then used by the neural network at the decoder as a conditioning signal to guide
the estimation of scaling factors at the decoder.
[0125] It will be appreciated that the trained artificial neural network may be implemented
in many different ways and that many different algorithms, structures, etc. for implementing
an artificial neural network are known and that any suitable approach may be used
without detracting from the invention.
[0126] Similarly, it will be appreciated that many different approaches for training one
or more artificial neural networks are known and that any suitable approach may be
used. Specifically, one or more of the artificial neural networks may be trained using
training data comprising a number of training multichannel audio signals and a cost
function dependent on a difference measure for the multichannel audio signal received
by the encoder apparatus and the multichannel audio signal generated by the decoder
apparatus. In many embodiments, a training set including a large number of multichannel
audio signals may be generated/provided and (repeatedly) be fed to a cascade of the
audio encoder apparatus and audio decoder apparatus that is to be trained. The resulting
multichannel audio signals at the output of the audio decoder apparatus are compared
to the original multichannel audio signal fed to the audio encoder apparatus and a
cost value depending on the differences may be generated and used to update and train
the artificial neural network(s).
[0127] The cost/loss function may specifically be dependent on a perceptual difference measure
and thus the differences may be weighted/evaluated by considering how perceptually
relevant they are. For example, a frequency and/or time dependent perceptual weighting
may be applied to the difference between the original training multichannel audio
signal and the regenerated multichannel audio signal generated by the decoder.
[0128] The trained artificial neural network may be implemented using different artificial
neural network architectures, models, and processes in different embodiments.
[0129] In some embodiments, the trained artificial neural network may be a convolutional
artificial neural network which for example may receive samples of the downmix audio
signal and may e.g. directly generate scale factors.
[0130] In other embodiments, the trained artificial neural network may include both convolutional
layers as well as fully connected layers to generate scale factors for samples of
a downmix audio signal.
[0131] Indeed, an artificial neural network as used in the described functions may be a
network of nodes arranged in layers and with each node holding a node value. FIG.
7 illustrates an example of a section of an artificial neural network of fully connected
nodes, i.e. each node in a specific layer weighs and sums activations from the previous
layer.
[0132] The node value for a given node may be calculated to include contributions from some
or often all nodes of a previous layer of the artificial neural network. Specifically,
the node value for a node may be calculated as a weighted summation of the node values
of all the nodes output of the previous layer. Typically, a bias may be added and
the result may be subjected to an activation function. The activation function provides
an essential part of each neuron by typically providing a non-linearity. Such non-linearities
and activation functions provide a significant effect in the learning and adaptation
process of the artificial neural network. Thus, the node value is generated as a function
of the node values of the previous layer.
[0133] The artificial neural network may as illustrated in FIG. 7 specifically comprise
an input layer 701 comprising a plurality of nodes receiving the input data values
for the artificial neural network. Thus, the node values for nodes of the input layer
may typically directly be the input data values to the artificial neural network and
thus may not be calculated from other node values.
[0134] The artificial neural network may further comprise none, one, or more hidden layers
703, 705 or processing layers. For each of such layers, the node values are typically
generated as a function of the node values of the nodes of the previous layer, and
specifically a weighted combination and added bias followed by an activation function
(such as a Sigmoid, ReLU, or Tanh function may be applied).
[0135] Specifically, as shown in FIG. 6, each node, which may also be referred to as a neuron,
may receive input values (from nodes of a previous layer) and therefrom calculate
a node value as a function of these values. Often, this includes first generating
a value as a linear combination of the input values with each of these weighted by
a weight:

where w refers to weights, x refers to the nodes of the previous layer and n is an
index referring to the different nodes of the previous layer.
[0137] Other often used functions include a sigmoid function or a tanh function. In many
embodiments, the node output or value may be calculated using a plurality of functions.
For example, both a ReLU and Sigmoid function may be combined using an activation
function such as:

[0138] Such operations may be performed by each node of the artificial neural network (except
for typically the input nodes).
[0139] The artificial neural network may further comprise an output layer 707 which provides
the output from the artificial neural network, i.e. the output data of the artificial
neural network is the node values of the output layer. As for the hidden/ processing
layers, the output node values are generated by a function of the node values of the
previous layer. However, in contrast to the hidden/ processing layers where the node
values are typically not accessible or used further, the node values of the output
layer are accessible and provide the result of the operation of the artificial neural
network. In many embodiments, a Sigmoid activation may be used for real scale factors,
for example, since the outputs of such an activation are constrained to be between
0 and 1 which is desired in many practical embodiments.
[0141] WaveNet is an architecture used for the synthesis of time domain signals using dilated
causal convolution, and has been successfully applied to audio signals. For WaveNet
the following activation function is commonly used:

where * denotes a convolution operator, ⊙ denotes an element-wise multiplication
operator, σ(·) is a sigmoid function, k is the layer index, f and g denote filter
and gate, respectively, and W represents the weights of the learned artificial neural
network. The filter product of the equation may typically provide a filtering effect
with the gating product providing a weighting of the result which may in many cases
effectively allow the contribution of the node to be reduced to substantially zero
(i.e. it may allow or "cutoff" the node providing a contribution to other nodes thereby
providing a "gate" function). In different circumstances, the gate function may result
in the output of that node being negligible, whereas in other cases it would contribute
substantially to the output. Such a function may substantially assist in allowing
the artificial neural network to effectively learn and be trained.
[0142] An artificial neural network may in some cases further be arranged to include additional
contributions that allow the artificial neural network to be dynamically adapted or
customized for a specific desired property or characteristics of the generated output.
For example, a set of values may be provided to adapt the artificial neural network.
These values may be included by providing a contribution to some nodes of the artificial
neural network. These nodes may be specifically input nodes but may typically be nodes
of a hidden or processing layer. Such adaptation values may for example be weighted
and added as a contribution to the weighted summation/ correlation value for a given
node. For example, for WaveNet such adaptation values may be included in the activation
function. For example, the output of the activation function may be given as:

where
y is a vector representing the adaptation values and V represents suitable weights
for these values.
[0143] The above description relates to an artificial neural network approach that may be
suitable for many embodiments and implementations. However, it will be appreciated
that many other types and structures of artificial neural network may be used. Indeed,
many different approaches to artificial neural networks have been, and are being,
developed including artificial neural networks using complex structures and processes
that differ from the ones described above. The approach is not limited to any specific
artificial neural network approach and any suitable approach may be used without detracting
from the invention.
[0144] In many embodiments, the artificial neural network may advantageously be a generative
model such as a diffusion model artificial neural network. An example of such an approach
is illustrated in FIG. 8.
[0145] In many embodiments, such as in the example of FIG. 4, the inputs to the trained
artificial neural network are shown as the subband domain downmix signal. In practice,
however, and depending on the available learning capacity of the neural network, the
subbands of the downmix signal
Xm(
k, m) can be grouped into the parameter-estimation bands producing the downmix signal
Xm(
b, m)
. In some embodiments, a subband-aggregation front-end can be also performed/learned
by the artificial neural network.
[0146] In some embodiments, the artificial neural network may be implemented using an encoder-decoder
architecture known as a U-net consisting of a downsampling branch and an upsampling
branch and it may perform a multi-resolution feature extraction. Optionally, the model
can preserve signal details at each level via skip connections. The final layer may
feature a convolutional layer followed by an activation function such as a linear
activation unit.
[0147] The shape of the output of the trained artificial neural network may typically be
a
B ×
M block to match the number of required estimation bands and timeslots by the parameter
estimation block. It will be appreciated that the analysis filterbank and parameter
estimation bands sub-block may also be implemented as either disparate neural networks
or a single neural network that translates the time-domain multichannel input (e.g.
2-channels) into a set of parameter estimation bands (
B bands).
[0148] In case the input to the model is the subband downmix signal
Xm(
k, m), the
K ×
M complex-valued spectrogram may first be split into real and imaginary parts with
the two resulting spectrogram channels then being stacked. The model can thus be implemented
using 2D convolutional layers as shown in FIG. 9, where the downsampling is performed
along both time and frequency dimensions, controlled by the stride parameters of the
convolutional layers. A stride of two along a given dimension downsamples the data
by two. The example of FIG. 9 illustrates an embodiment of a neural network with 2D
convolutional as well as a residual layers.
[0149] The batch normalization blocks are used to normalize the data after certain layers
to improve training. Batch normalization is a method to normalize data between network
layers to improve and speed up training with larger step sizes which could lead to
instabilities without normalization. The ReLU blocks are rectified linear unit activations.
The upper row of layers performs downsampling (along time and frequency), while the
lower set of layers performs the upsampling, restoring to the desired time-frequency
resolution.
[0150] In the upsampling branch, upsampling can be performed using transpose convolutional
layers (denoted as deconvolutional layers in 9). Details of the residual layer, which
can learn composite functions and improve training due to their bypass functionality
are shown in Fig. 10. Due to the input bypass, the 2D convolutional layers are required
to preserve the shape of the input using appropriate padding and stride parameters.
Furthermore, by stacking multiple residual layers, a functions of ranging complexity
can be learned.
[0151] A high-level of the trained artificial neural networks block's inputs and outputs,
assuming hQMF stereo inputs is shown in Fig. 11 along with their dimensions. In some
embodiments, the trained artificial neural network output may be of shape 2 ×
B ×
M and represent a complex window that includes phase information. FIG. 11 illustrates
an example of high level processing of hQMF 2D signals where a mono downmix's real
and imaginary parts are stacked, producing two channels. The artificial neural network
output consists of a single channel set of B windows each of length M timeslots.
[0152] Finally, each row of the resulting NN output may be processed by a Softmax layer
such that the resulting window within each estimation band sums to 1,

[0153] In some embodiments, the output of the final artificial neural network block layer
may already produce a positive output, e.g., when implemented using a rectified linear
unit (ReLU), for example. In this case, the output can be normalized by the sum of
the outputs of each row (along the time slots). In other embodiments a Sigmoid activation
may be applied to each value of
WNN.
[0154] Artificial neural networks are adapted to specific purposes by a training process
which is used to adapt/ tune/ modify the weights and other parameters (e.g. bias)
of the artificial neural network. It will be appreciated that many different training
processes and algorithms are known for training artificial neural networks. A training
setup that may be used for the described artificial neural networks is illustrated
in FIG. 12.
[0155] Typically, training is based on large training sets where a large number of examples
of input data are provided to the network. Further, the output of the artificial neural
network is typically (directly or indirectly) compared to an expected or ideal result.
A cost function may be generated to reflect the desired outcome of the training process.
In a typical scenario known as supervised learning, the cost (or loss) function often
represents the distance between the prediction and the ground truth for a particular
input data. Based on the cost function, the weights may be changed and by reiterating
the process for the modified weights, the artificial neural network may be adapted
towards a state for which the cost function is minimized.
[0156] Different approaches may be used to train the artificial neural networks of the audio
encoder apparatus and audio decoder apparatus, and in particular an overall training
may seek to result in the output of the audio apparatuses being a multichannel audio
signal that most closely corresponds to the original multichannel audio signal. Thus,
the trained artificial neural networks may be trained to provide scale factors that
most effectively result in accurate reconstruction of the multichannel audio signal.
In such a case, a cost function based on the difference between a generated multichannel
audio signal and an original training multichannel audio signal may be used. In other
cases, more specific training of an artificial neural network may be used.
[0157] In some embodiments, where the audio encoder apparatus and the audio decoder apparatus
use identical networks, the two networks may use the exact same weights on thus the
updating may be constrained to be identical for the two networks.
[0158] In more detail, during a training step an artificial neural network may have two
different flows of information from input to output (forward pass) and from output
to input (backward pass). In the forward pass, the data is processed by the artificial
neural network as described above while in the backward pass the weights are updated
to minimize the cost function. Typically, such a backward propagation follows the
gradient direction of the cost function landscape. In other words, by comparing the
predicted output with the ground truth for a batch of data input, one can estimate
the direction in which the cost function is minimized and propagate backward, by updating
the weights accordingly. Other approaches known for training artificial neural networks
include for example Levenberg-Marquardt algorithm, the conjugate gradient method,
and the Newton method etc.
[0159] In the present case, training may specifically include a training set comprising
a potentially large number of multichannel audio signals or corresponding downmix
audio signals. The training sets may include audio signals representing a number of
different audio sources including e.g. recording of videos, movies, telecommunications,
etc. In some embodiments, the training data may even include non-audio data such as
a training being performed in combination with training data from other sources, such
as text data etc.
[0160] In some embodiments, training data may be multichannel audio signals in time segments
corresponding to the processing time intervals of the trained artificial neural network
being trained, e.g. the number of samples in a training multichannel audio signal
may correspond to a number of samples corresponding to the input nodes of the artificial
neural network(s) being trained. Each training example may thus correspond to one
operation of the artificial neural network(s) being trained. Usually, however, a batch
of training samples is considered for each step to speed up and smoothen the training
process. Furthermore, many upgrades to gradient descent are possible also to speed
up convergence or avoid local minima in the cost function landscape.
[0161] For each training multichannel audio signal, a training processor may perform a downmix
operation to generate a downmix audio signal and corresponding upmix parametric data.
Thus, the encoding process that is applied to the multichannel audio signal during
normal operation may also be applied to the training multichannel audio signal thereby
generating a downmix and the upmix parametric data.
[0162] Based on the cost value, the training processor may adapt the weights of the artificial
neural networks. For example, a back-propagation approach may be used. In particular,
the training processor may adjust the weights of one or all of the artificial neural
networks based on the cost value. For example, given the derivative (representing
the slope) of the weights with respect to the cost function the weights values are
modified to go in the direction of the slope. For a simple/minima account one can
refer to the training of the perceptron (single neuron) in case of backward pass of
a single data input.
[0163] The process may be iterated until the artificial neural networks are considered to
be trained. For example, training may be performed for a predetermined number of iterations.
As another example, training may be continued until the weights change is less than
a predetermined amount. Also very common, a validation stop is implemented where the
network is tested again a validation metric and stopped when reaching the expected
outcome.
[0164] Specifically, for the trained artificial neural network implementing a diffusion
model, the training process includes randomly selecting timestamps or corresponding
noise scaling values for each of the training examples in a given batch. For each
training example in a batch, an instance of a standard Gaussian noise signal (can
also correspond to some other standard distribution) is generated with the same dimension
as the training example. Then after scaling each of the noise signals with their corresponding
noise scaling values, these are added to the examples in the training batch. The model
then processes this batch of noisy data with additional inputs of the noise scaling
factors and features to estimate the original Gaussian noise signals. Therefore the
loss between the estimated and real noise signals are back propagated through the
network to optimize their weights using an optimizer such as the Adam optimizer.
[0165] In many embodiments, the loss or cost function may be perceptually weighted loss/cost
function. A perceptual relevancy metric may be used to weigh each of the samples during
training so that the artificial neural network learns meaningful weighting values
to help improve the e.g. the binaural cue preservation properties of the system.
[0166] Starting with the relevancy metric, in some embodiments, the relevancy metric may
be used to weigh frames or time slots containing transients that have been panned
to one stereo channel or the other or differently than the background or diffuse component
of the stereo signal. Particularly when the background signal is diffuse (e.g., applause),
foreground components are more easily perceived when they display different binaural
properties from the background (this is related to the theory of binaural masking
level differences (BMLD), where the masking threshold for a maskee drops if its binaural
properties are different from its masker), i.e. such as panned foreground clapping.
[0167] For a metric that can emphasize relevant transient signal components, let
PL(
b, m) and
PR(
b, m) denote the (smoothed) left and right energy levels for time-frequency tile in timeslot
m and band b.
[0168] Next, a normalized IID is computed from the left and right channels per (
b, m) tile,

that lies between -1 and 1.
[0169] For noise-like background signals, the IID values will be distributed randomly over
frequency within a slot. Therefore, next, a summary IID is calculated by integrating
(and weighting) IID values over frequency, leading to a profile of which slots could
contain signals of relevance,

where the weights can be set based on the properties of human hearing and/or a particular
class of signal (transient), and where ∑
b wb = 1.
[0170] Finally, a relevancy metric for a given frame (set of contiguous time slots), c corresponds
to the maximum value of the absolute summary IID,

and lies between 0 and 1. The above relevancy metric can be computed for each frame
within a given training batch to weigh its contribution to the average loss.
[0171] An example of computing the summary IID and corresponding relevancy metric is shown
in FIG. 13.
[0172] In some embodiments, one or more of the trained artificial neural network(s) may
be trained using a perceptually weighted loss function. The loss function may be perceptually
weighted by applying a higher weight to time frequency intervals representing a perceptually
significant signal component, such as a transient. In some embodiments, the perceptual
weighted loss function may be generated by applying a higher weighting to frames/segments
comprising a perceptually significant signal component/event, such as a transient
or loud sound, relative to frames/segments that do not include perceptually significant
signal component/events.
[0173] Thus, for perceptually relevant components of the signal, the weighting of the error
for the corresponding samples/time-slots/ time-frequency tiles and/or frames/segments
may be increased when determining the loss function. This may focus the learning on
addressing the perceptually significant situations and may reduce the risk that the
neural network training will focus more on learning the more common components for
correct resynthesis (such as the background sound) potentially leading to subpar performance
for the perceptually relevant samples in the training data.
[0174] A loss function for a batch of I frames each with relevancy
ci may be given by:

where

is the loss attributed to the individual frames in the training set. The weight
ci may be adapted to provide desired perceptual weighting.
[0175] In other embodiments, instead of taking the maximum over all frames to compute the
relevancy metric, the summary IID may simply be used to scale the contribution of
each time slot within a frame when computing the average error for that frame first
before averaging over the batch of frames.
[0176] In further embodiments, a training multichannel/stereo dataset can be composed of
separate transient and non-transient components that are mixed according to various
mixing ratios. The relevancy metric can then be computed by first scaling both transient
and non-transient components such that the non-transient components have equal power
in both channels before combining with the transient signal. The resulting summary
IID and relevancy metric may then better follow the BMLD model.
[0177] The loss function can consist of a number of terms, including the mean square error
(MSE) between the stereo input to the PS encoder and the estimated output stereo signal
at the decoder, where

[0178] Another loss term ignores the phase between the complex subband-domain input and
predicted stereo output signals. Such a term is more correlated to the perceptual
differences between two audio signals since phase does not have such a large impact
for most types of audio signals,

where
α is a compression exponent term (
α > 0).
[0179] Finally, the loss function for sample i can be calculated as a weighted sum of the
terms
L1 and
L2,

with 0 ≤
β ≤ 1. In other embodiments, the weights do not have to sum to 1.0, particularly to
more effectively control the relative contributions of loss terms to the total loss.
[0180] Other loss components can be included depending on the specific neural network configurations
at encoder and decoder as described in the next section.
[0181] The audio apparatus(s) may specifically be implemented in one or more suitably programmed
processors. In particular, the artificial neural networks may be implemented in one
more such suitably programmed processors. The different functional blocks, and in
particular the artificial neural network, may be implemented in separate processors
and/or may e.g. be implemented in the same processor. An example of a suitable processor
is provided in the following.
[0182] FIG. 14 is a block diagram illustrating an example processor 1400 according to embodiments
of the disclosure. Processor 1400 may be used to implement one or more processors
implementing an apparatus as previously described or elements thereof (including in
particular one more artificial neural network). Processor 1400 may be any suitable
processor type including, but not limited to, a microprocessor, a microcontroller,
a Digital Signal Processor (DSP), a Field ProGrammable Array (FPGA) where the FPGA
has been programmed to form a processor, a Graphical Processing Unit (GPU), an Application
Specific Integrated Circuit (ASIC) where the ASIC has been designed to form a processor,
or a combination thereof.
[0183] The processor 1400 may include one or more cores 1402. The core 1402 may include
one or more Arithmetic Logic Units (ALU) 1404. In some embodiments, the core 1402
may include a Floating Point Logic Unit (FPLU) 1406 and/or a Digital Signal Processing
Unit (DSPU) 1408 in addition to or instead of the ALU 1404.
[0184] The processor 1400 may include one or more registers 1412 communicatively coupled
to the core 1402. The registers 1412 may be implemented using dedicated logic gate
circuits (e.g., flip-flops) and/or any memory technology. In some embodiments the
registers 1412 may be implemented using static memory. The register may provide data,
instructions and addresses to the core 1402.
[0185] In some embodiments, processor 1400 may include one or more levels of cache memory
1410 communicatively coupled to the core 1402. The cache memory 1410 may provide computer-readable
instructions to the core 1402 for execution. The cache memory 1410 may provide data
for processing by the core 1402. In some embodiments, the computer-readable instructions
may have been provided to the cache memory 1410 by a local memory, for example, local
memory attached to the external bus 1416. The cache memory 1410 may be implemented
with any suitable cache memory type, for example, Metal-Oxide Semiconductor (MOS)
memory such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM),
and/or any other suitable memory technology.
[0186] The processor 1400 may include a controller 1414, which may control input to the
processor 1400 from other processors and/or components included in a system and/or
outputs from the processor 1400 to other processors and/or components included in
the system. Controller 1414 may control the data paths in the ALU 1404, FPLU 1406
and/or DSPU 1408. Controller 1414 may be implemented as one or more state machines,
data paths and/or dedicated control logic. The gates of controller 1414 may be implemented
as standalone gates, FPGA, ASIC or any other suitable technology.
[0187] The registers 1412 and the cache 1410 may communicate with controller 1414 and core
1402 via internal connections 1420A, 1420B, 1420C and 1420D. Internal connections
may be implemented as a bus, multiplexer, crossbar switch, and/or any other suitable
connection technology.
[0188] Inputs and outputs for the processor 1400 may be provided via a bus 1416, which may
include one or more conductive lines. The bus 1416 may be communicatively coupled
to one or more components of processor 1400, for example the controller 1414, cache
1410, and/or register 1412. The bus 1416 may be coupled to one or more components
of the system.
[0189] The bus 1416 may be coupled to one or more external memories. The external memories
may include Read Only Memory (ROM) 1432. ROM 1432 may be a masked ROM, Electronically
Programmable Read Only Memory (EPROM) or any other suitable technology. The external
memory may include Random Access Memory (RAM) 1433. RAM 1433 may be a static RAM,
battery backed up static RAM, Dynamic RAM (DRAM) or any other suitable technology.
The external memory may include Electrically Erasable Programmable Read Only Memory
(EEPROM) 1435. The external memory may include Flash memory 1434. The External memory
may include a magnetic storage device such as disc 1436. In some embodiments, the
external memories may be included in a system.
[0190] The invention can be implemented in any suitable form including hardware, software,
firmware or any combination of these. The invention may optionally be implemented
at least partly as computer software running on one or more data processors and/or
digital signal processors. The elements and components of an embodiment of the invention
may be physically, functionally and logically implemented in any suitable way. Indeed
the functionality may be implemented in a single unit, in a plurality of units or
as part of other functional units. As such, the invention may be implemented in a
single unit or may be physically and functionally distributed between different units,
circuits and processors.
[0191] Although the present invention has been described in connection with some embodiments,
it is not intended to be limited to the specific form set forth herein. Rather, the
scope of the present invention is limited only by the accompanying claims. Additionally,
although a feature may appear to be described in connection with particular embodiments,
one skilled in the art would recognize that various features of the described embodiments
may be combined in accordance with the invention. In the claims, the term comprising
does not exclude the presence of other elements or steps.
[0192] Furthermore, although individually listed, a plurality of means, elements, circuits
or method steps may be implemented by e.g. a single circuit, unit or processor. Additionally,
although individual features may be included in different claims, these may possibly
be advantageously combined, and the inclusion in different claims does not imply that
a combination of features is not feasible and/or advantageous. Also the inclusion
of a feature in one category of claims does not imply a limitation to this category
but rather indicates that the feature is equally applicable to other claim categories
as appropriate. Furthermore, the order of features in the claims do not imply any
specific order in which the features must be worked and in particular the order of
individual steps in a method claim does not imply that the steps must be performed
in this order. Rather, the steps may be performed in any suitable order. In addition,
singular references do not exclude a plurality. Thus references to "a", "an", "first",
"second" etc do not preclude a plurality. Reference signs in the claims are provided
merely as a clarifying example shall not be construed as limiting the scope of the
claims in any way.