<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE ep-patent-document PUBLIC "-//EPO//EP PATENT DOCUMENT 1.7.1//EN" "ep-patent-document-v1-7-1.dtd">
<!-- This XML data has been generated under the supervision of the European Patent Office -->
<ep-patent-document id="EP25160554A1" file="EP25160554NWA1.xml" lang="en" country="EP" doc-number="4800688" kind="A1" date-publ="20260902" status="n" dtd-version="ep-patent-document-v1-7-1">
<SDOBI lang="en"><B000><eptags><B001EP>ATBECHDEDKESFRGBGRITLILUNLSEMCPTIESILTLVFIROMKCYALTRBGCZEEHUPLSKBAHRIS..MTNORSMESMMAKHTNMDGE........</B001EP><B005EP>J</B005EP><B007EP>0009012-RPUB02</B007EP></eptags></B000><B100><B110>4800688</B110><B120><B121>EUROPEAN PATENT APPLICATION</B121></B120><B130>A1</B130><B140><date>20260902</date></B140><B190>EP</B190></B100><B200><B210>25160554.9</B210><B220><date>20250227</date></B220><B250>en</B250><B251EP>en</B251EP><B260>en</B260></B200><B400><B405><date>20260902</date><bnum>202636</bnum></B405><B430><date>20260902</date><bnum>202636</bnum></B430></B400><B500><B510EP><classification-ipcr sequence="1"><text>G10L  19/008       20130101AFI20250523BHEP        </text></classification-ipcr><classification-ipcr sequence="2"><text>G10L  25/30        20130101ALI20250523BHEP        </text></classification-ipcr></B510EP><B520EP><classifications-cpc><classification-cpc sequence="1"><text>G10L  19/008       20130101 FI20250519BHEP        </text></classification-cpc><classification-cpc sequence="2"><text>G10L  25/30        20130101 LA20250519BHEP        </text></classification-cpc></classifications-cpc></B520EP><B540><B541>de</B541><B542>KODIERTES AUDIOSIGNAL FÜR MEHRKANAL-AUDIOSIGNAL</B542><B541>en</B541><B542>ENCODED AUDIO SIGNAL FOR MULTICHANNEL AUDIO SIGNAL</B542><B541>fr</B541><B542>SIGNAL AUDIO CODÉ POUR SIGNAL AUDIO MULTICANAL</B542></B540><B590><B598>1</B598></B590></B500><B700><B710><B711><snm>Koninklijke Philips N.V.</snm><iid>102120190</iid><irf>2024P00621EP</irf><adr><str>High Tech Campus 34</str><city>5656 AE Eindhoven</city><ctry>NL</ctry></adr></B711></B710><B720><B721><snm>KECHICHIAN, Patrick</snm><adr><city>Eindhoven</city><ctry>NL</ctry></adr></B721><B721><snm>SCHUIJERS, Erik Gosuinus Petrus</snm><adr><city>Eindhoven</city><ctry>NL</ctry></adr></B721></B720><B740><B741><snm>Philips Intellectual Property &amp; Standards</snm><iid>101808802</iid><adr><str>High Tech Campus 34</str><city>5656 AE Eindhoven</city><ctry>NL</ctry></adr></B741></B740></B700><B800><B840><ctry>AL</ctry><ctry>AT</ctry><ctry>BE</ctry><ctry>BG</ctry><ctry>CH</ctry><ctry>CY</ctry><ctry>CZ</ctry><ctry>DE</ctry><ctry>DK</ctry><ctry>EE</ctry><ctry>ES</ctry><ctry>FI</ctry><ctry>FR</ctry><ctry>GB</ctry><ctry>GR</ctry><ctry>HR</ctry><ctry>HU</ctry><ctry>IE</ctry><ctry>IS</ctry><ctry>IT</ctry><ctry>LI</ctry><ctry>LT</ctry><ctry>LU</ctry><ctry>LV</ctry><ctry>MC</ctry><ctry>ME</ctry><ctry>MK</ctry><ctry>MT</ctry><ctry>NL</ctry><ctry>NO</ctry><ctry>PL</ctry><ctry>PT</ctry><ctry>RO</ctry><ctry>RS</ctry><ctry>SE</ctry><ctry>SI</ctry><ctry>SK</ctry><ctry>SM</ctry><ctry>TR</ctry></B840><B844EP><B845EP><ctry>BA</ctry></B845EP></B844EP><B848EP><B849EP><ctry>GE</ctry></B849EP><B849EP><ctry>KH</ctry></B849EP><B849EP><ctry>MA</ctry></B849EP><B849EP><ctry>MD</ctry></B849EP><B849EP><ctry>TN</ctry></B849EP></B848EP></B800></SDOBI>
<abstract id="abst" lang="en">
<p id="pa01" num="0001">An audio apparatus comprises a downmixer (105) generating a downmix audio signal by downmixing a multichannel audio signal. A trained artificial neural network (111) determines scale factors for time frequency intervals of the multichannel audio signal from samples of the multichannel audio signal or the downmix audio signal. A signal generator (113) generates a modified multichannel audio signal by applying the first scale factors to samples of the time frequency intervals and a parameter circuit (115) generates spatial parameters representing relative properties for channels of the first modified multi-channel audio signal. An output circuit (109) generates am encoded audio signal to comprise the downmix audio signal and the spatial parameters. The encoded audio signal may also include a residual signal. A complementary apparatus may receive the encoded audio signal and comprise a trained artificial neural network (111) which from the downmix audio signal can generate corresponding scale factors. An output multichannel audio signal is generated by upmixing the downmix audio signal using these scale factors and the received spatial parameters.
<img id="iaf01" file="imgaf001.tif" wi="78" he="89" img-content="drawing" img-format="tif"/></p>
</abstract>
<description id="desc" lang="en"><!-- EPO <DP n="1"> -->
<heading id="h0001">FIELD OF THE INVENTION</heading>
<p id="p0001" num="0001">The invention relates to generation and/or processing of an encoded audio signal for a multichannel audio signals and in particular, but not exclusively, to generation and/or encoding of stereo signals.</p>
<heading id="h0002">BACKGROUND OF THE INVENTION</heading>
<p id="p0002" num="0002">Spatial audio applications have become numerous and widespread and increasingly form at least part of many audiovisual experiences. Indeed, new and improved spatial experiences and applications are continuously being developed which results in increased demands for audio processing and rendering.</p>
<p id="p0003" num="0003">For example, in recent years, Virtual Reality (VR) and Augmented Reality (AR) have received increasing interest, and a number of implementations and applications are reaching the consumer market. Indeed, equipment is being developed for both rendering the experience as well as for capturing or recording suitable data for such applications. For example, relatively low-cost equipment is being developed for allowing gaming consoles to provide a full VR experience. It is expected that this trend will continue and indeed will increase in speed with the market for VR and AR reaching a substantial size within a short time scale. In the audio domain, a prominent field explores the reproduction and synthesis of realistic and natural spatial audio. The ideal aim is to produce natural audio sources such that the user cannot recognize the difference between a synthetic or an original one.</p>
<p id="p0004" num="0004">A lot of research and development effort has focused on providing efficient and high-quality audio encoding and audio decoding for spatial audio. A frequently used spatial audio representation is multichannel audio representations, including stereo representation, and efficient encoding of such multichannel audio based on downmixing multichannel audio signals to downmix channels with fewer channels have been developed. One of the main advances in low bit-rate audio coding has been the use of parametric multichannel coding where a downmix audio signal is generated together with parametric data that can be used to upmix the downmix audio signal to recreate the multichannel audio signal.</p>
<p id="p0005" num="0005">In particular, instead of traditional mid-side or intensity coding, parametric multichannel audio coding uses a downmix of a multichannel input signal to a lower number of channels (e.g. two to one) and multichannel image (stereo) parameters are extracted. Then the downmix audio signal is encoded using a more traditional audio coder (e.g. a mono audio encoder). The data of the downmix is combined with the encoded multichannel parameter data to generate a suitable audio bitstream. This<!-- EPO <DP n="2"> --> bitstream is then transmitted to the decoder, where the process is inverted. First the downmix audio signal is decoded, after which the multichannel audio signal is reconstructed, guided by the encoded multichannel image spatial parameters.</p>
<p id="p0006" num="0006">An example of stereo coding is described in <nplcit id="ncit0001" npl-type="s"><text>E. Schuijers, W. Oomen, B. den Brinker, J. Breebaart, "Advances in Parametric Coding for High-Quality Audio", 114th AES Convention, Amsterdam, The Netherlands, 2003, Preprint 5852</text></nplcit>. In the described approach, the downmixed mono signal is parametrized by exploiting the natural separation of the signal into three components (objects): transients, sinusoids, and noise. In <nplcit id="ncit0002" npl-type="s"><text>E. Schuijers, J. Breebaart, H. Purnhagen, J. Engdegård, "Low Complexity Parametric Stereo Coding", 116th AES, Berlin, Germany, 2004, Preprint 6073</text></nplcit> more details are provided describing how parametric stereo was realized with a low (decoder) complexity when combining it with Spectral Band Replication (SBR).</p>
<p id="p0007" num="0007">However, although Parametric Stereo (PS) and similar multichannel downmix based encoding/ decoding approaches were a leap forward from traditional stereo and multichannel coding, the approach is not optimal in all scenarios. In particular, known approaches tend to introduce some distortion, changes, artefacts etc. that may introduce differences between the (original) multichannel audio signal input to the encoder and the multichannel audio signal recreated at the decoder. Typically, the audio quality may be degraded and imperfect recreation of the multichannel occurs. Further, the data rate may still be higher than desired and/or the complexity/ resource usage of the involved processing may be higher than preferred.</p>
<p id="p0008" num="0008">Hence, an improved approach would be advantageous. In particular, an approach allowing increased flexibility, improved adaptability, an improved performance, improved audio quality, improved audio quality to data rate trade-off, reduced complexity and/or resource usage, facilitated implementation, and/or an improved spatial audio experience would be advantageous.</p>
<heading id="h0003">SUMMARY OF THE INVENTION</heading>
<p id="p0009" num="0009">Accordingly, the invention seeks to preferably mitigate, alleviate or eliminate one or more of the above mentioned disadvantages singly or in any combination.</p>
<p id="p0010" num="0010">According to an aspect of the invention there is provided an apparatus for generating an encoded audio signal for a multichannel audio signal, the apparatus comprising: a receiver arranged to receive the multichannel audio signal; a downmixer arranged to generate a downmix audio signal by downmixing the multichannel audio signal; an encoder arranged to encode the downmix audio signal to generate encoded downmix audio data; a trained artificial neural network arranged to determine first scale factors for (sets, blocks, contiguous groups of) time frequency intervals of the multichannel audio signal, the trained artificial neural network having input nodes for receiving samples of the multichannel audio signal and/or the downmix audio signal, and output nodes providing the first scale factors; a signal generator arranged to generate a first modified multichannel audio signal by applying the first scale factors to samples of the time frequency intervals; a parameter circuit arranged to generate a first set of<!-- EPO <DP n="3"> --> spatial parameters for the downmix audio signal (from the modified multichannel audio signal), the first set of spatial parameters representing relative properties for channels of the first modified multi-channel audio signal; and an output circuit arranged to generate the encoded audio signal to comprise the encoded audio data and first set of spatial parameters.</p>
<p id="p0011" num="0011">The approach may provide an improved audio experience in many embodiments. For many signals and scenarios, the approach may provide improved representation of a multichannel audio signal allowing improved generation/ reconstruction of the multichannel audio signal with an improved perceived audio quality. The approach may provide a particularly advantageous arrangement which may in many embodiments and scenarios allow a facilitated and/or improved possibility of utilizing artificial neural networks in audio processing, including typically audio encoding and/or decoding. The approach may allow an advantageous employment of artificial neural network(s) in generating a multichannel audio signal from a downmix audio signal.</p>
<p id="p0012" num="0012">The approach may provide an efficient implementation and may in many embodiments allow a reduced complexity and/or resource usage.</p>
<p id="p0013" num="0013">The approach may in many cases allow improved representation of particular signal components with high perceptual impact, such as e.g. transients. The approach may in many cases in particular mitigate or avoid temporal smearing and/or frequency distortion thereby allowing improved quality of the resulting multichannel audio signal.</p>
<p id="p0014" num="0014">The set of spatial parameters may comprise parameter (values) relating properties of the downmix audio signal to properties of the multichannel audio signal. The set of spatial parameters may comprise data being indicative of relative properties between channels of the modified multichannel audio signal. The set of spatial parameters may comprise data being indicative of differences in properties between channels of the modified multichannel audio signal. The set of spatial parameters may comprise data being perceptually relevant for the synthesis of the multichannel audio signal. The properties may for example be differences in phase and/or intensity and/or timing and/or correlation. The set of spatial parameters may comprise data including at least one of interchannel intensity differences (IID), interchannel timing differences (ITD), interchannel correlations (ICC) and/or interchannel phase differences (IPD) for channels of the modified multichannel audio signal. In particular, the set of spatial parameters may include one or more of an ICC, IPD, IID parameter as known from Parametric Stereo encoding/decoding.</p>
<p id="p0015" num="0015">The artificial neural network(s) may be a trained artificial neural network(s) trained by training data including training multichannel audio signals; the training employing a cost function dependent on a (spectro-temporal) envelope/shape/ difference between the multichannel audio signal and a multichannel audio signal synthesized from the encoded audio signal. The cost function may be dependent on a correlation between the input multichannel audio signal and the synthesized multichannel audio signal, and specifically may provide a decreasing cost function for a decreasing correlation. The cost function may be dependent on a difference between statistical properties of the synthesized<!-- EPO <DP n="4"> --> multichannel audio signal and the original multichannel audio signal, and specifically may provide a decreasing cost function for a decreasing correlation. Depending on the preferences and the decorrelation calculation/measurements/metric, it may in some embodiments be appropriate for the cost function to be dependent on a correlation between the synthesized multichannel audio signal and the original multichannel audio signal, and specifically to provide an increasing cost function for a decreasing correlation.</p>
<p id="p0016" num="0016">The scale factors may be provided for individual time-frequency intervals/ tiles. <b>In</b> some cases, one scale factor may be provided for each time-frequency interval of the multichannel audio signal.</p>
<p id="p0017" num="0017">The multichannel audio signal may specifically be a stereo signal. The downmix audio signal may specifically be a mono downmix audio signal.</p>
<p id="p0018" num="0018">In some embodiments, one or more of the trained artificial neural network(s) may be a neural network trained using a perceptually weighted loss function. The loss function may be perceptually weighted by applying a higher weight to time frequency intervals representing a perceptually significant signal component, such as a transient. In some embodiments, the perceptual weighted loss function may be generated by applying a higher weighting to (time) frames/segments comprising a perceptually significant signal component/event, such as a transient or loud sound, relative to frames/segments that do not include perceptually significant signal component/events.</p>
<p id="p0019" num="0019">According to an optional feature of the invention, the apparatus further comprises a scale factor circuit arranged to determine second scale factors for the time frequency intervals of the multichannel audio signals; and the signal generator is arranged to generate a second modified multichannel audio signal by applying the second scale factors to samples of the time frequency intervals; the parameter circuit is arranged to generate a second set of spatial parameters for the downmix audio signal from the second modified multichannel audio signal, the second set of spatial parameters representing relative properties for channels of the second modified multi-channel audio signal; and the output circuit is arranged to generate the encoded audio signal to comprise the second set of spatial parameters.</p>
<p id="p0020" num="0020">This may provide a particularly efficient and high-performance operation in many scenarios and may typically result in substantially improved audio quality of a multichannel audio signal generated from the encoded audio signal.</p>
<p id="p0021" num="0021">According to an optional feature of the invention, the scale factor circuit is arranged to determine the second set of scale factors from the first set of scale factors.</p>
<p id="p0022" num="0022">This may provide a particularly advantageous and high performance operation and/or implementation. The second set of scale factors may be determined so that the first and second set of scale factors have a constant combined value (w<sub>2</sub>=1-w<sub>1</sub>).</p>
<p id="p0023" num="0023">According to an optional feature of the invention, the scale factor circuit comprises a further trained artificial neural network arranged to determine the second scale factors, the further trained<!-- EPO <DP n="5"> --> artificial neural network having input nodes for receiving samples of the multichannel audio signal and/or the downmix audio signal, and output nodes providing the second scale factors.</p>
<p id="p0024" num="0024">This may provide a particularly advantageous and high performance operation and/or implementation.</p>
<p id="p0025" num="0025">According to an optional feature of the invention, a data rate for the first set of spatial parameters is different than a data rate for the second set of spatial parameters.</p>
<p id="p0026" num="0026">This may provide a particularly advantageous and high performance operation and/or implementation.</p>
<p id="p0027" num="0027">According to an optional feature of the invention, the first scale factors are complex values.</p>
<p id="p0028" num="0028">This may provide a particularly advantageous and high performance operation and/or implementation.</p>
<p id="p0029" num="0029">According to an optional feature of the invention, the apparatus comprises an additional trained artificial neural network arranged to generate feature set values from the multichannel audio signal, the additional trained artificial neural network having input nodes for receiving samples of the multichannel audio signal, and the output circuit is arranged to include the feature set values in the encoded audio signal.</p>
<p id="p0030" num="0030">This may provide a particularly advantageous and high performance operation and/or implementation.</p>
<p id="p0031" num="0031">According to an aspect of the invention, there is provided an apparatus for generating a multichannel audio signal from an encoded audio signal, the apparatus comprising: a receiver arranged to receive the encoded audio signal comprising encoded downmix audio data for a downmix of a multi-channel audio signal and a first set of spatial parameters, the first set of spatial parameters representing relative properties for channels of a first modified multi-channel audio signal resulting from applying first scale factors to samples of time frequency intervals of the multichannel audio signal; a decoder arranged to decode the encoded audio data to generate the downmix audio signal; a decorrelator arranged to decorrelate the downmix audio signal to generate a decorrelated signal; a trained artificial neural network arranged to determine the first scale factors for time frequency intervals of the multichannel audio signal, the trained artificial neural network having input nodes for receiving samples of the downmix audio signal and output nodes providing the first scale factors; an upmixer arranged to generate the multichannel audio signal by upmixing the downmix audio signal and the decorrelated signal in dependence on the first set of spatial parameters and the first scale factors.</p>
<p id="p0032" num="0032">According to an optional feature of the invention, the upmixer is arranged to generate each channel of the multichannel audio signal as a linear combination of the downmix audio signal and the decorrelated signals where weights of the linear combination are dependent on the first set of scale factors and the first set of spatial parameters.<!-- EPO <DP n="6"> --></p>
<p id="p0033" num="0033">This may provide a particularly advantageous and high performance operation and/or implementation.</p>
<p id="p0034" num="0034">According to an optional feature of the invention, the encoded audio signal comprises a second set of spatial parameters, the second set of spatial parameters representing relative properties for channels of a second modified multi-channel audio signal resulting from applying second scale factors to samples of time frequency intervals of the multichannel audio signal; the apparatus comprises a circuit arranged to generate the second scale factors; and the upmixer is arranged to determine a first upmixed multichannel audio signal from the downmix audio signal and the decorrelated signal using the first set of spatial parameters and the first scale factors, to determine a second upmixed multichannel audio signal from the downmix audio signal and the decorrelated signal using the second set of spatial parameters and the second scale factors, and to include a combination of the first upmixed multichannel audio signal and the second upmixed multichannel audio signal in the multichannel audio.</p>
<p id="p0035" num="0035">This may provide a particularly advantageous and high performance operation and/or implementation.</p>
<p id="p0036" num="0036">According to an optional feature of the invention, the encoded audio signal comprises feature set values for a trained network and the trained artificial neural network has input nodes for receiving the feature set values.</p>
<p id="p0037" num="0037">This may provide a particularly advantageous and high performance operation and/or implementation.</p>
<p id="p0038" num="0038">According to an optional feature of the invention, the trained artificial neural network of an encoder apparatus is different from the trained artificial neural network of a decoder apparatus.</p>
<p id="p0039" num="0039">This may provide a particularly advantageous and high performance operation and/or implementation.</p>
<p id="p0040" num="0040">According to an optional feature of the invention, one or more of the trained artificial neural networks is trained using training data comprising a number of training multichannel audio signals and a cost function dependent on a perceptual difference measure for the multichannel audio signal received by the encoder apparatus and the multichannel audio signal generated by the decoder apparatus.</p>
<p id="p0041" num="0041">This may provide a particularly advantageous and high performance operation and/or implementation.</p>
<p id="p0042" num="0042">According to an aspect of the invention, there is provided a method of operation for an audio apparatus for generating an encoded audio signal for a multichannel audio signal, the method comprising: receiving the multichannel audio signal; generating a downmix audio signal by downmixing the multichannel audio signal; encoding the downmix audio signal to generate encoded downmix audio data; a trained artificial neural network determining first scale factors for time frequency intervals of the multichannel audio signal, the trained artificial neural network having input nodes for receiving samples of the multichannel audio signal and/or the downmix audio signal, and output nodes providing the first scale factors; generating a first modified multichannel audio signal by applying the first scale factors to<!-- EPO <DP n="7"> --> samples of the time frequency intervals; generating a first set of spatial parameters for the downmix audio signal, the first set of spatial parameters representing relative properties for channels of the first modified multi-channel audio signal; and generating an encoded audio signal to comprise the encoded audio data and first set of spatial parameters.</p>
<p id="p0043" num="0043">According to an aspect of the invention, there is provided a method of operation for an audio apparatus, the method comprising: receiving the encoded audio signal comprising encoded downmix audio data for a downmix of a multi-channel audio signal and a first set of spatial parameters, the first set of spatial parameters representing relative properties for channels of a first modified multi-channel audio signal resulting from applying first scale factors to samples of time frequency intervals of the multichannel audio signal; decoding the encoded audio data to generate the downmix audio signal; decorrelating the downmix audio signal to generate a decorrelated signal; a trained artificial neural network determining the first scale factors for time frequency intervals of the multichannel audio signal, the trained artificial neural network having input nodes for receiving samples of the downmix audio signal and output nodes providing the first scale factors; and generating the multichannel audio signal by upmixing the downmix audio signal and the decorrelated signal in dependence on the first set of spatial parameters and the first scale factors.</p>
<p id="p0044" num="0044">These and other aspects, features and advantages of the invention will be apparent from and elucidated with reference to the embodiment(s) described hereinafter.</p>
<heading id="h0004">BRIEF DESCRIPTION OF THE DRAWINGS</heading>
<p id="p0045" num="0045">Embodiments of the invention will be described, by way of example only, with reference to the drawings, in which
<ul id="ul0001" list-style="none">
<li><figref idref="f0001">FIG. 1</figref> illustrates some elements of an example of an audio apparatus in accordance with some embodiments of the invention;</li>
<li><figref idref="f0002">FIG. 2</figref> illustrates some elements of an example of an audio apparatus in accordance with some embodiments of the invention;</li>
<li><figref idref="f0003">FIG. 3</figref> illustrates some elements of an example of a time to frequency converter for an audio apparatus in accordance with some embodiments of the invention;</li>
<li><figref idref="f0004">FIG. 4</figref> illustrates some elements of an example of an audio apparatus in accordance with some embodiments of the invention;</li>
<li><figref idref="f0005">FIG. 5</figref> illustrates an example of a stereo signal including a transient;</li>
<li><figref idref="f0006">FIG. 6</figref> illustrates an example of a neuron for an artificial neural network;</li>
<li><figref idref="f0007">FIG. 7</figref> illustrates an example of a structure of an artificial neural network;</li>
<li><figref idref="f0008">FIG. 8</figref> illustrates an example of a structure of a diffusion model approach;</li>
<li><figref idref="f0009">FIG. 9</figref> illustrates an example of a structure of an artificial neural network;</li>
<li><figref idref="f0010">FIG. 10</figref> illustrates an example of a structure of an artificial neural network;</li>
<li><figref idref="f0011">FIG. 11</figref> illustrates an example of a structure of an artificial neural network arrangement;<!-- EPO <DP n="8"> --></li>
<li><figref idref="f0012">FIG. 12</figref> illustrates an example of elements of an arrangement for training an artificial neural network;</li>
<li><figref idref="f0013">FIG. 13</figref> illustrates an example of an inter-channel level difference; and</li>
<li><figref idref="f0014">FIG. 14</figref> illustrates some elements of a possible arrangement of a processor for implementing elements of an audio apparatus in accordance with some embodiments of the invention.</li>
</ul></p>
<heading id="h0005">DETAILED DESCRIPTION OF SOME EMBODIMENTS OF THE INVENTION</heading>
<p id="p0046" num="0046"><figref idref="f0001">FIG. 1</figref> illustrates some elements of an audio apparatus arranged to generate an encoded audio signal representing a multichannel audio signal and <figref idref="f0002">FIG. 2</figref> illustrates elements of an audio apparatus arranged to generate a multichannel audio signal from an encoded audio signal such as that generated by the audio apparatus of <figref idref="f0001">FIG. 1</figref>. An audio distribution/communication system may accordingly comprise the audio encoder apparatus of <figref idref="f0001">FIG. 1</figref> generating an encoded audio signal for a multichannel audio signal with the encoded audio signal being transmitted to an audio decoder apparatus of <figref idref="f0002">FIG. 2</figref> which may recreate a local replica of the multichannel audio signal from the received encoded audio signal.</p>
<p id="p0047" num="0047">The audio encoder apparatus comprises a receiver 101 which is arranged to receive a multichannel audio signal that is to be encoded and represented by the encoded audio signal. The multichannel audio signal may typically be a stereo signal, and the following description will focus on this specific example.</p>
<p id="p0048" num="0048">The processing by the audio encoder apparatus is typically performed predominantly in the frequency/ subband domain and the following description will focus on such embodiments. However, in some embodiments, some or all of the processing may be performed in the time domain using time domain processing and indeed in many embodiments the processing may include both processing of a time domain and frequency domain representations of the signals.</p>
<p id="p0049" num="0049">In many embodiments, the multichannel audio signal may be received as a time domain signal and a frequency representation may be generated and used in the processing.</p>
<p id="p0050" num="0050">In some embodiments, the frequency domain multichannel audio signal may be received from a remote source but in many other embodiments a remote source may generate a time domain representation of the audio signal and a local frequency transformer may be arranged to generate a (time-)frequency representation of the multichannel audio signal.</p>
<p id="p0051" num="0051">In particular, in the example of <figref idref="f0001">FIG. 1</figref>, the audio apparatus comprises a filter bank 103 which is arranged to generate a frequency (subband) representation of a received time domain multichannel audio signal. Typically, the audio apparatus may comprise a filter bank 103 that is applied to each channel signal of the multichannel audio signal such that this is divided into frequency subbands.</p>
<p id="p0052" num="0052">The filter bank may be a Quadrature Mirror Filter (QMF) bank or may e.g. be implemented by a Fast Fourier Transform (FFT), but it will be appreciated that many other filter banks and approaches for dividing an audio signal into a plurality of subband signals are known and may be<!-- EPO <DP n="9"> --> used. The filter-bank may specifically be a complex-exponential modulated pseudo QMF bank, resulting in e.g. 32 or 64 complex-valued sub-band signals.</p>
<p id="p0053" num="0053">The processing is furthermore typically performed in time segments or time slots. In most embodiments, the audio signal is divided into time intervals/segments with a conversion to the frequency/subband domain by applying e.g. an FFT or QMF filtering to the samples of each signal. For example, each channel of the downmix audio signal may be divided into time segments of e.g. 2048, 1024, or 512 samples. These signals may then be processed to generate samples for e.g. 64, 32 or 16 subbands. Thus, a set of samples may be determined for each subband of the downmix audio signal.</p>
<p id="p0054" num="0054">It should be noted that the number of time domain samples is not directly coupled to the number of subbands. Typically, for a so-called critically sampled filterbank of N bands, every N input samples will lead to N sub-band samples (one for every sub-band). An oversampled filterbank will produce more output samples. E.g. for every N input samples, it would generate k*N output samples, i.e., k consecutive samples for every band.</p>
<p id="p0055" num="0055">In some embodiments, the subbands are generated to have the same bandwidth but in other embodiments subbands are generated to have different bandwidths, e.g. reflecting the sensitivity of human hearing to different frequencies.</p>
<p id="p0056" num="0056"><figref idref="f0003">FIG. 3</figref> illustrates an example of an approach for a filterbank 103 where different bandwidths are generated by a hybrid filter bank approach.</p>
<p id="p0057" num="0057">In the example, this is realized by a combination of a complex-exponential modulated pseudo QMF bank 301 and a small filter bank 303 for the lower frequency bands to realize a higher frequency resolution, as desired for binaural perception of the human auditory system. The result is a hybrid filterbank with logarithmic filter band center-frequency spacings that follow that of human perception similar to equivalent rectangular bandwidths (ERBs). In order to compensate for the delay of the filtering by the small filter bank 303, a delay 305 is introduced for higher frequency subbands.</p>
<p id="p0058" num="0058">In the specific example, a time-domain signal <i>x</i>[<i>n</i>] is fed through a downsampled complex-exponential modulated QMF bank with <i>K</i> bands. Each frame of 64 time domain samples <i>x</i>[<i>n</i>] results in one slot of QMF samples <i>X</i>[<i>k, m</i>] with <i>k</i> = (0, ... <i>, K</i> - 1) at slot m. The lower slots are then filtered by additional complex-modulated filterbanks splitting the lower bands further. The higher slots are delayed ensuring that the filtered signals of the lower bands are in sync with the higher bands as the filtering introduces a delay. This finally results in a structure where for every 64 time-domain samples <i>x</i>[<i>n</i>], one slot m of hybrid QMF samples <i>Y</i>[<i>l, m</i>] is produced with <i>l</i> = (0, ..., <i>L</i> - 1) at slot <i>m</i>, e.g. with a total number of hybrid bands M = 77.</p>
<p id="p0059" num="0059">The receiver 101 is fed to a downmixer 105 which proceeds to generate a downmix audio signal and typically a mono downmix audio signal. In many embodiments and scenarios, the downmixer 105 may generate a mono downmix audio signal from a received stereo signal, e.g. simply by summing<!-- EPO <DP n="10"> --> the channels of the stereo signal in accordance with a standard Parametric Stereo (PS) downmix approach.</p>
<p id="p0060" num="0060">The downmixer 105 is coupled to an audio encoder 107 which is arranged to encode the downmix audio signal to generate encoded downmix audio data for the downmix audio signal. In many embodiments, a mono audio encoding may be used to encode a mono downmix audio signal. It will be appreciated that any suitable audio encoding approach may be used including in particular any suitable encoding Standard or algorithm for encoding an audio signal to generate audio data representing the audio signal. In particular, the mono downmix audio signal may be encoded using a standard encoding.</p>
<p id="p0061" num="0061">The audio encoder 107 is coupled to an output circuit 109 which is arranged to generate the encoded audio signal. The encoded audio signal is generated to include the encoded downmix audio data, and it will be appreciated that the encoded audio signal may be generated in accordance with any suitable structure, format etc. In many embodiments, the output circuit 109 may be arranged to generate the encoded audio signal to follow a Standard data format of audio signals.</p>
<p id="p0062" num="0062">The encoder audio apparatus of <figref idref="f0001">FIG. 1</figref> is further arranged to generate and include spatial (upmix) parameters in the encoded audio signal. However, rather than use a conventional approach, such as e.g. known from PS, where the spatial parameters are generated from the multichannel audio signal and are generated to reflect properties of the multichannel audio signal, the encoding audio apparatus of <figref idref="f0001">FIG. 1</figref> is arranged to perform specific processing to generate signal components of the multichannel audio signal and with the spatial parameters being generated to reflect these.</p>
<p id="p0063" num="0063">In particular, the audio encoder apparatus includes a first (encoder) trained artificial neural network 111 which determines scale factors, also referred to as the first (encoder) scale factors, for time frequency tiles of the multichannel audio signal. In many embodiments, the trained artificial neural network 111 may generate scale factors for individual time frequency intervals/times as generated by the filter banks 103. The first trained artificial neural network 111 is arranged to receive samples of the generated mono downmix audio signal and/or the multichannel audio signal and to generate the scale factors based on these values. The first trained artificial neural network 111 has input nodes that receive these samples and output nodes that provide scale factors for the time frequency tiles.</p>
<p id="p0064" num="0064">The first trained artificial neural network 111 may be arranged to generate scale factors that provide a filtering/windowing function which is applied to the multichannel audio signal to estimate/emphasize perceptually relevant parts of the channel signals of the multichannel audio signal. The scale factors may form a window function that may be a masking function between that weighs the relevancy of time-frequency tiles across each subband of (a frame of) the multichannel audio signal. In many cases, the scale factors may be determined as values within a given interval such as between 0 and 1. The first trained artificial neural network 111 may be trained using a perceptual-based loss function which assigns a relevancy weighting to training samples that include perceptually relevant signal components to help guide the scale factor/window estimation process to emphasize such components after<!-- EPO <DP n="11"> --> training is completed. The first trained artificial neural network 111 may thus be trained to generate scale factors that filter/mask/extract/estimate perceptually relevant parts of the multichannel audio signal.</p>
<p id="p0065" num="0065">It will be appreciated that, as described in more detail later, the first trained artificial neural network 111, may be implemented in many different ways and that any suitable approach or neural network implementation may be used. Similarly, as will be described later, different approaches for training the trained artificial neural network may be used in different applications and embodiments.</p>
<p id="p0066" num="0066">The first trained artificial neural network 111 is coupled to a signal generator 113 which is also coupled to the receiver 101. The signal generator 113 receives the multichannel audio signal and the first scale factors and is arranged to generate a first modified multichannel audio signal by applying the first scale factors to samples of the time frequency intervals of the multichannel audio signal. Specifically, in many embodiments, each frequency sample of the multichannel audio signal may be multiplied by the scale factor determined by the first trained artificial neural network 111 for the time frequency tile to which the frequency sample belongs. The signal generator 113 thus generates a modified multichannel audio signal that has been generated based on the scale factors determined by the first trained artificial neural network 111.</p>
<p id="p0067" num="0067">As will be described in more detail, a second modified multichannel audio signal may often be generated by the signal generator 113, and indeed the second modified multichannel audio signal may often be generated as a residual signal such that the first and second modified multichannel audio signal may form a decomposition of the original multichannel audio signal. In some embodiments, more than two modified multichannel audio signals may be generated and in particular the signal generator 113 may perform a signal decomposition of the multichannel audio signal into more than two (modified) multichannel audio signal components.</p>
<p id="p0068" num="0068">The modified multichannel audio signal(s) are fed to a parameter estimation circuit 115 which is coupled to the signal generator 113 and to the output circuit 109. The parameter circuit 115 is arranged to determine spatial parameters for the modified multichannel audio signal(s) where the spatial parameters are indicative of relative properties between the channels of the modified multichannel audio signal(s). Specifically, the parameter circuit 115 may determine a first set of spatial parameters for the first modified multichannel audio signal where the first set of spatial parameters represent relative properties of the first modified multichannel audio signal.</p>
<p id="p0069" num="0069">Typically, the upmix/spatial parameters may be indicative of time differences, phase differences, level/intensity differences and/or a measure of similarity, such as correlation, between channels of the modified multichannel audio signal(s). Typically, the spatial parameters are provided on a per time and per frequency basis (time frequency tiles). For example, new parameters may periodically be provided for a set of subbands. Parameters may specifically include Inter-channel phase difference (IPD), Overall phase difference (OPD), Inter-channel correlation (ICC), Inter-channel Intensity/Level difference parameters as known from Parametric Stereo encoding (as well as from higher channel encodings).<!-- EPO <DP n="12"> --></p>
<p id="p0070" num="0070">For stereo applications, the spatial parameters may specifically be spatial parameters as used for encoding a stereo signal using a Parametric Stereo (PS) encoding of the stereo signal, such as for example by a mono downmix audio signal given as: <maths id="math0001" num=""><math display="block"><mi>m</mi><mo>=</mo><mi>c</mi><mfenced separators=""><mi>l</mi><mo>+</mo><mi>r</mi></mfenced></math><img id="ib0001" file="imgb0001.tif" wi="21" he="5" img-content="math" img-format="tif"/></maths> where the parameter <i>c</i> is chosen such that the power of the stereo signal is preserved in the downmix, the power being defined using the 2-norm: <maths id="math0002" num=""><math display="block"><msup><mfenced open="‖" close="‖"><mi>m</mi></mfenced><mn>2</mn></msup><mo>=</mo><msup><mfenced open="‖" close="‖"><mi>l</mi></mfenced><mn>2</mn></msup><mo>+</mo><msup><mfenced open="‖" close="‖"><mi>r</mi></mfenced><mn>2</mn></msup></math><img id="ib0002" file="imgb0002.tif" wi="33" he="5" img-content="math" img-format="tif"/></maths> and thus e.g.: <maths id="math0003" num=""><math display="block"><mi>c</mi><mo>=</mo><msqrt><mfrac><mrow><msup><mfenced open="‖" close="‖"><mi>l</mi></mfenced><mn>2</mn></msup><mo>+</mo><msup><mfenced open="‖" close="‖"><mi>r</mi></mfenced><mn>2</mn></msup></mrow><msup><mfenced open="‖" close="‖" separators=""><mi>l</mi><mo>+</mo><mi>r</mi></mfenced><mn>2</mn></msup></mfrac></msqrt></math><img id="ib0003" file="imgb0003.tif" wi="33" he="15" img-content="math" img-format="tif"/></maths></p>
<p id="p0071" num="0071">The PS parameters are specifically an Inter-channel Intensity Difference IID, an Inter-channel Correlation ICC, and in some cases an Inter-channel Phase Difference IPD parameter. These may specifically be defined/determined as: <maths id="math0004" num=""><math display="block"><mi>IID</mi><mo>=</mo><mfrac><msup><mfenced open="‖" close="‖"><mi>l</mi></mfenced><mn>2</mn></msup><msup><mfenced open="‖" close="‖"><mi>r</mi></mfenced><mn>2</mn></msup></mfrac></math><img id="ib0004" file="imgb0004.tif" wi="20" he="10" img-content="math" img-format="tif"/></maths> <maths id="math0005" num=""><math display="block"><mi>ICC</mi><mo>=</mo><mfrac><mfenced open="|" close="|"><mfenced open="〈" close="〉"><mi>l</mi><mi>r</mi></mfenced></mfenced><msqrt><mrow><msup><mfenced open="‖" close="‖"><mi>l</mi></mfenced><mn>2</mn></msup><msup><mfenced open="‖" close="‖"><mi>r</mi></mfenced><mn>2</mn></msup></mrow></msqrt></mfrac></math><img id="ib0005" file="imgb0005.tif" wi="32" he="11" img-content="math" img-format="tif"/></maths> <maths id="math0006" num=""><math display="block"><mi>IPD</mi><mo>=</mo><mi>arg</mi><mfenced open="〈" close="〉"><mi>l</mi><mi>r</mi></mfenced></math><img id="ib0006" file="imgb0006.tif" wi="31" he="4" img-content="math" img-format="tif"/></maths> where the complex-valued inner product is defined as: <maths id="math0007" num=""><math display="block"><mfenced open="〈" close="〉"><mi mathvariant="normal">x</mi><mi mathvariant="normal">y</mi></mfenced><mo>=</mo><mstyle displaystyle="true"><munder><mo>∑</mo><mrow><mo>∀</mo><mi>i</mi></mrow></munder><msub><mi>x</mi><mi>i</mi></msub><mo>⋅</mo><msubsup><mi>y</mi><mi>i</mi><mo>∗</mo></msubsup></mstyle></math><img id="ib0007" file="imgb0007.tif" wi="33" he="9" img-content="math" img-format="tif"/></maths> and <maths id="math0008" num=""><math display="block"><msup><mfenced open="‖" close="‖"><mi>x</mi></mfenced><mn>2</mn></msup><mo>=</mo><mfenced open="〈" close="〉"><mi>x</mi><mi>x</mi></mfenced></math><img id="ib0008" file="imgb0008.tif" wi="29" he="5" img-content="math" img-format="tif"/></maths><!-- EPO <DP n="13"> --></p>
<p id="p0072" num="0072">It should be noted that in some embodiments, the input multichannel audio signal may be weighted by windows (time direction) and grouping/kernels (frequency direction). In some embodiments, this may be affected by the scale factors determined by the first trained artificial neural network 111 and the parameter estimation and/or downmixing may also take these into account.</p>
<p id="p0073" num="0073">The spatial parameters are typically provided for specific time frequency tiles, and thus specifically each parameter value is generated/provided for a given frequency subband and for a given time segment.</p>
<p id="p0074" num="0074">The parameter circuit 115 may determine spatial parameters that are indicative of relative properties of the channels (channel signals) of the stereo signal.</p>
<p id="p0075" num="0075">The spatial parameters may comprise sets of spatial parameters, each set of spatial parameters comprising at least one of: a level difference parameter indicative of a level difference between channels of the multichannel audio signal; a correlation parameter indicative of a coherence between channels of the multichannel audio signal; a timing difference parameter indicative of a timing difference between channels of the multichannel audio signal, and a phase difference parameter indicative of a phase difference between channels of the multichannel audio signal.</p>
<p id="p0076" num="0076">The output circuit 109 is arranged to include the generated spatial parameters in the encoded audio signal and thus the audio encoder apparatus is arranged to generate a representation of a multichannel audio signal by a (typically mono) downmix signal and associated spatial data which however is not providing data on relative properties of the channels of the multichannel audio signal but rather is indicative of relative properties of only a particular component/part of the channels of the multichannel audio signal with this component/part being determined/identified by a trained artificial neural network.</p>
<p id="p0077" num="0077">The audio decoder apparatus of <figref idref="f0002">FIG. 2</figref> comprises a receiver 201 which is arranged to receive an encoded audio signal and specifically it may receive the encoded audio signal from the audio encoder apparatus of <figref idref="f0001">FIG. 1</figref>. Thus, the receiver 201 may receive an encoded audio signal that comprises encoded downmix audio data for a downmix of a multi-channel audio signal and a first set of spatial parameters where the first set of spatial parameters represent relative properties for channels for the first modified multi-channel audio signal that is generated from applying first scale factors to samples of time frequency intervals of the multichannel audio signal. The first set of spatial parameters may be spatial parameters for a multichannel audio signal component of the multichannel audio signal where the multichannel audio signal component is related to the multichannel audio signal by a set of scale factors for time frequency intervals of the multichannel audio signal.</p>
<p id="p0078" num="0078">The audio decoder apparatus comprises a decoder 203 arranged to decode the encoded downmix audio data to generate the downmix audio signal <i>x<sub>m</sub></i>(<i>n</i>). The downmix audio signal is fed to a decorrelator 205 which is arranged to decorrelate the downmix audio signal to generate a decorrelated signal <i>x<sub>d</sub></i>(<i>n</i>) that is perceptually similar to, but uncorrelated with <i>x<sub>m</sub></i>(<i>n</i>). It will be appreciated that many<!-- EPO <DP n="14"> --> different algorithms are known for decorrelating an audio signal and that any suitable approach may be used by the audio decoder apparatus of <figref idref="f0002">FIG. 2</figref>.</p>
<p id="p0079" num="0079">The downmix audio signal, the decorrelated signal, and the received spatial parameters are fed to an upmixer (207) which is arranged to generate (a local copy of) the multichannel audio signal by upmixing the downmix audio signal and the decorrelated signal.</p>
<p id="p0080" num="0080">The audio decoder apparatus further comprises a trained artificial neural network, henceforth also referred to as the first decoder trained artificial neural network 209 which receives the downmix audio signal. The first decoder trained artificial neural network 209 is arranged and trained to determine scale factors, henceforth referred to as the first decoder scale factors, for time frequency intervals of the downmix audio signal. The first decoder trained artificial neural network 209 has input nodes for receiving samples of the downmix audio signal and output nodes providing the first scale factors. In many embodiments, the first decoder trained artificial neural network 209 may correspond directly to the first encoder trained artificial neural network 111 and indeed in many embodiments both of these trained networks may receive only the downmix audio signal for generating the scale factors and the first decoder trained artificial neural network 209 may typically generate scale factors that are identical or very close to those generated by the first encoder trained artificial neural network 111.</p>
<p id="p0081" num="0081">In many embodiments, the first encoder trained artificial neural network 111 at the audio encoder apparatus is arranged to generate the first encoder scale factors based (only) on the downmix audio signal which is also available at the audio decoder apparatus. Accordingly, by having the first decoder trained artificial neural network 209 determining the first decoder scale factors based on the locally available downmix audio signal, a similar processing can be performed by the trained artificial neural networks resulting in identical or similar scale factors being determined.</p>
<p id="p0082" num="0082">The upmixer 207 may specifically be arranged to generate a first multichannel audio signal by upmixing (at least) the downmix audio signal and the decorrelated signal based on the received first set of spatial parameters. The audio decoder apparatus comprises the first decoder trained artificial neural network 209 which may generate scale factors corresponding to those generated at the audio source device 101. Indeed, in many embodiments, the trained networks may be identical and both operate on the downmix audio signal and thus will generate (substantially) identical scale factors. The audio decoder apparatus can accordingly apply the scale factors resulting in it being possible to generate (a local replica) of the modified multichannel audio signal. The audio decoder apparatus may then generate an output multichannel audio signal directly corresponding to the modified multichannel audio signal, or may typically combine the recreated modified multichannel audio signal with other signal components, such as typically (a local replica of) a second modified signal representing a residual signal. Thus, in many embodiments, the output multichannel audio signal may be generated as a combination of a plurality of generated multichannel audio signals of which the local replica of the first modified multichannel audio signal is one.<!-- EPO <DP n="15"> --></p>
<p id="p0083" num="0083">In the approach, a multichannel audio signal may accordingly be represented by a downmix signal with spatial parameters that are representing relative properties for a signal component extracted from the multichannel audio signal rather than from the multichannel audio signal as a whole. The spatial parameters specifically reflects properties of a signal component of the multichannel audio signal that has been extracted/determined using a trained artificial neural network as will be described in more detail later. Further, at the decoder side, the signal component may be recreated from the mono downmix audio signal and a complementary artificial neural network may allow the scale factors to be determined allowing for the output signal to be generated while appropriately scaling/weighting the signal component.</p>
<p id="p0084" num="0084">In many embodiments, the audio encoder apparatus may include a scale factor circuit 117 which is arranged to further determine a second set of spatial parameters which describe properties of a second signal component. Specifically, the signal generator 113 may be arranged to generate a second modified multichannel audio signal by applying the second scale factors to samples of the time frequency intervals. In many embodiments, the second scale factors may be determined such that the second modified multichannel audio signal is a residual signal for the first modified multichannel audio signal. The second scale factors may for example be determined such that the sum of the first and second scale factors are constant (unity) for all time frequency tiles. For example, in many embodiments, second scale factors may be generated as 1-w1 where w1 is the first scale factor for the same time-frequency tile.</p>
<p id="p0085" num="0085">The parameter circuit 115 may accordingly also be fed the second modified multichannel audio signal, which specifically may be a residual signal, and may proceed to determine a second set of spatial parameters that represent/describe relative properties for channels of the second modified multi-channel audio signal. The same approaches as used for the first modified multichannel audio signal may be used and e.g. the second set of spatial parameters may include IID, ICC, and IPD values. The second set of spatial parameters are fed to the output circuit 109 which includes it in the encoded audio signal.</p>
<p id="p0086" num="0086">At the audio decoder apparatus side, the second set of spatial parameters are also fed to the upmixer 207 which is arranged to upmix the received mono downmix audio signal and the decorrelated signal based on the second set of spatial parameters to generate a second upmixed multichannel audio signal. The audio decoder apparatus may further proceed to generate second scale factors to correspond to the second decoder scale factors. For example, in the case where the second modified multichannel audio signal is a residual signal generated by the second scale factors and the first scale factors summing to a fixed, constant value, the second decoder scale factors may be generated by subtracting the first decoder scale factors determined by the first encoder trained artificial neural network 111 from the fixed, constant value.</p>
<p id="p0087" num="0087">The first and second upmixed multichannel audio signals may then be combined based on the first and second scale factors. For example, a linear combination of the frequency domain samples of the first and second upmixed multichannel audio signal with the weights of the first upmixed<!-- EPO <DP n="16"> --> multichannel audio signal corresponding to the first scale factors and the weights of the second upmixed multichannel audio signal corresponding to the second scale factors.</p>
<p id="p0088" num="0088">Thus, in some embodiments, trained artificial neural networks may be used to isolate and extract signal components for which individual spatial parameters are determined and communicated with the audio decoder apparatus performing differentiated upmixing based on the different spatial parameters sets. In particular, in some embodiments, the audio encoder apparatus may perform a signal decomposition operation on the multichannel audio signal and generate individual and separate spatial parameters for the different signal components. The different components may be separately upmixed at the decoder side with the resulting upmixed signals being combined. The approach has been found to provide a particularly advantageous operation in many embodiments, applications, and scenarios.</p>
<p id="p0089" num="0089">As a specific example, the multichannel audio signal may be a stereo signal and the audio decoder apparatus may employ the structure of <figref idref="f0004">FIG. 4</figref> where two complementary modified stereo signals may be generated as: <maths id="math0009" num=""><math display="block"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mfenced><mi>b</mi><mi>m</mi></mfenced><mo>=</mo><msub><mover accent="true"><mi>W</mi><mo>¯</mo></mover><mi mathvariant="italic">NN</mi></msub><mfenced><mi>b</mi><mi>m</mi></mfenced><mi>Y</mi><mfenced><mi>b</mi><mi>m</mi></mfenced><mo>,</mo><mspace width="1ex"/><mi>Y</mi><mo>∈</mo><mfenced open="{" close="}"><mi>L</mi><mi>R</mi></mfenced></math><img id="ib0009" file="imgb0009.tif" wi="75" he="5" img-content="math" img-format="tif"/></maths> <maths id="math0010" num=""><math display="block"><mover accent="true"><mi>Y</mi><mo>^</mo></mover><mfenced><mi>b</mi><mi>m</mi></mfenced><mo>=</mo><msub><mi>W</mi><mi mathvariant="italic">NN</mi></msub><mfenced><mi>b</mi><mi>m</mi></mfenced><mi>Y</mi><mfenced><mi>b</mi><mi>m</mi></mfenced><mo>,</mo><mspace width="1ex"/><mi>Y</mi><mo>∈</mo><mfenced open="{" close="}"><mi>L</mi><mi>R</mi></mfenced></math><img id="ib0010" file="imgb0010.tif" wi="75" he="5" img-content="math" img-format="tif"/></maths> where b and m are indexes representing time frequency times (frequency time band b and time interval/frame m), <i><o ostyle="single">W</o><sub>NN</sub></i> (<i>b, m</i>) represent the scale factor for time frequency tile b,m, and <i><o ostyle="single">W</o><sub>NN</sub></i> (<i>b</i>, <i>m</i>) = 1 <i>- W<sub>NN</sub></i>(<i>b, m</i>)<i>.</i></p>
<p id="p0090" num="0090">In some cases, a windowing function <i>W<sub>T</sub></i>(<i>m</i>) across frames may also be applied, such as e.g. a triangular windowing function: <maths id="math0011" num=""><math display="block"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mfenced><mi>b</mi><mi>m</mi></mfenced><mo>=</mo><msub><mi>W</mi><mi>T</mi></msub><mfenced><mi>m</mi></mfenced><msub><mover accent="true"><mi>W</mi><mo>¯</mo></mover><mi mathvariant="italic">NN</mi></msub><mfenced><mi>b</mi><mi>m</mi></mfenced><mi>Y</mi><mfenced><mi>b</mi><mi>m</mi></mfenced><mo>,</mo><mspace width="1ex"/><mi>Y</mi><mo>∈</mo><mfenced open="{" close="}"><mi>L</mi><mi>R</mi></mfenced></math><img id="ib0011" file="imgb0011.tif" wi="88" he="5" img-content="math" img-format="tif"/></maths> <maths id="math0012" num=""><math display="block"><mover accent="true"><mi>Y</mi><mo>^</mo></mover><mfenced><mi>b</mi><mi>m</mi></mfenced><mo>=</mo><msub><mi>W</mi><mi>T</mi></msub><mfenced><mi>m</mi></mfenced><msub><mi>W</mi><mi mathvariant="italic">NN</mi></msub><mfenced><mi>b</mi><mi>m</mi></mfenced><mi>Y</mi><mfenced><mi>b</mi><mi>m</mi></mfenced><mo>,</mo><mspace width="1ex"/><mi>Y</mi><mo>∈</mo><mfenced open="{" close="}"><mi>L</mi><mi>R</mi></mfenced></math><img id="ib0012" file="imgb0012.tif" wi="88" he="5" img-content="math" img-format="tif"/></maths></p>
<p id="p0091" num="0091">The parameter circuit 115 may then proceed to generate a set of spatial parameters for each of the modified stereo signals, and specifically the spatial parameters <maths id="math0013" num=""><math display="inline"><mover accent="true"><mi mathvariant="italic">ICC</mi><mo>^</mo></mover><mfenced><mi>b</mi></mfenced><mo>,</mo><mover accent="true"><mi mathvariant="italic">IID</mi><mo>^</mo></mover><mfenced><mi>b</mi></mfenced><mo>,</mo><mspace width="1ex"/><mover accent="true"><mi mathvariant="italic">IPD</mi><mo>^</mo></mover><mfenced><mi>b</mi></mfenced></math><img id="ib0013" file="imgb0013.tif" wi="42" he="6" img-content="math" img-format="tif" inline="yes"/></maths>, and those based on the complementary window <i><o ostyle="single">W</o><sub>NN</sub>, <o ostyle="single">ICC</o></i>(<i>b</i>), <i><o ostyle="single">IID</o></i>(<i>b</i>), <i><o ostyle="single">IPD</o></i>(<i>b</i>), may be determined. The spatial parameters are then included in the encoded audio signal.<!-- EPO <DP n="17"> --></p>
<p id="p0092" num="0092">The bit stream then includes the set of windowed and complementary PS parameters <maths id="math0014" num=""><math display="inline"><mfenced open="{" close="}" separators=""><mover accent="true"><mi mathvariant="italic">ICC</mi><mo>^</mo></mover><mfenced><mi>b</mi></mfenced><mo>,</mo><mover accent="true"><mi mathvariant="italic">IID</mi><mo>^</mo></mover><mfenced><mi>b</mi></mfenced><mo>,</mo><mover accent="true"><mi mathvariant="italic">IPD</mi><mo>^</mo></mover><mfenced><mi>b</mi></mfenced></mfenced></math><img id="ib0014" file="imgb0014.tif" wi="45" he="6" img-content="math" img-format="tif" inline="yes"/></maths> and {<i><o ostyle="single">ICC</o></i>(<i>b'</i>), <i><o ostyle="single">IID</o></i>(<i>b'</i>), <i><o ostyle="single">IPD</o></i>(<i>b</i>')}, where b' is used to emphasize different groupings of the complementary PS parameters.</p>
<p id="p0093" num="0093">Because the windowing function/scale factors <i><o ostyle="single">W</o><sub>NN</sub></i> are determined from the downmix signal <i>X<sub>m</sub>,</i> the window functions/scale factors themselves do not have to be included in the encoded audio signal as they can independently be estimated at the audio decoder apparatus.</p>
<p id="p0094" num="0094">At the audio decoder apparatus, the downmix signal is decoded by the mono decoder 203 and is by the first decoder trained artificial neural network 209 used to (re)estimate the set of scale factors/ windows, <i>W<sub>NN</sub>.</i> The estimated windows/scale factors, together with the mono downmix subband signal <i>X<sub>m</sub></i>(<i>k,</i> m) and the decorrelated signal <i>X<sub>d</sub></i>(<i>k,</i> m) generated by the decorrelator 205, are fed to the upmixer 207. The upmixer 207 may then perform upmixing to generate local replicas of the two modified stereo signals <i><o ostyle="single">Y</o></i>(<i>b,</i> m) and <i>Ŷ</i>(<i>b, m</i>)<i>.</i> These may then be combined into a single stereo output signal with the weighting of each being set to match the scale factors.</p>
<p id="p0095" num="0095">As the upmixing is typically linear (e.g. a matrix multiplication with coefficients dependent on the spatial parameters) and as the scale factors are the same for all channels of the multichannel audio signal (specifically both channels of a stereo signal), the order of the upmixing and scaling of the signals by the scale factors may be changed. Therefore, in some embodiments, the scale factors may be applied to the downmix audio signal and the decorrelated signal prior to the upmixing. Thus, in the specific example, the downmix audio signal and the decorrelated signal for each of the two signal components may be determined as: <maths id="math0015" num=""><math display="block"><msub><mover accent="true"><mi>X</mi><mo>^</mo></mover><mi>m</mi></msub><mfenced><mi>k</mi><mi>m</mi></mfenced><mo>=</mo><msub><mi>W</mi><mi>T</mi></msub><mfenced><mi>m</mi></mfenced><msub><mi>W</mi><mi mathvariant="italic">NN</mi></msub><mfenced><mi>k</mi><mi>m</mi></mfenced><msub><mi>X</mi><mi>m</mi></msub><mfenced><mi>k</mi><mi>m</mi></mfenced></math><img id="ib0015" file="imgb0015.tif" wi="72" he="5" img-content="math" img-format="tif"/></maths> <maths id="math0016" num=""><math display="block"><msub><mover accent="true"><mi>X</mi><mo>^</mo></mover><mi>d</mi></msub><mfenced><mi>k</mi><mi>m</mi></mfenced><mo>=</mo><msub><mi>W</mi><mi>T</mi></msub><mfenced><mi>m</mi></mfenced><msub><mi>W</mi><mi mathvariant="italic">NN</mi></msub><mfenced><mi>k</mi><mi>m</mi></mfenced><msub><mi>X</mi><mi>d</mi></msub><mfenced><mi>k</mi><mi>m</mi></mfenced></math><img id="ib0016" file="imgb0016.tif" wi="70" he="5" img-content="math" img-format="tif"/></maths> <maths id="math0017" num=""><math display="block"><msub><mover accent="true"><mi>X</mi><mo>¯</mo></mover><mi>m</mi></msub><mfenced><mi>k</mi><mi>m</mi></mfenced><mo>=</mo><msub><mi>W</mi><mi>T</mi></msub><mfenced><mi>m</mi></mfenced><mfenced separators=""><mn>1</mn><mo>−</mo><msub><mi>W</mi><mi mathvariant="italic">NN</mi></msub><mfenced><mi>k</mi><mi>m</mi></mfenced></mfenced><msub><mi>X</mi><mi>m</mi></msub><mfenced><mi>k</mi><mi>m</mi></mfenced></math><img id="ib0017" file="imgb0017.tif" wi="83" he="5" img-content="math" img-format="tif"/></maths> <maths id="math0018" num=""><math display="block"><msub><mover accent="true"><mi>X</mi><mo>¯</mo></mover><mi>d</mi></msub><mfenced><mi>k</mi><mi>m</mi></mfenced><mo>=</mo><msub><mi>W</mi><mi>T</mi></msub><mfenced><mi>m</mi></mfenced><mfenced separators=""><mn>1</mn><mo>−</mo><msub><mi>W</mi><mi mathvariant="italic">NN</mi></msub><mfenced><mi>k</mi><mi>m</mi></mfenced></mfenced><msub><mi>X</mi><mi>d</mi></msub><mfenced><mi>k</mi><mi>m</mi></mfenced></math><img id="ib0018" file="imgb0018.tif" wi="81" he="5" img-content="math" img-format="tif"/></maths></p>
<p id="p0096" num="0096">The upmixer 207 may then upmix each set of signals using the received spatial parameters for the individual signal component to produce the reconstructed left and right stereo signal <i>L</i>(<i>k, m</i>) and <i>R</i>(<i>k, m</i>)<i>.</i> This is achieved using upmix matrices according to: <maths id="math0019" num=""><math display="block"><mfenced open="[" close="]"><mtable><mtr><mtd><mover accent="true"><mi>L</mi><mo>^</mo></mover><mfenced><mi>k</mi><mi>m</mi></mfenced></mtd></mtr><mtr><mtd><mover accent="true"><mi>R</mi><mo>^</mo></mover><mfenced><mi>k</mi><mi>m</mi></mfenced></mtd></mtr></mtable></mfenced><mo>=</mo><mfenced open="[" close="]"><mtable equalrows="true" equalcolumns="true"><mtr><mtd><mrow><msub><mover accent="true"><mi>H</mi><mo>^</mo></mover><mn>11</mn></msub><mfenced><mi>k</mi></mfenced></mrow></mtd><mtd><mrow><msub><mover accent="true"><mi>H</mi><mo>^</mo></mover><mn>12</mn></msub><mfenced><mi>k</mi></mfenced></mrow></mtd></mtr><mtr><mtd><mrow><msub><mover accent="true"><mi>H</mi><mo>^</mo></mover><mn>21</mn></msub><mfenced><mi>k</mi></mfenced></mrow></mtd><mtd><mrow><msub><mover accent="true"><mi>H</mi><mo>^</mo></mover><mn>22</mn></msub><mfenced><mi>k</mi></mfenced></mrow></mtd></mtr></mtable></mfenced><mfenced open="[" close="]"><mtable><mtr><mtd><msub><mover accent="true"><mi>X</mi><mo>^</mo></mover><mi>m</mi></msub><mfenced><mi>k</mi><mi>m</mi></mfenced></mtd></mtr><mtr><mtd><msub><mover accent="true"><mi>X</mi><mo>^</mo></mover><mi>d</mi></msub><mfenced><mi>k</mi><mi>m</mi></mfenced></mtd></mtr></mtable></mfenced></math><img id="ib0019" file="imgb0019.tif" wi="76" he="11" img-content="math" img-format="tif"/></maths> and<!-- EPO <DP n="18"> --> <maths id="math0020" num=""><math display="block"><mfenced open="[" close="]"><mtable><mtr><mtd><mover accent="true"><mi>L</mi><mo>¯</mo></mover><mfenced><mi>k</mi><mi>m</mi></mfenced></mtd></mtr><mtr><mtd><mover accent="true"><mi>R</mi><mo>¯</mo></mover><mfenced><mi>k</mi><mi>m</mi></mfenced></mtd></mtr></mtable></mfenced><mo>=</mo><mfenced open="[" close="]"><mtable equalrows="true" equalcolumns="true"><mtr><mtd><mrow><msub><mover accent="true"><mi>H</mi><mo>¯</mo></mover><mn>11</mn></msub><mfenced><mi>k</mi></mfenced></mrow></mtd><mtd><mrow><msub><mover accent="true"><mi>H</mi><mo>¯</mo></mover><mn>12</mn></msub><mfenced><mi>k</mi></mfenced></mrow></mtd></mtr><mtr><mtd><mrow><msub><mover accent="true"><mi>H</mi><mo>¯</mo></mover><mn>21</mn></msub><mfenced><mi>k</mi></mfenced></mrow></mtd><mtd><mrow><msub><mover accent="true"><mi>H</mi><mo>¯</mo></mover><mn>22</mn></msub><mfenced><mi>k</mi></mfenced></mrow></mtd></mtr></mtable></mfenced><mfenced open="[" close="]"><mtable><mtr><mtd><msub><mover accent="true"><mi>X</mi><mo>¯</mo></mover><mi>m</mi></msub><mfenced><mi>k</mi><mi>m</mi></mfenced></mtd></mtr><mtr><mtd><msub><mover accent="true"><mi>X</mi><mo>¯</mo></mover><mi>d</mi></msub><mfenced><mi>k</mi><mi>m</mi></mfenced></mtd></mtr></mtable></mfenced></math><img id="ib0020" file="imgb0020.tif" wi="76" he="11" img-content="math" img-format="tif"/></maths> where the entries <maths id="math0021" num=""><math display="inline"><msub><mover accent="true"><mi>H</mi><mo>^</mo></mover><mi mathvariant="italic">ij</mi></msub><mfenced><mi>k</mi></mfenced><mo>=</mo><mi>f</mi><mfenced separators=""><mover accent="true"><mi mathvariant="italic">ICC</mi><mo>^</mo></mover><mfenced><mi>b</mi></mfenced><mo>,</mo><mover accent="true"><mi mathvariant="italic">IID</mi><mo>^</mo></mover><mfenced><mi>b</mi></mfenced><mo>,</mo><mover accent="true"><mi mathvariant="italic">IPD</mi><mo>^</mo></mover><mfenced><mi>b</mi></mfenced></mfenced></math><img id="ib0021" file="imgb0021.tif" wi="65" he="7" img-content="math" img-format="tif" inline="yes"/></maths> and <i><o ostyle="single">H</o><sub>ij</sub></i>(<i>k</i>) = <i>f</i>(<i><o ostyle="single">ICC</o></i>(<i>b</i>), <i><o ostyle="single">IID</o></i>(<i>b</i>), <i><o ostyle="single">IPD</o>(b)),</i> where <i>f</i>(·) is a function that translates the PS parameters into the corresponding upmix matrix entries per subband and is known in the prior art.</p>
<p id="p0097" num="0097">Finally, the two components may be combined and specifically summed to produce the reconstructed left and right stereo signals: <maths id="math0022" num=""><math display="block"><mi>L</mi><mo>′</mo><mfenced><mi>k</mi><mi>m</mi></mfenced><mo>=</mo><mover accent="true"><mi>L</mi><mo>^</mo></mover><mfenced><mi>k</mi><mi>m</mi></mfenced><mo>+</mo><mover accent="true"><mi>L</mi><mo>¯</mo></mover><mfenced><mi>k</mi><mi>m</mi></mfenced></math><img id="ib0022" file="imgb0022.tif" wi="52" he="5" img-content="math" img-format="tif"/></maths> <maths id="math0023" num=""><math display="block"><mi>R</mi><mo>′</mo><mfenced><mi>k</mi><mi>m</mi></mfenced><mo>=</mo><mover accent="true"><mi>R</mi><mo>^</mo></mover><mfenced><mi>k</mi><mi>m</mi></mfenced><mo>+</mo><mover accent="true"><mi>R</mi><mo>¯</mo></mover><mfenced><mi>k</mi><mi>m</mi></mfenced></math><img id="ib0023" file="imgb0023.tif" wi="54" he="5" img-content="math" img-format="tif"/></maths></p>
<p id="p0098" num="0098">A synthesis filterbank may then be used to reconstruct left and right time-domain signals, e.g. using an overlap-add process.</p>
<p id="p0099" num="0099"><b>In</b> contrast to e.g. traditional PS encoding, the current approach may provide a highly advantageous operation in many embodiments and may in particular in many embodiments provide an improved perceived audio quality.</p>
<p id="p0100" num="0100">In particular, in comparison to conventional approaches where the spatial parameters are typically generated to represent signal properties over an entire frame, the current approach may dynamically adapt to focus the parameters on particular perceptually relevant signal components.</p>
<p id="p0101" num="0101">Specifically, current PS encoding has a low update rate for spatial parameters in order to keep the bit rate low (typically one parameter set per frame) which results in the parameters reflecting average properties over a frame. Indeed, if e.g. some static windowing is introduced, this will tend to statically emphasize the middle of the frame thus averaging out the parameter estimation over frequency bands and biasing the estimate towards the most dominant component, which is usually the stationary background in this case.</p>
<p id="p0102" num="0102">For example, consider the depiction of a foreground transient component amid background noise as exemplified in <figref idref="f0005">FIG. 5</figref> where a single transient T is present in the right channel R and for only a short time interval of the frame. <figref idref="f0005">FIG. 5</figref> provides a depiction of left-panned transient spectrogram amid background applause. The checkered background noise is used to emphasize the random nature of the time-frequency distribution. The asterisk represents the binaural parameter estimation point. Transients can typically display wideband characteristics relative to background noise which has a fall-off given by room absorption/reflectivity. Here, ideally for low frequencies where the power of the background is larger than the transient, the PS parameters should reflect the background signal's properties, while for the foreground component, above a certain frequency, its parameters should dominate. However, given a static symmetric window centered at the asterisk (boundary between current<!-- EPO <DP n="19"> --> and previous frame), the transients PS parameters will be biased to the weak background (not shown here).</p>
<p id="p0103" num="0103">This results in the loss of specific binaural cues related to signal components such as transients which have a shorter time-span relative to the frame size. The described approach may employ a neural network-based approach that may be coupled to a relevancy weighting applied in the loss function during training, producing data-driven subband scale factors that can be used to generate a signal component that highlight perceptually relevant components and in particular with frequency selective emphasis on perceptually relevant components.</p>
<p id="p0104" num="0104">In some embodiments and scenarios where a plurality of sets of spatial parameters are included for different signal components, the data rate/size for the different sets may be set to be different. For example, the data rate/size for spatial parameters for the first modified multichannel audio signal determined by the scale factors from the first trained artificial neural network 111 may be substantially higher than for the spatial parameters for the residual signal. For example, the quantization (whether in frequency or of the actual value) may be different for different signal components and thus the encoding of the generated parameter values may be different. In many embodiments, the average number of bits allocated/used per parameter value may be different for (at least two) different sets of spatial parameters.</p>
<p id="p0105" num="0105">Thus, a different trade-off between accuracy and data rate may be implemented for the different sets of spatial parameters. This may allow improved audio quality versus data rate in many scenarios.</p>
<p id="p0106" num="0106">For example, for the PS example described above, the PS parameters for the signal component based on the complementary adaptive window <i><o ostyle="single">W</o><sub>NN</sub></i> may have a lower resolution and correspond to a single triplet of PS parameters (<i><o ostyle="single">ICC</o>, <o ostyle="single">IID</o>, <o ostyle="single">IPD</o></i>) than for the signal component based on the adaptive window <i>W<sub>NN</sub>.</i> This may typically result in a scenario where the data rate of the generated encoded audio signal is only negligibly higher than the data rate for a standard PS coder while allowing a better audio quality for the reproduced stereo signal.</p>
<p id="p0107" num="0107">In some embodiments, the scale factors may be generated as scale values and thus may simply be scalar gain values for frequency domain values of the multichannel audio signal(s). However, in most embodiments, the trained artificial neural networks are arranged to generate complex valued scale factors and the application to (typically complex) frequency domain values of the multichannel audio signal(s) are performed in the complex domain (specifically by performing complex multiplications). Thus, in most embodiments, the trained artificial neural networks are arranged to not only generate values that include an amplitude compensation but also to include a phase property and modification.</p>
<p id="p0108" num="0108">As described above, in some embodiments, the audio decoder apparatus may in some embodiments be arranged to generate the second set of scale factors for determining the second modified multichannel audio signal from the first set of scale factors. For example, for each frequency band, the audio decoder apparatus may determine a scale factor for the second set of scale factors as a function of a<!-- EPO <DP n="20"> --> scale factor from the first set of scale factors. In particular, in many embodiments, the second set of scale factors may as previously described be determined such that a sum of the first and second scale factors for a given frequency band add up to a fixed value. The fixed value may specifically be unity and thus <maths id="math0024" num=""><math display="block"><msub><mi>c</mi><mrow><mn>2</mn><mo>,</mo><mi>k</mi><mo>,</mo><mi>m</mi></mrow></msub><mo>=</mo><mn>1</mn><mo>−</mo><msub><mi>c</mi><mrow><mn>1</mn><mo>,</mo><mi>k</mi><mo>,</mo><mi>m</mi></mrow></msub></math><img id="ib0024" file="imgb0024.tif" wi="32" he="5" img-content="math" img-format="tif"/></maths> where c<sub>1,k,m</sub> and c<sub>2,k,m</sub> are the scale factors of respectively the first and second set of scale factors for frequency band k and time slot m. It will be appreciated that the audio decoder apparatus may proceed to perform the same operation to generate second scale factors and that these may be used in the upmixing/combination to generate the output multichannel audio signal as previously described.</p>
<p id="p0109" num="0109">In some embodiments, the audio encoder apparatus may be arranged to determine second scale factors for a second multichannel audio signal which represents a second perceptually relevant signal component that is not (necessarily) a residual signal.</p>
<p id="p0110" num="0110">In particular, in some embodiments, the audio encoder apparatus may include a second trained artificial neural network 117 which may generate second scale factors from the downmix audio signal and/or from the multichannel audio signal. The second trained artificial neural network 117 has input nodes for receiving samples of the multichannel audio signal and/or the downmix audio signal, and output nodes providing the second scale factors. The second trained artificial neural network 117 may typically be trained to extract a second modified multichannel audio signal that comprises perceptually relevant information that is not included in the first modified multichannel audio signal. For example, the second modified multichannel audio signal may represent a second transient, a dominant point source signal, etc.</p>
<p id="p0111" num="0111">The parameter circuit 115 may then proceed to determine a second set of spatial parameters for the second modified multichannel audio signal and the output circuit 109 may include the second set of spatial parameters in the encoded audio signal. Thus, the generated encoded audio signal includes spatial parameters for upmixing two (or more) intermediate multichannel audio signals that may represent different perceptually relevant properties/components of the original multichannel audio signal.</p>
<p id="p0112" num="0112">In many embodiments, the audio decoder apparatus may in addition to generating scale factors for two or more multichannel signal components also proceed to generate a residual signal comprising the signal components that are not included in the two or more multichannel signal components.</p>
<p id="p0113" num="0113">Correspondingly, the audio decoder apparatus may include additional trained artificial neural networks that determine the set of second scale factors. A corresponding intermediate multichannel audio signal may be generated by an upmixing based on the second spatial parameters and the resulting multichannel audio signal may be included in the combination to generate the output multichannel audio<!-- EPO <DP n="21"> --> signal with the weights being dependent on the generated second scale factors (or equivalently the weighting may be included in the upmix operation).</p>
<p id="p0114" num="0114">In many embodiments, the trained networks of the audio encoder apparatus and the audio decoder apparatus are substantially identical and may in particular be identical in terms of design, structure, dimensions etc. Further, in many embodiments, the trained artificial neural networks may be trained identically, and indeed by be initialized with values that result from the same training process. Indeed, in some embodiments, the encoded audio signal may include configuration data for the trained artificial neural network(s) of the audio decoder apparatus and the audio decoder apparatus may be arranged to initialize the trained artificial neural network with the received data values.</p>
<p id="p0115" num="0115">However, in other embodiments, the trained artificial neural networks of the audio encoder apparatus and the audio decoder apparatus may be different. In particular, in some embodiments, a trained artificial neural network of the audio decoder apparatus may have a different structure, size, dimension, or configuration than the corresponding trained artificial neural network of the audio encoder apparatus. For example, the audio decoder apparatus may in many embodiments employ a trained artificial neural network that is substantially less complex and resource demanding than for the audio encoder apparatus thereby allowing implementation in smaller and less capable devices, such as e.g. portable devices.</p>
<p id="p0116" num="0116">In such cases, the trained artificial neural networks may for example be trained by a joint training process simultaneously including both the audio encoder apparatus and the audio decoder apparatus and where a loss function is e.g. based on comparing the original multichannel audio signal to the replicated multichannel audio signal and with both weights/parameters of the trained artificial neural network of the audio encoder apparatus and the audio decoder apparatus being updated.</p>
<p id="p0117" num="0117">The approach may thus utilize a trained artificial neural network at both the audio encoder apparatus and the audio decoder apparatus. In some embodiments, the same neural network architecture and weights are used at both encoder and decoder, with the neural network weights being shared by the encoder and decoder during training.</p>
<p id="p0118" num="0118">In another embodiment of the invention, the same or different neural network architectures may be used at encoder and decoder with different weights being trained during training. In such a case, additional loss terms may be included in the total loss including the difference between the estimated windows of the encoder neural network and decoder neural network, <maths id="math0025" num=""><math display="block"><msub><mi>L</mi><mn>3</mn></msub><mfenced><msubsup><mi>W</mi><mi mathvariant="italic">NN</mi><mi mathvariant="italic">enc</mi></msubsup><msubsup><mi>W</mi><mi mathvariant="italic">NN</mi><mi mathvariant="italic">dec</mi></msubsup></mfenced><mo>=</mo><mfrac><mn>1</mn><mi mathvariant="italic">BM</mi></mfrac><mstyle displaystyle="true"><munder><mo>∑</mo><mi>b</mi></munder><munder><mo>∑</mo><mi>m</mi></munder><msup><mfenced open="[" close="]" separators=""><msubsup><mi>W</mi><mi mathvariant="italic">NN</mi><mi mathvariant="italic">enc</mi></msubsup><mfenced><mi>b</mi><mi>m</mi></mfenced><mo>−</mo><msubsup><mi>W</mi><mi mathvariant="italic">NN</mi><mi mathvariant="italic">dec</mi></msubsup><mfenced><mi>b</mi><mi>m</mi></mfenced></mfenced><mn>2</mn></msup></mstyle></math><img id="ib0025" file="imgb0025.tif" wi="106" he="12" img-content="math" img-format="tif"/></maths></p>
<p id="p0119" num="0119">In some cases, a non-neural network-based approach (e.g., based on signal processing and segmentation methods) could be used at the encoder side to estimate the subband windows <i><o ostyle="single">W</o><sub>NN</sub></i> and the<!-- EPO <DP n="22"> --> neural network at the decoder side may learn to adapt its weights to produce an identical set of windows at the decoder.</p>
<p id="p0120" num="0120">In some embodiments the neural network at the decoder side may not produce the same sets of subband windows <i><o ostyle="single">W</o><sub>NN</sub></i> as at the encoder side, due to reconstruction errors introduced by the mono downmix audio decoder.</p>
<p id="p0121" num="0121">In yet other embodiments, as an alternative to the approach where the encoder and decoder effectively both have the same artificial neural network to estimate the scale factors, it may be desirable to trade-off some of the complexity of the decoder by transmitting intermediate latent space variables (feature values) e.g. representing the scale factors.</p>
<p id="p0122" num="0122">Thus, in some embodiments, the audio encoder apparatus may further include a feature set trained artificial neural network 119 which is arranged to generate feature values that are fed to the output circuit 109 which includes the feature set in the encoded audio signal. Thus, in some embodiments, the audio encoder apparatus is arranged to generate the encoded audio signal to include feature set values for the trained artificial neural network of the audio decoder apparatus.</p>
<p id="p0123" num="0123">Correspondingly, the trained artificial neural network 209 of the audio decoder apparatus may have input nodes for receiving the feature set values and thus the generation of the scale factors may be dependent on dedicated data generated by a trained artificial neural network of the audio encoder apparatus. This may provide a highly efficient approach for the audio encoder apparatus to direct and assist the operation of the audio decoder apparatus leading to a typically improved determination of scale factors and thus an improved overall audio quality.</p>
<p id="p0124" num="0124">In practice, the feature values are typically also quantized at the encoder. The quantization may be achieved using a Vector Quantized Variational Auto-Encoder (VQ-VAE) neural network to learn a discrete set of codebook values during training. The feature set values are thus quantized at the encoder using a VQ-VAE. These discrete codewords are then used by the neural network at the decoder as a conditioning signal to guide the estimation of scaling factors at the decoder.</p>
<p id="p0125" num="0125">It will be appreciated that the trained artificial neural network may be implemented in many different ways and that many different algorithms, structures, etc. for implementing an artificial neural network are known and that any suitable approach may be used without detracting from the invention.</p>
<p id="p0126" num="0126">Similarly, it will be appreciated that many different approaches for training one or more artificial neural networks are known and that any suitable approach may be used. Specifically, one or more of the artificial neural networks may be trained using training data comprising a number of training multichannel audio signals and a cost function dependent on a difference measure for the multichannel audio signal received by the encoder apparatus and the multichannel audio signal generated by the decoder apparatus. In many embodiments, a training set including a large number of multichannel audio signals may be generated/provided and (repeatedly) be fed to a cascade of the audio encoder apparatus and audio decoder apparatus that is to be trained. The resulting multichannel audio signals at the output of<!-- EPO <DP n="23"> --> the audio decoder apparatus are compared to the original multichannel audio signal fed to the audio encoder apparatus and a cost value depending on the differences may be generated and used to update and train the artificial neural network(s).</p>
<p id="p0127" num="0127">The cost/loss function may specifically be dependent on a perceptual difference measure and thus the differences may be weighted/evaluated by considering how perceptually relevant they are. For example, a frequency and/or time dependent perceptual weighting may be applied to the difference between the original training multichannel audio signal and the regenerated multichannel audio signal generated by the decoder.</p>
<p id="p0128" num="0128">The trained artificial neural network may be implemented using different artificial neural network architectures, models, and processes in different embodiments.</p>
<p id="p0129" num="0129">In some embodiments, the trained artificial neural network may be a convolutional artificial neural network which for example may receive samples of the downmix audio signal and may e.g. directly generate scale factors.</p>
<p id="p0130" num="0130">In other embodiments, the trained artificial neural network may include both convolutional layers as well as fully connected layers to generate scale factors for samples of a downmix audio signal.</p>
<p id="p0131" num="0131">Indeed, an artificial neural network as used in the described functions may be a network of nodes arranged in layers and with each node holding a node value. <figref idref="f0007">FIG. 7</figref> illustrates an example of a section of an artificial neural network of fully connected nodes, i.e. each node in a specific layer weighs and sums activations from the previous layer.</p>
<p id="p0132" num="0132">The node value for a given node may be calculated to include contributions from some or often all nodes of a previous layer of the artificial neural network. Specifically, the node value for a node may be calculated as a weighted summation of the node values of all the nodes output of the previous layer. Typically, a bias may be added and the result may be subjected to an activation function. The activation function provides an essential part of each neuron by typically providing a non-linearity. Such non-linearities and activation functions provide a significant effect in the learning and adaptation process of the artificial neural network. Thus, the node value is generated as a function of the node values of the previous layer.</p>
<p id="p0133" num="0133">The artificial neural network may as illustrated in <figref idref="f0007">FIG. 7</figref> specifically comprise an input layer 701 comprising a plurality of nodes receiving the input data values for the artificial neural network. Thus, the node values for nodes of the input layer may typically directly be the input data values to the artificial neural network and thus may not be calculated from other node values.</p>
<p id="p0134" num="0134">The artificial neural network may further comprise none, one, or more hidden layers 703, 705 or processing layers. For each of such layers, the node values are typically generated as a function of the node values of the nodes of the previous layer, and specifically a weighted combination and added bias followed by an activation function (such as a Sigmoid, ReLU, or Tanh function may be applied).<!-- EPO <DP n="24"> --></p>
<p id="p0135" num="0135">Specifically, as shown in <figref idref="f0006">FIG. 6</figref>, each node, which may also be referred to as a neuron, may receive input values (from nodes of a previous layer) and therefrom calculate a node value as a function of these values. Often, this includes first generating a value as a linear combination of the input values with each of these weighted by a weight: <maths id="math0026" num=""><math display="block"><mi>k</mi><mo>=</mo><mstyle displaystyle="true"><munder><mo>∑</mo><mi>n</mi></munder><msub><mi>w</mi><mi>n</mi></msub><msub><mi>x</mi><mi>n</mi></msub></mstyle></math><img id="ib0026" file="imgb0026.tif" wi="24" he="11" img-content="math" img-format="tif"/></maths> where w refers to weights, x refers to the nodes of the previous layer and n is an index referring to the different nodes of the previous layer.</p>
<p id="p0136" num="0136">An activation function may then be applied to the resulting combination. For example, the node value l may be determined as: <maths id="math0027" num=""><math display="block"><mi>l</mi><mo>=</mo><mi>f</mi><mfenced><mi>k</mi></mfenced></math><img id="ib0027" file="imgb0027.tif" wi="15" he="5" img-content="math" img-format="tif"/></maths> where the function may for example be a Rectified Linear Unit (as described in <nplcit id="ncit0003" npl-type="s"><text>Xavier Glorot, Antoine Bordes, Yoshua Bengio Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, PMLR 15:315-323, 2011</text></nplcit>) function: <maths id="math0028" num=""><math display="block"><mi>f</mi><mfenced><mi>k</mi></mfenced><mo>=</mo><mi mathvariant="italic">ReLU</mi><mfenced><mi>k</mi></mfenced><mo>=</mo><mi>max</mi><mfenced><mn>0</mn><mi>k</mi></mfenced></math><img id="ib0028" file="imgb0028.tif" wi="54" he="5" img-content="math" img-format="tif"/></maths></p>
<p id="p0137" num="0137">Other often used functions include a sigmoid function or a tanh function. In many embodiments, the node output or value may be calculated using a plurality of functions. For example, both a ReLU and Sigmoid function may be combined using an activation function such as: <maths id="math0029" num=""><math display="block"><mi>f</mi><mfenced><mi>k</mi></mfenced><mo>=</mo><mi mathvariant="italic">ReLU</mi><mfenced><mi>k</mi></mfenced><mo>+</mo><mi>σ</mi><mfenced><mi>k</mi></mfenced></math><img id="ib0029" file="imgb0029.tif" wi="44" he="5" img-content="math" img-format="tif"/></maths></p>
<p id="p0138" num="0138">Such operations may be performed by each node of the artificial neural network (except for typically the input nodes).</p>
<p id="p0139" num="0139">The artificial neural network may further comprise an output layer 707 which provides the output from the artificial neural network, i.e. the output data of the artificial neural network is the node values of the output layer. As for the hidden/ processing layers, the output node values are generated by a function of the node values of the previous layer. However, in contrast to the hidden/ processing layers where the node values are typically not accessible or used further, the node values of the output layer are accessible and provide the result of the operation of the artificial neural network. In many<!-- EPO <DP n="25"> --> embodiments, a Sigmoid activation may be used for real scale factors, for example, since the outputs of such an activation are constrained to be between 0 and 1 which is desired in many practical embodiments.</p>
<p id="p0140" num="0140">A number of different networks structures and toolboxes for artificial neural networks have been developed and in many embodiments the artificial neural network may be based on adapting and customizing such a network. An example of a network architecture that may be suitable for the applications mentioned above is WaveNet by van den Oord et al which is described in <nplcit id="ncit0004" npl-type="s"><text>Oord, Aaron van den, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. "Wavenet: A generative model for raw audio." arXiv preprint arxiv:1609.03499 (2016</text></nplcit>).</p>
<p id="p0141" num="0141">WaveNet is an architecture used for the synthesis of time domain signals using dilated causal convolution, and has been successfully applied to audio signals. For WaveNet the following activation function is commonly used: <maths id="math0030" num=""><math display="block"><mi>z</mi><mo>=</mo><mi>tanh</mi><mfenced separators=""><msub><mi>W</mi><mrow><mi>f</mi><mo>,</mo><mi>k</mi></mrow></msub><mo>∗</mo><mi>x</mi></mfenced><mo>⊙</mo><mi>σ</mi><mfenced separators=""><msub><mi>W</mi><mrow><mi>g</mi><mo>,</mo><mi>k</mi></mrow></msub><mo>∗</mo><mi>x</mi></mfenced><mo>,</mo></math><img id="ib0030" file="imgb0030.tif" wi="60" he="5" img-content="math" img-format="tif"/></maths> where * denotes a convolution operator, ⊙ denotes an element-wise multiplication operator, σ(·) is a sigmoid function, k is the layer index, f and g denote filter and gate, respectively, and W represents the weights of the learned artificial neural network. The filter product of the equation may typically provide a filtering effect with the gating product providing a weighting of the result which may in many cases effectively allow the contribution of the node to be reduced to substantially zero (i.e. it may allow or "cutoff" the node providing a contribution to other nodes thereby providing a "gate" function). In different circumstances, the gate function may result in the output of that node being negligible, whereas in other cases it would contribute substantially to the output. Such a function may substantially assist in allowing the artificial neural network to effectively learn and be trained.</p>
<p id="p0142" num="0142">An artificial neural network may in some cases further be arranged to include additional contributions that allow the artificial neural network to be dynamically adapted or customized for a specific desired property or characteristics of the generated output. For example, a set of values may be provided to adapt the artificial neural network. These values may be included by providing a contribution to some nodes of the artificial neural network. These nodes may be specifically input nodes but may typically be nodes of a hidden or processing layer. Such adaptation values may for example be weighted and added as a contribution to the weighted summation/ correlation value for a given node. For example, for WaveNet such adaptation values may be included in the activation function. For example, the output of the activation function may be given as: <maths id="math0031" num=""><math display="block"><mi>z</mi><mo>=</mo><mi>tanh</mi><mspace width="1ex"/><mfenced separators=""><msub><mi>W</mi><mrow><mi>f</mi><mo>,</mo><mi>k</mi></mrow></msub><mo>⋅</mo><mi>x</mi><mo>+</mo><msub><mi>V</mi><mrow><mi>f</mi><mo>,</mo><mi>k</mi></mrow></msub><mo>∗</mo><mi>y</mi></mfenced><mo>⊙</mo><mi>σ</mi><mfenced separators=""><msub><mi>W</mi><mrow><mi>g</mi><mo>,</mo><mi>k</mi></mrow></msub><mo>∗</mo><mi>x</mi><mo>+</mo><msub><mi>V</mi><mrow><mi>g</mi><mo>,</mo><mi>k</mi></mrow></msub><mo>∗</mo><mi>y</mi></mfenced></math><img id="ib0031" file="imgb0031.tif" wi="101" he="5" img-content="math" img-format="tif"/></maths> where <b>y</b> is a vector representing the adaptation values and V represents suitable weights for these values.<!-- EPO <DP n="26"> --></p>
<p id="p0143" num="0143">The above description relates to an artificial neural network approach that may be suitable for many embodiments and implementations. However, it will be appreciated that many other types and structures of artificial neural network may be used. Indeed, many different approaches to artificial neural networks have been, and are being, developed including artificial neural networks using complex structures and processes that differ from the ones described above. The approach is not limited to any specific artificial neural network approach and any suitable approach may be used without detracting from the invention.</p>
<p id="p0144" num="0144">In many embodiments, the artificial neural network may advantageously be a generative model such as a diffusion model artificial neural network. An example of such an approach is illustrated in <figref idref="f0008">FIG. 8</figref>.</p>
<p id="p0145" num="0145">In many embodiments, such as in the example of <figref idref="f0004">FIG. 4</figref>, the inputs to the trained artificial neural network are shown as the subband domain downmix signal. In practice, however, and depending on the available learning capacity of the neural network, the subbands of the downmix signal <i>X<sub>m</sub></i>(<i>k, m</i>) can be grouped into the parameter-estimation bands producing the downmix signal <i>X<sub>m</sub></i>(<i>b, m</i>)<i>.</i> In some embodiments, a subband-aggregation front-end can be also performed/learned by the artificial neural network.</p>
<p id="p0146" num="0146">In some embodiments, the artificial neural network may be implemented using an encoder-decoder architecture known as a U-net consisting of a downsampling branch and an upsampling branch and it may perform a multi-resolution feature extraction. Optionally, the model can preserve signal details at each level via skip connections. The final layer may feature a convolutional layer followed by an activation function such as a linear activation unit.</p>
<p id="p0147" num="0147">The shape of the output of the trained artificial neural network may typically be a <i>B</i> × <i>M</i> block to match the number of required estimation bands and timeslots by the parameter estimation block. It will be appreciated that the analysis filterbank and parameter estimation bands sub-block may also be implemented as either disparate neural networks or a single neural network that translates the time-domain multichannel input (e.g. 2-channels) into a set of parameter estimation bands (<i>B</i> bands).</p>
<p id="p0148" num="0148">In case the input to the model is the subband downmix signal <i>X<sub>m</sub></i>(<i>k, m</i>), the <i>K</i> × <i>M</i> complex-valued spectrogram may first be split into real and imaginary parts with the two resulting spectrogram channels then being stacked. The model can thus be implemented using 2D convolutional layers as shown in <figref idref="f0009">FIG. 9</figref>, where the downsampling is performed along both time and frequency dimensions, controlled by the stride parameters of the convolutional layers. A stride of two along a given dimension downsamples the data by two. The example of <figref idref="f0009">FIG. 9</figref> illustrates an embodiment of a neural network with 2D convolutional as well as a residual layers.</p>
<p id="p0149" num="0149">The batch normalization blocks are used to normalize the data after certain layers to improve training. Batch normalization is a method to normalize data between network layers to improve and speed up training with larger step sizes which could lead to instabilities without normalization. The<!-- EPO <DP n="27"> --> ReLU blocks are rectified linear unit activations. The upper row of layers performs downsampling (along time and frequency), while the lower set of layers performs the upsampling, restoring to the desired time-frequency resolution.</p>
<p id="p0150" num="0150">In the upsampling branch, upsampling can be performed using transpose convolutional layers (denoted as deconvolutional layers in 9). Details of the residual layer, which can learn composite functions and improve training due to their bypass functionality are shown in <figref idref="f0010">Fig. 10</figref>. Due to the input bypass, the 2D convolutional layers are required to preserve the shape of the input using appropriate padding and stride parameters. Furthermore, by stacking multiple residual layers, a functions of ranging complexity can be learned.</p>
<p id="p0151" num="0151">A high-level of the trained artificial neural networks block's inputs and outputs, assuming hQMF stereo inputs is shown in <figref idref="f0011">Fig. 11</figref> along with their dimensions. In some embodiments, the trained artificial neural network output may be of shape 2 × <i>B</i> × <i>M</i> and represent a complex window that includes phase information. <figref idref="f0011">FIG. 11</figref> illustrates an example of high level processing of hQMF 2D signals where a mono downmix's real and imaginary parts are stacked, producing two channels. The artificial neural network output consists of a single channel set of B windows each of length M timeslots.</p>
<p id="p0152" num="0152">Finally, each row of the resulting NN output may be processed by a Softmax layer such that the resulting window within each estimation band sums to 1, <maths id="math0032" num=""><math display="block"><msub><mi>W</mi><mi mathvariant="italic">NN</mi></msub><mfenced><mi>b</mi><mi>m</mi></mfenced><mo>=</mo><mfrac><msup><mi>e</mi><mrow><mi>W</mi><mfenced><mi>b</mi><mi>m</mi></mfenced></mrow></msup><mstyle displaystyle="true"><msub><mo>∑</mo><mi>m</mi></msub><msup><mi>e</mi><mrow><mi>W</mi><mfenced><mi>b</mi><mi>m</mi></mfenced></mrow></msup></mstyle></mfrac></math><img id="ib0032" file="imgb0032.tif" wi="45" he="12" img-content="math" img-format="tif"/></maths></p>
<p id="p0153" num="0153">In some embodiments, the output of the final artificial neural network block layer may already produce a positive output, e.g., when implemented using a rectified linear unit (ReLU), for example. In this case, the output can be normalized by the sum of the outputs of each row (along the time slots). In other embodiments a Sigmoid activation may be applied to each value of <i>W<sub>NN</sub>.</i></p>
<p id="p0154" num="0154">Artificial neural networks are adapted to specific purposes by a training process which is used to adapt/ tune/ modify the weights and other parameters (e.g. bias) of the artificial neural network. It will be appreciated that many different training processes and algorithms are known for training artificial neural networks. A training setup that may be used for the described artificial neural networks is illustrated in <figref idref="f0012">FIG. 12</figref>.</p>
<p id="p0155" num="0155">Typically, training is based on large training sets where a large number of examples of input data are provided to the network. Further, the output of the artificial neural network is typically (directly or indirectly) compared to an expected or ideal result. A cost function may be generated to reflect the desired outcome of the training process. In a typical scenario known as supervised learning, the cost (or loss) function often represents the distance between the prediction and the ground truth for a particular input data. Based on the cost function, the weights may be changed and by reiterating the<!-- EPO <DP n="28"> --> process for the modified weights, the artificial neural network may be adapted towards a state for which the cost function is minimized.</p>
<p id="p0156" num="0156">Different approaches may be used to train the artificial neural networks of the audio encoder apparatus and audio decoder apparatus, and in particular an overall training may seek to result in the output of the audio apparatuses being a multichannel audio signal that most closely corresponds to the original multichannel audio signal. Thus, the trained artificial neural networks may be trained to provide scale factors that most effectively result in accurate reconstruction of the multichannel audio signal. In such a case, a cost function based on the difference between a generated multichannel audio signal and an original training multichannel audio signal may be used. In other cases, more specific training of an artificial neural network may be used.</p>
<p id="p0157" num="0157">In some embodiments, where the audio encoder apparatus and the audio decoder apparatus use identical networks, the two networks may use the exact same weights on thus the updating may be constrained to be identical for the two networks.</p>
<p id="p0158" num="0158">In more detail, during a training step an artificial neural network may have two different flows of information from input to output (forward pass) and from output to input (backward pass). In the forward pass, the data is processed by the artificial neural network as described above while in the backward pass the weights are updated to minimize the cost function. Typically, such a backward propagation follows the gradient direction of the cost function landscape. In other words, by comparing the predicted output with the ground truth for a batch of data input, one can estimate the direction in which the cost function is minimized and propagate backward, by updating the weights accordingly. Other approaches known for training artificial neural networks include for example Levenberg-Marquardt algorithm, the conjugate gradient method, and the Newton method etc.</p>
<p id="p0159" num="0159">In the present case, training may specifically include a training set comprising a potentially large number of multichannel audio signals or corresponding downmix audio signals. The training sets may include audio signals representing a number of different audio sources including e.g. recording of videos, movies, telecommunications, etc. In some embodiments, the training data may even include non-audio data such as a training being performed in combination with training data from other sources, such as text data etc.</p>
<p id="p0160" num="0160">In some embodiments, training data may be multichannel audio signals in time segments corresponding to the processing time intervals of the trained artificial neural network being trained, e.g. the number of samples in a training multichannel audio signal may correspond to a number of samples corresponding to the input nodes of the artificial neural network(s) being trained. Each training example may thus correspond to one operation of the artificial neural network(s) being trained. Usually, however, a batch of training samples is considered for each step to speed up and smoothen the training process. Furthermore, many upgrades to gradient descent are possible also to speed up convergence or avoid local minima in the cost function landscape.<!-- EPO <DP n="29"> --></p>
<p id="p0161" num="0161">For each training multichannel audio signal, a training processor may perform a downmix operation to generate a downmix audio signal and corresponding upmix parametric data. Thus, the encoding process that is applied to the multichannel audio signal during normal operation may also be applied to the training multichannel audio signal thereby generating a downmix and the upmix parametric data.</p>
<p id="p0162" num="0162">Based on the cost value, the training processor may adapt the weights of the artificial neural networks. For example, a back-propagation approach may be used. In particular, the training processor may adjust the weights of one or all of the artificial neural networks based on the cost value. For example, given the derivative (representing the slope) of the weights with respect to the cost function the weights values are modified to go in the direction of the slope. For a simple/minima account one can refer to the training of the perceptron (single neuron) in case of backward pass of a single data input.</p>
<p id="p0163" num="0163">The process may be iterated until the artificial neural networks are considered to be trained. For example, training may be performed for a predetermined number of iterations. As another example, training may be continued until the weights change is less than a predetermined amount. Also very common, a validation stop is implemented where the network is tested again a validation metric and stopped when reaching the expected outcome.</p>
<p id="p0164" num="0164">Specifically, for the trained artificial neural network implementing a diffusion model, the training process includes randomly selecting timestamps or corresponding noise scaling values for each of the training examples in a given batch. For each training example in a batch, an instance of a standard Gaussian noise signal (can also correspond to some other standard distribution) is generated with the same dimension as the training example. Then after scaling each of the noise signals with their corresponding noise scaling values, these are added to the examples in the training batch. The model then processes this batch of noisy data with additional inputs of the noise scaling factors and features to estimate the original Gaussian noise signals. Therefore the loss between the estimated and real noise signals are back propagated through the network to optimize their weights using an optimizer such as the Adam optimizer.</p>
<p id="p0165" num="0165">In many embodiments, the loss or cost function may be perceptually weighted loss/cost function. A perceptual relevancy metric may be used to weigh each of the samples during training so that the artificial neural network learns meaningful weighting values to help improve the e.g. the binaural cue preservation properties of the system.</p>
<p id="p0166" num="0166">Starting with the relevancy metric, in some embodiments, the relevancy metric may be used to weigh frames or time slots containing transients that have been panned to one stereo channel or the other or differently than the background or diffuse component of the stereo signal. Particularly when the background signal is diffuse (e.g., applause), foreground components are more easily perceived when they display different binaural properties from the background (this is related to the theory of binaural masking level differences (BMLD), where the masking threshold for a maskee drops if its binaural properties are different from its masker), i.e. such as panned foreground clapping.<!-- EPO <DP n="30"> --></p>
<p id="p0167" num="0167">For a metric that can emphasize relevant transient signal components, let <i>P<sub>L</sub></i>(<i>b, m</i>) and <i>P<sub>R</sub></i>(<i>b, m</i>) denote the (smoothed) left and right energy levels for time-frequency tile in timeslot m and band b.</p>
<p id="p0168" num="0168">Next, a normalized IID is computed from the left and right channels per (<i>b, m</i>) tile, <maths id="math0033" num=""><math display="block"><mi mathvariant="italic">IID</mi><mfenced><mi>b</mi><mi>m</mi></mfenced><mo>=</mo><mfrac><mrow><msub><mi>P</mi><mi>L</mi></msub><mfenced><mi>b</mi><mi>m</mi></mfenced><mo>−</mo><msub><mi>P</mi><mi>R</mi></msub><mfenced><mi>b</mi><mi>m</mi></mfenced></mrow><mrow><msub><mi>P</mi><mi>L</mi></msub><mfenced><mi>b</mi><mi>m</mi></mfenced><mo>+</mo><msub><mi>P</mi><mi>R</mi></msub><mfenced><mi>b</mi><mi>m</mi></mfenced></mrow></mfrac></math><img id="ib0033" file="imgb0033.tif" wi="59" he="11" img-content="math" img-format="tif"/></maths> that lies between -1 and 1.</p>
<p id="p0169" num="0169">For noise-like background signals, the IID values will be distributed randomly over frequency within a slot. Therefore, next, a summary IID is calculated by integrating (and weighting) IID values over frequency, leading to a profile of which slots could contain signals of relevance, <maths id="math0034" num=""><math display="block"><mover accent="true"><mi mathvariant="italic">IID</mi><mo>˜</mo></mover><mfenced><mi>m</mi></mfenced><mo>=</mo><mstyle displaystyle="true"><munder><mo>∑</mo><mi>b</mi></munder><msub><mi>w</mi><mi>b</mi></msub><mi mathvariant="italic">IID</mi><mfenced><mi>b</mi><mi>m</mi></mfenced></mstyle></math><img id="ib0034" file="imgb0034.tif" wi="49" he="11" img-content="math" img-format="tif"/></maths> where the weights can be set based on the properties of human hearing and/or a particular class of signal (transient), and where ∑<i><sub>b</sub> w<sub>b</sub></i> = 1.</p>
<p id="p0170" num="0170">Finally, a relevancy metric for a given frame (set of contiguous time slots), c corresponds to the maximum value of the absolute summary IID, <maths id="math0035" num=""><math display="block"><mi>c</mi><mo>=</mo><munder><mi>max</mi><mi>m</mi></munder><mover accent="true"><mi mathvariant="italic">IID</mi><mo>˜</mo></mover><mfenced><mi>m</mi></mfenced></math><img id="ib0035" file="imgb0035.tif" wi="29" he="7" img-content="math" img-format="tif"/></maths> and lies between 0 and 1. The above relevancy metric can be computed for each frame within a given training batch to weigh its contribution to the average loss.</p>
<p id="p0171" num="0171">An example of computing the summary IID and corresponding relevancy metric is shown in <figref idref="f0013">FIG. 13</figref>.</p>
<p id="p0172" num="0172">In some embodiments, one or more of the trained artificial neural network(s) may be trained using a perceptually weighted loss function. The loss function may be perceptually weighted by applying a higher weight to time frequency intervals representing a perceptually significant signal component, such as a transient. In some embodiments, the perceptual weighted loss function may be generated by applying a higher weighting to frames/segments comprising a perceptually significant signal component/event, such as a transient or loud sound, relative to frames/segments that do not include perceptually significant signal component/events.</p>
<p id="p0173" num="0173">Thus, for perceptually relevant components of the signal, the weighting of the error for the corresponding samples/time-slots/ time-frequency tiles and/or frames/segments may be increased<!-- EPO <DP n="31"> --> when determining the loss function. This may focus the learning on addressing the perceptually significant situations and may reduce the risk that the neural network training will focus more on learning the more common components for correct resynthesis (such as the background sound) potentially leading to subpar performance for the perceptually relevant samples in the training data.</p>
<p id="p0174" num="0174">A loss function for a batch of I frames each with relevancy <i>c<sub>i</sub></i> may be given by: <maths id="math0036" num=""><math display="block"><mi>L</mi><mo>=</mo><mstyle displaystyle="true"><munder><mo>∑</mo><mi>I</mi></munder><msub><mi>c</mi><mi>i</mi></msub><msub><mi>L</mi><mi>i</mi></msub></mstyle></math><img id="ib0036" file="imgb0036.tif" wi="22" he="12" img-content="math" img-format="tif"/></maths> where <img id="ib0037" file="imgb0037.tif" wi="5" he="5" img-content="character" img-format="tif" inline="yes"/> is the loss attributed to the individual frames in the training set. The weight <i>c<sub>i</sub></i> may be adapted to provide desired perceptual weighting.</p>
<p id="p0175" num="0175">In other embodiments, instead of taking the maximum over all frames to compute the relevancy metric, the summary IID may simply be used to scale the contribution of each time slot within a frame when computing the average error for that frame first before averaging over the batch of frames.</p>
<p id="p0176" num="0176">In further embodiments, a training multichannel/stereo dataset can be composed of separate transient and non-transient components that are mixed according to various mixing ratios. The relevancy metric can then be computed by first scaling both transient and non-transient components such that the non-transient components have equal power in both channels before combining with the transient signal. The resulting summary IID and relevancy metric may then better follow the BMLD model.</p>
<p id="p0177" num="0177">The loss function can consist of a number of terms, including the mean square error (MSE) between the stereo input to the PS encoder and the estimated output stereo signal at the decoder, where <maths id="math0037" num=""><math display="block"><msub><mi>L</mi><mn>1</mn></msub><mfenced separators=""><mi>S</mi><mo>,</mo><mi>S</mi><mo>′</mo></mfenced><mo>=</mo><mfrac><mn>1</mn><mi>M</mi></mfrac><mstyle displaystyle="true"><munder><mo>∑</mo><mrow><mi>m</mi><mo>,</mo><mi>k</mi></mrow></munder><msup><mfenced open="[" close="]" separators=""><mi>S</mi><mfenced><mi>m</mi><mi>k</mi></mfenced><mo>−</mo><mi>S</mi><mo>′</mo><mfenced><mi>m</mi><mi>k</mi></mfenced></mfenced><mn>2</mn></msup></mstyle></math><img id="ib0038" file="imgb0038.tif" wi="69" he="13" img-content="math" img-format="tif"/></maths></p>
<p id="p0178" num="0178">Another loss term ignores the phase between the complex subband-domain input and predicted stereo output signals. Such a term is more correlated to the perceptual differences between two audio signals since phase does not have such a large impact for most types of audio signals, <maths id="math0038" num=""><math display="block"><msub><mi>L</mi><mn>2</mn></msub><mfenced separators=""><mi>S</mi><mo>,</mo><mi>S</mi><mo>′</mo></mfenced><mo>=</mo><mfrac><mn>1</mn><mi>M</mi></mfrac><mstyle displaystyle="true"><munder><mo>∑</mo><mrow><mi>m</mi><mo>,</mo><mi>k</mi></mrow></munder><msup><mfenced open="[" close="]" separators=""><msup><mfenced open="|" close="|" separators=""><mi>S</mi><mfenced><mi>m</mi><mi>k</mi></mfenced></mfenced><mi>α</mi></msup><mo>−</mo><msup><mfenced open="|" close="|" separators=""><mi>S</mi><mo>′</mo><mfenced><mi>m</mi><mi>k</mi></mfenced></mfenced><mi>α</mi></msup></mfenced><mn>2</mn></msup></mstyle></math><img id="ib0039" file="imgb0039.tif" wi="79" he="13" img-content="math" img-format="tif"/></maths> where <i>α</i> is a compression exponent term (<i>α &gt;</i> 0).</p>
<p id="p0179" num="0179">Finally, the loss function for sample i can be calculated as a weighted sum of the terms <i>L</i><sub>1</sub> and <i>L</i><sub>2</sub>,<!-- EPO <DP n="32"> --> <maths id="math0039" num=""><math display="block"><msub><mi>L</mi><mi>i</mi></msub><mfenced><msub><mi>S</mi><mi>i</mi></msub><msubsup><mi>S</mi><mi>i</mi><mo>′</mo></msubsup></mfenced><mo>=</mo><msub><mi mathvariant="italic">βL</mi><mn>1</mn></msub><mfenced><msub><mi>S</mi><mi>i</mi></msub><msubsup><mi>S</mi><mi>i</mi><mo>′</mo></msubsup></mfenced><mo>+</mo><mfenced separators=""><mn>1</mn><mo>−</mo><mi>β</mi></mfenced><msub><mi>L</mi><mn>2</mn></msub><mfenced><msub><mi>S</mi><mi>i</mi></msub><msubsup><mi>S</mi><mi>i</mi><mo>′</mo></msubsup></mfenced></math><img id="ib0040" file="imgb0040.tif" wi="76" he="5" img-content="math" img-format="tif"/></maths> with 0 ≤ <i>β</i> ≤ 1. In other embodiments, the weights do not have to sum to 1.0, particularly to more effectively control the relative contributions of loss terms to the total loss.</p>
<p id="p0180" num="0180">Other loss components can be included depending on the specific neural network configurations at encoder and decoder as described in the next section.</p>
<p id="p0181" num="0181">The audio apparatus(s) may specifically be implemented in one or more suitably programmed processors. In particular, the artificial neural networks may be implemented in one more such suitably programmed processors. The different functional blocks, and in particular the artificial neural network, may be implemented in separate processors and/or may e.g. be implemented in the same processor. An example of a suitable processor is provided in the following.</p>
<p id="p0182" num="0182"><figref idref="f0014">FIG. 14</figref> is a block diagram illustrating an example processor 1400 according to embodiments of the disclosure. Processor 1400 may be used to implement one or more processors implementing an apparatus as previously described or elements thereof (including in particular one more artificial neural network). Processor 1400 may be any suitable processor type including, but not limited to, a microprocessor, a microcontroller, a Digital Signal Processor (DSP), a Field ProGrammable Array (FPGA) where the FPGA has been programmed to form a processor, a Graphical Processing Unit (GPU), an Application Specific Integrated Circuit (ASIC) where the ASIC has been designed to form a processor, or a combination thereof.</p>
<p id="p0183" num="0183">The processor 1400 may include one or more cores 1402. The core 1402 may include one or more Arithmetic Logic Units (ALU) 1404. In some embodiments, the core 1402 may include a Floating Point Logic Unit (FPLU) 1406 and/or a Digital Signal Processing Unit (DSPU) 1408 in addition to or instead of the ALU 1404.</p>
<p id="p0184" num="0184">The processor 1400 may include one or more registers 1412 communicatively coupled to the core 1402. The registers 1412 may be implemented using dedicated logic gate circuits (e.g., flip-flops) and/or any memory technology. In some embodiments the registers 1412 may be implemented using static memory. The register may provide data, instructions and addresses to the core 1402.</p>
<p id="p0185" num="0185">In some embodiments, processor 1400 may include one or more levels of cache memory 1410 communicatively coupled to the core 1402. The cache memory 1410 may provide computer-readable instructions to the core 1402 for execution. The cache memory 1410 may provide data for processing by the core 1402. In some embodiments, the computer-readable instructions may have been provided to the cache memory 1410 by a local memory, for example, local memory attached to the external bus 1416. The cache memory 1410 may be implemented with any suitable cache memory type, for example, Metal-Oxide Semiconductor (MOS) memory such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), and/or any other suitable memory technology.<!-- EPO <DP n="33"> --></p>
<p id="p0186" num="0186">The processor 1400 may include a controller 1414, which may control input to the processor 1400 from other processors and/or components included in a system and/or outputs from the processor 1400 to other processors and/or components included in the system. Controller 1414 may control the data paths in the ALU 1404, FPLU 1406 and/or DSPU 1408. Controller 1414 may be implemented as one or more state machines, data paths and/or dedicated control logic. The gates of controller 1414 may be implemented as standalone gates, FPGA, ASIC or any other suitable technology.</p>
<p id="p0187" num="0187">The registers 1412 and the cache 1410 may communicate with controller 1414 and core 1402 via internal connections 1420A, 1420B, 1420C and 1420D. Internal connections may be implemented as a bus, multiplexer, crossbar switch, and/or any other suitable connection technology.</p>
<p id="p0188" num="0188">Inputs and outputs for the processor 1400 may be provided via a bus 1416, which may include one or more conductive lines. The bus 1416 may be communicatively coupled to one or more components of processor 1400, for example the controller 1414, cache 1410, and/or register 1412. The bus 1416 may be coupled to one or more components of the system.</p>
<p id="p0189" num="0189">The bus 1416 may be coupled to one or more external memories. The external memories may include Read Only Memory (ROM) 1432. ROM 1432 may be a masked ROM, Electronically Programmable Read Only Memory (EPROM) or any other suitable technology. The external memory may include Random Access Memory (RAM) 1433. RAM 1433 may be a static RAM, battery backed up static RAM, Dynamic RAM (DRAM) or any other suitable technology. The external memory may include Electrically Erasable Programmable Read Only Memory (EEPROM) 1435. The external memory may include Flash memory 1434. The External memory may include a magnetic storage device such as disc 1436. In some embodiments, the external memories may be included in a system.</p>
<p id="p0190" num="0190">The invention can be implemented in any suitable form including hardware, software, firmware or any combination of these. The invention may optionally be implemented at least partly as computer software running on one or more data processors and/or digital signal processors. The elements and components of an embodiment of the invention may be physically, functionally and logically implemented in any suitable way. Indeed the functionality may be implemented in a single unit, in a plurality of units or as part of other functional units. As such, the invention may be implemented in a single unit or may be physically and functionally distributed between different units, circuits and processors.</p>
<p id="p0191" num="0191">Although the present invention has been described in connection with some embodiments, it is not intended to be limited to the specific form set forth herein. Rather, the scope of the present invention is limited only by the accompanying claims. Additionally, although a feature may appear to be described in connection with particular embodiments, one skilled in the art would recognize that various features of the described embodiments may be combined in accordance with the invention. In the claims, the term comprising does not exclude the presence of other elements or steps.</p>
<p id="p0192" num="0192">Furthermore, although individually listed, a plurality of means, elements, circuits or method steps may be implemented by e.g. a single circuit, unit or processor. Additionally, although<!-- EPO <DP n="34"> --> individual features may be included in different claims, these may possibly be advantageously combined, and the inclusion in different claims does not imply that a combination of features is not feasible and/or advantageous. Also the inclusion of a feature in one category of claims does not imply a limitation to this category but rather indicates that the feature is equally applicable to other claim categories as appropriate. Furthermore, the order of features in the claims do not imply any specific order in which the features must be worked and in particular the order of individual steps in a method claim does not imply that the steps must be performed in this order. Rather, the steps may be performed in any suitable order. In addition, singular references do not exclude a plurality. Thus references to "a", "an", "first", "second" etc do not preclude a plurality. Reference signs in the claims are provided merely as a clarifying example shall not be construed as limiting the scope of the claims in any way.</p>
</description>
<claims id="claims01" lang="en"><!-- EPO <DP n="35"> -->
<claim id="c-en-0001" num="0001">
<claim-text>An audio apparatus for generating an encoded audio signal for a multichannel audio signal, the apparatus comprising:
<claim-text>a receiver (101) arranged to receive the multichannel audio signal;</claim-text>
<claim-text>a downmixer (105) arranged to generate a downmix audio signal by downmixing the multichannel audio signal;</claim-text>
<claim-text>an encoder (107) arranged to encode the downmix audio signal to generate encoded downmix audio data;</claim-text>
<claim-text>a trained artificial neural network (111) arranged to determine first scale factors for time frequency intervals of the multichannel audio signal, the trained artificial neural network having input nodes for receiving samples of at least one of the multichannel audio signal and the downmix audio signal, and output nodes providing the first scale factors;</claim-text>
<claim-text>a signal generator (113) arranged to generate a first modified multichannel audio signal by applying the first scale factors to samples of the time frequency intervals;</claim-text>
<claim-text>a parameter circuit (115) arranged to generate a first set of spatial parameters for the downmix audio signal, the first set of spatial parameters representing relative properties for channels of the first modified multi-channel audio signal; and</claim-text>
<claim-text>an output circuit arranged (109) to generate the encoded audio signal to comprise the encoded downmix audio data and first set of spatial parameters.</claim-text></claim-text></claim>
<claim id="c-en-0002" num="0002">
<claim-text>The audio apparatus of claim 1 further comprising a scale factor circuit (117) arranged to determine second scale factors for the time frequency intervals of the multichannel audio signals; and wherein the signal generator (113) is arranged to generate a second modified multichannel audio signal by applying the second scale factors to samples of the time frequency intervals; the parameter circuit (115) is arranged to generate a second set of spatial parameters for the downmix audio signal from the second modified multichannel audio signal, the second set of spatial parameters representing relative properties for channels of the second modified multi-channel audio signal; and<br/>
the output circuit (109) is arranged to generate the encoded audio signal to comprise the second set of spatial parameters.</claim-text></claim>
<claim id="c-en-0003" num="0003">
<claim-text>The audio apparatus of claim 2 wherein the scale factor circuit (117) is arranged to determine the second scale factors from the first scale factors.<!-- EPO <DP n="36"> --></claim-text></claim>
<claim id="c-en-0004" num="0004">
<claim-text>The audio apparatus of claim 2 wherein the scale factor circuit (117) comprises a further trained artificial neural network arranged to determine the second scale factors, the further artificial trained artificial neural network having input nodes for receiving samples of at least one of the multichannel audio signal and the downmix audio signal, and output nodes providing the second scale factors.</claim-text></claim>
<claim id="c-en-0005" num="0005">
<claim-text>The apparatus of any of claims 2 to 4 wherein a data rate for the first set of spatial parameters is different than a data rate for the second set of spatial parameters.</claim-text></claim>
<claim id="c-en-0006" num="0006">
<claim-text>The audio apparatus of any previous claim wherein the first scale factors are complex values.</claim-text></claim>
<claim id="c-en-0007" num="0007">
<claim-text>The audio apparatus of any previous claim further comprising an additional trained artificial neural network arranged to generate feature set values from the multichannel audio signal, the additional trained artificial neural network having input nodes for receiving samples of the multichannel audio signal, and the output circuit (109) is arranged to include the feature set values in the encoded audio signal.</claim-text></claim>
<claim id="c-en-0008" num="0008">
<claim-text>An audio apparatus for generating a multichannel audio signal from an encoded audio signal, the apparatus comprising:
<claim-text>a receiver (201) arranged to receive the encoded audio signal comprising encoded downmix audio data for a downmix of a multi-channel audio signal and a first set of spatial parameters, the first set of spatial parameters representing relative properties for channels of a first modified multi-channel audio signal resulting from applying first scale factors to samples of time frequency intervals of the multichannel audio signal;</claim-text>
<claim-text>a decoder (203) arranged to decode the encoded audio data to generate the downmix audio signal;</claim-text>
<claim-text>a decorrelator (205) arranged to decorrelate the downmix audio signal to generate a decorrelated signal;</claim-text>
<claim-text>a trained artificial neural network (209) arranged to determine the first scale factors for time frequency intervals of the multichannel audio signal, the trained artificial neural network having input nodes for receiving samples of the downmix audio signal and output nodes providing the first scale factors;</claim-text>
<claim-text>an upmixer (207) arranged to generate the multichannel audio signal by upmixing the downmix audio signal and the decorrelated signal in dependence on the first set of spatial parameters and the first scale factors.</claim-text><!-- EPO <DP n="37"> --></claim-text></claim>
<claim id="c-en-0009" num="0009">
<claim-text>The audio apparatus of claim 8 wherein the upmixer (207) is arranged to generate each channel of the multichannel audio signal as a linear combination of the downmix audio signal and the decorrelated signals where weights of the linear combination are dependent on the first scale factors and the first set of spatial parameters.</claim-text></claim>
<claim id="c-en-0010" num="0010">
<claim-text>The audio apparatus of claim 8 wherein the encoded audio signal comprises a second set of spatial parameters, the second set of spatial parameters representing relative properties for channels of a second modified multi-channel audio signal resulting from applying second scale factors to samples of time frequency intervals of the multichannel audio signal; the apparatus comprises a circuit arranged to generate the second scale factors; and the upmixer (207) is arranged to determine a first upmixed multichannel audio signal from the downmix audio signal and the decorrelated signal using the first set of spatial parameters and the first scale factors, to determine a second upmixed multichannel audio signal from the downmix audio signal and the decorrelated signal using the second set of spatial parameters and the second scale factors, and to include a combination of the first upmixed multichannel audio signal and the second upmixed multichannel audio signal in the multichannel audio signal.</claim-text></claim>
<claim id="c-en-0011" num="0011">
<claim-text>The audio apparatus of any of the previous claims 8 to 10 wherein the encoded audio signal comprises feature set values for a trained network and the trained artificial neural network (209) has input nodes for receiving the feature set values.</claim-text></claim>
<claim id="c-en-0012" num="0012">
<claim-text>An audio system comprising an encoder apparatus in accordance with any of claims 1-7 and a decoder apparatus in accordance with any of claims 8-11.</claim-text></claim>
<claim id="c-en-0013" num="0013">
<claim-text>The audio system of claim 12 wherein the trained artificial neural network of the encoder apparatus is different from the trained artificial neural network of the decoder apparatus.</claim-text></claim>
<claim id="c-en-0014" num="0014">
<claim-text>The audio system of any previous claim 12 or 13 wherein the trained artificial neural network of at least one of the encoder apparatus and the decoder apparatus is trained using training data comprising a number of training multichannel audio signals and a cost function dependent on a perceptual difference measure for the multichannel audio signal received by the encoder apparatus and the multichannel audio signal generated by the decoder apparatus.</claim-text></claim>
<claim id="c-en-0015" num="0015">
<claim-text>A method of operation for an audio apparatus for generating an encoded audio signal for a multichannel audio signal, the method comprising:
<claim-text>receiving the multichannel audio signal;</claim-text>
<claim-text>generating a downmix audio signal by downmixing the multichannel audio signal;</claim-text>
<claim-text>encoding the downmix audio signal to generate encoded downmix audio data;<!-- EPO <DP n="38"> --></claim-text>
<claim-text>a trained artificial neural network (111) determining first scale factors for time frequency intervals of the multichannel audio signal, the trained artificial neural network having input nodes for receiving samples of at least one of the multichannel audio signal and the downmix audio signal, and output nodes providing the first scale factors;</claim-text>
<claim-text>generating a first modified multichannel audio signal by applying the first scale factors to samples of the time frequency intervals;</claim-text>
<claim-text>generating a first set of spatial parameters for the downmix audio signal, the first set of spatial parameters representing relative properties for channels of the first modified multi-channel audio signal; and</claim-text>
<claim-text>generating an encoded audio signal to comprise the encoded downmix audio data and first set of spatial parameters.</claim-text></claim-text></claim>
<claim id="c-en-0016" num="0016">
<claim-text>A method of operation for an audio apparatus, the method comprising:
<claim-text>receiving the encoded audio signal comprising encoded downmix audio data for a downmix of a multi-channel audio signal and a first set of spatial parameters, the first set of spatial parameters representing relative properties for channels of a first modified multi-channel audio signal resulting from applying first scale factors to samples of time frequency intervals of the multichannel audio signal;</claim-text>
<claim-text>decoding the encoded audio data to generate the downmix audio signal;</claim-text>
<claim-text>decorrelating the downmix audio signal to generate a decorrelated signal;</claim-text>
<claim-text>a trained artificial neural network (209) determining the first scale factors for time frequency intervals of the multichannel audio signal, the trained artificial neural network having input nodes for receiving samples of the downmix audio signal and output nodes providing the first scale factors; and</claim-text>
<claim-text>generating the multichannel audio signal by upmixing the downmix audio signal and the decorrelated signal in dependence on the first set of spatial parameters and the first scale factors.</claim-text></claim-text></claim>
</claims>
<drawings id="draw" lang="en"><!-- EPO <DP n="39"> -->
<figure id="f0001" num="1"><img id="if0001" file="imgf0001.tif" wi="165" he="186" img-content="drawing" img-format="tif"/></figure><!-- EPO <DP n="40"> -->
<figure id="f0002" num="2"><img id="if0002" file="imgf0002.tif" wi="165" he="137" img-content="drawing" img-format="tif"/></figure><!-- EPO <DP n="41"> -->
<figure id="f0003" num="3"><img id="if0003" file="imgf0003.tif" wi="141" he="179" img-content="drawing" img-format="tif"/></figure><!-- EPO <DP n="42"> -->
<figure id="f0004" num="4"><img id="if0004" file="imgf0004.tif" wi="98" he="234" img-content="drawing" img-format="tif"/></figure><!-- EPO <DP n="43"> -->
<figure id="f0005" num="5"><img id="if0005" file="imgf0005.tif" wi="165" he="109" img-content="drawing" img-format="tif"/></figure><!-- EPO <DP n="44"> -->
<figure id="f0006" num="6"><img id="if0006" file="imgf0006.tif" wi="109" he="113" img-content="drawing" img-format="tif"/></figure><!-- EPO <DP n="45"> -->
<figure id="f0007" num="7"><img id="if0007" file="imgf0007.tif" wi="132" he="150" img-content="drawing" img-format="tif"/></figure><!-- EPO <DP n="46"> -->
<figure id="f0008" num="8"><img id="if0008" file="imgf0008.tif" wi="165" he="97" img-content="drawing" img-format="tif"/></figure><!-- EPO <DP n="47"> -->
<figure id="f0009" num="9"><img id="if0009" file="imgf0009.tif" wi="125" he="231" img-content="drawing" img-format="tif"/></figure><!-- EPO <DP n="48"> -->
<figure id="f0010" num="10"><img id="if0010" file="imgf0010.tif" wi="110" he="214" img-content="drawing" img-format="tif"/></figure><!-- EPO <DP n="49"> -->
<figure id="f0011" num="11"><img id="if0011" file="imgf0011.tif" wi="165" he="104" img-content="drawing" img-format="tif"/></figure><!-- EPO <DP n="50"> -->
<figure id="f0012" num="12"><img id="if0012" file="imgf0012.tif" wi="165" he="123" img-content="drawing" img-format="tif"/></figure><!-- EPO <DP n="51"> -->
<figure id="f0013" num="13"><img id="if0013" file="imgf0013.tif" wi="165" he="161" img-content="drawing" img-format="tif"/></figure><!-- EPO <DP n="52"> -->
<figure id="f0014" num="14"><img id="if0014" file="imgf0014.tif" wi="113" he="237" img-content="drawing" img-format="tif"/></figure>
</drawings>
<search-report-data id="srep" lang="en" srep-office="EP" date-produced=""><doc-page id="srep0001" file="srep0001.tif" wi="160" he="240" type="tif"/><doc-page id="srep0002" file="srep0002.tif" wi="158" he="240" type="tif"/></search-report-data><search-report-data date-produced="20250520" id="srepxml" lang="en" srep-office="EP" srep-type="ep-sr" status="n"><!--
 The search report data in XML is provided for the users' convenience only. It might differ from the search report of the PDF document, which contains the officially published data. The EPO disclaims any liability for incorrect or incomplete data in the XML for search reports.
 -->

<srep-info><file-reference-id>2024P00621EP</file-reference-id><application-reference><document-id><country>EP</country><doc-number>25160554.9</doc-number></document-id></application-reference><applicant-name><name>Koninklijke Philips N.V.</name></applicant-name><srep-established srep-established="yes"/><srep-invention-title title-approval="yes"/><srep-abstract abs-approval="yes"/><srep-figure-to-publish figinfo="by-applicant"><figure-to-publish><fig-number>1</fig-number></figure-to-publish></srep-figure-to-publish><srep-info-admin><srep-office><addressbook><text>MN</text></addressbook></srep-office><date-search-report-mailed><date>20250530</date></date-search-report-mailed></srep-info-admin></srep-info><srep-for-pub><srep-fields-searched><minimum-documentation><classifications-ipcr><classification-ipcr><text>G10L</text></classification-ipcr></classifications-ipcr></minimum-documentation></srep-fields-searched><srep-citations><citation id="sr-cit0001"><nplcit id="sr-ncit0001" npl-type="s"><article><author><name>WU YULIN ET AL</name></author><atl>Adaptive subband partition encoding scheme for multiple audio objects using CNN and residual dense blocks mixture network</atl><serial><sertitle>EXPERT SYSTEMS WITH APPLICATIONS, ELSEVIER, AMSTERDAM, NL</sertitle><pubdate>20240124</pubdate><vid>247</vid><doi>10.1016/J.ESWA.2024.123323</doi><issn>0957-4174</issn></serial><refno>XP087496972</refno></article></nplcit><category>A</category><rel-claims>1-16</rel-claims><rel-passage><passage>* the whole document *</passage></rel-passage></citation><citation id="sr-cit0002"><patcit dnum="EP4339943A1" id="sr-pcit0001" url="http://v3.espacenet.com/textdoc?DB=EPODOC&amp;IDX=EP4339943&amp;CY=ep"><document-id><country>EP</country><doc-number>4339943</doc-number><kind>A1</kind><name>KONINKLIJKE PHILIPS NV [NL]</name><date>20240320</date></document-id></patcit><category>A</category><rel-claims>1-16</rel-claims><rel-passage><passage>* paragraph [0001] - paragraph [0049] *</passage><passage>* figure 1 *</passage></rel-passage></citation></srep-citations><srep-admin><examiners><primary-examiner><name>De Ceulaer, Bart</name></primary-examiner></examiners><srep-office><addressbook><text>Munich</text></addressbook></srep-office><date-search-completed><date>20250520</date></date-search-completed></srep-admin><!--							The annex lists the patent family members relating to the patent documents cited in the above mentioned European search report.							The members are as contained in the European Patent Office EDP file on							The European Patent Office is in no way liable for these particulars which are merely given for the purpose of information.							For more details about this annex : see Official Journal of the European Patent Office, No 12/82						--><srep-patent-family><patent-family><priority-application><document-id><country>EP</country><doc-number>4339943</doc-number><kind>A1</kind><date>20240320</date></document-id></priority-application><family-member><document-id><country>AU</country><doc-number>2023343374</doc-number><kind>A1</kind><date>20250424</date></document-id></family-member><family-member><document-id><country>CA</country><doc-number>3267259</doc-number><kind>A1</kind><date>20240321</date></document-id></family-member><family-member><document-id><country>CN</country><doc-number>119866523</doc-number><kind>A</kind><date>20250422</date></document-id></family-member><family-member><document-id><country>EP</country><doc-number>4339943</doc-number><kind>A1</kind><date>20240320</date></document-id></family-member><family-member><document-id><country>EP</country><doc-number>4588041</doc-number><kind>A1</kind><date>20250723</date></document-id></family-member><family-member><document-id><country>JP</country><doc-number>2025529995</doc-number><kind>A</kind><date>20250909</date></document-id></family-member><family-member><document-id><country>KR</country><doc-number>20250068699</doc-number><kind>A</kind><date>20250516</date></document-id></family-member><family-member><document-id><country>TW</country><doc-number>202429443</doc-number><kind>A</kind><date>20240716</date></document-id></family-member><family-member><document-id><country>WO</country><doc-number>2024056359</doc-number><kind>A1</kind><date>20240321</date></document-id></family-member></patent-family></srep-patent-family></srep-for-pub></search-report-data>
<ep-reference-list id="ref-list">
<heading id="ref-h0001"><b>REFERENCES CITED IN THE DESCRIPTION</b></heading>
<p id="ref-p0001" num=""><i>This list of references cited by the applicant is for the reader's convenience only. It does not form part of the European patent document. Even though great care has been taken in compiling the references, errors or omissions cannot be excluded and the EPO disclaims all liability in this regard.</i></p>
<heading id="ref-h0002"><b>Non-patent literature cited in the description</b></heading>
<p id="ref-p0002" num="">
<ul id="ref-ul0001" list-style="bullet">
<li><nplcit id="ref-ncit0001" npl-type="s"><article><author><name>E. SCHUIJERS</name></author><author><name>W. OOMEN</name></author><author><name>B. DEN BRINKER</name></author><author><name>J. BREEBAART</name></author><atl>Advances in Parametric Coding for High-Quality Audio</atl><serial><sertitle>114th AES Convention, Amsterdam, The Netherlands</sertitle><pubdate><sdate>20030000</sdate><edate/></pubdate></serial></article></nplcit><crossref idref="ncit0001">[0006]</crossref></li>
<li><nplcit id="ref-ncit0002" npl-type="s"><article><author><name>E. SCHUIJERS</name></author><author><name>J. BREEBAART</name></author><author><name>H. PURNHAGEN</name></author><author><name>J. ENGDEGÅRD</name></author><atl>Low Complexity Parametric Stereo Coding</atl><serial><sertitle>116th AES, Berlin, Germany</sertitle><pubdate><sdate>20040000</sdate><edate/></pubdate></serial></article></nplcit><crossref idref="ncit0002">[0006]</crossref></li>
<li><nplcit id="ref-ncit0003" npl-type="s"><article><author><name>XAVIER GLOROT</name></author><author><name>ANTOINE BORDES</name></author><author><name>YOSHUA BENGIO</name></author><atl>Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics</atl><serial><sertitle>PMLR</sertitle><pubdate><sdate>20110000</sdate><edate/></pubdate><vid>15</vid></serial><location><pp><ppf>315</ppf><ppl>323</ppl></pp></location></article></nplcit><crossref idref="ncit0003">[0136]</crossref></li>
<li><nplcit id="ref-ncit0004" npl-type="s"><article><author><name>OORD, AARON VAN DEN</name></author><author><name>SANDER DIELEMAN</name></author><author><name>HEIGA ZEN</name></author><author><name>KAREN SIMONYAN</name></author><author><name>ORIOL VINYALS</name></author><author><name>ALEX GRAVES</name></author><author><name>NAL KALCHBRENNER</name></author><author><name>ANDREW SENIOR</name></author><author><name>KORAY KAVUKCUOGLU</name></author><atl>Wavenet: A generative model for raw audio.</atl><serial><sertitle>arxiv:1609.03499</sertitle><pubdate><sdate>20160000</sdate><edate/></pubdate></serial></article></nplcit><crossref idref="ncit0004">[0140]</crossref></li>
</ul></p>
</ep-reference-list>
</ep-patent-document>
