FIELD OF THE INVENTION
[0001] The invention relates to an audio apparatus and method of operation therefor, and
specifically, but not exclusively, to rendering stereo signals using e.g. Parametric
Stereo based encoding.
BACKGROUND OF THE INVENTION
[0002] Spatial audio applications have become numerous and widespread and increasingly form
part of many audiovisual experiences. New and improved spatial experiences and applications
are continuously being developed which result in increased demands on the audio processing
and rendering.
[0003] A lot of research and development effort has focused on providing efficient and high
quality audio encoding and audio decoding for spatial audio. A frequently used spatial
audio representation is multichannel audio representations, including stereo representation,
and efficient encoding of such multichannel audio based on downmixing multichannel
audio signals to downmix channels with fewer channels have been developed. One of
the main advances in low bit-rate audio coding has been the use of parametric multichannel
coding where a downmix signal is generated together with parametric data that can
be used to upmix the downmix signal to recreate the multichannel audio signal.
[0004] In particular, instead of traditional mid-side or intensity coding, in parametric
multichannel audio coding, a multichannel input signal is downmixed to a lower number
of channels (e.g. two to one) and multichannel image (stereo) parameters are extracted.
Then the downmix signal is encoded using a more traditional audio coder (e.g. a mono
audio encoder). The bitstream of the downmix is multiplexed with the encoded multichannel
image parameter bitstream. This bitstream is then transmitted to the decoder, where
the process is inverted. First the downmix audio signal is decoded, after which the
multichannel audio signal is reconstructed guided by the encoded multichannel image/
upmix parameters.
[0005] An example of stereo coding is described in
E. Schuijers, W. Oomen, B. den Brinker, J. Breebaart, "Advances in Parametric Coding
for High-Quality Audio", 114th AES Convention, Amsterdam, The Netherlands, 2003, Preprint
5852. In the described approach, the downmixed mono signal is parametrized by exploiting
the natural separation of the signal into three components (objects): transients,
sinusoids, and noise. In
E. Schuijers, J. Breebaart, H. Pumhagen, J. Engdegård, "Low Complexity Parametric
Stereo Coding", 116th AES, Berlin, Germany, 2004, Preprint 6073 more details are provided describing how parametric stereo was realized with a low
(decoder) complexity when combining it with Spectral Band Replication (SBR).
[0006] Parametric Stereo (PS) is a technology which is widely used to efficiently code a
stereo signal as a mono downmix and a set of spatial parameters allowing an accurate
reconstruction of the stereo image. PS has been used to substantially improve the
compression efficiency for AAC and USAC at lower bit-rates, say 32kbps and below.
[0007] In addition to accurately reproducing a stereo signal, it has also been of interest
to create high quality binaural rendering of (encoded) stereo signals to emulate a
virtual loudspeaker playback.
[0008] Binaural rendering of content authored for multi-channel playback can be achieved
by the sum of convolutions of the input channel signals with left and right Head Related
Impulse Responses (HRIRs), where each HRIR pair corresponds to a measured/simulated
impulse response from a loudspeaker location to the ears. This can be expressed compactly
in the z-domain as:

where X
C(z) represents the z-transform of the time domain input signal x
c[n] with channel c, Y
L,R(z) represents the z-transform of the left and right time domain output signals l[n]
and r[n] respectively and

is the z-transform of the HRIR of the left and right channels h
l[n] and h
r[n] for the angle (and distance) corresponding to loudspeaker position φ
c. Approaches for rendering binaural stereo are disclosed in
WO2010/122455A1 and
WO2007/031896 A1.
[0010] However, whereas current approaches for audio rendering may provide acceptable performance
in many applications and scenarios, they tend to not be ideal and may exhibit suboptimal
behavior in some scenarios. In particular, it may result in suboptimal perceived quality
and/or a reduced user experience with e.g. perceived suboptimal spatial perception/audio
scene in some cases. Complexity and/or resource usage may also be higher than desired
and may in some case make the approach undesired or impractical for some implementations,
such as applications based on small and cheap portable devices.
[0011] Hence, an improved approach would be advantageous. In particular an approach allowing
increased flexibility, improved adaptability, improved performance, increased audio
quality, improved perceived quality, an improved rendering/generation of a stereo
signal, improved spatial perception, reduced complexity and/or resource usage, reduced
computational load, facilitated implementation, improved user experience, and/or an
improved spatial audio experience would be advantageous.
SUMMARY OF THE INVENTION
[0012] Accordingly, the Invention seeks to preferably mitigate, alleviate or eliminate one
or more of the above mentioned disadvantages singly or in any combination.
[0013] According to an aspect of the invention there is provided an audio apparatus comprising:
a receiver arranged to receive a data signal comprising encoded data for a mono downmix
audio signal (being a downmix of a stereo signal) and a set of spatial upmix parameters
for upmixing the mono downmix audio signal to a stereo signal, the set of spatial
upmix parameters being indicative of relative signal properties of channels of the
stereo signal; a store comprising directional transfer functions for different directions,
a directional transfer function for a given direction representing a mapping of a
mono audio signal to stereo channels such that the mono audio signal is positioned
in the given direction in a stereo image of the stereo channels; a decoder arranged
to generate the mono downmix audio signal by decoding the encoded data; a directional
circuit arranged to estimate a directional signal, the directional signal representing
a point source audio component of the stereo signal; a decorrelator arranged to generate
a decorrelated mono downmix audio signal from the mono downmix audio signal; a generator
arranged to generate a first estimated signal and a second estimated signal from the
mono downmix audio signal, at least the first estimated residual signal further being
generated from the first decorrelated signal and the directional signal, the first
estimated signal, and the second estimated signal forming a(n estimated) decomposition
of the stereo signal; a direction determining circuit arranged to determine a first
direction from the spatial parameters; a first renderer arranged to perform a first
rendering of the mono downmix audio signal to generate a first intermediate stereo
signal, the first rendering being a directional rendering using a first channel transfer
function retrieved from the store for the first direction; a second renderer arranged
to perform a second rendering being a rendering of the first estimated signal and
the second estimated signal to generate a second intermediate stereo signal; a combiner
arranged to combine at least the first intermediate stereo signal and the second intermediate
stereo signal to generate an output stereo signal; and an adapter arranged to control
the generator to adapt a first relative property being a relative property between
the first estimated signal and the second estimated signal in dependence on the spatial
parameters.
[0014] The approach may provide an improved audio experience in many embodiments. For many
signals and scenarios, the approach may provide improved rendering of a stereo audio
signal allowing improved generation/ reconstruction of a stereo audio signal with
an improved perceived audio quality. The approach may provide improved representation
of an audio scene by a stereo signal. The approach may in many embodiments allow an
improved and/or more attractive spatial audio experience from a stereo signal.
[0015] The approach may provide an efficient implementation and may in many embodiments
allow reduced complexity and/or resource usage. The approach may in many scenarios
allow a reduced computational burden while providing a perceived high quality rendering
of a stereo signal, and in particular based on an efficient representation of a stereo
signal using a downmix and spatial upmix parameters, such as specifically a PS encoded
stereo signal.
[0016] The processing may be in time frequency segments or tiles. Each time frequency segment/tile
may represent a frequency interval in a time interval. In many embodiments, the mono
downmix audio signal may be divided into time segments/intervals and a frequency representation
of the signal in the time segment/interval may be provided by signal values representing
different frequency segments of the signal in the time segment/interval. Some or all
of the processing may be performed in the frequency domain/ frequency subbands.
[0017] The spatial upmix parameters may comprise sets of upmix parameters, each set of upmix
parameters comprising at least one of: a level difference parameter indicative of
a level difference between channels of the multichannel audio signal; a correlation
parameter indicative of a coherence between channels of the multichannel audio signal;
a timing difference parameter indicative of a timing difference between channels of
the multichannel audio signal, and a phase difference parameter indicative of a phase
difference between channels of the multichannel audio signal.
[0018] The first direction may be a desired/target rendering direction. A direction may
be an angle and/or orientation from a listening position/in a stereo image.
[0019] The directional rendering may be a binaural rendering generating the first intermediate
stereo signal as a binaural stereo signal comprising a point source positioned in
the first direction, the binaural rendering comprising selecting binaural impulse
response values as values for a binaural impulse response for a sound source in the
first direction. The directional transfer functions may be parameterized transfer
functions, and specifically may be represented in the frequency domain as weights
for each of a plurality of subbands. A weight may be provided for each stereo channel.
The weights may typically be complex valued.
[0020] The binaural impulse response values may be parametric values and may be frequency
tile values. The binaural impulse response values may be values representing any suitable
binaural impulse response in any suitable way, including HRIR, HRTF, BRIR values etc.
[0021] In many embodiments, the encoded data and the spatial upmix parameters are part of
a Parametric Stereo encoding of the stereo signal.
[0022] In some embodiments, the second rendering is arranged to generate the second intermediate
stereo signal using a set of directional transfer functions retrieved from the store
for a set of predetermined directions.
[0023] In many scenarios and/or embodiments, the first and second estimated residual signals
are estimates of audio components of the mono downmix audio signal not included in
the directional signal. The first estimated signal may be a first estimated residual
signal. The second estimated signal may be a second estimated residual signal.
[0024] In many embodiments, the first and/or second rendering is a binaural rendering and
the directional transfer functions are binaural transfer functions.
[0025] In some embodiments, the second rendering is arranged to generate the second intermediate
stereo signal using a set of directional transfer functions retrieved from the store
for a set of predetermined directions.
[0026] This may provide an advantageous approach for many scenarios, including e.g. providing
an advantageous trade-off between complexity, computational resources, data rate and/or
the perceived audio quality of the generated output stereo signal.
[0027] In some embodiments, the set of predetermined directions consists of one predetermined
direction.
[0028] This may provide a particularly advantageous trade-off between complexity and user
experience/perceived audio quality in many scenarios and embodiments.
[0029] In some embodiments, the set of predetermined directions comprises a plurality of
predetermined directions.
[0030] This may provide a particularly advantageous trade-off between complexity and user
experience/perceived audio quality in many scenarios and embodiments.
[0031] In some embodiments, the spatial upmix parameters and the directional transfer functions
are provided for frequency subbands and the first renderer is arranged to generate
subband values for subbands of the first intermediate stereo signals from subband
values of the modified mono downmix audio signal based on spatial upmix parameters
and directional transfer functions for the subbands.
[0032] This may provide improved audio rendering in many embodiments. It may in many scenarios
and applications allow an improved and/or more attractive spatial audio experience
from a stereo signal.
[0033] In some embodiments, the direction determining circuit is arranged to determine a
point source direction in a stereo image of the stereo signal from the spatial upmix
parameters, and to determine the first direction by applying a mapping function to
the point source direction.
[0034] This may provide an advantageous approach for many scenarios, including e.g. providing
an advantageous trade-off between complexity, computational resources, data rate and/or
the perceived audio quality of the generated output stereo signal.
[0035] The mapping may be non-uniform. The mapping may be from one range to a different
range (e.g. from [0,90°] to [-30°,30°]). The mapping may be non-linear.
[0036] In some embodiments, the direction determining circuit may be arranged to determine
an indication of the point source direction as:

where IID is an interchannel intensity difference and ICC is an inter-channel cross-correlation.
The point source direction may be derived by interpreting y as a relative angle between
two loudspeakers (γ = 0° is one speaker, 90° is the other). In typical notation a
front-left speaker is typically at 30 degrees, and front-right at -30 degrees.
[0037] According to an optional feature of the invention, the first relative property includes
a level of the first estimated residual signal relative to the second estimated residual
signal.
[0038] This may provide improved audio rendering in many embodiments. It may in many scenarios
and applications allow an improved and/or more attractive spatial audio experience
from a stereo signal. It may provide an advantageous approach for many scenarios,
including e.g. providing an advantageous trade-off between complexity, computational
resources, data rate and/or the perceived audio quality of the generated output stereo
signal.
[0039] According to an optional feature of the invention, the first relative property includes
a correlation between the first estimated residual signal relative to the second estimated
residual signal.
[0040] This may provide improved audio rendering in many embodiments. It may in many scenarios
and applications allow an improved and/or more attractive spatial audio experience
from a stereo signal. It may provide an advantageous approach for many scenarios,
including e.g. providing an advantageous trade-off between complexity, computational
resources, data rate and/or the perceived audio quality of the generated output stereo
signal.
[0041] According to an optional feature of the invention, the adapter is arranged to adapt
a second relative property being a relative property between the directional signal
and the first estimated signal in dependence on the spatial parameters.
[0042] This may provide improved audio rendering in many embodiments. It may in many scenarios
and applications allow an improved and/or more attractive spatial audio experience
from a stereo signal.
[0043] According to an optional feature of the invention, the second relative property includes
a level of the directional signal relative to the first estimated residual signal.
[0044] The approach may allow an improved audio quality to data rate relationship in many
scenarios and applications.
[0045] According to an optional feature of the invention, the second relative property includes
a correlation between the directional signal relative to the first estimated residual
signal.
[0046] The approach may allow an improved audio quality in many scenarios and applications.
[0047] According to an optional feature of the invention, at least the directional signal
and the first estimated signal are each a linear combination of the mono downmix audio
signal and the decorrelated mono downmix audio signal.
[0048] This may provide improved audio rendering in many embodiments. It may in many scenarios
and applications allow an improved and/or more attractive spatial audio experience
from a stereo signal. It may provide an advantageous approach for many scenarios,
including e.g. providing an advantageous trade-off between complexity, computational
resources, data rate and/or the perceived audio quality of the generated output stereo
signal.
[0049] The second estimated signal may also typically be a linear combination of the mono
downmix audio signal and the decorrelated mono downmix audio signal.
[0050] According to an optional feature of the invention, the adapter is arranged to determine
weights of the linear combination as a function of the spatial parameters.
[0051] This may provide improved audio rendering in many embodiments. It may in many scenarios
and applications allow an improved and/or more attractive spatial audio experience
from a stereo signal. It may provide an advantageous approach for many scenarios,
including e.g. providing an advantageous trade-off between complexity, computational
resources, data rate and/or the perceived audio quality of the generated output stereo
signal.
[0052] In some embodiments, the directional circuit is arranged to generate the directional
signal by applying first frequency band weights to frequency band samples of the mono
downmix audio signal, the first frequency band weights being dependent on the spatial
parameters.
[0053] In many embodiments no other signals than the mono downmix audio signal are included
in the generation of the directional signal. In many embodiments, the first frequency
band weights are dependent only on the spatial parameters.
[0054] In some embodiments, the frequency band weights may be determined substantially as:

or

where IID is an interchannel intensity difference and ICC is an inter-channel cross-correlation
and make up part of the spatial parameters.
[0055] According to an optional feature of the invention, the directional signal, the first
estimated signal, and the second estimated signal are generated to be estimates of
a signal model for the stereo signal, the signal model representing the stereo signal
as a matrix multiplication of a signal vector by a matrix having a rank of two, the
signal vector comprising a directional signal and two diffuse signals.
[0056] The matrix may be a 2x3 matrix.
[0057] In some embodiments, the directional signal, the first estimated signal, and the
second estimated signal are generated to be estimates of a signal model for the stereo
signal, the signal model representing the stereo signal as a matrix multiplication
of a signal vector comprising a directional signal and two diffuse signals by a matrix
having a rank of two.
[0058] This may provide particularly advantageous operation and performance in many embodiments
and scenarios.
[0059] According to an optional feature of the invention, the adapter is arranged to adapt
the first relative property to match a corresponding relative property between the
two diffuse signals of the signal model.
[0060] This may provide particularly advantageous operation and performance in many embodiments
and scenarios.
[0061] According to an optional feature of the invention, the generator is arranged to generate
the first estimated signal by applying second frequency band weights to frequency
band samples of the decorrelated mono downmix audio signal, the second frequency band
weights being dependent on the spatial parameters.
[0062] This may provide particularly advantageous operation and performance in many embodiments
and scenarios.
[0063] In many embodiments no other signals than the decorrelated mono downmix audio signal
are included in the generation of the first estimated signal. In many embodiments,
the second frequency band weights are dependent only on the spatial parameters.
[0064] According to an optional feature of the invention, the generator is arranged to generate
the second estimated signal from the decorrelated mono downmix audio signal.
[0065] This may provide particularly advantageous operation and performance in many embodiments
and scenarios.
[0066] According to an optional feature of the invention, the generator is arranged to generate
the second estimated signal by applying third frequency band weights to frequency
band samples of the decorrelated mono downmix audio signal, the third frequency band
weights being dependent on the spatial parameters.
[0067] This may provide particularly advantageous operation and performance in many embodiments
and scenarios.
[0068] In many embodiments no other signals than the decorrelated mono downmix audio signal
are included in the generation of the second estimated signal. In many embodiments,
the third frequency band weights are dependent only on the spatial parameters.
[0069] According to another aspect of the invention, there is provided method of operation
for an audio apparatus, the method comprising: receiving a data signal comprising
encoded data for a mono downmix audio signal and a set of spatial upmix parameters
for upmixing the mono downmix audio signal to a stereo signal, the set of spatial
upmix parameters being indicative of relative signal properties of channels of the
stereo signal; providing directional transfer functions for different directions,
a directional transfer function for a given direction representing a mapping of a
mono audio signal to stereo channels such that the mono audio signal is positioned
in the given direction in a stereo image of the stereo channels; generating the mono
downmix audio signal by decoding the encoded data; estimating a directional signal,
the directional signal representing a point source audio component of the stereo signal;
generating a decorrelated mono downmix audio signal from the mono downmix audio signal;
generating a first estimated signal and a second estimated signal from the mono downmix
audio signal, at least the first estimated residual signal further being generated
from the first decorrelated signal and the directional signal, the first estimated
signal, and the second estimated signal forming a decomposition of the stereo signal;
determining a first direction from the spatial parameters; performing a first rendering
of the mono downmix audio signal to generate a first intermediate stereo signal, the
first rendering being a directional rendering using a first channel transfer function
retrieved from the store for the first direction; performing a second rendering being
a rendering of the first estimated signal and the second estimated signal to generate
a second intermediate stereo signal; combining at least the first intermediate stereo
signal and the second intermediate stereo signal to generate an output stereo signal;
and controlling the generator to adapt a first relative property being a relative
property between the first estimated signal and the second estimated signal in dependence
on the spatial parameters.
[0070] These and other aspects, features and advantages of the invention will be apparent
from and elucidated with reference to the embodiment(s) described hereinafter.
BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Embodiments of the invention will be described, by way of example only, with reference
to the drawings, in which
FIG. 1 illustrates some elements of an example of an audio distribution system;
FIG. 2 illustrates some elements of an example of an audio apparatus in accordance
with some embodiments of the invention;
FIG. 3 illustrates some elements of an example of an audio render apparatus in accordance
with some embodiments of the invention; and
FIG. 4 illustrates some elements of a possible arrangement of a processor for implementing
elements of an audio apparatus in accordance with some embodiments of the invention.
DETAILED DESCRIPTION OF SOME EMBODIMENTS OF THE INVENTION
[0072] FIG. 1 illustrates an example of an audio system wherein a stereo signal may be distributed/communicated
for remote rendering. In the system, an audio apparatus referred to as the audio source
device 101 generates an audio data signal including a representation of a stereo audio
signal. The stereo audio signal may be one captured at the audio source device 101,
may be received from another source, or indeed may e.g. be an artificially generated
stereo signal (e.g. it may be a virtual audio stereo signal).
[0073] The audio source device 101 may generate the data signal to include encoded data
that represents a mono downmix audio signal for the stereo signal. For example, the
mono downmix audio signal may be generated as a weighted combination, and specifically
as a weighted summation, of the channel signals of an input stereo signal. In many
cases, the weights may be fixed and may specifically be the same for the two channel
signals.
[0074] The generated mono downmix audio signal is encoded using a suitable mono audio encoding
algorithm/standard to generate encoded audio data representing the mono downmix audio
signal.
[0075] In addition to the mono downmix audio signal, the audio source device 101 generates
spatial upmix parameters for upmixing the mono downmix audio signal to recreate the
original stereo signal.
[0076] The spatial upmix parameters are generated to be indicative of/reflect relative properties
of the channel signals of the stereo audio signal. In particular, the spatial upmix
parameters may be indicated to include parameters that are indicative of at least
one of relative intensities/levels of the stereo channels, relative (frequency domain)
phases of the stereo channels, a relative time difference between the channels, and/or
a correlation between the channels. Specifically, the audio source device 101 may
generate spatial upmix parameters including one or more of an inter-channel intensity
difference, inter-channel level difference, inter-channel time difference, inter-channel
phase difference, and/or inter-channel correlation.
[0077] The data signal may specifically comprise a Parametric Stereo (PS) encoding of the
stereo signal.
[0078] A classical PS downmix is calculated as:

where the parameter
c is chosen such that the power of the stereo signal is preserved in the downmix, the
power being defined using the 2-norm:

and thus e.g.:

[0080] The spatial upmix parameters are typically generated for specific time frequency
tiles, and thus specifically each parameter value is generated/provided for a given
frequency subband/interval and for a given time segment/interval.
[0081] The audio source device 101 may accordingly encode a stereo signal as encoded data
representing a mono downmix audio signal of the stereo signal and associated spatial
upmix parameters that are indicative of relative properties of the channels (the channel
signals) of the stereo signal. Specifically, the audio source device 101 may be arranged
to generate a data signal comprising a conventional PS encoded stereo signal.
[0082] The system of FIG. 1 further comprises an audio apparatus which henceforth will be
referred to as the audio render apparatus 103. The audio render apparatus 103 is arranged
to receive the data signal generated by the audio source device 101. In the example,
the audio render apparatus 103 and audio source device 101 are both coupled to a network
105 through which the data signal can be communicated and specifically through which
it can be communicated from the audio source device 101 to the audio render apparatus
103. The network 105 may specifically be, or include, the Internet.
[0083] Thus, the receiver 101 receives a data signal which comprises encoded data for a
mono downmix audio signal and spatial upmix parameters for upmixing the mono downmix
audio signal to the stereo signal. The set of spatial upmix parameters comprises one
or more parameters indicative of relative signal properties of channels of the stereo
signal, and may specifically be indicative of a level/intensity difference between
the channels of the stereo signal, a cross-correlation between the channels of the
stereo signal, and/or a phase difference or time difference between the channels of
the stereo signals. In many cases, the data signal may comprise ICC, IID and/or IPD
parameters. The data signal may specifically comprise a PS (Parametric Stereo) encoded
stereo signal. The audio source device 101 may be arranged to generate a data signal
which comprises a stereo signal encoded in accordance with the Parametric Stereo (PS)
specifications/standard. The receiver 101 may accordingly receive a representation
of a stereo signal encoded by a mono downmix audio signal and spatial upmix parameters,
and specifically a PS encoded stereo signal.
[0084] The audio render apparatus 103 is arranged to process the data signal to render a
stereo signal. An audio render apparatus could render the stereo signal of the data
signal using a conventional PS rendering approach based on a PS upmixing of the mono
downmix audio signal using the PS spatial parameters. However, the audio render apparatus
103 of FIG. 1 uses a specific approach where point source components may be specifically
considered, and where in many embodiments different/parallel paths process the received
mono downmix audio signal in different ways to generate different stereo signal components
which are then combined to generate the output stereo signal.
[0085] The rendering by the audio render apparatus 103 may include a differentiated rendering
of different audio components with different properties, and in particular may seek
to render some components corresponding to point audio sources with a specific directional
property while other components may be rendered less spatially specific including
in particular as spatially diffuse components. The approach may be based on considering
a signal model for a stereo signal that may allow decomposition of the stereo signal
into different signal components.
[0086] A signal model may specifically consider the stereo signal to be a linear combination
of three signal components s, d
l, and d
r:

[0087] However, one problem with such an approach is that it is in principle not always
possible to determine the opposite relationship, i.e. to determine the three signal
components from the stereo signals (in principle it is not possible to determine three
independent components from two signal components). In other words, it is generally
not possible to perform a signal inversion as the inverse matrix
G of the signal matrix
F cannot be determined.
[0088] In order to perform a decomposition of a stereo signal into three components, it
may accordingly be required to include some assumptions or restrictions on the signal
model.
[0089] A possible signal model for a stereo signal may be represented by:

[0090] This signal model essentially represents a consideration that the stereo signal corresponds
to a combination of a directional signal and diffuse background audio. The directional
signal s is an audio component that corresponds/reflects audio that is considered
to originate from an audio source that has a point source property/characteristic
and which thus has a well-defined spatial origin. The audio of the directional signal
thus corresponds to audio that can/should be rendered to be perceived to reach the
listener from a specific direction. The directional signal may for each time frequency
tile represent a single audio point source and thus may for each time frequency tile
have a specific source position. The directional signal may in some cases/embodiments
correspond to different audio point sources in different time frequency tiles, i.e.
it is not required that all time frequency tiles of the directional source represent
audio from the same point source. In some cases, the directional signal may correspond
to different frequency components being reflected differently in the environment and
thus may (e.g. in some frequency bands) correspond to early reflections of the point
source within the acoustic environment.
[0091] In the signal model, the directional signal component
s is phase shifted using two parameters
ϕl and
ϕr, and is further panned/positioned in the stereo image of the original stereo channels
l and
r. The panning is to an angle represented by the panning angle
γ. Furthermore, a diffuse signal component is represented by diffuse residual signal
components
dl and
dr of the respective left and right channels.
[0092] Such a model may be used by a render apparatus to generate estimates of the three
signal decomposition signals. Further, these may be generated based on a single received
downmix
m and two decorrelated signals
d1 and
d2. In particular, a render apparatus may proceed to generate estimates of the decomposition
signals as:

[0093] The approach thus generates a main/directional signal
s' and two residual signals as decorrelated signals thereby allowing them to be generated
from the received mono signal. Further, the gains may be selected to seek to maintain
properties corresponding to the stereo signal. In particular, with ∥
d1∥
2 = ∥
d2∥
2 = ∥
m∥
2, the approach may reconstruct level properties of the underlying signal model, i.e.

[0094] In addition to the generated decomposition signals being decorrelated, the signal
levels of the generated residual signals are the same ∥
dl'∥
2 = ∥
dr∥
2.
[0095] However, whereas such an approach may provide an advantageous decomposition that
may allow improved rendering in many cases, it tends to not always provide optimum
audio quality.
[0096] A signal model for decomposition of a stereo signal (
l,
r) may be based on considering the linear combination:

but with a matrix
F having a rank of two. This allows a signal/matrix inversion allowing the decomposition
signals to be determined from the stereo signals:

[0097] Based on this consideration, a renderer may seek to generate local estimates/replicas
of the decomposition signals from the received mono downmix audio signal and a decorrelated
signal:

[0098] The matrix coefficients may be adapted to seek to maintain the properties of the
underlying signal model. Specifically, with ∥
d∥
2 = ∥
m∥
2 this may allow reconstruction of a range of properties:

[0099] It should be noted that in this case the correlation between the residual signals
(i.e. the last correlation) is typically not 0 (in contrast to the previous signal
model where it is an inherent and fundamental property/assumption of the model). It
is also noted that the levels of the residual signals are not (necessarily) the same
(in contrast to the previous signal model where it is an inherent and fundamental
property/assumption of the model).
[0100] Accordingly, whereas both models may reconstruct the properties of the signal components
of the underlying model, the Inventors have realized that improved decomposition into
specific signal components can be achieved based on the latter model considerations
and that this in particular is suitable for a particular and differentiated rendering
as will be described in detail later.
[0101] It is noted that the signal model description does not (typically/necessarily) refer
to a time-domain signal, but rather can alternatively or additionally refer to individual
(potentially relatively small) frequency subbands. For example, the described signal
model may individually apply to each of the frequency subbands for which separate
spatial upmix parameters are provided.
[0102] The audio render apparatus 103 may be based on a consideration of the signal model
as indicated above. In particular, the audio render apparatus 103 is arranged to receive
the data signal including the mono downmix audio signal and spatial parameters and
to generate local replicas/estimates of the directional signal s' and the residual
signals d'
l, d'
r from the mono downmix audio signal and the spatial parameters. These signal components
may then be separately rendered to generate a combined stereo signal that may provide
an improved spatial user experience in many embodiments.
[0103] The audio source device 101 of FIG. 1 is arranged to generate a data signal with
data representing an audio stereo signal. The audio source device 101 is specifically
arranged to generate the data signal to include encoded audio data for a mono downmix
audio signal which represents a downmix of the stereo signal. In addition, spatial
parameters that indicate relative properties between the channels of the stereo signal
are included. Such spatial data may be appropriate for upmixing the mono downmix audio
signal to recreate the downmix represented by the mono downmix audio signal.
[0104] As illustrated in FIG. 2, the audio source device 101 comprises a receiver 201 which
is arranged to receive audio components from which the mono downmix audio signal is
generated. In many embodiments, the receiver 201 may directly receive a stereo signal
which is to be represented by the data signal. In other embodiments, the receiver
201 may additionally or alternatively receive a number of audio components such as
audio objects, mono signals, single source signals, multichannel signals etc.
[0105] The audio components are fed to a downmixer 203 which proceeds to generate the mono
downmix audio signal. In many embodiments and scenarios, the downmixer 203 may generate
a mono downmix audio signal from a received stereo signal, e.g. simply by summing
the channels of the stereo signal in accordance with a standard PS downmix approach.
In other embodiments, the downmixer 203 may be arranged to generate a stereo signal
from received audio components, such as e.g. by generating an intermediate stereo
signal for each received audio component followed by a combination of the intermediate
stereo signals. For example, a multi-channel signal may be downmixed to an intermediate
stereo signal, an audio object may be panned to the stereo image of an intermediate
stereo signal based on position data provided for the audio object, etc.
[0106] The audio source device 101 further comprises a spatial parameter circuit 205 which
is arranged to determine the spatial parameters for the stereo signal represented
by the mono downmix audio signal. Specifically, the received or locally generated
stereo signal may be fed to the parameter circuit 205 which may determine the spatial
parameters from an analysis/processing of the stereo signal.
[0107] The spatial parameter circuit 205 may provide sets of frequency subband spatial parameters
for the stereo signal where the sets of frequency subband spatial parameters are indicative
of relative signal properties of the channels of the stereo signal. The frequency
subband spatial parameters are provided for individual subbands of the stereo signal.
[0108] The spatial parameters are indicative of/reflect relative properties of the channel
signals of the stereo audio signal. In particular, the spatial parameters may be indicated
to include parameters that are indicative of at least one of relative intensities/levels
of the stereo channels, relative (frequency domain) phases of the stereo channels,
a relative time difference between the channels, and/or a correlation between the
channels. Specifically, the spatial parameters may include one or more of an inter-channel
intensity difference, inter-channel level difference, inter-channel time difference,
inter-channel phase difference, and/or inter-channel correlation.
[0109] The spatial parameters may specifically be spatial parameters as used for encoding
a stereo signal using a Parametric Stereo (PS) encoding of the stereo signal, such
as for example by a mono downmix audio signal given as:

where the parameter
c is chosen such that the power of the stereo signal is preserved in the downmix, the
power being defined using the 2-norm:

and thus e.g.:

[0111] The spatial parameters are typically provided for specific time frequency tiles,
and thus specifically each parameter value is generated/provided for a given frequency
subband and for a given time segment.
[0112] The spatial parameter circuit 205 may determine spatial parameters that are indicative
of relative properties of the channels (channel signals) of the stereo signal.
[0113] In many embodiments, the spatial parameter circuit 205 may receive the input stereo
signal and process/analyze this to generate the spatial parameters. Specifically,
the spatial parameter circuit 205 may calculate the IID, ICC, and IPD values in accordance
with the formulas indicated above.
[0114] The spatial parameters may comprise sets of spatial parameters, each set of spatial
parameters comprising at least one of: a level difference parameter indicative of
a level difference between channels of the multichannel audio signal; a correlation
parameter indicative of a coherence between channels of the multichannel audio signal;
a timing difference parameter indicative of a timing difference between channels of
the multichannel audio signal, and a phase difference parameter indicative of a phase
difference between channels of the multichannel audio signal.
[0115] The audio source device 101 further comprises a data signal circuit 207 which generates
the audio data signal and specifically it generates the audio data signal to include
data representing the mono downmix audio signal and data representing the spatial
parameters. The data signal circuit 207 may specifically include an audio signal encoder
arranged to encode the mono downmix audio signal. It will be appreciated that any
suitable audio signal encoding algorithm and approach may be used, such as an Advanced
Audio Coding (AAC) mono encoding algorithm. Likewise, the spatial parameters may be
encoded using a suitable encoding approach.
[0116] The data signal circuit 207 may generate an output data signal in accordance with
any suitable format, and may specifically generate the output data signal to follow
a suitable standard for an audio signal.
[0117] FIG. 3 shows examples of elements of the audio render apparatus 103.
[0118] The audio render apparatus 103 comprises a receiver 301 which is arranged to receive
the data signal from the audio source device 101. Thus, the receiver 301 receives
a data signal comprising encoded data for a mono downmix audio signal of a stereo
signal. In addition, the data signal includes spatial upmix parameters for upmixing
the mono downmix audio signal to the stereo signal where the spatial upmix parameters
are indicative of relative signal properties of channels of the stereo signal. The
spatial upmix parameters may as mentioned specifically be inter-channel time, phase,
level, intensity differences and/or inter-channel correlation measures.
[0119] The receiver 301 is coupled to a decoder 303 which is arranged to receive the encoded
data representing the mono downmix audio signal and to decode this data to generate
the mono downmix audio signal. It will be appreciated that any suitable method for
encoding and decoding the mono downmix audio signal may be used and in particular
that any suitable standardized encoding format and algorithm may be used.
[0120] The receiver 301 may be arranged to receive a time domain audio signal and/or a frequency
domain audio signal version/representation of the mono downmix audio signal. In some
cases, the received data signal may include the mono downmix audio signal in only
one representation, i.e. the data signal may include only one of the frequency domain
audio signal and the time domain audio signal. In such cases, the received data signal
may be transformed to the other domain as appropriate. Thus, in some cases, a received
data signal may include a time domain audio signal being the time domain representation
of the mono downmix audio signal, and a time to frequency domain transformer may from
this generate the frequency domain audio signal for the mono downmix audio signal.
In some cases, a received data signal may include a frequency domain audio signal
being the frequency domain representation of the mono downmix audio signal and a frequency
to time domain transformer may from this generate the time domain audio signal for
the mono downmix audio signal if necessary.
[0121] In particular, in some embodiments, the receiver may comprise a filter bank which
is arranged to generate a frequency subband representation of a received time domain
mono downmix audio signal. The receiver 301 may comprise a filter bank that is applied
to the mono downmix audio signal such that it is divided into frequency subbands.
[0122] The filter bank may be Quadrature Mirror Filter (QMF) bank or may e.g. be implemented
by a Fast Fourier Transform (FFT), but it will be appreciated that many other filter
banks and approaches for dividing an audio signal into a plurality of subband signals
are known and may be used. The filterbank may specifically be a complex-valued pseudo
QMF bank, resulting in e.g. 32 or 64 complex-valued sub-band signals.
[0123] The processing is furthermore typically performed in time segments or time slots.
In most embodiments, the audio signal is divided into time intervals/segments with
a conversion to the frequency/subband domain by applying e.g. an FFT or QMF filtering
to the samples of each signal. For example, each channel of the downmix audio signal
may be divided into time segments of e.g. 2048, 1024, or 512 samples. These signals
may then be processed to generate samples for e.g. 64, 32 or 16 subbands. Thus, a
set of samples may be determined for each subband of the mono downmix audio signal.
[0124] It should be noted that the number of time domain samples is not directly coupled
to the number of subbands. Typically, for a so-called critically sampled filterbank
of N bands, every N input samples will lead to N sub-band samples (one for every sub-band).
An oversampled filterbank will produce more output samples. E.g. for every N input
samples, it would generate k*N output samples, i.e., k consecutive samples for every
band.
[0125] In some embodiments, the subbands are generated to have the same bandwidth but in
other embodiments subbands are generated to have different bandwidths, e.g. reflecting
the sensitivity of human hearing to different frequencies.
[0126] For example, the receiver 301 may employ a hybrid filterbank with logarithmic filter
band center-frequency spacings that follow that of human perception similar to equivalent
rectangular bandwidths (ERBs). In order to compensate for the delay of the filtering
by the small filter bank, a delay may be introduced for higher frequency subbands.
[0127] As a specific example, a time-domain signal
x[
n] may be fed through a downsampled complex-exponential modulated QMF bank with
K bands. Each frame of 64 time domain samples
x[
n] results in one slot of QMF samples
X[
k, l] with
k = (0, ... ,
K - 1) at slot
l. The lower slots may then be filtered by additional complex-modulated filterbanks
splitting the lower bands further. The higher slots are delayed ensuring that the
filtered mono downmix audio signals of the lower bands are in sync with the higher
bands as the filtering introduces a delay. This finally results in a structure where
for every 64 time-domain samples
x[
n], one slot m of hybrid QMF samples
Y[
k, l] is produced with
k = (0, ... ,
L - 1) at slot
l, e.g. with a total number of hybrid bands
M = 77.
[0128] Thus, in many embodiments, the signals and the processing may be performed in subbands
and for individual segments. Such blocks of a frequency interval/subband in a given
time interval/segment will also be referred to as time frequency segments/tiles.
[0129] The mono downmix audio signal is fed to a directional signal circuit 305 which is
arranged to generate a directional signal from the mono downmix audio signal. Thus,
the directional signal circuit 305 may generate a modified mono downmix audio signal
that seeks to represent/estimate a point source audio component of the stereo signal.
The directional signal may be generated to seek to represent parts of the stereo signal
which can be considered to have an origin with point source properties. Thus, the
directional signal is generated to include audio components that correspond to audio
that have a specific position/direction of origin.
[0130] The directional signal may in some cases be generated to represent/estimate audio
that has a specific position in the stereo image of the stereo signal, and thus which
corresponds to audio reaching the capture position (for the stereo signal) from a
single direction.
[0131] The directional signal may accordingly be generated to estimate audio that is linked
with audio sources that are spatially limited/small/localized and which specifically
are generated by point sources. It may thus seek to extract such audio components/parts
from the stereo signal to separate point source audio from audio with a spatial extension,
such as e.g. background audio, ambient audio, etc.
[0132] In some cases, the directional signal generator 305 may be arranged to estimate an
audio component from a single point source and specifically the directional signal
may be generated to represent/estimate a single point source. In such an approach,
audio from one specific direction may be identified/estimated and represented by the
directional signal.
[0133] In most embodiments, however, the directional signal is generated to represent/estimate
audio that has a point source origin but not necessarily from the same single point
source. The directional signal may accordingly in many embodiments represent/estimate
audio from a plurality of point sources.
[0134] The modified mono downmix audio signal/ directional signal is fed to a first renderer
307 which is arranged to render the modified mono downmix audio signal to generate
a first intermediate stereo signal. The rendering by the first renderer 307 (also
referred to as a first rendering) is a directional rendering which renders the first
intermediate stereo signal with a given direction/position in the stereo image of
the first intermediate stereo signal. The first rendering may specifically render
the mono downmix audio signal as a point source with a given direction/position in
the stereo image.
[0135] The first renderer 307 is coupled to a direction determining circuit 309 which is
arranged to determine a direction γ' which is fed to the first renderer 307 resulting
in this rendering the modified mono downmix audio signal from this position/direction.
Thus, the first rendering is specifically such that the modified mono downmix audio
signal in the first intermediate stereo signal is perceived as a point audio source
positioned in the direction corresponding to the direction γ' determined by the direction
determining circuit 309. The direction γ' will also be referred to as the rendering
direction or rendering angle.
[0136] The direction determining circuit 309 is arranged to determine the direction from
the received spatial upmix parameters. The spatial upmix parameters provide information
on the relationship between the channels of the stereo signal that is downmixed and
as such provide information of the position/orientation of the audio, and specifically
of a dominant signal component in the stereo image of the stereo signal. For example,
for a PS encoded signal, the spatial upmix parameters provide information of the position
of the dominant signal component in the stereo signal, and specifically it provides
information of an orientation angle for the dominant signal.
[0137] The direction determining circuit 309 may specifically determine the rendering direction
γ' from the spatial upmix parameters. The rendering direction will typically be determined
on a frequency tile basis, and specifically in frequency subbands and time segments
matching those for which the spatial upmix parameters are provided.
[0138] The first renderer 307 may accordingly proceed to render the modified mono downmix
audio signal such that is perceived from the given direction and it specifically achieves
this directional rendering by applying a channel transfer function to the mono downmix
audio signal with the channel transfer function generating the intermediate stereo
signal from the mono downmix audio signal. The channel transfer function may specifically
include a sub-transfer function for each channel, i.e. it may include one (sub)transfer
function for generating a left channel signal and one (sub)transfer function for generating
the right channel signal.
[0139] In many cases, the transfer function may be provided as a set of complex weights
for the different subbands of a frequency representation of the modified mono downmix
audio signal. The audio apparatus may perform many or all of the operations in the
frequency domain and thus the transfer function may also be expressed and applied
in the frequency domain. For example, for each frequency subband of the representation
of the mono downmix audio signal, the transfer function may provide a complex weight
for each of the output channels and a frequency representation of the first intermediate
stereo signal may be generated by applying/multiplying the subband samples of the
mono downmix audio signal by these weights to generate the subband samples of the
first intermediate stereo signal.
[0140] The first transfer function is determined to correspond to the desired direction,
i.e. it reflects the mapping from the mono downmix audio signal to the channels of
the first intermediate stereo signal such that it is perceived as/corresponds to an
audio source at a position in the stereo image corresponding the rendering direction/angle.
[0141] For example, in some cases, the transfer function for a given direction may correspond
to a panning of the mono downmix audio signal to the given direction in the stereo
image.
[0142] In many embodiments, the first rendering may be a binaural rendering and the first
intermediate stereo signal may be a binaural stereo signal providing an enhanced spatial
experience/perception when heard through headphones. Thus, the first renderer 307
may specifically be a binaural audio renderer which generates binaural audio signals
for the left and right ear of a user. Binaural audio signals are generated to provide
a desired spatial experience and are typically reproduced by headphones or earphones
that specifically may be part of a headset worn by a user (the headset typically also
comprises left and right eye displays).
[0143] Thus, in many embodiments, the audio rendering by the first renderer 307 is a binaural
render process using suitable binaural transfer functions to provide the desired spatial
effect for a user wearing a headphone. For example, the first renderer 307 may be
arranged to generate an audio component to be perceived to arrive from a specific
position using binaural processing.
[0144] Binaural processing is known to be used to provide a spatial experience by virtual
positioning of sound sources using individual signals for the listener's ears. With
an appropriate binaural rendering processing, the signals required at the eardrums
in order for the listener to perceive sound from any desired direction can be calculated,
and the signals can be rendered such that they provide the desired effect. These signals
are then recreated at the eardrum using either headphones or a crosstalk cancelation
method (suitable for rendering over closely spaced speakers). Binaural rendering can
be considered to be an approach for generating signals for the ears of a listener
resulting in tricking the human auditory system into perceiving that a sound is coming
from the desired positions.
[0145] The binaural rendering is based on binaural transfer functions which vary from person
to person due to the acoustic properties of the head, ears and reflective surfaces,
such as the shoulders. Binaural transfer functions may therefore be personalized for
an optimal binaural experience. For example, binaural filters can be used to create
a binaural recording simulating multiple sources at various locations. This can be
realized by convolving each sound source with the pair of e.g., Head Related Impulse
Responses (HRIRs) that correspond to the position of the sound source.
[0146] A well-known method to determine binaural transfer functions is binaural recording.
It is a method of recording sound that uses a dedicated microphone arrangement and
is intended for replay using headphones. The recording is made by either placing microphones
in the ear canal of a subject or using a dummy head with built-in microphones, a bust
that includes pinnae (outer ears). The use of such dummy head including pinnae provides
a very similar spatial impression as if the person listening to the recordings was
physically present during the recording.
[0147] By measuring e.g., the responses from a sound source at a specific location in 2D
or 3D space to microphones placed in or near the human ears, the appropriate binaural
filters can be determined. Based on such measurements, binaural filters reflecting
the acoustic transfer functions to the user's ears can be generated. The binaural
filters can be used to create a binaural recording simulating multiple sources at
various locations. This can be realized e.g., by convolving each sound source with
the pair of measured impulse responses for a desired position of the sound source.
In order to create the illusion that a sound source is moving around the listener,
a large number of binaural filters is typically required with a certain spatial resolution,
e.g., 10 degrees.
[0148] The head related binaural transfer functions may be represented e.g., as Head Related
Impulse Responses (HRIR), or equivalently as Head Related Transfer Functions (HRTFs)
or, Binaural Room Impulse Responses (BRIRs). The (e.g., estimated or assumed) transfer
function from a given position to the listener's ears (or eardrums) may for example
be represented in the frequency domain in which case it is typically referred to as
an HRTF or BRTF, or in the time domain in which case it is typically referred to as
a HRIR or BRIR. In some scenarios, the head related binaural transfer functions are
determined to include aspects or properties of the acoustic environment and specifically
of the environment in which the measurements are made, whereas in other examples only
the user characteristics are considered. Examples of the first type of functions are
the BRIRs and BRTFs.
[0149] The audio render apparatus 103 comprises a store 311 which stores directional transfer
functions for different directions. The directional transfer function for a given
direction represents the mapping of a mono audio signal to stereo channels such that
the mono audio signal is positioned in the given direction in a stereo image of the
stereo channels. Thus, applying the directional transfer function for a given direction
to the mono downmix audio signal may generate a stereo signal representing the mono
downmix audio signal as an audio source positioned in the given direction. The mapping
may in some cases be a time domain mapping (such as a gain, filter or other transfer
function) or may in many cases be a frequency domain mapping, such as a set of parameter
values/scale values (typically complex values) for different subbands. In the latter
case, a frequency domain intermediate stereo signal may be generated by for each subband
multiplying the subband sample of the mono downmix audio signal with respectively
a complex value for that subband for a first channel of the intermediate stereo signal
and with a complex value for that subband for a second channel of the intermediate
stereo signal.
[0150] For example, in examples where a panning is performed in the horizontal 2D plane,
the store 311 may comprise panning parameters for different directions. For example,
panning parameters for azimuth angles in a 0-360° interval may be provided for each
1° angle increment. The first renderer 307 may be coupled to the store 311 and be
arranged to extract the directional transfer function for the rendering direction
and then proceed to perform the rendering using the extracted directional transfer
function. The rendering of the modified mono downmix audio signal may accordingly
be rendered such that it is positioned/perceived in the stereo image to arrive from
the rendering position.
[0151] It will be appreciated that the store 311 may not have directional transfer function
stored for the desired rendering direction. In such cases, the first renderer 307
may be arranged to retrieve the nearest directional transfer function from the store
311 and use this for rendering. In such cases, the rendering direction may be considered
to correspond to the direction for the retrieved directional transfer function, i.e.
the rendered direction may be a quantized value γ' of the desired rendering direction
determined by the direction determining circuit 309.
[0152] In other embodiments, the first renderer 307 may be arranged to estimate a desired
directional transfer function for a desired rendering direction by interpolating between
two directional transfer functions from the store 311 corresponding to the two rendering
angles nearest to the desired rendering direction determined by the direction determining
circuit 309.
[0153] In most embodiments, the first renderer 307 is as mentioned arranged to perform a
binaural rendering and the directional transfer functions stored in the store 311
are binaural transfer functions. Thus, the store may store data describing binaural
transfer functions for different directions. The binaural transfer functions may for
example be HRTFs, BRIRs, or HRIRs. The store 311 may specifically store frequency
subband complex values for each channel for each frequency subband for a range of
different frequencies. The first renderer 307 may thus perform the binaural rendering
by multiplying the subband samples of the mono downmix audio signal with the corresponding
subband coefficients/complex values of the selected binaural transfer function to
generate subband sample values of the intermediate binaural stereo signal.
[0154] It will be appreciated that in many embodiments, the directional transfer functions
may be stored as a plurality of functions linked with different directions. For example,
the store 311 may be a look-up table which can receive the rendering direction as
an index and provide a set of values of the directional transfer function for that
direction. The directional transfer function may for example be represented by individual
subband values/coefficients, or may e.g. in other embodiments be represented by e.g.
parameter values defining the directional transfer function operation (e.g. coefficients
for the transfer function), a mathematical description/function from which suitable
values of the transfer function can be generated etc.
[0155] Thus, the audio render apparatus 103 comprises a processing path which generates
an intermediate stereo signal comprising the modified mono downmix audio signal represented
as an audio source at a specific position in the spatial image of the first intermediate
stereo signal. The mono downmix audio signal may typically be represented as a point
audio source at the given direction. The rendering is adaptive with the direction
being given by the spatial parameters and thus is dynamically adapted to reflect the
characteristics of the stereo signal.
[0156] Further, the spatially definite/well defined and directional rendering is specifically
of a directional signal representing point source audio. Thus, an improved spatial
rendering of point source audio can be achieved with point source audio typically
being rendered with a higher audio quality.
[0157] The generated first intermediate stereo signal is fed to an output circuit 313 which
is arranged to generate an output stereo signal that includes the first intermediate
stereo signal. As will be described more in the following, the output stereo signal
may be generated to include other signal components, including audio components representing
non-point source audio, such as background or ambient sounds.
[0158] In many embodiments, the rendering of the mono downmix audio signal includes extracting
a modified mono downmix audio signal/directional signal s' which may specifically
represent a point source audio component of the stereo signal.
[0159] It will be appreciated that different approaches, algorithms, and functions may be
used to estimate the directional signal from the mono downmix audio signal using the
spatial parameters.
[0160] For example, a rudimentary estimate of the directional signal from the mono signal
may follow from:
x' = ICC · m. This example may follow the consideration that for fully decorrelated signals all
of the signal power of the original stereo signal is captured by the directional signal,
whereas for fully decorrelated signals none of the signal power of the original stereo
signal is captured by the directional signal.
[0161] In a particular approach, the directional signal is estimated as a spectro-temporally
shaped version of the received mono downmix audio signal, where the shaping is done
in such a way that it approximates the signal power of the directional signal of the
signal model as previously described, i.e. where the original stereo signal is considered
to represent a main component s and residual/more diffuse components d
1, d
2. The directional signal can specifically be generated by a spectral shaping wherein
frequency band samples of the mono downmix audio signal are multiplied by subband
weights determined from the spatial parameters.
[0162] A particular approach can be determined from the signal model of:

where x can be considered to correspond to a directional/point source signal.
[0163] The audio render apparatus 103/first renderer 307 may then seek to generate an estimate
of the directional signal component x by scaling of the mono downmix audio signal
(with the scaling typically being separate for different frequency band). The estimated
direct signal component may be represented by:

where the resulting approximation
x' is the estimate of the directional signal in a given subband and g
x is determined from the spatial parameters, and specifically with

[0164] The spatial parameters are indicative of the relative properties of the channel signals
for the stereo signal and, as has been realized by the inventors, this may also provide
information on the relative levels for respectively a directional component (and for
a non-directional component). The first renderer 307 may specifically determine a
scaling factor that scales the effective gain of subbands of the mono downmix audio
signal to generate the directional signal estimate.
[0165] In many embodiments, the first renderer 307 is accordingly arranged to generate the
gains dependent on the spatial parameters. In particular, the gain for the directional
rendering may be set in dependence on a relative power level of a directional component
of the stereo signal relative to a power level of the mono downmix audio signal. The
spatial upmix parameters provide information on the relative properties of the channel
signals of the stereo signal, and specifically may provide information on both the
interchannel levels/intensity differences as well as on the interchannel correlation.
Accordingly, the spatial upmix parameters can be considered to provide information
on the directional signal component and the diffuse/residual components. Accordingly,
the gains
gx may be estimated/calculated from the received spatial parameters.
[0166] The gains may in many embodiments be determined to ensure power preservation. For
the indicated signal model of the stereo signal, the gains for an actual point source
with no other audio being present may be 1 and for a completely diffuse signal, the
gains would be 0.
[0167] The gains for estimating the directional signal may in particular in many embodiments
advantageously be determined in line with one or more of the following:

where IID is an interchannel intensity difference and ICC is an inter-channel cross-correlation,
and specifically

and

and:

[0168] Thus, in many embodiments, one or more of the gains may be determined based on received
interchannel intensity differences, and interchannel correlations (being part of the
spatial upmix parameters).
[0169] Different approaches for determining the rendering direction from the spatial upmix
parameters may be used in different embodiments. In particular, the signal model as
indicated above is based on directional component
x being at a direction γ in the stereo image of the stereo signal, henceforth also
referred to as the orientation direction. In many embodiments, the direction determining
circuit 309 may determine the orientation direction y and then determine the (desired)
rendering direction γ' from the orientation direction y. Indeed, in some embodiments
or scenarios, the rendering direction γ' may simply be set equal to the orientation
direction y.
[0170] The determination of the orientation direction y may be based on the signal model
indicated above. The spatial upmix parameters provide information on the relative
properties of the channel signals of the stereo signal and specifically they may provide
information on both the interchannel levels/intensity differences as well as on the
interchannel correlation. Accordingly, the spatial upmix parameters can be considered
to provide information on the directional signal component x and on the position of
this in the stereo image of the stereo signal, i.e. the spatial upmix parameters provide
information on the orientation direction y allowing this to be determined from the
provided parameter values.
[0171] The direction determining circuit 309 may determine the orientation direction y as
a direction to a directional signal component in a stereo image of the stereo signal
from the spatial upmix parameters, and to map this to a direction in a stereo image
of the output stereo signal. The directional signal component may be a dominant signal
component. The direction determining circuit 309 may be arranged to determine the
orientation direction y as a direction of a dominant sound source in the stereo signal
where the direction of the dominant sound source is represented by the spatial upmix
parameters.
[0172] The directional signal component may specifically be a signal component (estimated/determined)
to originate from a point source. Specifically, the direction determining circuit
309 may be arranged to determine the orientation direction y as a direction for which
a single point source audio source will result in spatial upmix parameter values matching
the spatial upmix parameters of the data signal.
[0173] In some embodiments, the direction determining circuit may be arranged to determine
the first direction in line with:

where IID is an interchannel intensity difference and ICC is an inter-channel cross-correlation,
and specifically with these given by the equations provided above in connection with
the equations for determining gains. The determination of the orientation direction
y above results from an assumption that the signals
x, nl and
nr of the signal model are mutually decorrelated, and that the power of the left and
right diffuse signals
nl and
nr are equal.
[0174] The direction determining circuit 309 may, as previously mentioned, in some embodiments
be used directly as the rendering direction y', i.e. γ = γ'. However, in many embodiments,
a mapping may be included which for at least some values of the orientation direction
y may result in a different rendering direction γ'.
[0175] Thus, in many embodiments, the direction determining circuit 309 may be arranged
to apply a mapping function to the orientation direction y to determine the rendering
direction γ'.
[0176] For example, the mapping may map the position in the stereo image of the original
stereo signal as represented by the orientation direction y to a desired position
in the stereo image of the output stereo signal as represented by the rendering direction
γ'. In many cases, where the output stereo signal is a binaural signal, the mapping
may include a consideration/ determination of a distance to the audio sources. For
example, a range of the orientation direction y in the interval of [0,180°] may be
mapped to a location between two virtual stereo speakers in the audio scene created
by the binaural rendering. Such speakers may for example be positioned at angles of
-30° and +30° relative to a center direction for the binaural signal. Thus, in such
situations, the direction determining circuit 309 may include a mapping between an
orientation direction y in the range of [0,180°] to a rendering direction γ' in the
range of [-30°,+30°].
[0177] Thus, in some embodiments, the directional component (the mono downmix audio signal)
may be rendered to a virtual angle in the range of a virtual loudspeaker angle range
generated by a binaural rendering. The rendered directional component may be combined
with a diffuse rendering of the residual signal(s).
[0178] In many embodiments, the direction determining circuit 309 may be arranged to map
an orientation direction γ representing an angle in one interval/range to a rendering
direction γ' representing an angle in a different interval/range.
[0179] The audio render apparatus 103 comprises a second processing path which generates
a second intermediate stereo signal. The mono downmix audio signal is fed to a decorrelator
315 which is arranged to apply a decorrelation to the mono downmix audio signal to
generate a decorrelated mono downmix audio signal. It will be appreciated that a large
number of different algorithms and functions for decorrelating an audio signal is
known to the skilled person, and that any suitable approach or algorithm may be used
without detracting from the invention.
[0180] The decorrelated mono downmix audio signal is fed to a signal generator 317 which
is typically also fed the mono downmix audio signal and which is arranged to generate
a first and second estimated signals based on the decorrelated mono downmix audio
signal and the mono downmix audio signal. The first and second estimated signals are
generated/estimated to, together with the directional signal, correspond to a decomposition
of the stereo signal. The first and second estimated signals may be estimated signals
representing signal components of the stereo signal that are not represented by the
directional signal. The first and second estimated signal may be estimated to represent
audio that is more diffuse such as ambient or background audio for the scene. The
directional signal and the first and second estimated signals may be generated/estimated
to represent respectively point source audio and non-point source audio/diffuse/ambient
audio.
[0181] In some cases, each of the two signals may be generated by a filtering of the decorrelated
mono downmix audio signal. The two filters may be different and generate different
estimated signals from the decorrelated mono downmix audio signal.
[0182] The signal generator 317 is coupled to a second renderer 319 which is arranged to
perform a second rendering being a rendering of the first estimated signal and the
second estimated signal to generate a second intermediate stereo signal. However,
in contrast to the first rendering process, the second rendering process is typically
a predetermined rendering which is not dependent on the spatial parameters, and which
typically is not depending on properties of the stereo signal. The second rendering
may typically be a diffuse rendering seeking to generate the second intermediate stereo
signal to provide a perception of a more diffuse and spatially less definite audio
source. The second rendering is specifically a predetermined rendering employing a
predetermined mapping of the first and second estimated signals to channel signals
of the second intermediate stereo signal.
[0183] As a specific example, the second rendering may typically generate the second intermediate
stereo signal by simply mapping the first estimated signals and the second estimated
signal to two phase-inverse signals, i.e. the second intermediate stereo signal may
be generated with the first and second estimated signals being mapped to the two stereo
channels. For example, in some embodiments, the first estimated signal may be mapped
to the right signal of the second intermediate stereo signal and the second estimated
signal may be mapped to the left signals of the second intermediate stereo signal.
[0184] The first renderer 307 and the second renderer 319 are coupled to a combiner 313
which is arranged to combine at least the first intermediate stereo signal and the
second intermediate stereo signal to generate an output stereo signal. In many embodiments,
the combiner 313 may be arranged to combine/sum the samples/values of the individual
channels of the first and second intermediate stereo signals to generate the samples/values
of the output stereo signal. In many cases, the combination may be performed by combining/summing
subband values of the intermediate stereo signals. In other embodiments, the combination
may be performed in the time domain by combining/summing time domain values of the
intermediate stereo signals.
[0185] In many embodiments, the combination of the intermediate stereo signals may be by
a (possibly weighted) combination/summation of corresponding channel signals for the
first intermediate stereo signal and the second intermediate stereo signal.
[0186] In some embodiments where binaural processing is used, each of the estimated residual
signals may be rendered from a specific position, such as each estimated residual
signal being rendered from a different virtual position, such as for example from
different virtual positions.
[0187] In some embodiments, the rendering for the first and/or second estimated signal,
may position the signal at a specific position.
[0188] In many embodiments, the rendering of the estimated signals may be performed by the
renderer 319 retrieving a set of directional transfer functions from the store 311
and rendering the estimated signals using the retrieved transfer functions.
[0189] In many embodiments, the audio render apparatus 103 may be arranged to extract a
directional transfer function for a single predetermined direction and render the
first and/or second estimated signal using this directional transfer function. Accordingly,
(each of) the estimated signals may be rendered from one predetermined direction/position,
such as a direction/position corresponding to a virtual speaker position.
[0190] An example of subband parametric rendering may e.g. result in left and right signals:

where
Gl, Gr, ϕl, ϕr form the parametric HRIRs,
f(
γ) is a mapping function converting the estimated angles (orientation direction y)
to HRIR direction angles,
βl and
βr are two pre-determined angles and
H1{.} and
H2{.} represent the processing generating the first and second estimated signals respectively,
m is the received mono downmix audio signal and
md is the directional signal generated by the directional signal generator 305, and
gs and
gn are suitable scaling factors.
[0191] In some embodiments, the renderer 319 may retrieve directional transfer functions
for a plurality of predetermined directions and it may use multiple directional transfer
functions in performing the predetermined rendering. For example, different directional
transfer functions may be used for different frequency subbands. This may provide
a more diffuse perception with the audio being generated such that it is perceived
from different directions for different subbands thereby resulting in a perception
of a more distributed and spread audio source.
[0192] In the latter case, the sets of predetermined directions for the different estimated
signals are different in order to enhance the perceived diffuseness of the non-directional
audio components.
[0193] The renderer 319 may generate the second intermediate stereo signal using a first
set of directional transfer functions retrieved from the store 311 for a first set
of predetermined directions, and may generate the third intermediate stereo signal
using a second set of directional transfer functions retrieved from the store 311
for a second set of predetermined directions where the first set of set of predetermined
directions is different from the second set of predetermined directions.
[0194] In particular, the directional transfer functions may be binaural transfer functions
and the renderer 319 may be arranged to perform binaural rendering to generate the
first intermediate stereo signal using binaural impulse response values for a first
set of predetermined directions and may be arranged to perform binaural rendering
to generate the first intermediate stereo signal using binaural impulse response values
for a second set of predetermined directions where the first set of predetermined
directions are different from the second set of predetermined directions.
[0195] In many cases, the use of multiple directional transfer functions may be achieved
by using directional transfer functions for different directions in different frequency
subbands.
[0196] Thus, instead of rendering the diffuse/ non-directional estimated signals using fixed
angles, e.g. mimicking a virtual stereo speaker setup, the diffuse signals may also
be rendered using composite, e.g. pre-calculated HRIRs for many sources/directions,
e.g. spread over a (part of a) circle, or (part of) a sphere.

where e.g.:

with
Bl being a set of angles at which the left diffuse signal is to be rendered, B
r a set of angles at which the right diffuse signal is to be rendered, and
gnorm a normalisation factor.
[0197] In some embodiments, the diffuse signal component may be directly rendered onto left
and right channels without any HRIR processing.
[0198] The audio apparatus may accordingly be arranged to generate an output stereo signal,
and often an output binaural stereo signal from the received mono downmix audio signal
and spatial upmix parameters. The audio apparatus specifically implements two different
rendering paths with one being a directional (binaural) rendering of a directional
(e.g. a dominant) signal component which is generated from the mono downmix audio
signal and the difference signal. The other rendering path may employ a predetermined
rendering/mapping of estimated signals generated from the mono downmix audio signal
and a decorrelation thereof. The rendering of the output stereo signal is typically
not a conventional adaptive upmixing of the received and decorrelated mono signals,
and is specifically not a conventional 2x2 matrix upmixing of the mono signal and
a decorrelated signal, but rather is a direct generation of a stereo signal by parallel
processing of respectively the mono downmix audio signal and estimated signals generated
from the mono downmix audio signal and decorrelated versions of the mono downmix audio
signal, with the former rendering being directional dependent on the spatial upmix
parameters and the latter rendering being a predetermined rendering.
[0199] The processing seeks to render direct/dominant/directional point source components
using a direct rendering with a direction that is given by the spatial parameters.
The rendering employs a directionally dependent transfer function to the left and
right stereo output signal for that purpose. The approach further seeks to render
a residual/remaining signal component as more diffuse audio, and specifically it may
use a predetermined rendering where a decorrelated signal is mapped directly to the
channels of the output binaural signal using a transfer function. The mapping may
be predetermined and may specifically be such that it allows a more diffuse and non-directional
perception of this signal component. The rendering process thus uses fundamentally
different approaches to provide different signal components in the output binaural
signal, but does so without specifically decomposing the mono downmix audio signal
into a dominant and diffuse/residual signal component with these subsequently being
individually rendered. Rather a direct rendering of respectively the mono downmix
audio signal and estimated signals generated using a decorrelated version of the mono
downmix audio signal is performed to generate the binaural output signal.
[0200] The approach allows low complexity and computationally efficient rendering of a stereo
signal encoded as a downmix and spatial upmix parameters, such as a PS encoded signal.
It may further allow high performance rendering with a perceived improved audio quality.
In many cases, a substantially improved spatial perception and user experience may
be achieved, and indeed, can be provided using headphones. The approach may in many
cases provide a user perception of an audio scene where individual/dominant audio
sources are well defined at specific directions/positions in a stereo image whereas
other sources (e.g. ambient or background audio sources) are perceived more diffuse.
[0201] The approach may typically allow a very efficient operation and rendering with reduced
complexity. A particular advantage of the approach is that it does not require a decomposition
of the received mono downmix audio signal into different components with different
specific properties.
[0202] The audio render apparatus 103 accordingly generates an output stereo signal which
is the combination of a directional rendering putting point source audio at a specific
and spatially definite position determined from the received spatial upmix parameters,
and of a typically predetermined rendering providing a more diffuse and decorrelated
perception of the corresponding audio source. The approach provides two parallel rendering
processes/paths for the mono downmix audio signal with the rendered results being
combined to generate the output stereo signal.
[0203] However, further, the audio render apparatus 103 of FIG. 3 is arranged to control
the generation of the signals such that relative properties between the two estimated
signals are controlled based on the spatial parameters. Thus, rather than merely generating
the two signal components representing non-point source audio (typically ambient/
diffuse sound) by two signal components that are decorrelated with respect to each
other and with the same signal level, the audio render apparatus 103 of FIG. 3 includes
an adapter 321 which is specifically arranged to dynamically control the relative
property, and specifically the correlation between and/or relative level of the estimated
signals.
[0204] The adapter 321 receives the spatial parameters and proceeds to control the signal
generator 317 to adapt the generation of the estimated signals such that they have
appropriate relative properties, and specifically such that they have appropriate
(cross)correlation and signal levels as indicated by the spatial properties. The adapter
321 and signal generator 317 accordingly operate to generate estimated signals that
have varying and signal dependent (via the spatial parameters) correlation and signal
level properties. The estimated signals may typically be residual signal estimated
to represent audio components of the stereo signal that are not represented/included
in the directional signal.
[0205] As mentioned, the directional signal circuit 305 may generate the directional signal
as

[0206] Further the signal generator 317 may generate the two estimated signals as linear
combinations of the mono downmix audio signal and the decorrelated mono downmix audio
signal, e.g.:

[0207] Accordingly the overall decomposition can be represented by:

[0208] Thus, the audio render apparatus 103 may generate the directional signal and the
estimated signals to correspond to a signal composition of the stereo signal in accordance
with the previously introduced signal model. The signals may be generated as linear
combinations of the mono downmix audio signal and the decorrelated mono downmix audio
signal.
[0209] The adapter 321 may proceed to adapt the (matrix) coefficients to seek to maintain
the properties of the underlying signal model. Specifically, by imposing that the
decorrelator signal power equals the mono downmix power, ∥
d∥
2 = ∥
m∥
2 this may allow reconstruction of a range of properties:

[0210] Accordingly, the adapter 321 may by determining suitable weights/scale factors/gains/
(matrix) coefficients proceed to generate the estimated signals such that they have
a relative signal level and a cross correlation that match those of a decomposition
of the original stereo signal according to the signal model. The estimated signals
are thus not merely generated as decorrelated signals with the same signal level but
are generated to be signals having relative/cross properties that correspond more
closely to a decomposition according to a suitable signal model for the stereo signals.
In addition, the relative properties, and specifically the relative levels and the
correlations, between the difference signal and respectively the first and second
estimated signals may be adapted to correspond to those of the underlying signal model.
[0211] The directional signal and the estimated signals are each a linear combination of
the mono downmix audio signal and the decorrelated mono downmix audio signal and the
weights of the linear combination may be determined as a function of the spatial parameters.
The adapter 321 may specifically apply functions that result in one or more of the
properties indicated above being the same for the generated signals as for a decomposition
in accordance with the signal model.
[0212] The linear combinations and the determination of suitable weights based on the spatial
parameters are typically performed in the frequency domain and may be individual and
separate for each frequency subband/ time frequency tile.
[0213] In some embodiments, as in particular in the example above, no other signal than
the mono downmix audio signal is included in the generation of the directional signal.
Similarly, no other signals than the mono downmix audio signal and the decorrelated
mono downmix audio signal are included in the generation of the estimated signals.
[0214] It will be appreciated that different detailed approaches can be used to determine
suitable weights and coefficients from the spatial parameters, and that different
approaches may be used in different embodiments.
[0215] For example, using the overall decomposition matrix:

[0216] If the parameter
gm,l is chosen to be set to 0, the correlation between the generated signal pair (
s',
d'l) becomes 0, under the condition that the correlation between the mono signal
m and the decorrelated signal
d was 0. The same holds for the for the parameter
gm,r and the generated signal pair (
s',
d'r). Similarly, if the parameter
gd,l is chosen to be set to 0, the (absolute) correlation between the generated signal
pair (
s',
d'l) becomes 1. Likewise, if the parameter
gd,r is chosen to be set to 0, the (absolute) correlation between the generated signal
pair (
s',
d'r) becomes 1. By balancing
gm,l and
gd,l an arbitrary correlation between the signal pair (
s',
d'l) can be generated. By balancing
gm,r and
gd,r an arbitrary correlation between the signal pair
(s', d'r) can be generated. The relative levels of the signal pair (
s',
d'l) can be realized by weighting
gs to the combined weights of
gm,l and
gd,l (typically as

). For example, if the weights of
gm,l and
gd,l are relatively close to 0, compared to the weight
gs a large relative level is realized. If the weight of
gs is relatively close to 0, compared to the weights of
gm,l and
gd,l, a small relative level is realized. Similarly, the relative levels of the signal
(
s', d'r) can be realized by weighting
gs to the combined weights of
gm,r and
gd,r (typically as

).
[0217] It can be shown that for two signals
a,
b that are a linear combination of the left and right signals according to:

that:

[0218] This means that if there are given relationships between the coefficients
a1,
a2,
b1,
b2 and the PS parameters, i.e.,
a1 = f1(
IID, ICC, IPD), a2 =
f2(
IID, ICC, IPD), b1 = f3 (IID, ICC, IPD), b2 = f4(
IID, ICC, IPD), it is possible to calculate the cross-correlation between the resulting two signals
a and
b.
[0219] Similarly, it can be shown that the intensity ratio can be determined as:

[0220] This approach may be generalized as following. Suppose a set of
n signals
x1, ... ,
xn are generated as a linear combination of a left and right signal using a matrix
G, where each element is a function of the PS parameters (
gi,j = fi,j (
IID, ICC, IPD):

then, the resulting normalized covariance matrix between each pair of the linearly
combined signals
xi,
xj is a function of the PS parameters and the underlying matrix functions
gi,j:

[0221] In many embodiments, the adapter 321 may accordingly adapt the generation of the
estimated signals to have desired relative properties, and specifically to have a
desired correlation and signal level properties.
[0222] In many embodiments, the adapter 321 may further be arranged to adapt a relative
property between the directional signal and the first and/or second estimated signal
in dependence on the spatial parameters. The relative property may specifically be
the level difference, and specifically the IID, between the directional signal and
one of the estimated signals, and/or a correlation, and specifically an ICC, between
the directional signal and one of the estimated signals.
[0223] For example, using the overall decomposition matrix:

the IID between the directional signal and the left and right estimated signals are
determined as:

[0224] If we want to reinstate these IID values, without affecting the correlation, we can
set
gm,l =
gm,r = 0.
gs is equal to
9x as defined before. Then the desired
IIDs,dl and
IIDs,dr, which can be calculated from the linear equations of the decomposition model, can
be realized by setting the values of
gd,l and
gd,r accordingly.
[0225] The ICC between the direct and the left estimated signal using the overall decomposition
matrix is given by:

where
gs is again equal to
gx defined before, resulting in two equations, one for the
IIDs,dl and one for
ICCs,dl and two unknowns,
gm,l and
gd,l.
[0226] As a particular example, the processing may be based on a Principal Component Analysis
(PCA) and the directional signal s may be considered as the principal component determined
by a PCA process.
[0227] The general decomposition of a stereo signal into a directional signal s, and a left
and right residual signal d
l and d
r, can be determined as:

where the matrix
H is a 3x2 matrix, where each element is a function of the PS parameters:

[0228] An example of such matrix is given by a PCA decomposition, for which:

where:

and:

[0230] It is noted that
IIDdl,dr follows from the other IID values.
[0231] For the PCA decomposition it can be shown that:

[0232] By filling in the expressions for
wl and
wr as provided above, it can be shown that:

which is well-known property of the PCA decomposition. It realizes an orthogonalization
of the resulting signals.
[0233] In a similar way, it can be shown that

[0234] Finally, using the same method as described above, it is possible to show that:

[0235] Using the overall decomposition matrix, it becomes obvious that
gm,l and
gm,r can be set to 0, as no correlation is required between the directional signal and
one of the estimated signals:

[0237] It is noted that in the matrix
G, the phase angle IPD can be spread differently.
[0238] Thus, in the specific case where the decomposition is based on a PCA approach, the
directional signal can be generated by scaling the mono downmix audio signal by suitable
frequency subband weights and the first and second estimated signals can be generated
by scaling the decorrelated mono downmix audio signal by suitable frequency subband
weights that are different for the two estimated signals.
[0239] In many embodiments, the signal generator 317 may be arranged to generate the first
estimated signal by applying frequency band weights to frequency band samples of the
decorrelated mono downmix audio signal where the frequency band weights are dependent
on the spatial parameters. Specifically, it may be generated as/from:

[0240] Similarly, the signal generator 317 may be arranged to generate the second estimated
signal by applying frequency band weights to frequency band samples of the decorrelated
mono downmix audio signal where the frequency band weights being dependent on the
spatial parameters. Specifically, it may be generated as/from:

[0241] The processing may be performed in subbands and may be performed in time segments.
The processing in each subband may for some (any) or all steps be performed separately/independently
in each subband (with respect to the processing in other subbands). The processing
in each time segment may for some (any) or all steps be performed separately/independently
in each time segment (with respect to the processing in other time segments).
[0242] The processing may be time interval/segment based with all processing being performed
for each time segment. Equivalently, the signal(s) for each segment may be considered
a signal (and in particular signals of different time segments, may be considered
different signals).
[0243] The audio apparatus(s) may specifically be implemented in one or more suitably programmed
processors. An example of a suitable processor is provided in the following.
[0244] FIG. 4 is a block diagram illustrating an example processor 400 according to embodiments
of the disclosure. Processor 400 may be used to implement one or more processors implementing
an apparatus as previously described or elements thereof (including in particular
one more artificial neural network). Processor 400 may be any suitable processor type
including, but not limited to, a microprocessor, a microcontroller, a Digital Signal
Processor (DSP), a Field ProGrammable Array (FPGA) where the FPGA has been programmed
to form a processor, a Graphical Processing Unit (GPU), an Application Specific Integrated
Circuit (ASIC) where the ASIC has been designed to form a processor, or a combination
thereof.
[0245] The processor 400 may include one or more cores 402. The core 402 may include one
or more Arithmetic Logic Units (ALU) 404. In some embodiments, the core 402 may include
a Floating Point Logic Unit (FPLU) 406 and/or a Digital Signal Processing Unit (DSPU)
408 in addition to or instead of the ALU 404.
[0246] The processor 400 may include one or more registers 412 communicatively coupled to
the core 402. The registers 412 may be implemented using dedicated logic gate circuits
(e.g., flip-flops) and/or any memory technology. In some embodiments the registers
412 may be implemented using static memory. The register may provide data, instructions
and addresses to the core 402.
[0247] In some embodiments, processor 400 may include one or more levels of cache memory
410 communicatively coupled to the core 402. The cache memory 410 may provide computer-readable
instructions to the core 402 for execution. The cache memory 410 may provide data
for processing by the core 402. In some embodiments, the computer-readable instructions
may have been provided to the cache memory 410 by a local memory, for example, local
memory attached to the external bus 416. The cache memory 410 may be implemented with
any suitable cache memory type, for example, Metal-Oxide Semiconductor (MOS) memory
such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), and/or
any other suitable memory technology.
[0248] The processor 400 may include a controller 414, which may control input to the processor
400 from other processors and/or components included in a system and/or outputs from
the processor 400 to other processors and/or components included in the system. Controller
414 may control the data paths in the ALU 404, FPLU 406 and/or DSPU 408. Controller
414 may be implemented as one or more state machines, data paths and/or dedicated
control logic. The gates of controller 414 may be implemented as standalone gates,
FPGA, ASIC or any other suitable technology.
[0249] The registers 412 and the cache 410 may communicate with controller 414 and core
402 via internal connections 420A, 420B, 420C and 420D. Internal connections may be
implemented as a bus, multiplexer, crossbar switch, and/or any other suitable connection
technology.
[0250] Inputs and outputs for the processor 400 may be provided via a bus 416, which may
include one or more conductive lines. The bus 416 may be communicatively coupled to
one or more components of processor 400, for example the controller 414, cache 410,
and/or register 412. The bus 416 may be coupled to one or more components of the system.
[0251] The bus 416 may be coupled to one or more external memories. The external memories
may include Read Only Memory (ROM) 432. ROM 432 may be a masked ROM, Electronically
Programmable Read Only Memory (EPROM) or any other suitable technology. The external
memory may include Random Access Memory (RAM) 433. RAM 433 may be a static RAM, battery
backed up static RAM, Dynamic RAM (DRAM) or any other suitable technology. The external
memory may include Electrically Erasable Programmable Read Only Memory (EEPROM) 435.
The external memory may include Flash memory 434. The External memory may include
a magnetic storage device such as disc 436. In some embodiments, the external memories
may be included in a system.
[0252] The invention can be implemented in any suitable form including hardware, software,
firmware, or any combination of these. The invention may optionally be implemented
at least partly as computer software running on one or more data processors and/or
digital signal processors. The elements and components of an embodiment of the invention
may be physically, functionally and logically implemented in any suitable way. Indeed,
the functionality may be implemented in a single unit, in a plurality of units or
as part of other functional units. As such, the invention may be implemented in a
single unit or may be physically and functionally distributed between different units,
circuits and processors.
[0253] Although the present invention has been described in connection with some embodiments,
it is not intended to be limited to the specific form set forth herein. Rather, the
scope of the present invention is limited only by the accompanying claims. Additionally,
although a feature may appear to be described in connection with particular embodiments,
one skilled in the art would recognize that various features of the described embodiments
may be combined in accordance with the invention. In the claims, the term comprising
does not exclude the presence of other elements or steps.
[0254] Furthermore, although individually listed, a plurality of means, elements, circuits
or method steps may be implemented by e.g. a single circuit, unit or processor. Additionally,
although individual features may be included in different claims, these may possibly
be advantageously combined, and the inclusion in different claims does not imply that
a combination of features is not feasible and/or advantageous. Also, the inclusion
of a feature in one category of claims does not imply a limitation to this category
but rather indicates that the feature is equally applicable to other claim categories
as appropriate. Furthermore, the order of features in the claims do not imply any
specific order in which the features must be worked and in particular the order of
individual steps in a method claim does not imply that the steps must be performed
in this order. Rather, the steps may be performed in any suitable order. In addition,
singular references do not exclude a plurality. Thus references to "a", "an", "first",
"second" etc. do not preclude a plurality. Reference signs in the claims are provided
merely as a clarifying example shall not be construed as limiting the scope of the
claims in any way.