<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE ep-patent-document PUBLIC "-//EPO//EP PATENT DOCUMENT 1.4//EN" "ep-patent-document-v1-4.dtd">
<ep-patent-document id="EP10156530A1" file="EP10156530NWA1.xml" lang="en" country="EP" doc-number="2372707" kind="A1" date-publ="20111005" status="n" dtd-version="ep-patent-document-v1-4">
<SDOBI lang="en"><B000><eptags><B001EP>ATBECHDEDKESFRGBGRITLILUNLSEMCPTIESILTLVFIROMKCYALTRBGCZEEHUPLSKBAHRIS..MTNORSMESM..................</B001EP><B005EP>J</B005EP><B007EP>DIM360 Ver 2.15 (14 Jul 2008) -  1100000/0</B007EP></eptags></B000><B100><B110>2372707</B110><B120><B121>EUROPEAN PATENT APPLICATION</B121></B120><B130>A1</B130><B140><date>20111005</date></B140><B190>EP</B190></B100><B200><B210>10156530.7</B210><B220><date>20100315</date></B220><B250>en</B250><B251EP>en</B251EP><B260>en</B260></B200><B400><B405><date>20111005</date><bnum>201140</bnum></B405><B430><date>20111005</date><bnum>201140</bnum></B430></B400><B500><B510EP><classification-ipcr sequence="1"><text>G10L  21/02        20060101AFI20100730BHEP        </text></classification-ipcr></B510EP><B540><B541>de</B541><B542>Adaptive Spektralumwandlung für akustische Sprechsignale</B542><B541>en</B541><B542>Adaptive spectral transformation for acoustic speech signals</B542><B541>fr</B541><B542>Transformation spectrale adaptative pour signaux vocaux acoustiques</B542></B540><B590><B598>1</B598></B590></B500><B700><B710><B711><snm>Svox AG</snm><iid>100767722</iid><irf>WP-2259-EP</irf><adr><str>Baslerstrasse 30</str><city>8048 Zürich</city><ctry>CH</ctry></adr></B711></B710><B720><B721><snm>Withopf, Jochen</snm><adr><str> SVOX Deutschland GmbH
Götzenberg 1</str><city>97941 Tauberbischofsheim</city><ctry>DE</ctry></adr></B721><B721><snm>Hannon, Patrick</snm><adr><str> SVOX Deutschland GmbH
Wagnerstr. 60</str><city>89077 Ulm</city><ctry>DE</ctry></adr></B721><B721><snm>Krini, Mohamed</snm><adr><str>Elchinger Weg 10</str><city>89075 Ulm</city><ctry>DE</ctry></adr></B721><B721><snm>Schmidt, Gerhard Uwe</snm><adr><str> SVOX Deutschland GmbH
Wiesenhof 4</str><city>24248 Mönkeberg</city><ctry>DE</ctry></adr></B721></B720><B740><B741><snm>Stocker, Kurt</snm><iid>100041439</iid><adr><str>Büchel, von Révy &amp; Partner 
Zedernpark 
Bronschhoferstrasse 31</str><city>9500 Wil</city><ctry>CH</ctry></adr></B741></B740></B700><B800><B840><ctry>AT</ctry><ctry>BE</ctry><ctry>BG</ctry><ctry>CH</ctry><ctry>CY</ctry><ctry>CZ</ctry><ctry>DE</ctry><ctry>DK</ctry><ctry>EE</ctry><ctry>ES</ctry><ctry>FI</ctry><ctry>FR</ctry><ctry>GB</ctry><ctry>GR</ctry><ctry>HR</ctry><ctry>HU</ctry><ctry>IE</ctry><ctry>IS</ctry><ctry>IT</ctry><ctry>LI</ctry><ctry>LT</ctry><ctry>LU</ctry><ctry>LV</ctry><ctry>MC</ctry><ctry>MK</ctry><ctry>MT</ctry><ctry>NL</ctry><ctry>NO</ctry><ctry>PL</ctry><ctry>PT</ctry><ctry>RO</ctry><ctry>SE</ctry><ctry>SI</ctry><ctry>SK</ctry><ctry>SM</ctry><ctry>TR</ctry></B840><B844EP><B845EP><ctry>AL</ctry></B845EP><B845EP><ctry>BA</ctry></B845EP><B845EP><ctry>ME</ctry></B845EP><B845EP><ctry>RS</ctry></B845EP></B844EP></B800></SDOBI>
<abstract id="abst" lang="en">
<p id="pa01" num="0001">The present invention relates to a method for adaptive transformation of the frequency spectrum of a windowed speech signal. According to the present invention, frequency compression is achieved by applying frequency compression functions, which are dependent upon characteristics of the current frame. Formant boosting takes place using formant-dependent functions to increase the contrast between the formants and the non-formant frequencies in the frequency spectrum of the current frame. Generally, frequencies below a defined threshold are compressed linearly, corresponding to no compression when the slope is equal to 1, while the compression rate of higher frequencies varies over time and frequency depending on the current frame features. Thus features of the input signal located above the frequency threshold are moved into the low-pass range and are audible in the transmitted output signal. The feature based processing is designed to enhance the understandability of speech in the transmitted bandwidth and to increase recognition rates in an automatic speech recognition module.
<img id="iaf01" file="imgaf001.tif" wi="99" he="82" img-content="drawing" img-format="tif"/></p>
</abstract><!-- EPO <DP n="1"> -->
<description id="desc" lang="en">
<heading id="h0001"><b><u>Technical Field</u></b></heading>
<p id="p0001" num="0001">The present invention generally relates to speech synthesis technology.</p>
<heading id="h0002"><b><u>Background of the invention</u></b></heading>
<p id="p0002" num="0002">In most telecommunication systems, speech signals are not transmitted with their full analog bandwidth in order to use the available transmission channel more efficiently. For telephone communication, the signal bandwidth is typically limited to less than 4 kHz, even though there are signal components up to 8 kHz and higher in the original speech signal. This band limitation has little or no effect on the intelligibility of voiced sounds, but fricatives such as /s/, /sh/, /ch/, /z/ or /f/ may be lost completely. In <figref idref="f0001">Fig. 1</figref> a spectrogram of the utterance "Sauerkraut is served" is depicted. The phoneme /s/ is spoken at approximately 0, 0.8, and 1.4 seconds. At these times almost 100% of the signal energy is above 4 kHz. With a bandwidth limited to less than 4 kHz the speech intelligibility may still be sufficient because the listener is often able to predict missing phonemes from syntax and the context of what is being said. However, errors arise if such a prediction is not possible, e.g., because names or unknown words of a foreign language are transmitted. Furthermore, the phoneme /s/ is important in the English language to indicate the plural and possessive pronouns.</p>
<p id="p0003" num="0003">Various spectral compression schemes exist that aim to present signals from one frequency range in another more useful range of the frequency spectrum. The side-effects of the schemes vary according to the limitations of the candidate identification techniques, artifacts resulting from the signal processing, as well as high computational complexity and delay in the signal analysis stage.</p>
<p id="p0004" num="0004"><nplcit id="ncit0001" npl-type="s"><text>P. Patrick, R. Steele, and C. Xydeas, Frequency compression of 7.6 kHz speech into 3.3 kHz bandwidth, 31 (5):692-701, May 1983</text></nplcit> describes a system for retaining information under frequency compression. This system consists of frequency mapping at the transmitter and demapping at the receiver side. According to a frequency compression factor c, every cth sample is retained in the compressed magnitude spectrum, so that it occupies only the frequency range that is available for transmission. At the receiver side, the received frequency components are spaced out to their correct locations and the magnitude of the missing components is found by linear interpolation. The phase is chosen randomly.</p>
<p id="p0005" num="0005">The compression factor c can also be set frequency dependently to give higher or lower compression in certain regions. Three different mapping laws, which are tailored for certain<!-- EPO <DP n="2"> --> phonemes, and one that simply corresponds to a band limitation, are used in the system. For switching between these laws, the signal is first classified as voiced or unvoiced speech based on its autocorrelation. Mapping is only applied for unvoiced speech. Then the mapping and demapping procedure is applied according to the same laws.</p>
<p id="p0006" num="0006">It is important to note that this system needs to transmit side information about the mapping laws that have been used to ensure correct demapping. Additionally, the receiver must be able to handle this side information and to perform the demapping.</p>
<p id="p0007" num="0007"><nplcit id="ncit0002" npl-type="s"><text>D. A. Heide and G. S. Kang, Speech enhancement for bandlimited speech. In Proc. IEEE International Conference on Acoustics, Speech and Signal Processing, volume 1, pages 393-396, May 12-15, 1998</text></nplcit> describes a method that works without transmitting any side information. A classification of the speech sound is based on the spectral centroid of the spectrum, <i>f<sub>c</sub></i>. For voiced speech, <i>f<sub>c</sub></i> ≈ 2 kHz usually holds.</p>
<p id="p0008" num="0008">When the spectral centroid is greater, portions of the spectrum above 4 kHz are added into the 3 - 4 kHz bands according to the following rules:
<ul id="ul0001" list-style="none">
<li>fc = 2 kHz: No spectral translation is performed.</li>
<li>fc &gt; 3 kHz: The 4 - 5 kHz band is added into the 3 - 4 kHz band.</li>
<li>fc &gt; 4 kHz: The 5 - 6 kHz band is also added into the 3 - 4 kHz band.</li>
<li>fc &gt; 5 kHz: The 6 - 7 kHz band is also added into the 3 - 4 kHz band.</li>
</ul></p>
<p id="p0009" num="0009">Uncontrolled artifacts are resulting from this signal processing. There is no general improvement of the speech intelligibility.</p>
<p id="p0010" num="0010">Another application area of frequency mapping methods is hearing aids. People with hearing impairment usually suffer from a bad perception of high frequency sounds. The traditional approach is to strongly amplify these critical frequency regions. However, for some people, hearing sensitivity is so poor at high frequencies that sufficient gain for achieving audibility cannot be provided.</p>
<p id="p0011" num="0011"><nplcit id="ncit0003" npl-type="s"><text>Andrea Simpson, Adam A. Hersbach, and Hugh J. McDermott, Improvements in speech perception with an experimental nonlinear frequency compression hearing device, International Journal of Audiology, 44:5:281-292, 2006</text></nplcit> discloses further method of frequency mapping. The methods don't improve sound quality but enable the hearing-impaired to receive at least some information contained in the high frequency speech components.<!-- EPO <DP n="3"> --></p>
<p id="p0012" num="0012">Some further approaches are:
<ul id="ul0002" list-style="bullet" compact="compact">
<li>to shift the complete spectrum to lower frequencies if high frequency sounds are detected (<nplcit id="ncit0004" npl-type="s"><text>H. J. McDermott, V. P. Dorkos, M. R. Dean, and T. Y. Ching. Improvements in speech perception with use of the avr transonic frequency-transposing hearing aid. J Speech Hear Lang Res, 42(6):1323-1335, 1999</text></nplcit>). Problems with this scheme are the reliable detection of high frequency signals under noisy conditions and artifacts that occur during the on/off switching.</li>
<li>to detect a range of high energy and to shift that region downwards (<nplcit id="ncit0005" npl-type="s"><text>Kuk, Korhonen, Peeters, Jessen, and Andersen. Linear frequency transposition: Extending the audibility of high frequency information. The Hearing Review, 2006</text></nplcit>). Shifting means here, to overlap the identified frequency interval with a lower band and to add the two spectra.</li>
</ul></p>
<p id="p0013" num="0013">This can lead to blurring of vowel sounds.
<ul id="ul0003" list-style="bullet" compact="compact">
<li>to compress frequencies above a threshold frequency (described above). The compression is always switched on, regardless of the current speech sound.</li>
</ul></p>
<p id="p0014" num="0014"><patcit id="pcit0001" dnum="US5771299A"><text>US 5 771 299</text></patcit> aims to compress or expand the spectral envelope of an input signal. The goal is to move signal information from a region of the spectrum that is outside of the audible range of hearing aid users into another range that is still audible for the user. The system uses an LPC analysis filter and transmits the filter coefficients to the synthesis filter directly. Then, with the help of all-pass filters non-integer delays are introduced in the analysis and/or the synthesis filters. Delays larger than 1 compress the spectral envelope while delays smaller than 1 expand the spectral envelope.</p>
<p id="p0015" num="0015">The main problem with this system is that the voice specific characteristics such as formant positions are not preserved, since they are also compressed/expanded as part of the signal processing. This has the effect of enhancing audibility for the hearing impaired, but not of preserving the audio quality of the original input speech signal. Further disadvantages are delays of the output signal.</p>
<p id="p0016" num="0016"><patcit id="pcit0002" dnum="US20090074197A"><text>US 2009 0074197</text></patcit> discloses a similar transposition but with user-dependent information for spatial hearing. The goal is to perform a frequency transposition that moves frequency regions of the input signal to user specific ranges that are measured using built-in mechanisms of the hearing aid.</p>
<p id="p0017" num="0017">A method according to <patcit id="pcit0003" dnum="EP1333700A2"><text>EP 1 333 700 A2</text></patcit> applies a perception-based nonlinear transposition function to the input spectrum. The same transposition function is applied over the whole signal so that artifacts resulting from switching between non-transposition and transposition processing are avoided. However, the use of a single function over the entire speech signal<!-- EPO <DP n="4"> --> ignores time-varying phoneme-specific characteristics in the signal. Also, transforming the input signal to and then from the perception based scale to apply the transposition function reduces spectral resolution. As a result, characteristics of parts of the spectrum, which would otherwise only be subjected to a linear section of the transposition function, are not accurately preserved.</p>
<p id="p0018" num="0018"><patcit id="pcit0004" dnum="US20060226016A"><text>US 2006 0226016</text></patcit> aims to correct the phase of a transpositioned spectrum, wherein the transposition system is always active. Such processing of voiced segments has a negative effect on the phase and harmonic characteristics of the signal.</p>
<p id="p0019" num="0019">According to <patcit id="pcit0005" dnum="US20090226016A"><text>US 2009 0226016</text></patcit> the input signal is subjected to a high pass (or band pass) filter, after which the spectral envelope of the high pass signal is estimated using an all-pole model. A warping function is then applied to the all-pole model, translating the poles to lower frequencies using both non-linear and linear warping factors. An excitation signal (the prediction error signal) is then applied to the newly transposed all-pole model and shaped according to this transposed spectral envelope to create a synthesized transposed signal. The synthesized signal with an optional amplification and the original low-pass signal are then summed to create a signal containing compressed and transposed segments of the original spectrum. The LPC analysis and voice synthesis step is costly and the quality of such signals usually sounds quite artificial without proper postprocessing. This postprocessing is also expensive and this method is therefore not suitable for a system focusing on improved audio quality compared to systems for the hearing-impaired.</p>
<p id="p0020" num="0020"><patcit id="pcit0006" dnum="US6577739B"><text>US 6 577 739</text></patcit> describes a proportional compression factor to the frequency spectrum generated by an FFT of the input signal. The factors are set to be between 0.5 and 0.99, and are applied using different methods. Using linear interpolation and assuming a compression factor of 0.5, the contribution of two input FFT bins are calculated and applied to one output IFFT pin. Another method would use an IFFT length double that of the input FFT length. By additionally placing the contributions from the input FFT bins in a higher or lower position in the output IFFT vector, a frequency transposition can be applied as well. This method again places a compressed frequency range of an input signal in another range in the output signal, which is still available to the communication channel or the end user of a hearing aid device. This system is however quite limited in scope by requiring FFT processing of the input and output signals. The additional time domain trimming of signals required by using FFT and IFFT vectors of different lengths also presents problems for time-domain synchronization and variability of the compression factors is not easily implemented.</p>
<p id="p0021" num="0021"><patcit id="pcit0007" dnum="CA2569221"><text>CA 2 569 221</text></patcit> describes an attempt to retain information in frequency ranges outside of a band pass threshold. The input signal is converted to the frequency domain via the FFT<!-- EPO <DP n="5"> --> algorithm or a polyphase filterbank. A compression function is then applied. Additionally, an amplification is applied in accordance to the energy level of the uncompressed portion of the input signal. The system can also comprise sending the compressed signal to an automatic speech recognition module. This method can have negative compression effects by giving portions of the input spectrum either too much or too little emphasis.</p>
<p id="p0022" num="0022"><patcit id="pcit0008" dnum="WO20080123886A"><text>WO 2008/0123886</text></patcit> first defines a pass band for the output signal, and then defines a threshold, generally lower than the pass band, above which the frequency compression is applied. In the frequency region above the threshold, the highest frequency still of interest is identified and the appropriate compression is then applied. After the compression, the new peak power is normalized in a manner proportional to the compression, or is simply halved by -3 dB. Additionally, this method provides for the ability to expand a received signal, compressed or otherwise and synthetically reconstruct high frequency portions of the signal that may or may not have been present in the original transmitted signal. This can have a negative effect on the audio quality of the signal. This method still neglects the time-varying speaker dependent characteristics of speech sounds and applies the same compression function to each frame.</p>
<heading id="h0003"><b><u>Summary of the Invention</u></b></heading>
<p id="p0023" num="0023">In view of the foregoing, the need exists for improved solutions that preserve and enhance voice specific characteristics. The object of the present invention is to improve at least one out of controllability, precision, signal quality, processing load, and computational complexity.</p>
<p id="p0024" num="0024">The inventive method for adaptive spectral transformation for acoustic speech signals comprises the steps of<br/>
receiving at least one spectral input representation corresponding to at least one window of a time domain input signal of acoustic speech,<br/>
selecting of the spectral input representations at least one selected spectral representation to be transformed,<br/>
assigning the at least one selected spectral representation to one of a set of cluster centres, wherein<br/>
the cluster centres are defined on the bases of spectral representations of windowed acoustic speech segments of a speech corpus by a clustering algorithm,<br/>
spectral class representations are assigned to the cluster centres and are elements of a code book and<br/>
<!-- EPO <DP n="6"> -->the code book links to each spectral class representation at least one spectral transformation which enhances the corresponding spectral class representation,<br/>
transforming each selected spectral representation to a spectral output representation, wherein the applied transformation corresponds to the at least one spectral transformation linked to the cluster centre which is assigned to the respective selected spectral representation, and<br/>
providing the spectral output representations to synthesize an acoustic speech signal.</p>
<p id="p0025" num="0025">Selected spectral representations are classified by assigning them to spectral class representation. The spectral transformation applied to a selected spectral representation is adapted to this selected spectral representation because the applied spectral transformation enhances the assigned spectral class representation. The adaptations to enhance the spectral class representations are made in a setup procedure and include heuristic steps in order to find transformations which enhance the spectral class representations. The adapted advantageous transformations for each cluster centre respectively for each class of the code book ensure enhancement of the spectral representation of all the spectral representations assigned to this class of the code book. The controllability, precision and signal quality are enhanced while the processing load and computational complexity are reduced.</p>
<p id="p0026" num="0026">Best results can easily be found by the use of an appropriate code book, respectively by selecting an appropriate speech corpus and an appropriate clustering algorithm. Clustering of spectral representation with sufficient energy in the 4 to 8 kHz band allows specific enhancement of fricatives such as /s/, /sh/, /ch/, /z/ or /f/. The step of finding transformations which enhance fricatives of different classes allows linking of very specific transformations to the classes of the code book.</p>
<p id="p0027" num="0027">An English speech corpus can for example be taken from the TIMIT acoustic-phonetic continuous speech data base. This speech data base has been designed by the Massachusetts Institute of Technology (MIT), SRI International (SRI) and Texas Instruments, Inc. (TI) for acoustic-phonetic studies. It contains utterances of 630 male and female speakers in American English, divided into eight dialect regions. From each person, ten phonetically rich sentences have been recorded and tagged with phonetic information. The training data can be reduced to 70 speakers from all dialect regions and pauses can been removed with the help of phoneme tags. The reduced training data has a duration of approximately 23 minutes. For validation purposes, a second data set can been extracted that consists for example of 10 minutes speech data from 30 speakers. The sampling frequency is preferably fs = 16 kHz.<!-- EPO <DP n="7"> --></p>
<p id="p0028" num="0028">There are two different kinds of preferred transformations, frequency compression and formant boosting. Frequency compression is compressing the bandwidth for example from a bandwidth of 0 to 8 kHz to an bandwidth of 0 to 4 kHz preferably with a compression only at the upper or lower end of the bandwidth, optionally linear at least in the middle frequency range, corresponding to no compression when the slope is equal to 1. Formant boosting takes place using formant-dependent functions to increase the contrast between the formants and the non-formant frequencies in the frequency spectrum. The formant boosting gain function linked to each spectral class representation of the code book is amplifying at least one preferably two or three of the formants at low frequencies.</p>
<p id="p0029" num="0029">Assigning the at least one selected spectral representation to one of a set of cluster centres includes<br/>
calculating distance measures between the selected spectral representation and all the<br/>
spectral class representations of the code book, and<br/>
assigning the at least one selected spectral representation to the cluster centre with the shortest distance measures between the selected spectral representation and the spectral class representations of the cluster centre.</p>
<p id="p0030" num="0030">Calculating distance measures becomes very simple by first calculating feature vectors to the spectral representations. The distance measures are just distances between the feature vectors. The feature vectors are calculated from spectral envelope representations by a filtering transformation, preferably with a mel-filterbank, wherein the mel-filterbank optionally uses overlapping triangular windows with widths variable with frequency. With such feature vectors it can be sufficient to use a code book including at least eight (for example for fricative enhancement), preferably thirty-two, optionally 128 cluster centres. With higher class numbers more different enhancement problems can be solved each in a different way by the linked at least one spectral transformations.</p>
<p id="p0031" num="0031">The spectral class representations for the cluster centres are preferably averaged spectral representations averaged over spectral representations of corresponding cluster elements. A reduction to spectral class representations of special interest can be made by applying the clustering algorithm on the bases of a preselected sub-corpus of the speech corpus. The sub-corpus can for example be reduced to spectral representations of the speech corpus which have a spectral centroid lying above a given threshold frequency, preferably of 3 kHz.</p>
<p id="p0032" num="0032">Transformations are only needed for spectral representation which can be improved. A selection can be made by calculating the spectral centroid of each spectral input<!-- EPO <DP n="8"> --> representation and selecting spectral input representations which have a spectral centroid lying above a threshold, frequency preferably above 3 kHz. Such a selection fits to a code book base on a sub-corpus of the speech corpus with the same threshold frequency. A selection can also be made by detecting at least one of speech activity and background noise level. Spectral input representations with speech to be transformed will be selected with appropriate selection criteria.</p>
<p id="p0033" num="0033">The method for adaptive spectral transformation for acoustic speech signals can be implemented in different fields. In some applications the acoustic speech signal is windowed and of each window a spectral representation is deduced. There are also applications where the spectral representations of windows of an acoustic speech signal are provided by a system. Therefore the method of this invention starts with the step of receiving spectral input representations, which can be provided by the same system or by another system. At least some of the received spectral input representations are selected and enhanced by adapted transformations and the enhanced spectral representations are provided in the form of at least one spectral output representation. The combination of untransformed and transformed spectral representations allow synthesizing an enhanced acoustic speech signal.</p>
<p id="p0034" num="0034">The invention can be implemented in a computer program comprising program code means for performing all the steps of the disclosed methods when said program is run on a computer.</p>
<p id="p0035" num="0035">The described solutions preserve and enhance voice specific characteristics such as formants and their respective positions. This has the effect of enhancing audibility, while preserving the audio quality of the original input speech signal. The current invention also avoids unnecessary delays of the output signal. Various compression functions are applied to the speech signal, which takes into account the time-varying phoneme-specific characteristics in the signal and allows for spectral sharpening through adaptive formant boosting. Also, unnecessary transformations of the input signal, which could reduce spectral resolution are avoided. The effect of phase distortion on the harmonic section of the input signal can be avoided in the current invention by limiting the compression processing to only unvoiced segments of speech.</p>
<p id="p0036" num="0036">A focus on efficient algorithms and output signal synthesis allows the current invention to also avoid costly postprocessing to achieve high quality audio output. Additionally, transforming the input signal from the time domain to the frequency domain is not bound to the FFT (Fast Fourier Transform) algorithm. Compression functions are intended to be designed such that they are able to be efficiently stored in memory and do not inadvertently give portions of the input spectrum undesired emphasis.<!-- EPO <DP n="9"> --></p>
<heading id="h0004"><b><u>Brief description of the figures</u></b></heading>
<p id="p0037" num="0037">
<ul id="ul0004" list-style="none">
<li><figref idref="f0001">Fig. 1</figref> spectrogram of the utterance "Sauerkraut is served once a week". The plots below show the percentage of signal energy above 4 kHz and the amplitude of the signal</li>
<li><figref idref="f0001">Fig. 2</figref> example of the LBG-algorithm for clusters of two dimensional feature vectors with (c) k = <i>K</i> = 8 classes, x- and y-axis correspond to the value ranges of the first and the second features</li>
<li><figref idref="f0002">Fig. 3</figref> results of clustering with K = 8 in classes after pre-classification with a threshold of <i>fT</i> = 3kHz. The line with more details shows the average spectrum in dB, the line with less details shows a filtered spectral representation of the code book entry and the vertical line the spectral centroid fc</li>
<li><figref idref="f0003">Fig. 4</figref> a block diagram of the adaptive transformation method</li>
<li><figref idref="f0003">Fig. 5 and 6</figref>, compression functions</li>
<li><figref idref="f0004">Fig. 7</figref> a formant boosting gain function and spectral representations (cluster mean and filtered cluster mean), where the line with more details shows the average spectrum in dB, the line with less details shows a filtered spectral representation of the code book entry</li>
<li><figref idref="f0004">Fig. 8</figref> a block diagram of the adaptive transformation method using feature vectors</li>
<li><figref idref="f0004">Fig. 9</figref> examples for processing of phonemes: (a) spectral compression with compression characteristic below, (b) formant boosting with gain <i>g</i>(Ω<i><sub>µ</sub></i>) between + -10dB. The processed signals are band limited to 4 kHz (shorter curve)</li>
<li><figref idref="f0005">Fig. 10</figref> example of the sauerkraut utterance processed with the 128 class scheme (input signal above, processed signal below)</li>
<li><figref idref="f0005">Fig. 11</figref> time series auf a noise reduction filter and its derivative</li>
<li><figref idref="f0006">Fig. 12</figref> block diagram of the onset sharpening using an algorithm according to <maths id="math0001" num=""><math display="inline"><msubsup><mi>G</mi><mi mathvariant="italic">rec</mi><mfenced separators=""><mi mathvariant="italic">mod</mi><mspace width="1em"/><mn>1</mn></mfenced></msubsup><mfenced separators=""><msub><mi mathvariant="normal">Ω</mi><mi>μ</mi></msub><mo>⁢</mo><mi>l</mi></mfenced></math><img id="ib0001" file="imgb0001.tif" wi="26" he="8" img-content="math" img-format="tif" inline="yes"/></maths> of the recursive Wiener filter</li>
</ul></p>
<heading id="h0005"><b><u>Detailed description of the invention</u></b></heading>
<p id="p0038" num="0038"><figref idref="f0001">Fig. 1</figref> shows for the acoustic speech "Sauerkraut is served once a week" the energy distribution in time and frequency, a time series of the energy above 4 kHz and a time series of the amplitude. The fricatives "s", "ce" and "k" have quite some energy in the frequency range from 4 to 8 kHz.<!-- EPO <DP n="10"> --></p>
<p id="p0039" num="0039">A time-domain input signal, as for example "Sauerkraut is served once a week", is windowed before being transferred to the frequency domain via a fast Fourier transform (FFT) algorithm, discrete Fourier transform (DFT), discrete cosine transform (DCT), polyphase filterbank, wavelet transform, or some other time to frequency domain transformation. In the frequency domain the windowed time-domain signal has the form of spectral representations. By filtering for example with a mel-filterbank the spectral representations can be reduced to feature vectors.</p>
<p id="p0040" num="0040"><figref idref="f0001">Fig. 2</figref> shows an example of clusters of two dimensional feature vectors. With the LBG-algorithm for (c) k = <i>K</i> = 8 classes eight cluster centers + are found. The procedure is similar for feature vectors with higher dimensionality. The spectral representations or the feature vectors of the cluster centres are defined on the bases of spectral representations of windowed acoustic speech segments of a speech corpus by the clustering algorithm. A code book includes the spectral class representations for the cluster centres, which are averages over the elements of the cluster.</p>
<p id="p0041" num="0041"><figref idref="f0002">Fig. 3</figref> shows results of clustering with K = 8 in classes after pre-classification with a threshold of <i>fT</i> = 3kHz in (b). The line with more details shows the average spectrum in dB, the line with less details shows a filtered spectral representation of the code book entry and the vertical line the spectral centroid <i>fc</i> . The filtered spectral representations of the code book entries are the spectral class representations for the cluster centres, which are averages over the elements of the cluster.</p>
<p id="p0042" num="0042">The spectral representations or feature vectors of windowed frames of an input signal are subjected to a classification technique from the field of pattern recognition, such as the minimum mean squared error estimator. The input frames are classified by finding the class corresponding to the smallest value of a cost function, d(<b>v,c</b><sub>k</sub>) in Equation 1, where v<sub>d</sub> is the D-dimensional set of features of the current frame and c<sub>dk</sub> is the set of features of a codebook entry k. <maths id="math0002" num="equation 1"><math display="block"><mi>d</mi><mfenced separators=""><mi mathvariant="normal">v</mi><mo>⁢</mo><msub><mi mathvariant="normal">c</mi><mi>k</mi></msub></mfenced><mo>=</mo><mstyle displaystyle="true"><munderover><mo>∑</mo><mrow><mi>d</mi><mo>=</mo><mn>1</mn></mrow><mi>D</mi></munderover></mstyle><msup><mfenced separators=""><msub><mi>ν</mi><mi>d</mi></msub><mo>-</mo><msub><mi>c</mi><mi mathvariant="italic">dk</mi></msub></mfenced><mn>2</mn></msup></math><img id="ib0002" file="imgb0002.tif" wi="57" he="31" img-content="math" img-format="tif"/></maths></p>
<p id="p0043" num="0043">In <figref idref="f0003">Fig. 4</figref>, this setup is illustrated using K=128 as a possible number of entries for such a codebook. The codebook can be trained using feature vectors from a training set and any number of vector quantization algorithms, such as k-means [MacQueen, 1967] or the Linde-Buzo-Gray (LBG) algorithm[Linde, Buzo, &amp; Gray, 1980]. Linear discriminant analysis (LDA), principal component analysis (PCA), support vector machines (SVM), and other class<!-- EPO <DP n="11"> --> enhancing algorithms can also be applied to decrease the intraclass variances and/or increase the interclass distances, which improves separability.</p>
<p id="p0044" num="0044">The feature vectors for training and testing are in this case a set of perception based features in the mel scale, although the feature vectors must not be limited to these specific features. The features can also be designed to emphasize signal characteristics in or near the region of interest in the frequency spectrum. The key concept here is that the feature vector has a reduced dimensionality with respect to the input signal frequency spectrum in order to reduce computational costs.</p>
<p id="p0045" num="0045">A spectral compression function and/or a formant boosting function are chosen from the relevant codebook entry.</p>
<p id="p0046" num="0046">To avoid possible switching artifacts, a smoothing procedure can be applied to the chosen compression function in the frequency domain but in the time direction. Using for example, an IIR-smoothing function, the time-variance of the applied compression functions can be reduced.</p>
<p id="p0047" num="0047">The spectral compression functions are functions- in the preferred case continuous and nonlinear - that apply compression rates in the frequency domain. Both the compression functions and the frequency spectrum can be either linearly or nonlinearly scaled, reflecting the occasional advantages of processing in more perception based scales like the logarithmic scale.</p>
<p id="p0048" num="0048">The operation of spectral compression maps a frequency interval of the input signal into a smaller frequency interval of the output. When working on a discrete representation of the spectrum, this is equivalent to mapping a set of frequency pins from the input spectrum X(e<sup>jΩµ</sup>) into one frequency pin of the output spectrum Y (e<sup>jΩµ</sup>). Mathematically, an operation that reduces the input frequency bandwidth by approximately 0.5 can be expressed by <maths id="math0003" num="equation 2"><math display="block"><mtable><mtr><mtd><mi>Y</mi><mfenced><msup><mi>e</mi><mrow><mi>j</mi><mo>⁢</mo><msub><mi mathvariant="normal">Ω</mi><mi>ν</mi></msub></mrow></msup></mfenced><mo>=</mo><mfrac><mrow><mi>Y</mi><mfenced><msup><mi>e</mi><mrow><mi>j</mi><mo>⁢</mo><msub><mi mathvariant="normal">Ω</mi><mi>ν</mi></msub></mrow></msup></mfenced><mspace width="3em"/><mn>1</mn></mrow><mrow><mfenced open="|" close="|" separators=""><mi>X</mi><mfenced><msup><mi>e</mi><mrow><mi>j</mi><mo>⁢</mo><msub><mi mathvariant="normal">Ω</mi><mi>ν</mi></msub></mrow></msup></mfenced></mfenced><mo>⁢</mo><msubsup><mi>μ</mi><mi>u</mi><mfenced><mi>ν</mi></mfenced></msubsup><mo>-</mo><msubsup><mi>μ</mi><mi>l</mi><mfenced><mi>ν</mi></mfenced></msubsup><mo>-</mo><mn>1</mn></mrow></mfrac><mstyle displaystyle="true"><munderover><mo>∑</mo><mrow><mi>μ</mi><mo>=</mo><msubsup><mi>μ</mi><mi>l</mi><mfenced><mi>ν</mi></mfenced></msubsup></mrow><msubsup><mi>μ</mi><mi>u</mi><mfenced><mi>ν</mi></mfenced></msubsup></munderover></mstyle><mfenced open="|" close="|" separators=""><mi>X</mi><mfenced><msup><mi>e</mi><mrow><mi>j</mi><mo>⁢</mo><msub><mi mathvariant="normal">Ω</mi><mi>μ</mi></msub></mrow></msup></mfenced></mfenced></mtd><mtd><mi>ν</mi><mo>=</mo><mn>0</mn><mo>,</mo><mo>…</mo><mo>,</mo><mfrac><mi>N</mi><mn>4</mn></mfrac><mo>+</mo><mn>1.</mn></mtd></mtr></mtable></math><img id="ib0003" file="imgb0003.tif" wi="148" he="31" img-content="math" img-format="tif"/></maths></p>
<p id="p0049" num="0049">The upper and lower boundaries of the interval that is projected onto output frequency pin <i>v</i> are represented by the variables µ<sub>u</sub><sup>(v)</sup>and µ<sub>l</sub><sup>(v)</sup>, respectively. These boundaries are actually functions of <i>v</i>, allowing a variable amount of compression for different frequencies. Compression with this equation derives the magnitude of Y (e<sup>jΩv</sup>) as the mean value of the<!-- EPO <DP n="12"> --> magnitude of X(e<sup>jΩµ</sup>) for µ = µ<sub>u</sub><sup>(v)</sup>, ... , µ<sub>l</sub><sup>(v)</sup> while retaining the phase of input component <i>v</i> for output component <i>v</i>.</p>
<p id="p0050" num="0050"><figref idref="f0003">Fig. 5</figref> illustrates an example for the relationship between input frequency <i>f</i><sub>in</sub> and output frequency <i>f</i><sub>out</sub>. The curved line shows the compression characteristic, the horizontal and vertical lines indicate the frequency interval that is mapped. The frequency in Hz can be converted to the pin-index with the following relationship <maths id="math0004" num="equation 3"><math display="block"><mi>μ</mi><mo>=</mo><mo>⌊</mo><mfrac><mi>f</mi><msub><mi>f</mi><mn>3</mn></msub></mfrac><mo>⁢</mo><msub><mi>N</mi><mi>FFT</mi></msub><mo>⌉</mo></math><img id="ib0004" file="imgb0004.tif" wi="33" he="30" img-content="math" img-format="tif"/></maths></p>
<p id="p0051" num="0051">Equation 2 also performs energy normalization with respect to the amount of frequency compression. Energy normalization can take many forms, however, and could also be applied in a fashion based on the momentary broadband or frequency localized SNR. Another option is to maintain the form of the precompression estimated spectral envelope and apply a gain or dampening factor to correct the energy level to correspond to the slope of the envelope in the region of interest.</p>
<p id="p0052" num="0052">The equation also preserves the phase of the original uncompressed signal, but other phase corrections and adjustments are also possible, such as using a random phase.</p>
<p id="p0053" num="0053">The compression functions can be continuous in nature and have various compression rates over frequency. Preferably, the compression function is linear in lower frequencies, corresponding to no compression when the slope is equal to 1. In the higher frequencies of interest, the compression rate increases gradually, with the rate and degree of increase dependent upon the feature-based classification.</p>
<p id="p0054" num="0054">In <figref idref="f0003">Fig. 6</figref> various examples of continuous nonlinear compression functions are shown. Those displayed with solid lines are linear with a slope of, or near, 1 in the lower frequency range, hence performing little to no compression. With increasing input frequency, the compression rate increases differently for each of the curves. The dashed curves perform little to no compression in the middle frequency range, rather than the low frequencies. These curves compress frequencies in the lower and higher frequency extremes, and can be better understood as performing a transposition of the middle frequencies down into a lower region with almost no compression.</p>
<p id="p0055" num="0055">The appropriate compression function for each class of the code book is to be defined either manually, e.g., using subjective listening criteria, or automatically, e.g., applying unsupervised learning methods to find a preferred mapping to an output characteristic.<!-- EPO <DP n="13"> --> Different sets of class function assignments are generally possible, since the ideal output for improved audio quality does not always coincide with that providing improved speech recognition rates.</p>
<p id="p0056" num="0056">Speech can be made more accentuated, if formants are amplified. Especially in situations with background noise, we can expect a better localized SNR in these frequency regions, so the broadband SNR can be improved by applying a frequency dependent gain factor to the input spectrum. Ideally, the amplification should only be applied during speech activity. Otherwise, background noise will be amplified during pauses.</p>
<p id="p0057" num="0057">The current invention can receive information about speech activity from an external module, such as a noise detection and cancelation module.</p>
<p id="p0058" num="0058">The formant-boosting functions are also functions - (non)continuous, (non)linear, and on a (non)linear scale - designed such that a variable gain factor is applied to frequency ranges around formants in the spectrum. <figref idref="f0004">Fig. 7</figref> shows a formant boosting gain function and spectral representations (cluster mean and filtered cluster mean), where the line with more details shows the average spectrum in dB, the line with less details shows a filtered spectral representation of the code book entry. A curve can be used to determine the percentage of the gain factor that is applied to the formant frequency range. Alternatively individual gain and dampening curves can be stored in a codebook and applied to those frequency ranges identified as formants in voiced speech.</p>
<p id="p0059" num="0059">For some spaces between formants, the gain factor just described can be transformed into a dampening factor using another set of curves. The boosting curves can be designed to have positive values for amplifying the formant frequencies and negative values for the dampening curves to reduce the magnitude of the valleys between formants.</p>
<p id="p0060" num="0060">When compression is applied an additional curve can be used in the frequency range of the compression results. This amplifies or dampens the spectrum in the region of the frequency compression and so can be used to amplify or dampen the effects of the compression.</p>
<p id="p0061" num="0061">A decision can be made for each codebook entry of the spectral compression, formant boosting signal processing about the order of the steps. In some instances it can be necessary to first apply spectral compression and then the boosting function. At other times it should be possible to first boost (or dampen) regions of the spectrum and then to apply the spectral compression. This could be important when the compression is designed such that one of the formants is located in the compressed frequency range. The prior boosting of the formant would serve to retain the formant shape even after compression.<!-- EPO <DP n="14"> --></p>
<p id="p0062" num="0062">An overall block diagram of one possible incarnation of the system is seen in <figref idref="f0004">Fig. 8</figref>. Windows of a signal in the time domain are transformed to spectral representations by an analysis filter-bank. By extracting feature vectors a classification in relation to elements of a code book a signal processing is realized with the transformations linked to the code book elements. The transformed spectral representations are transformed to windows of a signal in the time domain by a synthesis filter-bank.</p>
<p id="p0063" num="0063"><figref idref="f0004">Fig. 9</figref> shows examples for processing of phonemes:
<ol id="ol0001" compact="compact" ol-style="">
<li>(a) with a spectral compression according to the shown compression characteristic ,</li>
<li>(b) with formant boosting according to the shown gain <i>g</i>(Ω<i><sub>µ</sub></i>) between + -10dB. The processed signals are band limited to 4 kHz.</li>
</ol></p>
<p id="p0064" num="0064"><figref idref="f0005">Fig. 10</figref> shows an example of the sauerkraut utterance processed with the 128 class scheme. The energy of the fricatives "s", "ce" and "k" is transformed below 4 kHz.</p>
<p id="p0065" num="0065">The system is also capable of receiving information from other signal processing modules, such as silence/unvoiced/voiced decisions from the noise estimation and reduction module. Furthermore the module should be capable of sending information to other modules, especially an ASR module, which can use the enhanced signal to improve recognition rates. The ASR module could be retrained with the compressed output data. However the compressed signal alone achieves a change in recognition rates that can be useful.</p>
<p id="p0066" num="0066">A focus on efficient algorithms and output signal synthesis allows the current invention to also avoid costly postprocessing to achieve high quality audio output. Additionally, transforming the input signal from the time domain to the frequency domain is not bound to the FFT algorithm. Compression functions in the current invention are intended to be designed such that they are able to be efficiently stored in memory and do not inadvertently give portions of the input spectrum undesired emphasis.</p>
<heading id="h0006"><b><u>Onset Sharpening</u></b></heading>
<p id="p0067" num="0067">Another invention is disclosed, which is new and inventive independent of the independent claims. This further invention is related to onset sharpening, which can be used independently but is of course advantageously combinable with the previously described speech enhancement methods.</p>
<p id="p0068" num="0068">The onset sharpening method performs an onset sharpening, i.e., to introduce attenuation immediately before and/or amplification after speech onsets in order to make them more accentuated. This is improving speech quality and intelligibility, especially for speech signals corrupted with background noise that are to be enhanced with noise reduction methods.<!-- EPO <DP n="15"> --> Noise reduction filters and their derivatives (<figref idref="f0005">Fig. 11</figref>) tend to react too slow and therefore do remove desired signal components during speech onsets.</p>
<p id="p0069" num="0069">Before any onset-dependent signal enhancement can be performed, the speech onsets need to be found first. This is done based on a recursive Wiener noise reduction filter. There are no real-time constraints, so all of the following steps can be applied to the entire signal x(t) giving the necessary data for the next step.</p>
<p id="p0070" num="0070">An analysis filter bank is needed to transform the input signal x(t) into the frequency domain. The result is a function <i>X</i>(<i>e</i><sup><i>j</i>Ω<i>µ</i></sup>, l) with µ= 0, ... , N<sub>DFT</sub>=2 and I = 0, ... , M-1, where M is the number of signal frames. This could also be interpreted as a (N<sub>DFT</sub>=2+1 ) x M matrix which is constituting a spectrogram.</p>
<p id="p0071" num="0071">Based on the spectral signal representation of the previous step, the attenuation factors of the recursive Wiener filter characteristic can be computed. The following, non-frequency dependent, parameters have been used as filter parameters:
<ul id="ul0005" list-style="bullet" compact="compact">
<li>Maximum attenuation G<sub>min</sub>(<sup>Ω</sup>µ, l) = G<sub>min</sub> = 0.25 (corresponding to 20 log<sub>10</sub>(G<sub>min</sub>) = -20 dB)</li>
<li>Overestimation factor <i>β̃</i>(Ω<sub>µ</sub>, l) = <i>β̃</i> = 5</li>
</ul></p>
<p id="p0072" num="0072">The result G<sub>rec</sub>(Ωµ, l) can again be seen as a (N<sub>DFT</sub>=2+1 ) x M matrix containing the attenuation factor for each sub-band for all time instances.</p>
<p id="p0073" num="0073">In order to smooth the filter coefficients in a perceptually meaningful manner over time, a mel-filterbank of 32 bands is applied to G<sub>rec</sub>(Ω<sub>µ</sub>, l), resulting in the <maths id="math0005" num=""><math display="inline"><mn>32</mn><mspace width="1em"/><mi>x M matrix</mi><mspace width="1em"/><msubsup><mi>G</mi><mi mathvariant="italic">rec</mi><mfenced><mi mathvariant="italic">mel</mi></mfenced></msubsup><mfenced separators=""><mi mathvariant="normal">m</mi><mo>⁢</mo><mi mathvariant="normal">l</mi></mfenced><mn>.</mn></math><img id="ib0005" file="imgb0005.tif" wi="47" he="9" img-content="math" img-format="tif" inline="yes"/></maths> The actual onset detection is performed within each mel-band of <maths id="math0006" num=""><math display="inline"><msubsup><mi>G</mi><mi mathvariant="italic">rec</mi><mfenced><mi mathvariant="italic">mel</mi></mfenced></msubsup><mfenced separators=""><mi mathvariant="normal">m</mi><mo>⁢</mo><mi mathvariant="normal">l</mi></mfenced><mn>.</mn></math><img id="ib0006" file="imgb0006.tif" wi="23" he="9" img-content="math" img-format="tif" inline="yes"/></maths>. Here, the moments when <maths id="math0007" num=""><math display="inline"><msubsup><mi>G</mi><mi mathvariant="italic">rec</mi><mfenced><mi mathvariant="italic">mel</mi></mfenced></msubsup><mfenced separators=""><mi mathvariant="normal">m</mi><mo>⁢</mo><mi mathvariant="normal">l</mi></mfenced></math><img id="ib0007" file="imgb0007.tif" wi="24" he="9" img-content="math" img-format="tif" inline="yes"/></maths> changes its value from G<sub>min</sub> to 1 (or close to 1) are of interest, i.e., the points when the filter opens. These time instances can be found by taking the numerical derivative <maths id="math0008" num=""><math display="block"><mfrac><mrow><msubsup><mi mathvariant="italic">dG</mi><mi mathvariant="italic">rec</mi><mfenced><mi mathvariant="italic">mel</mi></mfenced></msubsup><mfenced separators=""><mi>m</mi><mo>⁢</mo><mi>l</mi></mfenced></mrow><mi mathvariant="italic">dl</mi></mfrac><mo>≈</mo><msubsup><mi mathvariant="italic">G</mi><mi mathvariant="italic">rec</mi><mfenced><mi mathvariant="italic">mel</mi></mfenced></msubsup><mfenced separators=""><mi>m</mi><mo>⁢</mo><mi>l</mi></mfenced><mo>-</mo><msubsup><mi mathvariant="italic">G</mi><mi mathvariant="italic">rec</mi><mfenced><mi mathvariant="italic">mel</mi></mfenced></msubsup><mo>⁢</mo><mfenced separators=""><mi>m</mi><mo>,</mo><mi>l</mi><mo>-</mo><mn>1</mn></mfenced></math><img id="ib0008" file="imgb0008.tif" wi="88" he="18" img-content="math" img-format="tif"/></maths><br/>
and comparing the resulting value with a threshold <maths id="math0009" num=""><math display="block"><mfrac><mrow><mi>d</mi><mo>⁢</mo><msubsup><mi mathvariant="italic">G</mi><mi mathvariant="italic">rec</mi><mfenced><mi mathvariant="italic">mel</mi></mfenced></msubsup><mfenced separators=""><mi>m</mi><mo>⁢</mo><mi>l</mi></mfenced></mrow><mi mathvariant="italic">dl</mi></mfrac><mo>&gt;</mo><mi>γ</mi></math><img id="ib0009" file="imgb0009.tif" wi="36" he="17" img-content="math" img-format="tif"/></maths></p>
<p id="p0074" num="0074">The mel-band m for time instance l is labeled as a speech onset. The derivative of the noise reduction coefficient lies in the range of d <maths id="math0010" num=""><math display="inline"><msubsup><mi>G</mi><mi mathvariant="italic">rec</mi><mfenced><mi mathvariant="italic">mel</mi></mfenced></msubsup><mfenced separators=""><mi mathvariant="normal">m</mi><mo>⁢</mo><mi mathvariant="normal">l</mi></mfenced><mo>/</mo><mi>dl</mi><mo>∈</mo><mfenced open="[" close="]" separators=""><msub><mi mathvariant="normal">G</mi><mi>min</mi></msub><mo mathvariant="normal">-</mo><mn mathvariant="normal">1</mn><mo mathvariant="normal">,</mo><mn mathvariant="normal">1</mn><mo mathvariant="normal">-</mo><msub><mi mathvariant="normal">G</mi><mi>min</mi></msub></mfenced></math><img id="ib0010" file="imgb0010.tif" wi="57" he="9" img-content="math" img-format="tif" inline="yes"/></maths> and positive<!-- EPO <DP n="16"> --> values indicate times when the filter is opening. A threshold = 0.2 has proved to give good detection results for various speech signals and SNRs. Because it could happen that the derivative is greater than for several consecutive frames, also a sliding time window of 100ms duration is applied. Within this time window, only one detection is allowed. Furthermore, if an onset has been detected in a certain mel-band, the neighboring mel-band towards lower frequencies of the same frame l is also marked as a speech onset. The clean speech is mixed with background noise recorded in a car driving at a speed of 160 km/h to form an SNR of 1 dB during speech activity. Of course, the detection can be made more sensitive by taking a lower threshold, e.g. y= 0.1.</p>
<p id="p0075" num="0075">Based on the detections made with the method of the previous section, the onset sharpening can be performed. It consists of placing attenuation immediately before a detected speech onset and boosting the signal for a short time interval afterwards. The shape of the attenuation/amplification is determined by a prototype function that has been chosen to be <maths id="math0011" num=""><math display="block"><mi>f</mi><mfenced><mi>x</mi></mfenced><mo>=</mo><mrow><mo>{</mo><mtable><mtr><mtd><mo>-</mo><mi>ϱ</mi><mo>⁢</mo><mfrac><mi>e</mi><mi>σ</mi></mfrac><mo>⁢</mo><msup><mi mathvariant="italic">xe</mi><mfrac><mi>x</mi><mi>σ</mi></mfrac></msup><mo>⁢</mo><mi mathvariant="italic">for</mi><mo>-</mo><mn>20</mn><mo>≤</mo><mi>x</mi><mo>&lt;</mo><mn>0</mn></mtd></mtr><mtr><mtd><mfrac><mi>e</mi><mi>σ</mi></mfrac><mo>⁢</mo><msup><mi mathvariant="italic">xe</mi><mfrac><mrow><mo>-</mo><mi>x</mi></mrow><mi>σ</mi></mfrac></msup><mo>⁢</mo><mi mathvariant="italic">for</mi><mspace width="1em"/><mn>0</mn><mo>≤</mo><mi>x</mi><mo>≤</mo><mn>20</mn></mtd></mtr><mtr><mtd><mn>0</mn><mmultiscripts><mi mathvariant="italic">otherwise</mi><mprescripts/><mspace width="1em"/><none/></mmultiscripts></mtd></mtr></mtable></mrow></math><img id="ib0011" file="imgb0011.tif" wi="72" he="29" img-content="math" img-format="tif"/></maths><br/>
with the attenuation <img id="ib0012" file="imgb0012.tif" wi="4" he="5" img-content="character" img-format="tif" inline="yes"/> = 0.5 and σ = 3. The term e/ σ is used to normalize f (x) to a maximum value of 1. Out of this prototype function, the onset sharpening gain function g<sub>os</sub>(l) can be sampled. When deriving onset sharpening gain function, three parameters can be set:
<ol id="ol0002" compact="compact" ol-style="">
<li>1. The width of the (negative) attenuation and the (positive) boosting part, defined by the variable τ<sub>os</sub> in ms.</li>
<li>2. The offset τ<sub>offs</sub> in ms that defines at which time instance the gain function is placed. For τ<sub>offs</sub> = 0 ms, the zero crossing is exactly at the detection point, for negative offset values the zero crossing will be earlier. This is desirable because then the noise reduction filter is forced to open earlier.</li>
<li>3. A gain parameter α<sub>os</sub> that controls the amount of attenuation and boosting. It is multiplied with the onset sharpening gain function g<sub>os</sub>(l) which is interpreted as dB values.</li>
</ol></p>
<p id="p0076" num="0076">Several other prototype functions could be devised for onset sharpening, e.g., a sinusoid. The prototype function has been chosen with the stated parameters because it decays smoothly towards the ends and offers a steep slope around the zero crossing. This prototype function is now placed at each detected speech onset time instance in the corresponding mel-band, giving the onset sharpening matrix for mel-bands <maths id="math0012" num=""><math display="inline"><msubsup><mi>g</mi><mi mathvariant="italic">os</mi><mfenced><mi mathvariant="italic">mel</mi></mfenced></msubsup><mfenced separators=""><mi>m</mi><mo>⁢</mo><mi>l</mi></mfenced></math><img id="ib0013" file="imgb0013.tif" wi="21" he="8" img-content="math" img-format="tif" inline="yes"/></maths>,<!-- EPO <DP n="17"> --> which then is expanded into the onset sharpening matrix for the subbands g<sub>os</sub>(m,l). The parameters that have been used for the onset sharpening gain functions are τ<sub>os</sub> = 75 ms, τ<sub>offs</sub> = -10ms and <i>α</i><sub>os</sub> = 3, of course other parameters are possible as well. Using the notation of the mel-filterbank matrix A the expansion from melbands to subbands can be expressed as <maths id="math0013" num=""><math display="block"><msub><mi>g</mi><mi mathvariant="italic">os</mi></msub><mfenced separators=""><msub><mi mathvariant="normal">Ω</mi><mi>μ</mi></msub><mo>⁢</mo><mi>l</mi></mfenced><mo>=</mo><mstyle displaystyle="true"><munderover><mo>∑</mo><mrow><mi>m</mi><mo>=</mo><mn>0</mn></mrow><mi>D</mi></munderover></mstyle><mfrac><msub><mi>A</mi><mi mathvariant="italic">mμ</mi></msub><mrow><msub><mi mathvariant="italic">max</mi><mi>μ</mi></msub><mfenced open="{" close="}"><msub><mi>A</mi><mi mathvariant="italic">mμ</mi></msub></mfenced></mrow></mfrac><mo>*</mo><msubsup><mi>g</mi><mi mathvariant="italic">os</mi><mfenced><mi mathvariant="italic">mel</mi></mfenced></msubsup><mfenced separators=""><mi>m</mi><mo>⁢</mo><mi>l</mi></mfenced><mn>.</mn></math><img id="ib0014" file="imgb0014.tif" wi="78" he="20" img-content="math" img-format="tif"/></maths></p>
<p id="p0077" num="0077">The weighting of the filters contained in A that gives triangles of a broader bandwidth a lower amplitude is removed by the normalization containing the maximum operation.</p>
<p id="p0078" num="0078"><figref idref="f0006">Fig. 12</figref> discloses an onset sharpening algorithm with the following recursive Wiener noise reduction filter being modified by the onset sharpening gain function. The simplest way is by multiplication <maths id="math0014" num=""><math display="block"><mtable columnalign="left"><mtr><mtd><msubsup><mi>G</mi><mi mathvariant="italic">rec</mi><mfenced separators=""><mi mathvariant="italic">mod</mi><mspace width="1em"/><mn>1</mn></mfenced></msubsup><mfenced separators=""><msub><mi mathvariant="normal">Ω</mi><mi>μ</mi></msub><mo>⁢</mo><mi>l</mi></mfenced></mtd><mtd><mo>=</mo><msubsup><mi>g</mi><mi mathvariant="italic">os</mi><mi mathvariant="italic">lin</mi></msubsup><mfenced separators=""><msub><mi mathvariant="normal">Ω</mi><mi>μ</mi></msub><mo>⁢</mo><mi>l</mi></mfenced><mo>⋅</mo><msub><mi>G</mi><mi mathvariant="italic">rec</mi></msub><mfenced separators=""><msub><mi mathvariant="normal">Ω</mi><mi>μ</mi></msub><mo>⁢</mo><mi>l</mi></mfenced></mtd></mtr><mtr><mtd><mspace width="1em"/></mtd><mtd><mo>=</mo><msubsup><mi>g</mi><mi mathvariant="italic">os</mi><mi mathvariant="italic">lin</mi></msubsup><mfenced separators=""><msub><mi mathvariant="normal">Ω</mi><mi>μ</mi></msub><mo>⁢</mo><mi>l</mi></mfenced><mo>⋅</mo><mi>max</mi><mfenced open="{" close="}" separators=""><msub><mi>G</mi><mi mathvariant="italic">min</mi></msub><mfenced separators=""><msub><mi mathvariant="normal">Ω</mi><mi>μ</mi></msub><mo>⁢</mo><mi>l</mi></mfenced><mo>,</mo><mn>1</mn><mo>-</mo><mfrac><mrow><mover><mi>β</mi><mo>˜</mo></mover><mfenced separators=""><msub><mi mathvariant="normal">Ω</mi><mi>μ</mi></msub><mo>⁢</mo><mi>l</mi></mfenced></mrow><mrow><mi>G</mi><mo>⁢</mo><mfenced separators=""><msub><mi mathvariant="normal">Ω</mi><mi>μ</mi></msub><mo>,</mo><mi>l</mi><mo>-</mo><mn>1</mn></mfenced></mrow></mfrac><mo>⋅</mo><mfrac><msup><mfenced open="|" close="|" separators=""><mover><mi>B</mi><mo>^</mo></mover><mfenced separators=""><msup><mi>e</mi><mrow><mi>j</mi><mo>⁢</mo><msub><mi mathvariant="normal">Ω</mi><mi>μ</mi></msub></mrow></msup><mo>⁢</mo><mi>l</mi></mfenced></mfenced><mn>2</mn></msup><msup><mfenced open="|" close="|" separators=""><mover><mi>X</mi><mo>^</mo></mover><mfenced separators=""><msup><mi>e</mi><mrow><mi>j</mi><mo>⁢</mo><msub><mi mathvariant="normal">Ω</mi><mi>μ</mi></msub></mrow></msup><mo>⁢</mo><mi>l</mi></mfenced></mfenced><mn>2</mn></msup></mfrac></mfenced></mtd></mtr></mtable></math><img id="ib0015" file="imgb0015.tif" wi="145" he="28" img-content="math" img-format="tif"/></maths><br/>
where <maths id="math0015" num=""><math display="block"><msubsup><mi>g</mi><mi mathvariant="italic">os</mi><mi mathvariant="italic">lin</mi></msubsup><mfenced separators=""><msub><mi mathvariant="normal">Ω</mi><mi>μ</mi></msub><mo>⁢</mo><mi>l</mi></mfenced><mo>=</mo><msup><mn>10</mn><mrow><msub><mi>α</mi><mi mathvariant="italic">dB</mi></msub><mn>.</mn><msub><mi>g</mi><mi mathvariant="italic">os</mi></msub><mfenced separators=""><msub><mi mathvariant="normal">Ω</mi><mi>μ</mi></msub><mo>⁢</mo><mi>l</mi></mfenced><mo>/</mo><mn>20</mn></mrow></msup></math><img id="ib0016" file="imgb0016.tif" wi="62" he="12" img-content="math" img-format="tif"/></maths><br/>
is the onset sharpening function in linear values. Since the noise reduction filter is applied multiplicative to the input spectrum, this filter modification could also be interpreted as a multiplication of the signal spectrum <i>X</i>(<i>e</i><sup><i>j</i>Ωµ</sup><i>, l</i>) with the onset sharpening gain before (or after) applying the noise reduction filter.</p>
<p id="p0079" num="0079">In a second modification, the onset sharpening gain is built into the noise reduction characteristic to modify the spectral floor and the maximum gain of the filter (which is set to 1 in the recursive Wiener filter): <maths id="math0016" num=""><math display="block"><msubsup><mi mathvariant="normal">G</mi><mi>rec</mi><mfenced separators=""><mi>mod</mi><mspace width="1em"/><mn mathvariant="normal">2</mn></mfenced></msubsup><mfenced separators=""><msub><mi mathvariant="normal">Ω</mi><mi mathvariant="normal">μ</mi></msub><mo>⁢</mo><mi mathvariant="normal">l</mi></mfenced><mo mathvariant="normal">=</mo><mi>max</mi><mfenced open="{" close="}" separators=""><msub><mi mathvariant="normal">G</mi><mi>min</mi></msub><mfenced separators=""><msub><mi mathvariant="normal">Ω</mi><mi mathvariant="normal">μ</mi></msub><mo>⁢</mo><mi mathvariant="normal">l</mi></mfenced><mo mathvariant="normal">⋅</mo><msubsup><mi mathvariant="normal">g</mi><mi>os</mi><mi>att</mi></msubsup><mfenced separators=""><msub><mi mathvariant="normal">Ω</mi><mi mathvariant="normal">μ</mi></msub><mo>⁢</mo><mi mathvariant="normal">l</mi></mfenced><mo mathvariant="normal">,</mo><msubsup><mi mathvariant="normal">g</mi><mi>os</mi><mi>amp</mi></msubsup><mfenced separators=""><msub><mi mathvariant="normal">Ω</mi><mi mathvariant="normal">μ</mi></msub><mo>⁢</mo><mi mathvariant="normal">l</mi></mfenced><mo mathvariant="normal">-</mo><mfrac><msup><mfenced open="|" close="|" separators=""><mover><mi mathvariant="normal">B</mi><mo mathvariant="normal">^</mo></mover><mfenced separators=""><msup><mi mathvariant="normal">e</mi><msub><mi mathvariant="normal">jΩ</mi><mi mathvariant="normal">μ</mi></msub></msup><mo>⁢</mo><mi mathvariant="normal">l</mi></mfenced></mfenced><mn mathvariant="normal">2</mn></msup><msup><mfenced open="|" close="|" separators=""><mover><mi mathvariant="normal">X</mi><mo mathvariant="normal">^</mo></mover><mfenced separators=""><msup><mi mathvariant="normal">e</mi><msub><mi mathvariant="normal">jΩ</mi><mi mathvariant="normal">μ</mi></msub></msup><mo>⁢</mo><mi mathvariant="normal">l</mi></mfenced></mfenced><mn mathvariant="normal">2</mn></msup></mfrac></mfenced><mn mathvariant="normal">.</mn></math><img id="ib0017" file="imgb0017.tif" wi="128" he="19" img-content="math" img-format="tif"/></maths></p>
<p id="p0080" num="0080">For this description, the gain function has to be separated into the part responsible for the attenuation before speech onsets <maths id="math0017" num=""><math display="block"><msubsup><mi mathvariant="normal">g</mi><mi>os</mi><mi>att</mi></msubsup><mfenced separators=""><msub><mi mathvariant="normal">Ω</mi><mi mathvariant="normal">μ</mi></msub><mo>⁢</mo><mi mathvariant="normal">l</mi></mfenced><mo>=</mo><mrow><mo>{</mo><mtable><mtr><mtd><msup><mn>10</mn><mrow><msub><mi>α</mi><mi mathvariant="italic">dB</mi></msub><mn>.</mn><msub><mi>g</mi><mi mathvariant="italic">os</mi></msub><mfenced separators=""><msub><mi mathvariant="normal">Ω</mi><mi>μ</mi></msub><mo>⁢</mo><mi>l</mi></mfenced><mo>/</mo><mn>20</mn></mrow></msup></mtd><mtd columnalign="left"><msub><mi mathvariant="italic">for g</mi><mi>os</mi></msub><mfenced separators=""><msub><mi mathvariant="normal">Ω</mi><mi mathvariant="normal">μ</mi></msub><mo>⁢</mo><mi mathvariant="normal">l</mi></mfenced><mo>&lt;</mo><mn>0</mn></mtd></mtr><mtr><mtd><mn>1</mn></mtd><mtd columnalign="left"><mi>otherwise</mi></mtd></mtr></mtable></mrow></math><img id="ib0018" file="imgb0018.tif" wi="95" he="24" img-content="math" img-format="tif"/></maths><br/>
<!-- EPO <DP n="18"> -->and the part for amplification after speech onsets <maths id="math0018" num=""><math display="block"><msubsup><mi mathvariant="normal">g</mi><mi>os</mi><mi>amp</mi></msubsup><mfenced separators=""><msub><mi mathvariant="normal">Ω</mi><mi mathvariant="normal">μ</mi></msub><mo>⁢</mo><mi mathvariant="normal">l</mi></mfenced><mo>=</mo><mrow><mo>{</mo><mtable><mtr><mtd><msup><mn>10</mn><mrow><msub><mi>α</mi><mi mathvariant="italic">dB</mi></msub><mn>.</mn><msub><mi>g</mi><mi mathvariant="italic">os</mi></msub><mfenced separators=""><msub><mi mathvariant="normal">Ω</mi><mi>μ</mi></msub><mo>⁢</mo><mi>l</mi></mfenced><mo>/</mo><mn>20</mn></mrow></msup></mtd><mtd columnalign="left"><msub><mi mathvariant="italic">for g</mi><mi>os</mi></msub><mfenced separators=""><msub><mi mathvariant="normal">Ω</mi><mi mathvariant="normal">μ</mi></msub><mo>⁢</mo><mi mathvariant="normal">l</mi></mfenced><mo>≥</mo><mn>0</mn></mtd></mtr><mtr><mtd><mn>1</mn></mtd><mtd columnalign="left"><mi>otherwise</mi></mtd></mtr></mtable></mrow></math><img id="ib0019" file="imgb0019.tif" wi="92" he="20" img-content="math" img-format="tif"/></maths></p>
<p id="p0081" num="0081">A third possibility is to modify also the overestimation factor: <maths id="math0019" num=""><math display="block"><msubsup><mi mathvariant="normal">G</mi><mi>rec</mi><mfenced separators=""><mi>mod</mi><mspace width="1em"/><mn mathvariant="normal">3</mn></mfenced></msubsup><mfenced separators=""><msub><mi mathvariant="normal">Ω</mi><mi mathvariant="normal">μ</mi></msub><mo>⁢</mo><mi mathvariant="normal">l</mi></mfenced><mo mathvariant="normal">=</mo><mi>max</mi><mfenced open="{" close="}" separators=""><msub><mi mathvariant="normal">G</mi><mi>min</mi></msub><mfenced separators=""><msub><mi mathvariant="normal">Ω</mi><mi mathvariant="normal">μ</mi></msub><mo>⁢</mo><mi mathvariant="normal">l</mi></mfenced><mo mathvariant="normal">⋅</mo><msubsup><mi mathvariant="normal">g</mi><mi>os</mi><mi>att</mi></msubsup><mfenced separators=""><msub><mi mathvariant="normal">Ω</mi><mi mathvariant="normal">μ</mi></msub><mo>⁢</mo><mi mathvariant="normal">l</mi></mfenced><mo mathvariant="normal">,</mo><mn mathvariant="normal">1</mn><mo mathvariant="normal">-</mo><mfrac><mrow><mover><mi mathvariant="italic">β</mi><mo mathvariant="normal">˜</mo></mover><mfenced separators=""><msub><mi mathvariant="normal">Ω</mi><mi>μ</mi></msub><mo>⁢</mo><mi>l</mi></mfenced></mrow><mrow><msubsup><mi mathvariant="normal">g</mi><mi>os</mi><mi>amp</mi></msubsup><mfenced separators=""><msub><mi mathvariant="normal">Ω</mi><mi mathvariant="normal">μ</mi></msub><mo>⁢</mo><mi mathvariant="normal">l</mi></mfenced><mo mathvariant="normal">⋅</mo><mi mathvariant="normal">G</mi><mo>⁢</mo><mfenced separators=""><msub><mi mathvariant="normal">Ω</mi><mi mathvariant="normal">μ</mi></msub><mo mathvariant="normal">,</mo><mi mathvariant="normal">l</mi><mo mathvariant="normal">-</mo><mn mathvariant="normal">1</mn></mfenced></mrow></mfrac><mo>⁢</mo><mfrac><msup><mfenced open="|" close="|" separators=""><mover><mi mathvariant="normal">B</mi><mo mathvariant="normal">^</mo></mover><mfenced separators=""><msup><mi mathvariant="normal">e</mi><msub><mi mathvariant="normal">jΩ</mi><mi mathvariant="normal">μ</mi></msub></msup><mo>⁢</mo><mn mathvariant="normal">1</mn></mfenced></mfenced><mn mathvariant="normal">2</mn></msup><msup><mfenced open="|" close="|" separators=""><mover><mi mathvariant="normal">X</mi><mo mathvariant="normal">^</mo></mover><mfenced separators=""><msup><mi mathvariant="normal">e</mi><msub><mi mathvariant="normal">jΩ</mi><mi mathvariant="normal">μ</mi></msub></msup><mo>⁢</mo><mn mathvariant="normal">1</mn></mfenced></mfenced><mn mathvariant="normal">2</mn></msup></mfrac></mfenced></math><img id="ib0020" file="imgb0020.tif" wi="161" he="24" img-content="math" img-format="tif"/></maths></p>
<p id="p0082" num="0082">Inspection of noise reduction filters from the first and second modification shows that there are only small differences between the characteristics. The parameter settings for τ<sub>os</sub>, τ<sub>offs</sub>, α<sub>os</sub> and the choice of the gain prototype function are of much greater influence. Therefore, the simpler modification <maths id="math0020" num=""><math display="inline"><msubsup><mi>G</mi><mi mathvariant="italic">rec</mi><mfenced separators=""><mi mathvariant="italic">mod</mi><mspace width="1em"/><mn>1</mn></mfenced></msubsup><mfenced separators=""><msub><mi mathvariant="normal">Ω</mi><mi>μ</mi></msub><mo>⁢</mo><mi>l</mi></mfenced></math><img id="ib0021" file="imgb0021.tif" wi="28" he="9" img-content="math" img-format="tif" inline="yes"/></maths>) will be used for the evaluation.</p>
<p id="p0083" num="0083">For the evaluation of the speech onset enhancement method, the recursive Wiener filter with <maths id="math0021" num=""><math display="inline"><msubsup><mi>G</mi><mi mathvariant="italic">rec</mi><mfenced separators=""><mi mathvariant="italic">mod</mi><mspace width="1em"/><mn>1</mn></mfenced></msubsup><mfenced separators=""><msub><mi mathvariant="normal">Ω</mi><mi>μ</mi></msub><mo>⁢</mo><mi>l</mi></mfenced></math><img id="ib0022" file="imgb0022.tif" wi="26" he="8" img-content="math" img-format="tif" inline="yes"/></maths> has been compared to a standard recursive Wiener filter <maths id="math0022" num=""><math display="inline"><msubsup><mi>G</mi><mi mathvariant="italic">rec</mi><mspace width="1em"/></msubsup><mfenced separators=""><msub><mi mathvariant="normal">Ω</mi><mi>μ</mi></msub><mo>⁢</mo><mi>l</mi></mfenced></math><img id="ib0023" file="imgb0023.tif" wi="19" he="8" img-content="math" img-format="tif" inline="yes"/></maths><i>.</i> This has mainly been done on the basis of a logarithmic spectral distance (LSD) measure. Comparison of the noise reduction filter coefficients for several characteristics gives a qualitative impression about the opening/closing properties of a filter. Finally, listening tests give a useful criterion that help to judge intelligibility and the amount of artifacts such as musical tones.</p>
<p id="p0084" num="0084">The idea in using an LSD measure is to create a signal x(t) = s(t)+ b(t) corrupted with background noise, where the speech component s(t) and the noise b(t) are known. A noise reduction filter is computed for the disturbed signal x(t) and the two signal components are passed through this filter separately. Then, the distortion measures LSD<sub>speech</sub> and LSD<sub>noise</sub> can be calculated between the original and the filtered signal. Ideally, the speech component is passed through the filter unchanged, leading to a distance close to zero, whereas large distortions can be expected for the noise component. Based on these two measures, it is possible to judge on the noise suppression ability and on how much speech components are affected. The LSD between two time varying spectra S(<i>e</i><sup><i>j</i>Ω<i>µ</i></sup>, <i>l</i>) and Ŝ(<i>e</i><sup><i>j</i>Ω<i>µ</i></sup>, <i>l</i>) is defined as <maths id="math0023" num=""><math display="block"><mi mathvariant="italic">LSD</mi><mo>=</mo><mfrac><mn>10</mn><mi>D</mi></mfrac><mstyle displaystyle="true"><munderover><mo>∑</mo><mrow><mi>l</mi><mo>=</mo><mn>0</mn></mrow><mi>L</mi></munderover></mstyle><msqrt><mstyle displaystyle="true"><munderover><mo>∑</mo><mrow><mi>μ</mi><mo>=</mo><mn>0</mn></mrow><msub><mi>N</mi><mrow><mi mathvariant="italic">DFT</mi><mo>⁢</mo><msub><mo>/</mo><mn>2</mn></msub></mrow></msub></munderover></mstyle><mfrac><msub><mi>K</mi><mrow><mi>μ</mi><mo>,</mo><mi>l</mi></mrow></msub><msub><mover><mi>K</mi><mo>‾</mo></mover><mi>l</mi></msub></mfrac><mo>⁢</mo><msubsup><mi mathvariant="italic">log</mi><mn>10</mn><mn>2</mn></msubsup><mfenced open="{" close="}"><mfrac><mrow><mi mathvariant="italic">max</mi><mfenced open="{" close="}" separators=""><msup><mfenced open="|" close="|" separators=""><mi>S</mi><mfenced separators=""><msup><mi>e</mi><mrow><mi>j</mi><mo>⁢</mo><msub><mi mathvariant="normal">Ω</mi><mi>μ</mi></msub></mrow></msup><mo>⁢</mo><mi>l</mi></mfenced></mfenced><mn>2</mn></msup><mo>⁢</mo><msub><mi>δ</mi><mi>S</mi></msub></mfenced></mrow><mrow><mi mathvariant="italic">max</mi><mfenced open="{" close="}" separators=""><msup><mfenced open="|" close="|" separators=""><mover><mi>S</mi><mo>^</mo></mover><mfenced separators=""><msup><mi>e</mi><mrow><mi>j</mi><mo>⁢</mo><msub><mi mathvariant="normal">Ω</mi><mi>μ</mi></msub></mrow></msup><mo>⁢</mo><mi>l</mi></mfenced></mfenced><mn>2</mn></msup><mo>⁢</mo><msub><mi>δ</mi><mover><mi>S</mi><mo>^</mo></mover></msub></mfenced></mrow></mfrac></mfenced></msqrt></math><img id="ib0024" file="imgb0024.tif" wi="100" he="26" img-content="math" img-format="tif"/></maths><!-- EPO <DP n="19"> --></p>
<p id="p0085" num="0085">The two variables <maths id="math0024" num=""><math display="block"><msub><mi>δ</mi><mi>S</mi></msub><mo>=</mo><msup><mn>10</mn><mrow><mo>-</mo><mn>5</mn></mrow></msup><mo>⁢</mo><munder><mi>max</mi><mrow><mi>μ</mi><mo>,</mo><mi>l</mi></mrow></munder><mfenced open="{" close="}"><msup><mfenced open="|" close="|" separators=""><mi>S</mi><mfenced separators=""><msup><mi>e</mi><mrow><mi>j</mi><mo>⁢</mo><msub><mi mathvariant="normal">Ω</mi><mi>μ</mi></msub></mrow></msup><mo>⁢</mo><mi>l</mi></mfenced></mfenced><mn>2</mn></msup></mfenced></math><img id="ib0025" file="imgb0025.tif" wi="61" he="14" img-content="math" img-format="tif"/></maths> <maths id="math0025" num=""><math display="block"><msub><mi>δ</mi><mover><mi>S</mi><mo>^</mo></mover></msub><mo>=</mo><msup><mn>10</mn><mrow><mo>-</mo><mn>5</mn></mrow></msup><mo>⁢</mo><munder><mi>max</mi><mrow><mi>μ</mi><mo>,</mo><mi>l</mi></mrow></munder><mfenced open="{" close="}"><msup><mfenced open="|" close="|" separators=""><mover><mi>S</mi><mo>^</mo></mover><mfenced separators=""><msup><mi>e</mi><mrow><mi>j</mi><mo>⁢</mo><msub><mi mathvariant="normal">Ω</mi><mi>μ</mi></msub></mrow></msup><mo>⁢</mo><mi>l</mi></mfenced></mfenced><mn>2</mn></msup></mfenced></math><img id="ib0026" file="imgb0026.tif" wi="61" he="14" img-content="math" img-format="tif"/></maths><br/>
give lower bounds for the values that enter the measure. These components are selected by the binary mask <maths id="math0026" num=""><math display="block"><msub><mi>K</mi><mrow><mi>μ</mi><mo>,</mo><mi>l</mi></mrow></msub><mo>=</mo><mrow><mo>{</mo><mtable><mtr><mtd><mn>1</mn><mspace width="1em"/><mi mathvariant="italic">if</mi><mspace width="1em"/><msup><mfenced open="|" close="|" separators=""><mi>S</mi><mfenced separators=""><msup><mi>e</mi><mrow><mi>j</mi><mo>⁢</mo><msub><mi mathvariant="normal">Ω</mi><mi>μ</mi></msub></mrow></msup><mo>⁢</mo><mi>l</mi></mfenced></mfenced><mn>2</mn></msup><mo>≥</mo><msub><mi>δ</mi><mi>S</mi></msub></mtd></mtr><mtr><mtd><mn>0</mn><mspace width="1em"/><mi mathvariant="italic">otherwise</mi></mtd></mtr></mtable></mrow></math><img id="ib0027" file="imgb0027.tif" wi="58" he="18" img-content="math" img-format="tif"/></maths></p>
<p id="p0086" num="0086">The normalization factor <i><o ostyle="single">K̅</o><sub>l</sub></i> counts the number of components that are used for the distance measure in each time frame. In order to avoid a division by zero, it is defined as <maths id="math0027" num=""><math display="block"><msub><mover><mi>K</mi><mo>‾</mo></mover><mi>l</mi></msub><mo>=</mo><mi>max</mi><mfenced open="{" close="}" separators=""><mstyle displaystyle="true"><munderover><mo>∑</mo><mrow><mi>μ</mi><mo>=</mo><mn>0</mn></mrow><msub><mi>N</mi><mrow><mi mathvariant="italic">DFT</mi><mo>⁢</mo><msub><mo>/</mo><mn>2</mn></msub></mrow></msub></munderover></mstyle><msub><mi>K</mi><mrow><mi>μ</mi><mo>,</mo><mi>l</mi></mrow></msub><mo>,</mo><mn>0.1</mn></mfenced></math><img id="ib0028" file="imgb0028.tif" wi="55" he="24" img-content="math" img-format="tif"/></maths></p>
<p id="p0087" num="0087">The variable D gives the number of signal frames that are used for the calculation of the LSD, i.e., the number of signal frames with <i><o ostyle="single">K̅</o><sub>l</sub></i> &gt; 0.</p>
<p id="p0088" num="0088">For evaluating the performance of the modified recursive Wiener filter <maths id="math0028" num=""><math display="inline"><msubsup><mi>G</mi><mi mathvariant="italic">rec</mi><mfenced separators=""><mi mathvariant="italic">mod</mi><mspace width="1em"/><mn>1</mn></mfenced></msubsup><mfenced separators=""><msub><mi mathvariant="normal">Ω</mi><mi>μ</mi></msub><mo>⁢</mo><mi>l</mi></mfenced><mo>,</mo></math><img id="ib0029" file="imgb0029.tif" wi="28" he="9" img-content="math" img-format="tif" inline="yes"/></maths><i>,</i> a set of 616 filters has been computed with all possible combinations of the parameters τ<sub>os</sub> ∈ [50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100] ms<br/>
τ <sub>offs</sub> ∈ [0, -5, -10, -15, -20, -25, -30, -35] ms<br/>
α <sub>os</sub> ∈ [0, 2, 4, 6, 8, 10, 12] .</p>
<p id="p0089" num="0089">The signal that has been used is the "Sauerkraut is serve once a week" utterance used throughout this text, mixed with background noise from a car driving at a constant speed of 160km/h. The SNR d ring speech activity is adjusted to 1 dB. Then, only the speech and only the noise components have been processed with these filters and the distances <maths id="math0029" num=""><math display="inline"><msubsup><mi mathvariant="italic">LSD</mi><mi mathvariant="italic">speech</mi><mfenced><mi mathvariant="italic">os</mi></mfenced></msubsup></math><img id="ib0030" file="imgb0030.tif" wi="19" he="8" img-content="math" img-format="tif" inline="yes"/></maths> and <maths id="math0030" num=""><math display="inline"><msubsup><mi mathvariant="italic">LSD</mi><mi mathvariant="italic">noise</mi><mfenced><mi mathvariant="italic">os</mi></mfenced></msubsup></math><img id="ib0031" file="imgb0031.tif" wi="17" he="8" img-content="math" img-format="tif" inline="yes"/></maths> have been calculated. As reference for the comparison, a recursive Wiener filter has been designed and the measures <maths id="math0031" num=""><math display="inline"><msubsup><mi mathvariant="italic">LSD</mi><mi mathvariant="italic">speech</mi><mi mathvariant="italic">rw</mi></msubsup></math><img id="ib0032" file="imgb0032.tif" wi="19" he="8" img-content="math" img-format="tif" inline="yes"/></maths> and <maths id="math0032" num=""><math display="inline"><msubsup><mi mathvariant="italic">LSD</mi><mi mathvariant="italic">noise</mi><mfenced><mi mathvariant="italic">rw</mi></mfenced></msubsup></math><img id="ib0033" file="imgb0033.tif" wi="17" he="8" img-content="math" img-format="tif" inline="yes"/></maths> have been evaluated.</p>
<p id="p0090" num="0090">As mentioned earlier, a good filter is characterized by a small LSD for speech and a large value for noise. Obviously, these two requirements are difficult to meet at the same time and<!-- EPO <DP n="20"> --> an increase in one category usually falls together with an increase in the other one. For better comparison of the two filter types, the <maths id="math0033" num=""><math display="inline"><msubsup><mi mathvariant="italic">LSD</mi><mi mathvariant="italic">speech</mi><mfenced><mi mathvariant="italic">rw</mi></mfenced></msubsup><mo>-</mo><msubsup><mi mathvariant="italic">LSD</mi><mi mathvariant="italic">speech</mi><mfenced><mi mathvariant="italic">os</mi></mfenced></msubsup></math><img id="ib0034" file="imgb0034.tif" wi="37" he="8" img-content="math" img-format="tif" inline="yes"/></maths> and <maths id="math0034" num=""><math display="inline"><msubsup><mi mathvariant="italic">LSD</mi><mi mathvariant="italic">noise</mi><mfenced><mi mathvariant="italic">os</mi></mfenced></msubsup><mo>-</mo><msubsup><mi mathvariant="italic">LSD</mi><mi mathvariant="italic">noise</mi><mfenced><mi mathvariant="italic">rw</mi></mfenced></msubsup></math><img id="ib0035" file="imgb0035.tif" wi="36" he="9" img-content="math" img-format="tif" inline="yes"/></maths> could be calculated They are defined such, that a positive value means that the onset sharpening approach gives better results in the LSD sense. This is apparently not always the case and again it can be seen that an improvement in noise suppression leads to worse results for speech and vice versa. However, there are some combinations where a positive value can be achieved in both disciplines, e.g., for the parameter combination<br/>
τ <sub>os</sub> = 100ms<br/>
τ <sub>offs</sub> = -30 ms.</p>
<p id="p0091" num="0091">The gain factor α<sub>os</sub> basically only scales the distance measure. For the special case of T<sub>os</sub> = 0, the modified filter reduces to the standard recursive Wiener filter and thus gives the same LSD.</p>
<p id="p0092" num="0092">Comparing the noise reduction filter coefficients for a Wiener filter, a recursive Wiener filter and a recursive Wiener filter modified by onset detection for several subbands and different filter parameters gives the following result. The same signal as for the LSD measures has been used for the design of these filters.</p>
<p id="p0093" num="0093">The parameters that have been used are <maths id="math0035" num=""><math display="block"><mfenced open="[" close="]" separators=""><msub><mi>τ</mi><mi mathvariant="italic">os</mi></msub><mo>⁢</mo><msub><mi>τ</mi><mi mathvariant="italic">offs</mi></msub><mo>⁢</mo><msub><mi>α</mi><mi mathvariant="italic">os</mi></msub></mfenced><mo>=</mo><mfenced open="[" close="]" separators=""><mn>75</mn><mo>⁢</mo><mi>ms</mi><mo>,</mo><mo>-</mo><mn>10</mn><mo>⁢</mo><mi>ms</mi><mo>,</mo><mn>3</mn></mfenced></math><img id="ib0036" file="imgb0036.tif" wi="65" he="14" img-content="math" img-format="tif"/></maths></p>
<p id="p0094" num="0094">But of course other parameters are possible, too.</p>
<p id="p0095" num="0095">The Wiener filter opens more often, which potentially results in musical tones. It has also been seen that the recursive Wiener filter opens a bit later, which can be corrected by the onset sharpening modification.</p>
<p id="p0096" num="0096">First of all it should be noticed that the offset τ<i><sub>offs</sub></i> = -30ms is fairly large compared to the duration of <i>τ<sub>os</sub></i> = 100 ms. This causes that the modified filter coefficient near a speech onset increases, then decreases for a few frames and finally grows again with the opening of the recursive Wiener filter. Apparently, this is no optimal behavior even though this parameter combination gave the best results in the LSD. This gives rise to the assumption, that a different prototype function should be used. Good candidates would be more flat around their maximum, avoiding the "crash down" near 0.18 and 0.5 seconds. At any rate, a filter that is opening earlier seems to give benefits in terms of the LSD measure.<!-- EPO <DP n="21"> --></p>
<p id="p0097" num="0097">However, for larger gains α<sub>os</sub>, the noise floor was decreased on the expense of increasing filtering artifacts.</p>
<p id="p0098" num="0098"><figref idref="f0005">Fig. 11</figref> shows a procedure where the attenuation factor and its derivative together with the threshold are shown for mel-band 25 (corresponding to frequencies between 3.6 and 4.4 kHz). Because it could happen that the derivative is greater than y for several consecutive frames, also a sliding time window of 100ms duration is applied. Within this time window, only one detection is allowed. Furthermore, if an onset has been detected in a certain mel-band, the neighboring mel-band towards lower frequencies of the same frame I is also marked as a speech onset.</p>
<p id="p0099" num="0099">The signal that has been used is the utterance "Sauerkraut is served once a week" from the TIMIT database that has also been used before. The clean speech is mixed with background noise recorded in a car driving at a speed of 160 km/h to form an SNR of 1 dB during speech activity.</p>
</description><!-- EPO <DP n="22"> -->
<claims id="claims01" lang="en">
<claim id="c-en-0001" num="0001">
<claim-text>A method for adaptive spectral transformation for acoustic speech signals comprising the steps of<br/>
receiving at least one spectral input representation corresponding to at least one window of a time domain input signal of acoustic speech,<br/>
selecting of the spectral input representations at least one selected spectral representation to be transformed,<br/>
assigning the at least one selected spectral representation to one of a set of cluster centres, wherein<br/>
the cluster centres are defined on the bases of spectral representations of windowed acoustic speech segments of a speech corpus by a clustering algorithm,<br/>
spectral class representations are assigned to the cluster centres and are elements of a code book and<br/>
the code book links to each spectral class representation at least one spectral transformation which enhances the corresponding spectral class representation,<br/>
transforming each selected spectral representation to a spectral output representation, wherein the applied transformation corresponds to the at least one spectral transformation linked to the cluster centre which is assigned to the respective selected spectral representation, and<br/>
providing the one spectral output representations to synthesize an acoustic speech signal.</claim-text></claim>
<claim id="c-en-0002" num="0002">
<claim-text>Method as claimed in claim 1, wherein assigning the at least one selected spectral representation to one of a set of cluster centres includes<br/>
calculating distance measures between the selected spectral representation and all the spectral class representations of the code book, and<br/>
assigning the at least one selected spectral representation to the cluster centre with the shortest distance measures between the selected spectral representation and the spectral class representations of the cluster centre.</claim-text></claim>
<claim id="c-en-0003" num="0003">
<claim-text>Method as claimed in claim 2, wherein calculating distance measures includes calculating feature vectors for the spectral representations and the distances measures are distances between the feature vectors.<!-- EPO <DP n="23"> --></claim-text></claim>
<claim id="c-en-0004" num="0004">
<claim-text>Method as claimed in claim 3, wherein the feature vectors are calculated from the spectral representations by a filtering transformation, preferably with a mel-filterbank, wherein the mel-filterbank optionally uses overlapping triangular windows with widths variable with frequency.</claim-text></claim>
<claim id="c-en-0005" num="0005">
<claim-text>Method as claimed in one of claims 1 to 4, wherein the code book includes at least eight, preferably thirty-two, optionally 128 cluster centres.</claim-text></claim>
<claim id="c-en-0006" num="0006">
<claim-text>Method as claimed in one of claims 1 to 5, wherein the spectral class representations for the cluster centres are averaged spectral representations averaged over spectral representations of corresponding cluster elements respectively classes.</claim-text></claim>
<claim id="c-en-0007" num="0007">
<claim-text>Method as claimed in one of claims 1 to 6, wherein the definition of cluster centres by the clustering algorithm is made on the bases of a preselected sub-corpus of the speech corpus.</claim-text></claim>
<claim id="c-en-0008" num="0008">
<claim-text>Method as claimed in claim 7, wherein preselecting a sub-corpus includes the reduction to spectral representations of the speech corpus which have a spectral centroid lying above a threshold frequency, preferably above 3 kHz.</claim-text></claim>
<claim id="c-en-0009" num="0009">
<claim-text>Method as claimed in one of claims 1 to 8, wherein one of the at least one spectral transformation linked to each spectral class representation of the code book is a spectral compression transformation mapping a frequency interval of the selected spectral representation to a smaller frequency interval of the spectral output representation.</claim-text></claim>
<claim id="c-en-0010" num="0010">
<claim-text>Method as claimed in claim 9, wherein the spectral compression transformation linked to each spectral class representation of the code book is compressing an bandwidth of 0 to 8 kHz to an bandwidth of 0 to 4 kHz preferably with a compression only at the upper or lower end of the bandwidth, optionally linear at least in the middle frequency range, corresponding to no compression when the slope is equal to 1.</claim-text></claim>
<claim id="c-en-0011" num="0011">
<claim-text>Method as claimed in one of claims 1 to 7, wherein one of the at least one spectral transformation linked to each spectral class representation of the code book is a formant boosting gain function.<!-- EPO <DP n="24"> --></claim-text></claim>
<claim id="c-en-0012" num="0012">
<claim-text>Method as claimed in claim 11, wherein the formant boosting gain function linked to each spectral class representation of the code book is amplifying at least one preferably two or three of the formants at low frequencies.</claim-text></claim>
<claim id="c-en-0013" num="0013">
<claim-text>Method as claimed in claim 8, wherein selecting at least one selected spectral representation to be transformed includes calculating the spectral centroid of each spectral input representation and selecting spectral input representations which have a spectral centroid lying above a threshold, frequency preferably above 3 kHz.</claim-text></claim>
<claim id="c-en-0014" num="0014">
<claim-text>Method as claimed in one of claims 1 to 13, wherein selecting at least one selected spectral representation to be transformed includes detecting at least one of speech activity and background noise level and selecting spectral input representations with speech to be transformed.</claim-text></claim>
<claim id="c-en-0015" num="0015">
<claim-text>A computer program comprising program code means for performing all the steps of any one of the claims 1 to 14 when said program is run on a computer.</claim-text></claim>
</claims><!-- EPO <DP n="25"> -->
<drawings id="draw" lang="en">
<figure id="f0001" num="1,2"><img id="if0001" file="imgf0001.tif" wi="165" he="231" img-content="drawing" img-format="tif"/></figure><!-- EPO <DP n="26"> -->
<figure id="f0002" num="3"><img id="if0002" file="imgf0002.tif" wi="165" he="231" img-content="drawing" img-format="tif"/></figure><!-- EPO <DP n="27"> -->
<figure id="f0003" num="4,5,6"><img id="if0003" file="imgf0003.tif" wi="163" he="233" img-content="drawing" img-format="tif"/></figure><!-- EPO <DP n="28"> -->
<figure id="f0004" num="7,8,9"><img id="if0004" file="imgf0004.tif" wi="163" he="233" img-content="drawing" img-format="tif"/></figure><!-- EPO <DP n="29"> -->
<figure id="f0005" num="10,11"><img id="if0005" file="imgf0005.tif" wi="163" he="233" img-content="drawing" img-format="tif"/></figure><!-- EPO <DP n="30"> -->
<figure id="f0006" num="12"><img id="if0006" file="imgf0006.tif" wi="165" he="97" img-content="drawing" img-format="tif"/></figure>
</drawings>
<search-report-data id="srep" lang="en" srep-office="EP" date-produced=""><doc-page id="srep0001" file="srep0001.tif" wi="159" he="233" type="tif"/><doc-page id="srep0002" file="srep0002.tif" wi="159" he="233" type="tif"/></search-report-data>
<ep-reference-list id="ref-list">
<heading id="ref-h0001"><b>REFERENCES CITED IN THE DESCRIPTION</b></heading>
<p id="ref-p0001" num=""><i>This list of references cited by the applicant is for the reader's convenience only. It does not form part of the European patent document. Even though great care has been taken in compiling the references, errors or omissions cannot be excluded and the EPO disclaims all liability in this regard.</i></p>
<heading id="ref-h0002"><b>Patent documents cited in the description</b></heading>
<p id="ref-p0002" num="">
<ul id="ref-ul0001" list-style="bullet">
<li><patcit id="ref-pcit0001" dnum="US5771299A"><document-id><country>US</country><doc-number>5771299</doc-number><kind>A</kind></document-id></patcit><crossref idref="pcit0001">[0014]</crossref></li>
<li><patcit id="ref-pcit0002" dnum="US20090074197A"><document-id><country>US</country><doc-number>20090074197</doc-number><kind>A</kind></document-id></patcit><crossref idref="pcit0002">[0016]</crossref></li>
<li><patcit id="ref-pcit0003" dnum="EP1333700A2"><document-id><country>EP</country><doc-number>1333700</doc-number><kind>A2</kind></document-id></patcit><crossref idref="pcit0003">[0017]</crossref></li>
<li><patcit id="ref-pcit0004" dnum="US20060226016A"><document-id><country>US</country><doc-number>20060226016</doc-number><kind>A</kind></document-id></patcit><crossref idref="pcit0004">[0018]</crossref></li>
<li><patcit id="ref-pcit0005" dnum="US20090226016A"><document-id><country>US</country><doc-number>20090226016</doc-number><kind>A</kind></document-id></patcit><crossref idref="pcit0005">[0019]</crossref></li>
<li><patcit id="ref-pcit0006" dnum="US6577739B"><document-id><country>US</country><doc-number>6577739</doc-number><kind>B</kind></document-id></patcit><crossref idref="pcit0006">[0020]</crossref></li>
<li><patcit id="ref-pcit0007" dnum="CA2569221"><document-id><country>CA</country><doc-number>2569221</doc-number></document-id></patcit><crossref idref="pcit0007">[0021]</crossref></li>
<li><patcit id="ref-pcit0008" dnum="WO20080123886A"><document-id><country>WO</country><doc-number>20080123886</doc-number><kind>A</kind></document-id></patcit><crossref idref="pcit0008">[0022]</crossref></li>
</ul></p>
<heading id="ref-h0003"><b>Non-patent literature cited in the description</b></heading>
<p id="ref-p0003" num="">
<ul id="ref-ul0002" list-style="bullet">
<li><nplcit id="ref-ncit0001" npl-type="s"><article><author><name>P. PATRICK</name></author><author><name>R. STEELE</name></author><author><name>C. XYDEAS</name></author><atl/><serial><sertitle>Frequency compression of 7.6 kHz speech into 3.3 kHz bandwidth</sertitle><pubdate><sdate>19830500</sdate><edate/></pubdate><vid>31</vid><ino>5</ino></serial><location><pp><ppf>692</ppf><ppl>701</ppl></pp></location></article></nplcit><crossref idref="ncit0001">[0004]</crossref></li>
<li><nplcit id="ref-ncit0002" npl-type="s"><article><author><name>D. A. HEIDE</name></author><author><name>G. S. KANG</name></author><atl>Speech enhancement for bandlimited speech</atl><serial><sertitle>Proc. IEEE International Conference on Acoustics, Speech and Signal Processing</sertitle><pubdate><sdate>19980512</sdate><edate/></pubdate><vid>1</vid></serial><location><pp><ppf>393</ppf><ppl>396</ppl></pp></location></article></nplcit><crossref idref="ncit0002">[0007]</crossref></li>
<li><nplcit id="ref-ncit0003" npl-type="s"><article><author><name>ANDREA SIMPSON</name></author><author><name>ADAM A. HERSBACH</name></author><author><name>HUGH J. MCDERMOTT</name></author><atl>Improvements in speech perception with an experimental nonlinear frequency compression hearing device</atl><serial><sertitle>International Journal of Audiology</sertitle><pubdate><sdate>20060000</sdate><edate/></pubdate><vid>44</vid><ino>5</ino></serial><location><pp><ppf>281</ppf><ppl>292</ppl></pp></location></article></nplcit><crossref idref="ncit0003">[0011]</crossref></li>
<li><nplcit id="ref-ncit0004" npl-type="s"><article><author><name>H. J. MCDERMOTT</name></author><author><name>V. P. DORKOS</name></author><author><name>M. R. DEAN</name></author><author><name>T. Y. CHING</name></author><atl>Improvements in speech perception with use of the avr transonic frequency-transposing hearing aid</atl><serial><sertitle>J Speech Hear Lang Res</sertitle><pubdate><sdate>19990000</sdate><edate/></pubdate><vid>42</vid><ino>6</ino></serial><location><pp><ppf>1323</ppf><ppl>1335</ppl></pp></location></article></nplcit><crossref idref="ncit0004">[0012]</crossref></li>
<li><nplcit id="ref-ncit0005" npl-type="s"><article><author><name>KUK</name></author><author><name>KORHONEN</name></author><author><name>PEETERS</name></author><author><name>JESSEN</name></author><author><name>ANDERSEN</name></author><atl>Linear frequency transposition: Extending the audibility of high frequency information</atl><serial><sertitle>The Hearing Review</sertitle><pubdate><sdate>20060000</sdate><edate/></pubdate></serial></article></nplcit><crossref idref="ncit0005">[0012]</crossref></li>
</ul></p>
</ep-reference-list>
</ep-patent-document>
