<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE ep-patent-document PUBLIC "-//EPO//EP PATENT DOCUMENT 1.4//EN" "ep-patent-document-v1-4.dtd">
<ep-patent-document id="EP03762154B1" file="EP03762154NWB1.xml" lang="en" country="EP" doc-number="1518096" kind="B1" date-publ="20140423" status="n" dtd-version="ep-patent-document-v1-4">
<SDOBI lang="en"><B000><eptags><B001EP>......DE....FRGB..........SE............FI..........................................................</B001EP><B003EP>*</B003EP><B005EP>J</B005EP><B007EP>DIM360 Ver 2.40 (30 Jan 2013) -  2100000/0</B007EP><B010EP><B011EP><date>20130410</date><dnum><text>01</text></dnum><ctry>DE</ctry><ctry>FI</ctry><ctry>FR</ctry><ctry>GB</ctry><ctry>SE</ctry></B011EP></B010EP></eptags></B000><B100><B110>1518096</B110><B120><B121>EUROPEAN PATENT SPECIFICATION</B121></B120><B130>B1</B130><B140><date>20140423</date></B140><B190>EP</B190></B100><B200><B210>03762154.7</B210><B220><date>20030627</date></B220><B240><B241><date>20040301</date></B241><B242><date>20130301</date></B242></B240><B250>en</B250><B251EP>en</B251EP><B260>en</B260></B200><B300><B310>186862</B310><B320><date>20020701</date></B320><B330><ctry>US</ctry></B330></B300><B400><B405><date>20140423</date><bnum>201417</bnum></B405><B430><date>20050330</date><bnum>200513</bnum></B430><B450><date>20140423</date><bnum>201417</bnum></B450><B452EP><date>20131030</date></B452EP></B400><B500><B510EP><classification-ipcr sequence="1"><text>G10L  25/69        20130101AFI20130816BHEP        </text></classification-ipcr></B510EP><B540><B541>de</B541><B542>ÄUSSERUNGSABHÄNGIGE-AUSSPRACHEKOMPENSATION FÜR SPRACHQUALITÄTSBEWERTUNG</B542><B541>en</B541><B542>COMPENSATION FOR UTTERANCE DEPENDENT ARTICULATION FOR SPEECH QUALITY ASSESSMENT</B542><B541>fr</B541><B542>COMPENSATION DESTINEE A L'ARTICULATION DEPENDANTE DE L'ENONCIATION POUR L'EVALUATION DE LA QUALITE VOCALE</B542></B540><B560><B561><text>EP-A- 1 187 100</text></B561><B561><text>US-A- 4 352 182</text></B561><B561><text>US-A- 5 794 188</text></B561></B560></B500><B700><B720><B721><snm>KIM, Doh-Suk</snm><adr><str>42 Huntington Road</str><city>Basking Ridge, NJ 07920</city><ctry>US</ctry></adr></B721></B720><B730><B731><snm>Alcatel Lucent</snm><iid>101311164</iid><irf>Kim 3 (D)-EP</irf><adr><str>3, Avenue Octave Gréard</str><city>75007 Paris</city><ctry>FR</ctry></adr></B731></B730><B740><B741><snm>Wetzel, Emmanuelle</snm><sfx>et al</sfx><iid>100818093</iid><adr><str>Alcatel Lucent 
Intellectual Property &amp; Standards</str><city>70430 Stuttgart</city><ctry>DE</ctry></adr></B741></B740></B700><B800><B840><ctry>DE</ctry><ctry>FI</ctry><ctry>FR</ctry><ctry>GB</ctry><ctry>SE</ctry></B840><B860><B861><dnum><anum>US2003020354</anum></dnum><date>20030627</date></B861><B862>en</B862></B860><B870><B871><dnum><pnum>WO2004003499</pnum></dnum><date>20040108</date><bnum>200402</bnum></B871></B870></B800></SDOBI>
<description id="desc" lang="en"><!-- EPO <DP n="1"> -->
<heading id="h0001"><u>Field of the Invention</u></heading>
<p id="p0001" num="0001">The present invention relates generally to communications systems and, in particular, to speech quality assessment.</p>
<heading id="h0002"><u>Background of the Related Art</u></heading>
<p id="p0002" num="0002">Performance of a wireless communication system can be measured, among other things, in terms of speech quality. In the current art, there are two techniques of speech quality assessment. The first technique is a subjective technique (hereinafter referred to as "subjective speech quality assessment"). In subjective speech quality assessment, human listeners are used to rate the speech quality of processed speech, wherein processed speech is a transmitted speech signal which has been processed at the receiver. This technique is subjective because it is based on the perception of the individual human, and human assessment of speech quality typically takes into account phonetic contents, speaking styles or individual speaker differences. Subjective speech quality assessment can be expensive and time consuming.</p>
<p id="p0003" num="0003">The second technique is an objective technique (hereinafter referred to as "objective speech quality assessment"). Objective speech quality assessment is not based on the perception of the individual human. Most objective speech quality assessment techniques are based on known source speech or reconstructed source speech estimated from processed speech. However, these objective techniques do not account for phonetic contents, speaking styles or individual speaker differences.</p>
<p id="p0004" num="0004">Accordingly, there exists a need for assessing speech quality objectively which takes into account phonetic contents, speaking styles or individual speaker differences.<!-- EPO <DP n="2"> --></p>
<p id="p0005" num="0005"><patcit id="pcit0001" dnum="EP1187100A1"><text>EP 1187 100 A1</text></patcit> discloses a method for objective speech quality assessment without reference signal.</p>
<p id="p0006" num="0006"><patcit id="pcit0002" dnum="US4352182A"><text>US 4,352,182</text></patcit> discloses a method for testing the quality of digital speech-transmission equipment.</p>
<heading id="h0003"><u>Summary of the Invention</u></heading>
<p id="p0007" num="0007">The present invention is a method for objective speech quality assessment that accounts for phonetic contents, speaking styles or individual speaker differences by distorting speech signals under speech quality assessment. By using a<!-- EPO <DP n="3"> --> distorted version of a speech signal, it is possible to compensate for different phonetic contents, different individual speakers and different speaking styles when assessing speech quality. The amount of degradation in the objective speech quality assessment by distorting the speech signal is maintained similarly for different speech signals, especially when the amount of distortion of the distorted version of speech signal is severe. Objective speech quality assessment for the distorted speech signal and the original undistorted speech signal are compared to obtain a speech quality assessment compensated for utterance dependent articulation. In one embodiment, the comparison corresponds to a difference between the objective speech quality assessments for the distorted and undistorted speech signals.</p>
<heading id="h0004"><u>Brief Description of the Drawings</u></heading>
<p id="p0008" num="0008">The features, aspects, and advantages of the present invention will become better understood with regard to the following description, appended claims, and accompanying drawings where:
<ul id="ul0001" list-style="none" compact="compact">
<li><figref idref="f0001">Fig. 1</figref> depicts an objective speech quality assessment arrangement which compensates for utterance dependent articulation in accordance with the present invention;</li>
<li><figref idref="f0002">Fig. 2</figref> depicts an embodiment of an objective speech quality assessment module employing an auditory-articulatory analysis module in accordance with the present invention.;</li>
<li><figref idref="f0002">Fig. 3</figref> depicts a flowchart for processing, in an articulatory analysis module, the plurality of envelopes a<sub>i</sub>(t) in accordance with one embodiment of the invention; and</li>
<li><figref idref="f0003">Fig. 4</figref> depicts an example illustrating a modulation spectrum A<sub>i</sub>(m,f) in terms of power versus frequency.</li>
</ul></p>
<heading id="h0005"><u>Detailed Description</u></heading>
<p id="p0009" num="0009">The present invention is a method for objective speech quality assessment that accounts for phonetic contents, speaking styles or individual speaker differences by distorting processed speech. Objective speech quality assessment tend to yield different values for different speech signals which have same subjective<!-- EPO <DP n="4"> --> speech quality scores. The reason these values differ is because of different distributions of spectral contents in the modulation spectral domain. By using a distorted version of a processed speech signal, it is possible to compensate for different phonetic contents, different individual speakers and different speaking styles. The amount of degradation in the objective speech quality assessment by distorting the speech signal is maintained similarly for different speech signals, especially when the distortion is severe. Objective speech quality assessment for the distorted speech signal and the original undistorted speech signal are compared to obtain a speech quality assessment compensated for utterance dependent articulation.</p>
<p id="p0010" num="0010"><figref idref="f0001">Fig. 1</figref> depicts an objective speech quality assessment arrangement 10 which compensates for utterance dependent articulation in accordance with the present invention. Objective speech quality assessment arrangement 10 comprises a plurality of objective speech quality assessment modules 12, 14, a distortion module 16 and a compensation utterance-specific bias module 18. Speech signal s(t) is provided as inputs to distortion module 16 and objective speech quality assessment module 12. In distortion module 16, speech signal s(t) is distorted to produce a modulated noise reference unit (MNRU) speech signal s'(t). In other words, distortion module 16 produces a noisy version of input signal s(t). MNRU speech signal s'(t) is then provided as input to objective speech quality assessment module 14.</p>
<p id="p0011" num="0011">In objective speech quality assessment modules 12, 14, speech signal s(t) and MNRU speech signal s'(t) are processed to obtain objective speech quality assessments SQ(s(t) and SQ(s'(t)). Objective speech quality assessment modules 12, 14 are essentially identical in terms of the type of processing performed to any input speech signals. That is, if both objective speech quality assessment modules 12, 14 receive the same input speech signal, the output signals of both modules 12, 14 would be approximately identical. Note that, in other embodiments, objective speech quality assessment modules 12, 14 may process speech signals s(t) and s'(t) in a manner different from each other. Objective speech quality assessment modules are well-known in the art. An example of such a module will be described later herein.</p>
<p id="p0012" num="0012">Objective speech quality assessments SQ(s(t) and SQ(s'(t)) are then compared to obtain speech quality assessment SQ<sub>compensated</sub>, which compensates for<!-- EPO <DP n="5"> --> utterance dependent articulation. In one embodiment, speech quality assessment SQ<sub>compensated</sub> is determined using the difference between objective speech quality assessments SQ(s(t) and SQ(s'(t)). For example, SQ<sub>compensated</sub> is equal to SQ(s(t) minus SQ(s'(t)), or vice-versa. In another embodiment, speech quality assessment SQ<sub>compensated</sub> is determined based on a ratio between objective speech quality assessments SQ(s(t) and SQ(s'(t)). For example, <maths id="math0001" num=""><math display="block"><msub><mi>SQ</mi><mi>compensated</mi></msub><mo>=</mo><mfrac><mrow><mi>SQ</mi><mfenced separators=""><mi mathvariant="normal">s</mi><mfenced><mi mathvariant="normal">t</mi></mfenced></mfenced><mo>+</mo><mi mathvariant="normal">μ</mi></mrow><mrow><mi>SQ</mi><mfenced separators=""><mi mathvariant="normal">sʹ</mi><mfenced><mi mathvariant="normal">t</mi></mfenced></mfenced><mo>+</mo><mi mathvariant="normal">μ</mi></mrow></mfrac><mspace width="1em"/><mi>or</mi><mspace width="1em"/><msub><mi>SQ</mi><mi>compensated</mi></msub><mo>=</mo><mfrac><mrow><mi>SQ</mi><mfenced separators=""><mi mathvariant="normal">sʹ</mi><mfenced><mi mathvariant="normal">t</mi></mfenced></mfenced><mo>+</mo><mi mathvariant="normal">μ</mi></mrow><mrow><mi>SQ</mi><mfenced separators=""><mi mathvariant="normal">s</mi><mfenced><mi mathvariant="normal">t</mi></mfenced></mfenced><mo>+</mo><mi mathvariant="normal">μ</mi></mrow></mfrac></math><img id="ib0001" file="imgb0001.tif" wi="104" he="15" img-content="math" img-format="tif"/></maths> where µ is a small constant value.</p>
<p id="p0013" num="0013">As mentioned earlier, objective speech quality assessment modules 12, 14 are well known in the art. <figref idref="f0002">Fig. 2</figref> depicts an embodiment 20 of an objective speech quality assessment module 12, 14 employing an auditory-articulatory analysis module in accordance with the present invention. As shown in <figref idref="f0002">Fig. 2</figref>, objective quality assessment module 20 comprises of cochlear filterbank 22, envelope analysis module 24 and articulatory analysis module 26. In objective quality assessment module 20, speech signal s(t) is provided as input to cochlear filterbank 22. Cochlear filterbank 22 comprises a plurality of cochlear filters h<sub>i</sub>(t) for processing speech signal s(t) in accordance with a first stage of a peripheral auditory system, where i=1,2,...,N<sub>c</sub> represents a particular cochlear filter channel and N<sub>c</sub> denotes the total number of cochlear filter channels. Specifically, cochlear filterbank 22 filters speech signal s(t) to produce a plurality of critical band signals s<sub>i</sub>(t), wherein critical band signal s<sub>i</sub>(t) is equal to s(t)*h<sub>i</sub>(t).</p>
<p id="p0014" num="0014">The plurality of critical band signals s<sub>i</sub>(t) is provided as input to envelope analysis module 24. In envelope analysis module 24, the plurality of critical band signals s<sub>i</sub>(t) is processed to obtain a plurality of envelopes a<sub>i</sub>(t), wherein <maths id="math0002" num=""><math display="inline"><msub><mi mathvariant="normal">a</mi><mi mathvariant="normal">i</mi></msub><mfenced><mi mathvariant="normal">t</mi></mfenced><mo>=</mo><msqrt><msubsup><mi mathvariant="normal">s</mi><mi mathvariant="normal">i</mi><mn mathvariant="normal">2</mn></msubsup><mfenced><mi mathvariant="normal">t</mi></mfenced><mo>+</mo><msubsup><mover><mi mathvariant="normal">s</mi><mo>^</mo></mover><mi mathvariant="normal">i</mi><mn mathvariant="normal">2</mn></msubsup><mfenced><mi mathvariant="normal">t</mi></mfenced></msqrt></math><img id="ib0002" file="imgb0002.tif" wi="34" he="9" img-content="math" img-format="tif" inline="yes"/></maths> and ŝ<sub>i</sub>(t) is the Hilbert transform of s<sub>i</sub>(t).</p>
<p id="p0015" num="0015">The plurality of envelopes a<sub>i</sub>(t) is then provided as input to articulatory analysis module 26. In articulatory analysis module 26, the plurality of envelopes a<sub>i</sub>(t) is processed to obtain a speech quality assessment for speech signal s(t). Specifically, articulatory analysis module 26 does a comparison of the power associated with signals generated from the human articulatory system (hereinafter referred to as "articulation power P<sub>A</sub>(m,i)") with the power associated with signals not<!-- EPO <DP n="6"> --> generated from the human articulatory system (hereinafter referred to as "non-articulation power P<sub>NA</sub>(m,i)"). Such comparison is then used to make a speech quality assessment.</p>
<p id="p0016" num="0016"><figref idref="f0002">Fig. 3</figref> depicts a flowchart 300 for processing, in articulatory analysis module 26, the plurality of envelopes a<sub>i</sub>(t) in accordance with one embodiment of the invention. In step 310, Fourier transform is performed on frame m of each of the plurality of envelopes a<sub>i</sub>(t) to produce modulation spectrums A<sub>i</sub>(m,f), where f is frequency.</p>
<p id="p0017" num="0017"><figref idref="f0003">Fig. 4</figref> depicts an example 40 illustrating modulation spectrum A<sub>i</sub>(m,f) in terms of power versus frequency. In example 40, articulation power P<sub>A</sub>(m,i) is the power associated with frequencies 2∼12.5 Hz, and non-articulation power P<sub>NA</sub>(m,i) is the power associated with frequencies greater than 12.5 Hz. Power P<sub>No</sub>(m,i) associated with frequencies less than 2 Hz is the DC-component of frame m of critical band signal a<sub>i</sub>(t). In this example, articulation power P<sub>A</sub>(m,i) is chosen as the power associated with frequencies 2∼12.5 Hz based on the fact that the speed of human articulation is 2∼12.5 Hz, and the frequency ranges associated with articulation power P<sub>A(</sub>m,i) and non-articulation power P<sub>NA</sub>(m,i) (hereinafter referred to respectively as "articulation frequency range" and "non-articulation frequency range") are adjacent, non-overlapping frequency ranges. It should be understood that, for purposes of this application, the term "articulation power P<sub>A</sub>(m,i)" should not be limited to the frequency range of human articulation or the aforementioned frequency range 2∼12.5 Hz. Likewise, the term "non-articulation power P<sub>NA</sub>(m,i)" should not be limited to frequency ranges greater than the frequency range associated with articulation power P<sub>A</sub>(m,i). The non-articulation frequency range may or may not overlap with or be adjacent to the articulation frequency range. The non-articulation frequency range may also include frequencies less than the lowest frequency in the articulation frequency range, such as those associated with the DC-component of frame m of critical band signal a<sub>i</sub>(t).</p>
<p id="p0018" num="0018">In step 320, for each modulation spectrum A<sub>i</sub>(m,f), articulatory analysis module 26 performs a comparison between articulation power P<sub>A(</sub>m,i) and non-articulation power P<sub>NA</sub>(m,i). In this embodiment of articulatory analysis module 26, the comparison between articulation power P<sub>A</sub>(m,i) and non-articulation power<!-- EPO <DP n="7"> --> P<sub>NA</sub>(m,i) is an articulation-to-non-articulation ratio ANR(m,i). The ANR is defined by the following equation <maths id="math0003" num="equation (1)"><math display="block"><mi>ANR</mi><mfenced><mi mathvariant="normal">m</mi><mi mathvariant="normal">i</mi></mfenced><mo>=</mo><mfrac><mrow><msub><mi mathvariant="normal">P</mi><mi mathvariant="normal">A</mi></msub><mfenced><mi mathvariant="normal">m</mi><mi mathvariant="normal">i</mi></mfenced><mo>+</mo><mi mathvariant="normal">ε</mi></mrow><mrow><msub><mi mathvariant="normal">P</mi><mi>NA</mi></msub><mfenced><mi mathvariant="normal">m</mi><mi mathvariant="normal">i</mi></mfenced><mo>+</mo><mi mathvariant="normal">ε</mi></mrow></mfrac></math><img id="ib0003" file="imgb0003.tif" wi="104" he="15" img-content="math" img-format="tif"/></maths> where ε is some small constant value. Other comparisons between articulation power P<sub>A</sub>(m,i) and non-articulation power P<sub>NA</sub>(m,i) are possible. For example, the comparison may be the reciprocal of equation (1), or the comparison may be a difference between articulation power P<sub>A(</sub>m,i) and non-articulation power P<sub>NA</sub>(m,i). For ease of discussion, the embodiment of articulatory analysis module 26 depicted by flowchart 300 will be discussed with respect to the comparison using ANR(m,i) of equation (1). This should not, however, be construed to limit the present invention in any manner.</p>
<p id="p0019" num="0019">In step 330, ANR(m,i) is used to determine local speech quality LSQ(m) for frame m. Local speech quality LSQ(m) is determined using an aggregate of the articulation-to-non-articulation ratio ANR(m,i) across all channels i and a weighing factor R(m,i) based on the DC-component power P<sub>No</sub>(m,i). Specifically, local speech quality LSQ(m) is determined using the following equation <maths id="math0004" num="equation (2)"><math display="block"><mi>LSQ</mi><mfenced><mi mathvariant="normal">m</mi></mfenced><mo>=</mo><mi>log</mi><mfenced open="[" close="]" separators=""><mstyle displaystyle="true"><munderover><mo>∑</mo><mrow><mi mathvariant="normal">i</mi><mo>=</mo><mn mathvariant="normal">1</mn></mrow><msub><mi mathvariant="normal">N</mi><mi mathvariant="normal">c</mi></msub></munderover></mstyle><mi>ANR</mi><mfenced><mi mathvariant="normal">m</mi><mi mathvariant="normal">i</mi></mfenced><mo>⁢</mo><mi mathvariant="normal">R</mi><mfenced><mi mathvariant="normal">m</mi><mi mathvariant="normal">i</mi></mfenced></mfenced></math><img id="ib0004" file="imgb0004.tif" wi="105" he="18" img-content="math" img-format="tif"/></maths> where <maths id="math0005" num="equation (3)"><math display="block"><mi mathvariant="normal">R</mi><mfenced><mi mathvariant="normal">m</mi><mi mathvariant="normal">i</mi></mfenced><mo>=</mo><mfrac><mrow><mi>log</mi><mrow><mo>(</mo><mn mathvariant="normal">1</mn><mo>+</mo><msub><mi mathvariant="normal">P</mi><mi>No</mi></msub><mfenced><mi mathvariant="normal">m</mi><mi mathvariant="normal">i</mi></mfenced></mrow></mrow><mrow><mstyle displaystyle="true"><munderover><mo>∑</mo><mrow><mi mathvariant="normal">k</mi><mo>=</mo><mn mathvariant="normal">1</mn></mrow><mi>Nc</mi></munderover></mstyle><mi>log</mi><mrow><mo>(</mo><mn mathvariant="normal">1</mn><mo>+</mo><msub><mi mathvariant="normal">P</mi><mi>No</mi></msub><mfenced><mi mathvariant="normal">m</mi><mi mathvariant="normal">k</mi></mfenced></mrow></mrow></mfrac></math><img id="ib0005" file="imgb0005.tif" wi="106" he="23" img-content="math" img-format="tif"/></maths> and k is a frequency index.</p>
<p id="p0020" num="0020">In step 340, overall speech quality SQ for speech signal s(t) is determined using local speech quality LSQ(m) and a log power P<sub>s</sub>(m) for frame m. Specifically, speech quality SQ is determined using the following equation <maths id="math0006" num="equation (4)"><math display="block"><mi>SQ</mi><mo>=</mo><mi>L</mi><mo>⁢</mo><msubsup><mfenced open="{" close="}" separators=""><msub><mi mathvariant="normal">P</mi><mi mathvariant="normal">s</mi></msub><mfenced><mi mathvariant="normal">m</mi></mfenced><mo>⁢</mo><mi>LSQ</mi><mfenced><mi mathvariant="normal">m</mi></mfenced></mfenced><mrow><mi mathvariant="normal">m</mi><mo>=</mo><mn mathvariant="normal">1</mn></mrow><mi mathvariant="normal">T</mi></msubsup><mo>⁢</mo><msup><mfenced open="[" close="]" separators=""><mstyle displaystyle="true"><munderover><mo>∑</mo><mtable columnalign="left"><mtr><mtd><mi mathvariant="normal">m</mi><mo>=</mo><mn mathvariant="normal">1</mn></mtd></mtr><mtr><mtd><msub><mi mathvariant="normal">P</mi><mi mathvariant="normal">s</mi></msub><mo>&gt;</mo><msub><mi mathvariant="normal">P</mi><mi>th</mi></msub></mtd></mtr></mtable><mi mathvariant="normal">T</mi></munderover></mstyle><msubsup><mi mathvariant="normal">P</mi><mi mathvariant="normal">s</mi><mi mathvariant="normal">λ</mi></msubsup><mfenced><mi mathvariant="normal">m</mi></mfenced><mo>⁢</mo><msup><mi>LSQ</mi><mi mathvariant="normal">λ</mi></msup><mfenced><mi mathvariant="normal">m</mi></mfenced></mfenced><mmultiscripts><msub><mo>/</mo><mi mathvariant="normal">λ</mi></msub><mprescripts/><none/><mn mathvariant="normal">1</mn></mmultiscripts></msup></math><img id="ib0006" file="imgb0006.tif" wi="130" he="25" img-content="math" img-format="tif"/></maths><!-- EPO <DP n="8"> --> where <maths id="math0007" num=""><math display="inline"><msub><mi mathvariant="normal">P</mi><mi mathvariant="normal">s</mi></msub><mfenced><mi mathvariant="normal">m</mi></mfenced><mo>=</mo><mi>log</mi><mfenced open="[" close="]" separators=""><mstyle displaystyle="true"><munder><mo>∑</mo><mrow><mi mathvariant="normal">t</mi><mo>⁢</mo><mover><mi mathvariant="normal">I</mi><mo>^</mo></mover><mo>⁢</mo><mi mathvariant="normal">m</mi></mrow></munder></mstyle><msup><mi mathvariant="normal">s</mi><mn mathvariant="normal">2</mn></msup><mfenced><mi mathvariant="normal">t</mi></mfenced></mfenced><mo>,</mo></math><img id="ib0007" file="imgb0007.tif" wi="39" he="15" img-content="math" img-format="tif" inline="yes"/></maths> <i>L</i> is L<sub>p</sub>-norm, T is the total number of frames in speech signal s(t), λ is any value, and P<sub>th</sub> is a threshold for distinguishing between audible signals and silence. In one embodiment, λ is preferably an odd integer value.</p>
<p id="p0021" num="0021">The output of articulatory analysis module 26 is an assessment of speech quality SQ over all frames m. That is, speech quality SQ is a speech quality assessment for speech signal s(t).</p>
<p id="p0022" num="0022">Although the present invention has been described in considerable detail with reference to certain embodiments, other versions are possible. Therefore, the scope of the present invention should not be limited to the description of the embodiments contained herein.</p>
</description>
<claims id="claims01" lang="en"><!-- EPO <DP n="9"> -->
<claim id="c-en-01-0001" num="0001">
<claim-text>A method of assessing speech quality comprising the steps of:
<claim-text>determining a first and second speech quality assessment for first and second speech quality signals (SQ(s(t)), SQ(s'(t))) based, respectively, on first and second speech signal, the second speech signals being a processed speech signal, and the first speech signal being a distorted (MNRU) version of the second speech signal; and</claim-text>
<claim-text>obtaining a compensated speech SQ compensated quality assessment by comparing the first and second speech quality assessments.</claim-text></claim-text></claim>
<claim id="c-en-01-0002" num="0002">
<claim-text>The method of claim 1 comprising the additional steps of<br/>
prior to determining the first and second speech quality assessments, distorting the second speech signal to produce the first speech signal.</claim-text></claim>
<claim id="c-en-01-0003" num="0003">
<claim-text>The method of claim 1, wherein the first and second speech qualities are assessed using an identical technique for objective speech quality assessment.</claim-text></claim>
<claim id="c-en-01-0004" num="0004">
<claim-text>The method of claim 1, wherein the compensated speech quality assessment corresponds to a difference between the first and second speech qualities.</claim-text></claim>
<claim id="c-en-01-0005" num="0005">
<claim-text>The method of claim 1, wherein the compensated speech quality assessment corresponds to a ratio between the first and second speech qualities.</claim-text></claim>
<claim id="c-en-01-0006" num="0006">
<claim-text>The method of claim 1, wherein the first and second speech qualities are assessed using auditory-articulatory analysis.</claim-text></claim>
<claim id="c-en-01-0007" num="0007">
<claim-text>The method of claim 1, wherein the step assessing the second or first speech quality comprises the steps of;<br/>
<!-- EPO <DP n="10"> -->comparing articulation power and non-articulation power for the speech signal or distorted speech signal, wherein articulation P<sub>A</sub> and non-articulation P<sub>NA</sub> powers are powers associated with articulation and non-articulation frequencies of the speech signal or distorted speech signal; and<br/>
and assessing the second or first speech quality based on the comparison.</claim-text></claim>
<claim id="c-en-01-0008" num="0008">
<claim-text>The method of claim 7, wherein the articulation frequencies are approximately 2∼12.5 Hz.</claim-text></claim>
<claim id="c-en-01-0009" num="0009">
<claim-text>The method of claim 7, wherein the articulation frequencies correspond approximately to a speed of human articulation.</claim-text></claim>
<claim id="c-en-01-0010" num="0010">
<claim-text>The method of claim 7, wherein the non-articulation frequencies are approximately greater than the articulation frequencies.</claim-text></claim>
<claim id="c-en-01-0011" num="0011">
<claim-text>The method of claim 7, wherein the comparison between the articulation power and non-articulation power is a ratio between the articulation power and non-articulation power.</claim-text></claim>
<claim id="c-en-01-0012" num="0012">
<claim-text>The method of claim 10, wherein the ratio includes a denominator and numerator, the numerator including the articulation power and a small constant, the denominator including the non-articulation power plus the small constant.</claim-text></claim>
<claim id="c-en-01-0013" num="0013">
<claim-text>The method of claim 7, wherein the comparison between the articulation power and non-articulation power is a difference between the articulation power and non-articulation power.</claim-text></claim>
<claim id="c-en-01-0014" num="0014">
<claim-text>The method of claim 7, wherein the step of assessing the first or second speech quality includes the step of:
<claim-text>determining a local speech quality using the comparison.</claim-text><!-- EPO <DP n="11"> --></claim-text></claim>
<claim id="c-en-01-0015" num="0015">
<claim-text>The method of claim 7, wherein the local speech quality is further determined using a weighing factor based on a DC-component power P<sub>NO</sub>.</claim-text></claim>
<claim id="c-en-01-0016" num="0016">
<claim-text>The method of claim 9, wherein the first or second speech quality is determined using the local speech quality.</claim-text></claim>
<claim id="c-en-01-0017" num="0017">
<claim-text>The method of claim 7, wherein the step of comparing articulation power and non-articulation power includes the step of:
<claim-text>performing a Fourier transform on each of a plurality of envelopes obtained from a plurality of critical band signals.</claim-text></claim-text></claim>
<claim id="c-en-01-0018" num="0018">
<claim-text>The method of claim 7, wherein the step of comparing articulation power and non-articulation power includes the step of:
<claim-text>filtering the speech signal to obtain a plurality of critical band signals.</claim-text></claim-text></claim>
<claim id="c-en-01-0019" num="0019">
<claim-text>The method of claim 18, wherein the step of comparing articulation power and non-articulation power includes the step of:
<claim-text>performing an envelope analysis on the plurality of critical band signals to obtain a plurality of modulation spectrums.</claim-text></claim-text></claim>
<claim id="c-en-01-0020" num="0020">
<claim-text>The method of claim 18, wherein the step of comparing articulation power and non-articulation power includes the step of:
<claim-text>performing a Fourier transform on each of the plurality of modulation spectrums.</claim-text></claim-text></claim>
</claims>
<claims id="claims02" lang="de"><!-- EPO <DP n="12"> -->
<claim id="c-de-01-0001" num="0001">
<claim-text>Verfahren zur Bewertung der Sprachqualität, die folgenden Schritte umfassend:
<claim-text>Bestimmen einer ersten und einer zweiten Sprachqualitätsbewertung für erste und zweite Sprachqualitätssignale, jeweils ein (SQ(s(t), SQ(s'(t))-basiertes erstes und zweites Sprachsignal, wobei das zweite Sprachsignal ein verarbeitetes Sprachsignal ist und das erste Sprachsignal eine verzerrte (MNRU)-Version des zweiten Sprachsignals ist; und</claim-text>
<claim-text>Gewinnen einer SQ-kompensierten Sprachqualitätsbewertung durch Vergleichen der ersten und der zweiten Sprachqualitätsbewertungen.</claim-text></claim-text></claim>
<claim id="c-de-01-0002" num="0002">
<claim-text>Verfahren nach Anspruch 1, den folgenden zusätzlichen Schritt umfassend:
<claim-text>Vor Bestimmen der ersten und zweiten Sprachqualitätsbewertungen, Verzerren des zweiten Sprachsignals, um das erste Sprachsignal zu erzeugen.</claim-text></claim-text></claim>
<claim id="c-de-01-0003" num="0003">
<claim-text>Verfahren nach Anspruch 1, wobei die erste und die zweite Sprachqualität unter Verwendung einer identischen Technik für objektive Sprachqualitätsbewertung bewertet werden.</claim-text></claim>
<claim id="c-de-01-0004" num="0004">
<claim-text>Verfahren nach Anspruch 1, wobei die kompensierte Sprachqualitätsbewertung einer Differenz zwischen der ersten und der zweiten Sprachqualität entspricht.</claim-text></claim>
<claim id="c-de-01-0005" num="0005">
<claim-text>Verfahren nach Anspruch 1, wobei die kompensierte Sprachqualitätsbewertung einem Verhältnis zwischen der ersten und der zweiten Sprachqualität entspricht.</claim-text></claim>
<claim id="c-de-01-0006" num="0006">
<claim-text>Verfahren nach Anspruch 1, wobei die erste und die zweite Sprachqualität unter Verwendung einer auditorisch-artikulatorischen Analyse bewertet werden.</claim-text></claim>
<claim id="c-de-01-0007" num="0007">
<claim-text>Verfahren nach Anspruch 1, wobei der Schritt des Bewertens der zweiten und der ersten Sprachqualität die folgenden Schritte umfasst:
<claim-text>Vergleichen der Lautbildungsleistung und der Nicht-Lautbildungsleistung für das Sprachsignal oder das verzerrte Sprachsignal, wobei die Lautbildungs- (P<sub>A</sub>) und die Nicht-Lautbildungs- (P<sub>NA</sub>)-Leistungen mit Lautbildungs- und Nicht-Lautbildungsfrequenzen des Sprachsignals oder des verzerrten Sprachsignals assoziiert werden; und<!-- EPO <DP n="13"> --></claim-text>
<claim-text>Bewerten der zweiten oder der ersten Sprachqualität auf der Basis des Vergleichs.</claim-text></claim-text></claim>
<claim id="c-de-01-0008" num="0008">
<claim-text>Verfahren nach Anspruch 7, wobei die Lautbildungsfrequenzen im Bereich von ca. 2 ∼ 12.5 Hz liegen.</claim-text></claim>
<claim id="c-de-01-0009" num="0009">
<claim-text>Verfahren nach Anspruch 7, wobei die Lautbildungsfrequenzen in etwa der Geschwindigkeit der menschlichen Lautbildung entsprechen.</claim-text></claim>
<claim id="c-de-01-0010" num="0010">
<claim-text>Verfahren nach Anspruch 7, wobei die Nicht-Lautbildungsfrequenzen in etwa höher als die Lautbildungsfrequenzen sind.</claim-text></claim>
<claim id="c-de-01-0011" num="0011">
<claim-text>Verfahren nach Anspruch 7, wobei der Vergleich der Lautbildungsleistung und der Nicht-Lautbildungsleistung ein Verhältnis zwischen der Lautbildungsleistung und der Nicht-Lautbildungsleistung ist.</claim-text></claim>
<claim id="c-de-01-0012" num="0012">
<claim-text>Verfahren nach Anspruch 10, wobei das Verhältnis einen Nenner und einen Zähler umfasst, wobei der Zähler die Lautbildungsleistung und eine kleine Konstante einschließt, und wobei der Nenner die Nicht-Lautbildungsleistung plus die kleine Konstante einschließt.</claim-text></claim>
<claim id="c-de-01-0013" num="0013">
<claim-text>Verfahren nach Anspruch 7, wobei der Vergleich zwischen der Lautbildungsleistung und der Nicht-Lautbildungsleistung eine Differenz zwischen der Lautbildungsleistung und der Nicht-Lautbildungsleistung ist.</claim-text></claim>
<claim id="c-de-01-0014" num="0014">
<claim-text>Verfahren nach Anspruch 7, wobei der Schritt des Bewertens der ersten oder der zweiten Sprachqualität den folgenden Schritt umfasst:
<claim-text>Bestimmen einer lokalen Sprachqualität und Verwendung des Vergleichs.</claim-text></claim-text></claim>
<claim id="c-de-01-0015" num="0015">
<claim-text>Verfahren nach Anspruch 7, wobei die lokale Sprachqualität weiterhin unter Verwendung eines Gewichtungsfaktors auf der Basis einer Gleichstromkomponente-Leistung (P<sub>NO</sub>) bestimmt wird.<!-- EPO <DP n="14"> --></claim-text></claim>
<claim id="c-de-01-0016" num="0016">
<claim-text>Verfahren nach Anspruch 9, wobei die erste oder die zweite Sprachqualität unter Verwendung der lokalen Sprachqualität bestimmt wird.</claim-text></claim>
<claim id="c-de-01-0017" num="0017">
<claim-text>Verfahren nach Anspruch 7, wobei der Schritt des Vergleichens der Lautbildungsleistung und der Nicht-Lautbildungsleistung den folgenden Schritt umfasst:
<claim-text>Durchführen einer Fourier-Transformation auf einer jeden der Vielzahl von aus einer Vielzahl von Signalen kritischer Bänder gewonnenen Hüllkurven.</claim-text></claim-text></claim>
<claim id="c-de-01-0018" num="0018">
<claim-text>Verfahren nach Anspruch 7, wobei der Schritt des Vergleichens der Lautbildungsleistung und der Nicht-Lautbildungsleistung den folgenden Schritt umfasst:
<claim-text>Filtern des Sprachsignals, um eine Vielzahl von Signalen kritischer Bänder zu erhalten.</claim-text></claim-text></claim>
<claim id="c-de-01-0019" num="0019">
<claim-text>Verfahren nach Anspruch 18, wobei der Schritt des Vergleichens der Lautbildungsleistung und der Nicht-Lautbildungsleistung den folgenden Schritt umfasst:
<claim-text>Durchführen einer Hüllkurvenanalyse auf der Vielzahl von Signalen kritischer Bänder, um eine Vielzahl von Modulationsspektren zu erhalten.</claim-text></claim-text></claim>
<claim id="c-de-01-0020" num="0020">
<claim-text>Verfahren nach Anspruch 18, wobei der Schritt des Vergleichens der Lautbildungsleistung und der Nicht-Lautbildungsleistung den folgenden Schritt umfasst:
<claim-text>Durchführen einer Fourier-Transformation auf einem jeden der Vielzahl von Modulationsspektren.</claim-text></claim-text></claim>
</claims>
<claims id="claims03" lang="fr"><!-- EPO <DP n="15"> -->
<claim id="c-fr-01-0001" num="0001">
<claim-text>Procédé d'évaluation de la qualité vocale, comprenant les étapes suivantes :
<claim-text>déterminer une première et une deuxième évaluations de la qualité vocale pour des premier et deuxième signaux de qualité vocale (SQ(s(t)), SQ(s'(t))) sur la base, respectivement, de premier et deuxième signaux vocaux, le deuxième signal vocal étant un signal vocal traité, et le premier signal vocal étant une version déformée (MNRU) du deuxième signal vocal ; et</claim-text>
<claim-text>obtenir une évaluation de la qualité vocale SQ compensée en comparant la première et la deuxième évaluations de la qualité vocale.</claim-text></claim-text></claim>
<claim id="c-fr-01-0002" num="0002">
<claim-text>Procédé selon la revendication 1 comprenant les étapes supplémentaires suivantes<br/>
avant de déterminer la première et la deuxième évaluations de la qualité vocale, déformer le deuxième signal vocal pour produire le premier signal vocal.</claim-text></claim>
<claim id="c-fr-01-0003" num="0003">
<claim-text>Procédé selon la revendication 1, dans lequel la première et la deuxième qualités vocales sont évaluées en utilisant une technique identique pour l'évaluation objective de la qualité vocale.</claim-text></claim>
<claim id="c-fr-01-0004" num="0004">
<claim-text>Procédé selon la revendication 1, dans lequel l'évaluation de la qualité vocale compensée correspond à une différence entre la première et la deuxième qualités vocales.</claim-text></claim>
<claim id="c-fr-01-0005" num="0005">
<claim-text>Procédé selon la revendication 1, dans lequel l'évaluation de la qualité vocale compensée correspond à un rapport entre la première et la deuxième qualités vocales.</claim-text></claim>
<claim id="c-fr-01-0006" num="0006">
<claim-text>Procédé selon la revendication 1, dans lequel la première et la deuxième qualités vocales sont évaluées en utilisant une analyse articulatoire d'audition.</claim-text></claim>
<claim id="c-fr-01-0007" num="0007">
<claim-text>Procédé selon la revendication 1, dans lequel l'étape d'évaluation de la deuxième ou de la première qualité vocale comprend les étapes suivantes ;<br/>
comparer la puissance d'articulation et la puissance de non-articulation pour le signal vocal ou le signal vocal déformé, dans lequel les puissances d'articulation P<sub>A</sub> et de non-articulation P<sub>NA</sub> sont des puissances associées aux fréquences d'articulation et de non-articulation du signal vocal ou du signal vocal déformé ; et<br/>
<!-- EPO <DP n="16"> -->évaluer la deuxième ou la première qualité vocale sur la base de la comparaison.</claim-text></claim>
<claim id="c-fr-01-0008" num="0008">
<claim-text>Procédé selon la revendication 7 ; dans lequel les fréquences d'articulation sont environ comprises entre 2 et 12,5 Hz.</claim-text></claim>
<claim id="c-fr-01-0009" num="0009">
<claim-text>Procédé selon la revendication 7, dans lequel les fréquences d'articulation correspondent environ à une vitesse de l'articulation humaine.</claim-text></claim>
<claim id="c-fr-01-0010" num="0010">
<claim-text>Procédé selon la revendication 7, dans lequel les fréquences de non-articulation sont environ supérieures aux fréquences d'articulation.</claim-text></claim>
<claim id="c-fr-01-0011" num="0011">
<claim-text>Procédé selon la revendication 7, dans lequel la comparaison entre la puissance d'articulation et la puissance de non-articulation est un rapport entre la puissance d'articulation et la puissance de non-articulation.</claim-text></claim>
<claim id="c-fr-01-0012" num="0012">
<claim-text>Procédé selon la revendication 10, dans lequel le rapport comprend un dénominateur et un numérateur, le numérateur comprenant la puissance d'articulation et une faible constante, le dénominateur comprenant la puissance de non-articulation plus la faible constante.</claim-text></claim>
<claim id="c-fr-01-0013" num="0013">
<claim-text>Procédé selon la revendication 7, dans lequel la comparaison entre la puissance d'articulation et la puissance de non-articulation est une différence entre la puissance d'articulation et la puissance de non-articulation.</claim-text></claim>
<claim id="c-fr-01-0014" num="0014">
<claim-text>Procédé selon la revendication 7, dans lequel l'étape d'évaluation de la première ou de la deuxième qualité vocale comprend l'étape suivante :
<claim-text>déterminer une qualité vocale locale en utilisant la comparaison.</claim-text></claim-text></claim>
<claim id="c-fr-01-0015" num="0015">
<claim-text>Procédé selon la revendication 7, dans lequel la qualité vocale locale est en outre déterminée en utilisant un facteur de pondération sur la base d'une puissance de composante CC P<sub>NO</sub>.</claim-text></claim>
<claim id="c-fr-01-0016" num="0016">
<claim-text>Procédé selon la revendication 9, dans lequel la première ou la deuxième qualité vocale est déterminée en utilisant la qualité vocale locale.<!-- EPO <DP n="17"> --></claim-text></claim>
<claim id="c-fr-01-0017" num="0017">
<claim-text>Procédé selon la revendication 7, dans lequel l'étape de comparaison de la puissance d'articulation et de la puissance de non-articulation comprend l'étape suivante :
<claim-text>effectuer une transformation de Fourier sur chaque enveloppe parmi une pluralité d'enveloppes obtenues à partir d'une pluralité de signaux de bande critique.</claim-text></claim-text></claim>
<claim id="c-fr-01-0018" num="0018">
<claim-text>Procédé selon la revendication 7, dans lequel l'étape de comparaison de la puissance d'articulation et de la puissance de non-articulation comprend l'étape suivante :
<claim-text>filtrer le signal vocal pour obtenir une pluralité de signaux de bande critique.</claim-text></claim-text></claim>
<claim id="c-fr-01-0019" num="0019">
<claim-text>Procédé selon la revendication 18, dans lequel l'étape de comparaison de la puissance d'articulation et de la puissance de non-articulation comprend l'étape suivante :
<claim-text>effectuer une analyse d'enveloppes sur la pluralité de signaux de bande critique pour obtenir une pluralité de spectres de modulation.</claim-text></claim-text></claim>
<claim id="c-fr-01-0020" num="0020">
<claim-text>Procédé selon la revendication 18, dans lequel l'étape de comparaison de la puissance d'articulation et de la puissance de non-articulation comprend l'étape suivante :
<claim-text>effectuer une transformation de Fourier sur chaque spectre parmi la pluralité de spectres de modulation.</claim-text></claim-text></claim>
</claims>
<drawings id="draw" lang="en"><!-- EPO <DP n="18"> -->
<figure id="f0001" num="1"><img id="if0001" file="imgf0001.tif" wi="140" he="101" img-content="drawing" img-format="tif"/></figure><!-- EPO <DP n="19"> -->
<figure id="f0002" num="2,3"><img id="if0002" file="imgf0002.tif" wi="155" he="189" img-content="drawing" img-format="tif"/></figure><!-- EPO <DP n="20"> -->
<figure id="f0003" num="4"><img id="if0003" file="imgf0003.tif" wi="95" he="167" img-content="drawing" img-format="tif"/></figure>
</drawings>
<ep-reference-list id="ref-list">
<heading id="ref-h0001"><b>REFERENCES CITED IN THE DESCRIPTION</b></heading>
<p id="ref-p0001" num=""><i>This list of references cited by the applicant is for the reader's convenience only. It does not form part of the European patent document. Even though great care has been taken in compiling the references, errors or omissions cannot be excluded and the EPO disclaims all liability in this regard.</i></p>
<heading id="ref-h0002"><b>Patent documents cited in the description</b></heading>
<p id="ref-p0002" num="">
<ul id="ref-ul0001" list-style="bullet">
<li><patcit id="ref-pcit0001" dnum="EP1187100A1"><document-id><country>EP</country><doc-number>1187100</doc-number><kind>A1</kind></document-id></patcit><crossref idref="pcit0001">[0005]</crossref></li>
<li><patcit id="ref-pcit0002" dnum="US4352182A"><document-id><country>US</country><doc-number>4352182</doc-number><kind>A</kind></document-id></patcit><crossref idref="pcit0002">[0006]</crossref></li>
</ul></p>
</ep-reference-list>
</ep-patent-document>
