Background of the Invention
[0001] The invention relates generally to communicating voice and video information over
a channel having a fixed capacity, such as a telephone communication channel.
[0002] Video conferencing systems typically transmit both voice and video information over
the same channel. A portion of the channel's bandwidth is typically dedicated to voice
information and the remaining bandwidth is allocated to video information.
[0003] The amount of video and voice information varies with time. For example, at certain
moments in time, a person at one end of the system may be silent. Thus, if the system
includes a variable capacity voice encoder, little information needs to be transmitted
during such moments of silence.
[0004] Similarly, the video signal may have little or no change between frames as, for example,
when all objects within the field of view are still. If the system includes a variable
capacity video encoder, little information needs to be transmitted during such moments
of inactivity. At the other extreme, during times of great activity, the amount of
video information may exceed the channel capacity allocated to video information.
Accordingly, the system transmits as much video information as possible, discarding
the remainder.
[0005] Typically, the video encoder accords priority to the most noticeable features of
the video signal. Thus, high priority information is transmitted first and the less
noticeable low priority information is temporarily discarded if the channel lacks
sufficient capacity. Accordingly, it is desirable to have available as much video
bandwidth as possible.
[0006] US-A-4949383 (Koh et al) discloses a bit allocation method in which bits are allocated
to channels of a sub-band coder or to the coefficients of a transformer coder using
a fixed number of bits. Only the selection of the channels to which the available
bits are assigned is varied. Thus, different frequency bands can be encoded separately
and differently.
[0007] US-A-4965830 (Burham et al) discloses a bit allocation method attempting to maintain
a predetermined average quantization distortion level.
[0008] EP-A-309974 discloses a method for communicating a digital signal, in which data
compression is obtained by encoding the sufficiently large samples of the signal,
and then encoding a residual signal.
[0009] The invention relates to a method and apparatus for allocating transmission bits
for use in transmitting samples of a digital signal. An aggregate allowable quantization
distortion value is selected representing an allowable quantization distortion error
for a frame of samples of the digital signal. A set of samples are selected from the
frame of samples such that a plurality of the selected samples are greater than a
noise threshold. For each sample of the set, a sample quantization distortion value
is computed which represents an allowable quantization distortion error for the sample.
The sum of all sample quantization distortion values is approximately equal to the
aggregate allowable quantization distortion value. For each sample of the set, a quantization
step size is selected which yields a quantization distortion error approximately equal
to the sample's corresponding quantization distortion value. Each sample is then quantized
using its quantization step size.
[0010] In preferred embodiments, the digital signal includes a noise component and a signal
component. A signal index is prepared representing, for at least one sample of the
frame, the magnitude of the signal component relative to the magnitude of the noise
component. The aggregate allowable quantization distortion is selected based on the
signal index.
[0011] The sample quantization distortion value is computed by dividing the aggregate allowable
quantization distortion value by a number of samples in the frame of samples to form
a first sample distortion value. A tentative set of samples of the digital signal
is selected wherein each sample of the tentative set is greater than a noise threshold
determined at least in part by the value of the first sample distortion value. The
first sample distortion value is then adjusted by an amount determined by the difference
between the first sample distortion value and at least one sample excluded from the
tentative set (i.e, a "noisy sample"). Based on the adjusted distortion value, the
process is repeated to identify any noisy samples of the tentative set; remove the
noisy samples, if any, from the tentative set; and again adjust the first sample distortion
value by an amount determined by the difference between the first sample distortion
value and a noisy sample. The process is repeated until an adjusted first sample distortion
value is reached for which no additional noisy samples of the tentative set are found
or until the process has been repeated a maximum number of times.
[0012] After the adjustment is terminated, the number of bits required to transmit all samples
of the tentative set is estimated. The estimated bit number is compared to a maximum
bit number. If the estimated bit number is less than or equal to the maximum bit number,
a final noise threshold is selected based on the adjusted first sample distortion
value.
[0013] If the estimated bit number exceeds the maximum bit number, a second sample distortion
value, is prepared. A second tentative set of samples of the digital signal is then
selected wherein each of a plurality of samples of the second tentative set have a
magnitude above the second sample distortion value. The number of bits required to
transmit all samples of the second tentative set is then estimated. The estimated
bit number is compared to the maximum bit number. If it is greater than the maximum
bit number, the second sample distortion value is increased and the second tentative
set of samples is re-selected based on the adjusted second sample distribution value.
The number of bits required to transmit the second tentative set is again estimated.
This process is repeated until a second sample distortion value is reached for which
the estimated bit number is less than or equal to the maximum bit number.
[0014] The sample distortion value is then calculated from the adjusted first sample distortion
value and the second sample distortion value. A final set of samples of the digital
signal, is selected wherein each of a plurality of samples of the final set have a
magnitude above a final threshold determined by the sample's corresponding sample
distortion value.
[0015] In another aspect, the invention relates to a method and apparatus for communicating
a digital signal which includes a noise component and a signal component. An estimation
signal is prepared which is representative of the digital signal but has fewer samples
than the digital signal. A signal index is prepared which represents, for at least
one sample of the estimation signal, the magnitude of the signal component relative
to the magnitude of the noise component. Based on the signal index, samples of the
digital signal are selected which have a sufficiently large signal component. The
selected samples of the digital signal and the samples of the digital estimation signal
are both transmitted to a remote device. The remote device reconstructs the digital
signal from the transmitted selected samples and estimation samples.
[0016] In preferred embodiments, the digital signal is a frequency domain speech signal
representative of voice information to be communicated, and each sample of the estimation
signal is a spectral estimate of the frequency domain signal in a corresponding band
of frequencies. To reconstruct the frequency domain signal, a random number is generated
for each nonselected sample of the frequency domain speech signal. A noise estimate
of the magnitude of a noise component of the nonselected sample is prepared from at
least one spectral estimate. Based on the noise estimate, a scaling factor is generated.
The random number is then scaled according to the scaling factor to produce a reconstructed
sample representative of the nonselected sample.
[0017] The estimation signal and the frequency domain speech signal each include a series
of frames. Each frame represents the voice information over a specified window of
time. To prepare a first noise estimate for a current frame, an initial noise estimate
is first prepared for a prior frame of the estimation signal. The initial noise estimate
is prepared from at least one spectral estimate representative of a band of frequencies
of the prior frame. Based on the magnitude of the signal index, a rise time constant
t
r and a fall time constant t
f are then selected. The selected rise time constant is added to the initial noise
estimate to form an upper threshold, and the selected fall time constant is subtracted
from the initial noise estimate to form a lower threshold.
[0018] A current spectral estimate of the current frame, representative of the same band
of frequencies, is then compared to the upper and lower thresholds. If the current
spectral estimate is between the thresholds, the current noise estimate is set equal
to the current spectral estimate. If it is below the lower threshold, the current
noise estimate is set equal to the lower threshold. If the current spectral estimate
is above the upper threshold, the current noise estimate is set equal to the upper
threshold.
[0019] To generate the scaling factor, a noise coefficient and a voice coefficient are first
selected based on the value of the signal index. The current noise estimate is then
multiplied by the noise index. Similarly, the current spectral estimate is multiplied
by the signal index. The products of the multiplication are added together and the
scaling factor is then formed from the sum.
[0020] The methods and apparatus described in this specification have as an objective, a
reduction in the amount of channel bandwidth allocated to audio information whenever
there is little audio information required to be sent. The remaining portion of bandwidth
is allocated to video information. Thus, average, a lower bit rate is provided for
audio information and a higher bit rate is provided for video information.
[0021] The invention will now be described by way of example with reference to the drawings,
a brief description of which follows.
Brief Description of the Drawings
[0022] Fig. 1(a) is a block diagram of the near end of a video conferencing system.
[0023] Fig. 1(b) is a block diagram of a far end of a video conferencing system.
[0024] Figs. 2(a-c) are diagrams of a set of frequency coefficients and two types of estimates
of the frequency coefficients.
[0025] Figs. 3(a) and 3(b) are a flow chart of a process for computing spectral estimates.
[0026] Fig. 4 is a diagram of a set of spectral estimates arranged in broad bands.
[0027] Fig. 5 is a block diagram of a speech detector.
[0028] Fig. 6 is a flow chart of a process for estimating the amount of voice in a frame
of the microphone signal.
[0029] Fig. 7 is a block diagram of a bit rate estimator.
[0030] Figs. 8(a) and 8(b) are a flow chart of a process for computing a first estimate
of the allowable distortion per spectral estimate.
[0031] Fig. 9 is a flow chart of a process for computing a second estimate of the allowable
distortion per spectral estimate.
[0032] Fig. 10 is a flow chart of a process for computing quantization steps sizes.
[0033] Fig. 11 is a diagram illustrating the interpolation between band quantization step
sizes to form coefficient quantization step sizes.
[0034] Fig. 12 is a block diagram of a coefficient fill-in module.
System Overview
[0035] Referring to Figs. 1(a), and 1(b), a video conferencing system includes a near end
microphone 12 which receives the voice of a person speaking at the near end and generates
an electronic microphone signal m(t) representative of the voice, where t is a variable
representing time. Similarly, a camera 14 is focused on the person speaking to generate
video signal v(t) representative of the image of the person.
[0036] The microphone signal is applied to an audio encoder 16 which digitizes and encodes
the microphone signal m(t). As will be explained more fully below, the encoder 16
generates a set of digitally encoded signals F
h(k), S
h(j), g
i(F) and A
i(F) which collectively represent the microphone signal, where, as explained more fully
below, k and j are integers representing values of frequency, and F is an integer
identifying a "frame" of the microphone signal. Similarly, the video signal v(t) is
applied to a video encoder 18 which digitizes and encodes the video signal.
[0037] The encoded video signal V
e and the encoded set of microphone signals are applied to a bit stream controller
20 which merges the signals into a serial bit stream for transmission over a communication
channel 22.
[0038] Referring to Fig. 1(b), a bit stream controller 24 at the far end receives the bit
stream from communication channel 22 and separates it into the video and audio components.
The received video signal V
e is decoded by a video decoder 26 and applied to a display device 27 to recreate a
representation of the near end camera image. At the same time, an audio decoder 28
decodes the set of received microphone signals to generate a loudspeaker signal L(t)
representative of the near end microphone signal m(t). In response to the loudspeaker
signal, a loudspeaker 29 reproduces the near end voice.
[0039] Referring to Fig. 1(a), encoder 16 includes an input signal conditioner 30 for digitizing
and filtering the microphone signal m(t) to produce a digitized microphone signal
m(n), where n is an integer representing a moment in time. The digitized microphone
signal, m(n), is provided to a windowing module 32 which separates m(n) into groups
m
F(n), in the illustrated embodiment, of 512 consecutive samples referred to as a frame.
The frames are selected to overlap. More specifically, each frame includes the last
16 samples from the previous group and 496 new samples.
[0040] Each frame of samples is then applied to a normalization module 34 which calculates
the average energy E
av of all microphone samples in the frame:

It then computes a normalizing gain g(F), for each frame F, equal to the square root
of the frame's average energy E
av:

[0041] As explained more fully below, the normalizing gain g is used at the near end to
scale the microphone samples in each frame. At the far end, the normalizing gain is
used to restore the decoded microphone samples to their original scale. Accordingly,
the same gain used to scale the samples at the near end must be transmitted to the
far end for use in rescaling the microphone signal. Since the normalization gain g
calculated by module 34 has a relatively large number of bits, normalization module
34 supplies the gain g to a quantizer 35 which generates a gain quantization index
g
i which represents the gain using fewer bits than the number of bits specifying gain
g. The gain quantization index g
i is then provided to bit steam controller 20 for transmission to the far end.
[0042] The far end audio decoder 28 reconstructs the gain from the transmitted gain quantization
index g
i and uses the reconstructed gain g
q to restore the microphone signal to its original scale. Since the reconstructed gain
g
q typically differs slightly from the original gain g, audio encoder 16 at the near
end normalizes the microphone signal using the same gain g
q as used at the far end. More specifically, an inverse quantizer 37 reconstructs the
gain g
q from the gain quantization index g
i in the same manner as the far end audio decoder 28.
[0043] The quantized gain g
q and the frame of microphone signals m
F(n) are forwarded to a discrete cosine transform module (DCT) 36 which divides each
microphone sample m
F(n) by the gain g
q to yield a normalized sample m'(n). It then converts the group of normalized microphone
signals m'(n) to the frequency domain using the well known discrete cosine transform
algorithm. DCT 36 thus generates 512 frequency coefficients F(k) representing samples
of the frame's frequency spectrum, where k is an integer representing a discrete frequency
in the frame's spectrum.
[0044] The frequency coefficients F(k) are encoded (using an entropy adaptive transfer coder
described below) and transmitted to the far end. To reduce the number of bits necessary
to transmit these coefficients, the encoder 16 estimates the relative amounts of signal
(e.g., voice) and noise in various regions of the frequency spectrum, and chooses
not to transmit coefficients for frequency bands having a relatively large amount
of noise.
Further, for each coefficient selected for transmission, encoder 16 selects the number
of bits required to represent the coefficient based on the amount of noise present
in the frequency band which includes the coefficient's frequency. More specifically,
it is known that humans can tolerate more corrupting noise in regions of the audio
spectrum having relatively large amounts of audio signal, because the audio signal
tends to mask the noise. Accordingly, encoder 16 coarsely quantizes coefficients having
relatively large amounts of audio signal since the audio signal masks the quantization
distortion introduced through the coarse quantization. Thus, the encoder minimizes
the number of bits needed to represent each coefficient by selecting a quantization
step size tailored to the amount of audio signal represented by the coefficient.
[0045] To estimate the relative amounts of noise and signal in various regions of the spectrum,
the frequency coefficients F(k) are first supplied to a spectrum estimation module
38. As will be explained more fully below, estimation module 38 reduces the frequency
coefficients F(k) to a smaller set of spectral estimates S(j) which represent the
frequency spectrum of the frame with less detail, wherein j an integer representing
a band of frequencies in the spectrum. A speech detection module 40 processes the
spectral estimates of each frame to estimate the energy component of the microphone
signal due to noise in the near end room. It then provides a signal index A
i(F) for each frame F which approximates the percentage of the microphone signal in
the frame which is attributable to voice information.
[0046] For each frame F, the signal index A
i(F) and the spectral estimates S(j) are both used in quantizing and encoding the frequency
coefficients F(k) for transmission to the far end. Accordingly, they are both needed
at the far end for use in reconstructing the frequency coefficients. The signal index
A
i(F) has only three bits and accordingly is applied directly to the bit stream controller
20 for transmission. The spectral estimates S(j) however are converted to logarithmic
spectral estimates Log
2 S
2(j) and are encoded using a well known differential pulse code modulation encoder
(herein "DPCM") 39 to reduce the number of bits to be transmitted. (The DPCM is preferably
a first order DPCM with a unity gain predictor coefficient.) The encoded logarithmic
spectral estimates S
e(j) are further encoded using a Huffman encoder 49 to further reduce the number of
bits to be transmitted. The resultant Huffman codes S
h(j) are provided to bit stream controller 20 for transmission.
[0047] Far end audio decoder 28 reconstructs the logarithmic spectral estimates Log
2 S
2(j) from the Huffman codes S
h(j). However, due to the operation of the DPCM encoder 39, the reconstructed spectral
estimates Log
2 S
q 2(j) are not identical to the original estimates Log
2 S
2(j). Accordingly, a decoder 41 in near end encoder 16 decodes the encoded estimates
S
e(j) in the same manner as performed at the far end, and uses the thus decoded estimates
Log
2 S
q 2(j) in quantizing and encoding the frequency coefficients F(k). Thus, (as explained
more fully below) the audio encoder 16 encodes the frequency coefficients using the
identical estimates Log
2 S
q 2(j) as used by the far end audio decoder 28 in reconstructing the signal Log
2 S
q 2(j) at the far end.
[0048] Based on the values of the decoded spectral estimates Log
2 S
q 2(j) and the signal index A
i(F), a bit rate estimator 42 selects groups of frequency coefficients F(k) which collectively
have a sufficient amount of voice information to merit being transmitted to the far
end. Next, the bit rate estimator selects for each coefficient to be transmitted,
a quantization step size which determines the number of bits to be used in quantizing
the coefficients for transmission. For each selected group of frequency coefficients,
bit rate estimator 42 first computes a group quantization step size Q(j) and a "class"
C(j) (to be described below) for each frequency band j. The group quantization step
sizes are then interpolated to yield a coefficient quantization step size Q(k) for
each frequency coefficient F(k). Q(k) is then applied to a coefficient quantizer 44
which, based on the assigned step size, quantizes each frequency coefficient in the
band to provide a corresponding quantization index I(k).
[0049] In response to each quantization index I(k), a limit controller 45 generates a Huffman
index I
h(k). The Huffman indices I
h(k) provided by the limit controller are further encoded by Huffman encoder 47 to
yield Huffman codes F
h(k). The Huffman codes are applied to bit stream controller 20 for transmission over
communication channel 22.
[0050] The limit controller begins the above described encoding process starting with the
low frequency quantization indices (i.e., k=0) and continuing until all indices are
encoded or until the limit controller 45 concludes that the number of encoded bits
exceeds the capacity of channel 22 allocated to the microphone signal. If the limit
controller 45 concludes that the capacity has been exceeded, it discards the remaining
indices I
h(k) and provides a unique Huffman index I
h(k) indicating that the frame's remaining frequency coefficients will not be coded
for transmission. (As explained more fully below, the far end decoder estimates the
uncoded coefficients from the transmitted spectral estimates).
Spectrum Estimator for Calculating Spectral Estimates S(j)
[0051] Referring to Figs. 2(a-c) and 3(a-b), the following describes the spectrum estimation
module 38 in further detail. Module 38 first separates the frequency coefficients
F(k) into bands of L adjacent coefficients, where L is preferably equal to 10. (Step
110). For each band, j, the module prepares a first approximation sample, representative
of the entire band, by computing the energy of the spectrum in the band. More specifically,
module 38 squares each frequency coefficient F(k) in the band and sums the squares
of all coefficients in the band. (Steps 112, 114) (See also Figs. 2(b), 2(c))
[0052] This approximation may provide a poor representation of the spectrum if the spectrum
includes a high energy tone at the border between adjacent bands. For example, the
spectrum shown Fig. 2(b) includes a tone 50 at the border between band 52 and band
54. Interpolating between approximation sample 58 (representing the sum of the squares
of all coefficients in band 54) and approximation sample 56 (representing the sum
of the square of all coefficients in band 52) yields a value 59 which does not accurately
reflect the presence of tone 50.
[0053] Accordingly, module 38 also employs a second approximation technique wherein the
spectral estimates for both bands 52 and 54 will reflect the presence of a tone 50
near the border between the bands.
[0054] More specifically, module 38 derives a second approximation for each band using the
ten samples in the band and five neighboring samples from each adjacent band. (See
Fig. 2(a)) Module 38 performs the logical inclusive "OR" operation on the binary values
of all twenty samples. (Step 116). This operation provides a computationally inexpensive
estimate of the magnitude of the largest sample in the set. More specifically, each
digit of the binary result of an "OR" operation is set to one if any of the operand
binary values have a one in the digit (e.g., 0110 "OR" 0011 = 0111). Accordingly,
the result of the "OR" operation is at a minimum equal to the magnitude of the largest
sample in the set and at most equal to twice its magnitude. The result of the OR operation
is then doubled to form the second approximation. (Step 117).
[0055] As shown in Fig. 2(a), the second approximation in band 52 yields a relatively large
approximation value 60 for band 52 since the approximation includes tone 50 from band
54. The second approximation 60 more accurately reflects the presence of tone 50 than
the first approximation 56. Accordingly, for each band, module 38 compares the second
approximation to the first approximation (Step 118) and selects the larger of the
two as the squared spectral estimate S
2(j). (Steps 120-122). Finally, module 38 computes the logarithm of the squared spectral
estimate Log
2 S
2(j)). (Step 224).
Speech Detector for Calculating The Signal index Ai
[0056] Referring to Fig. 4, speech detector 40 computes a signal index A
i(F) representative of the relative amount of voice energy in the frame. Toward this
end, detector 40 groups the samples of spectral estimates into broad bands of frequencies.
The broad bands have varying widths, nonuniformly distributed across the spectrum.
For example, the first broad band may be larger than the second broad band which is
smaller than the third.
[0057] Referring to Fig. 5, the speech detector estimates the amount of background noise
S
n in each broad band. For each broad band, speech detector 40 forms an aggregate estimate
S
a(F) as follows:

where F is an integer identifying the current frame, and Y and X are numbers representing
the upper and lower frequencies respectively of the broad band. (Step 210). As will
be explained in more detail below, the speech detector then compares the aggregate
estimate of the current frame against a pair of noise thresholds (derived from prior
band aggregate estimates from prior frames) to determine the amount of the noise in
the band. (Step 211). In general, a frame having a relatively low aggregate estimate
likely includes little voice energy in the broad band. Accordingly, the aggregate
estimate in such a frame provides a reasonable estimate of the background noise in
the broad band.
[0058] To compare aggregate estimates against the noise thresholds, the speech detector
must first unnormalize each aggregate estimate. Otherwise, the estimates from the
present frame will be on different scales than those of the prior frames and the comparison
will not be meaningful. Accordingly, the speech detector unnormalizes the aggregate
estimate by first computing Log
2(g
q 2) to place the normalization gain in the same logarithmic scale as the aggregate estimates.
It then adds the scaled normalization gain Log
2(g
q 2) to the aggregate estimate to unnormalize the estimate. (Step 212).
[0059] This unnormalized estimate is then compared to an upper threshold S
r(F) and a lower threshold S
f(F) where F identifies the current frame. (Step 214). As explained more fully below,
the thresholds are computed for each frame based on the value of a noise estimate
from the prior frame. Since the first frame lacks a predecessor, the upper threshold
for the first frame, S
r(0), is initialized to a value "r" and the lower threshold S
f(0) is initialized to -f, where f is substantially greater than r (e.g., r=1, and
f=10).
[0060] If the aggregate estimate S
a(F) is between the thresholds, the speech detector sets the noise estimate for the
frame equal to the aggregate estimate (Step 216):

[0061] If the aggregate estimate is greater than the upper threshold, the noise estimate
is set equal to the upper threshold (Step 218):

[0062] Finally, if the aggregate estimate is below the lower threshold, the noise estimate
is set equal to the lower threshold (Step 220):

[0063] Before computing the noise estimate for the next frame, the speech detector adjusts
the thresholds for the next frame S
f(F+1), S
r(F+1) such that they straddle the noise estimate for the current frame S
n(F). (Step 222) More specifically, for the next frame, F = F + 1, the speech detector
calculates an upper noise threshold from the current frame's noise estimate as follows:

Similarly, it calculates a lower threshold S
f(F + 1) = S
n(F) - f.
Thus, for each new frame, the speech detector calculates a noise estimate and adjusts
the upper and lower noise thresholds to straddle the noise estimate.
[0064] This technique adaptively adjusts the noise estimates over time such that the noise
estimate in a given broad band of a current frame approximately equals the most recent
minimum aggregate estimate for the broad band. For example, if a series of frames
arrive having no voice component in a broad band, the aggregate estimates will be
relatively small since they reflect only the presence of background noise. Thus, if
these aggregate values are below the lower threshold, S
f, the above technique will quickly reduce the noise estimate in relatively large increments,
f, until the noise estimate equals a value of the relatively low aggregate estimate.
[0065] Once frames having a voice component begin to arrive, the noise estimate remains
relatively low, stepping upward in relatively small increments r. By allowing the
noise level to increment upward, the speech detector is able to detect increases in
the background noise. However, since a relatively small increment "r" is used, the
noise estimate tends to remain near the most recent minimum aggregate estimate. This
prevents the speech detector from mistaking transient voice energy for background
noise.
[0066] After calculating the noise estimates for a frame, the speech detector subtracts
the noise estimate S
n(F) from the frame's aggregate estimate S
a(F) to obtain an estimate of the voice signal in the broad band. (Step 224). It also
subtracts a threshold constant S
T from the aggregate estimate to generate a signal S
out representative of the extent to which the voice signal exceeds the specified threshold
S
T. (Step 224). Finally, the speech detector computes the index A
i(F) for the entire frame from the collection of S
out signals and the normalization gain g
q. (Step 226).
[0067] Referring to Fig. 6, the following describes in further detail the calculation of
A
i(F) from the collection of S
out signals. The speech detector first selects the largest S
out from all broad bands. (Step 246). If the selected value, S
max, is less than or equal to zero, all broad bands likely have no voice component. (Step
248). Accordingly, the speech detector sets the index A
i(F) to zero indicating that the broad bands contain only noise. (Step 250).
[0068] If S
max is greater than zero, (Step 248) the speech detector calculates the index A
i(F) from the value of S
max. Toward this end, it scales S
max by a fixed gain G
s i.e., S'
max = S
max * G
s where G
s preferably is approximately 0.2734. (Step 252). It next computes a correspondingly
attenuated representation g
o of the normalization gain as follows:

where T
g and G
g are predefined constants e.g., T
g = 4096 and G
g = 0.15625. (Step 254). If g
o is greater than S'
max, the speech detector assumes that S
max is less than the voice energy in the frame. Accordingly, it selects g
o as the index A
i(F). (Step 256). Otherwise, it selects S'
max as the index A
i(F). (Step 256). Finally, the speech detector compares the selected index to a maximum
index i
max. (Step 258). If the selected index exceeds the maximum index, the speech detector
sets the index A
i(F) equal to its maximum value, i
max. (Step 262). Otherwise, the selected value is used as the index. (Step 260).
Bit Rate Estimator for Calculating Step Sizes Q(k) and Class Information C(j)
[0069] Referring to Fig. 1(a), the bit rate estimator 42 receives the index A
i(F) and the quantized log spectral estimates Log
2S
q 2(j). In response, it computes a step size Q(k) for each frequency coefficient and
a class indicator c(j) for each band j of frequency coefficients. The quantization
step size Q(k) is used by quantizer 44 to quantize the frequency coefficients. The
class indicator c(j) is used by Huffman encoder 47 to select the appropriate coding
tables.
[0070] The following describes in further detail the procedure used by bit rate estimator
42 to compute the step size and class information. Referring to Fig. 7, the bit estimator
includes a first table 70 which contains a predetermined aggregate allowable distortion
value D
T for each value of the signal index A
i(F). Each value D
T represents an aggregate quantization distortion error for the entire frame. A second
table 72 contains, for each value of index A
i(F), a predetermined maximum number of bits r
max initially allocated to the frame. (As explained below, more bits may be allocated
to the frame if necessary.) The stored bit rates r
max increase with A
i(F). To the contrary, the stored distortion values D
T decrease with each increase in A
i(F). For example, for A
i(F) equals zero, (i.e., 100% noise) the tables provide a small bit rate and a high
allowable distortion. For A
i(F) equal seven, (i.e., 100% audio) the tables provide a high bit rate and a low allowable
distortion. Based on the value of index A
i(F), the first and second tables select an aggregate allowable quantization distortion
D
T and an allowable maximum number of bits r
max. This allowable distortion D
T is provided to a first distortion approximation module 74. Module 74 then computes
a first allowable sample distortion value d
l representative of the allowable distortion per estimate. The sample distortion value
d
l is then used in deriving c(j) and a block quantization step size Q(k).
Calculation of a First Estimate of the Allowable Distortion, d1
[0071] Referring to Figs. 8(a) and 8(b), module 74 first computes an initial value of the
sample distortion value d
1 by dividing the allowable aggregate distortion D
T by the number of spectral estimates P. (Step 310). As will be explained more fully
below, only frequency coefficients from a band whose squared spectral estimate is
sufficiently greater than a final quantization distortion value d
w will be coded for transmission. Thus, the allowable quantization value operates as
a noise threshold for determining which coefficients will be transmitted. Accordingly,
module 74 performs an inverse logarithm operation on the log spectral estimates Log
2 S
q 2(j) to form the square of the quantized spectral estimates S
q 2(j). Module 74 tentatively assumes that if the squared spectral estimate S
q 2(j) is less than or equal to d
1, the spectral estimate's constituent frequency coefficients (referred to herein as
"noisy samples") will not be coded. (Step 312). Module 74 accordingly increases the
sample distortion value d
1 to reflect the fact that such constituent coefficients will not be coded. More specifically,
it computes the sum D
NT of all squared spectral estimates which are less than or equal to d
1 (Step 314):

It then subtracts the sum from the aggregate allowable distortion D
T (Step 316) and divides the result by the number of remaining spectral estimates N
to compute an adjusted sample distortion value d
1 (Step 318):

where N is the number of squared spectral estimates above the initial distortion
value d
1. Since d
1 may now be greater than its initial value, module 74 compares each squared spectral
estimate to the new d
1 to determine if any other coefficients will not be coded. (Step 320). If so, the
estimator repeats the process to compute an adjusted sample distortion value d
1. (Steps 322, 324). The search terminates when no additional squared spectral coefficients
are less than or equal to an adjusted sample distortion value d
1 (Step 320). It also terminates after a maximum number of allowable iterations. (Steps
322). The resultant sample distortion value d
1 is then provided to a bit rate comparator 76 (Fig. 7). (Step 326).
[0072] Referring again to Fig. 7, based on the first sample distortion value d
1 and the log spectral estimate Log
2 S
q 2(j), comparator 76 computes a tentative number of bits per frame "r" as follows:

[0073] It then compares the estimated number of bits to the maximum allowable number of
bits per frame, r
max. If r is less than the maximum, r
max, comparator 76 signals module 78 to compute step sizes for the frequency coefficients
based on the first sample distortion value d
1. However, if r exceeds the maximum r
max, comparator 76 assumes that more distortion per estimate must be tolerated to keep
the number of bits below r
max.
Accordingly, it signals a second distortion approximation module 80 to begin an iterative
search for a new distortion value d
2 which will yield a bit rate below r
max.
Second Estimate of the Allowable Distortion, d2
[0074] Referring to Fig. 9, approximation module 80 initially computes a distortion increment
value D
i which satisfies the relation (Step 410):

The distortion increment value D
i is an estimate of a necessary increment in the first sample distortion value to reduce
the bit rate below the maximum R
max. Accordingly, the approximation module computes a new distortion value d
2 which satisfies the relation (Step 412):

[0075] In the same manner described above, it then compares each squared spectral estimate
S
q 2(j) to d
2 to determine which estimates are less than or equal to the distortion, thereby predicting
which frequency coefficients will not be coded. (Step 414). Based on this prediction,
module 80 again computes the total number of bits r required for the frame according
to the following equation (Step 416):

Module 80 again compares the bit rate r to the maximum r
max to determine if the new distortion value d
2 yields a bit rate below the maximum. (Step 418). If so, module 80 provides d
2 to module 78 and notifies it to calculate the quantization step sizes Q(j) based
on both d
2 and d
1. (Step 422). If not, module 80 performs another iteration of the process in an attempt
to find a distortion d
2 which will yield a sufficiently low bit rate r. (Steps 418, 420-428). However, if
a maximum number of iterations have been tried without finding such a distortion estimate,
the search is terminated and the most recent value of d
2 is supplied to module 78. (Steps 420, 422).
Calculation of Step Sizes and Class Information From the Distortion Estimates
[0076] Referring to Fig. 10, the module 78 combines the two estimates of allowable sample
distortion values d
1, d
2 to form a weighted distortion d
w satisfying the following relation:

where a is a constant weighting factor. (Step 510). (Note: If no value d
2 is calculated, the weighted distortion d
w is set equal to the first estimate d
1). Based on the weighted distortion estimate d
w, module 78 computes the class parameter c'(j) for each spectral estimate S(j) as
follows (Step 512):

[0077] The class parameter c'(j) is then rounded upward to the nearest integer value in
the range of zero to eight to form a class integer c(j). (Step 520). A class value
of zero indicates that all coefficients in the class should not be coded. Accordingly,
the class value c(j) for each spectral estimate is provided to the quantizer to indicate
which coefficients should be coded.
[0078] Class values greater than or equal to one are used to select huffman tables used
in encoding the quantized coefficients. Accordingly, the class value c(j) for each
spectral estimate is provided to the huffman encoder 47.
[0079] Module 78 next computes a band step size Q(j) for each spectral estimate, based on
the value of the weighted distortion d
w and the value of the class parameter c (j) for the estimate. (Steps 514-518). More
specifically, for spectral estimates whose class values are less than 7.5, the step
size Q(j) is calculated to satisfy the following relation:

where z is a constant offset value, e.g., - 1.47156 (Step 516). For spectral coefficients
whose class parameter c'(j) is greater than or equal to 7.5, the step size is chosen
to satisfy the following relation (Step 518):

[0080] The band step sizes Q(j) for each band j, are then interpolated to derive a step
size Q(k) for each frequency coefficient k within the band j. First, each band step
size is scaled downward. In this regard, recall that the spectral estimate S(j) for
each band j was computed as the greater of 1) the sum of all squared coefficients
in the band and 2) twice the logical OR of all coefficients in the block and of the
ten neighboring samples. (See Fig. 2(a)-2(c)). Accordingly, the selected spectral
estimate roughly approximates the aggregate energy of the entire band. However, the
quantization step size for each coefficient should be chosen based on the average
energy per coefficient within each band.
[0081] Accordingly, the band step size Q(j) is scaled downward by dividing it by the number
of coefficients used in computing the spectral estimate for the band. (i.e., by either
ten or twenty depending on the technique chosen to calculate S(j)). Next, bit rate
estimator 42 linearly interpolates between the log band step sizes Log
2Q(j) to compute the logarithm of coefficient step sizes log
2 Q(k). (See Fig. 11). Finally, the reciprocal of the coefficient step size is derived
as follows:

Quantization and Encoding of the Frequency Coefficients
[0082] Referring to Figure 1(a), the class integers c(j) and the quantization steps sizes
Q(k) are provided to the coefficient quantizer 44. Coefficient quantizer 44 is a mid-tread
quantizer which quantizes each frequency coefficient F(k) using its associated inverse
step size l/Q(k) to produce an index I(k). The indices I(k) are encoded for transmission
by the combined operation of a limit controller 45 and a Huffman encoder 47.
[0083] Huffman encoder 47 includes, for each class c(j), a Huffman table containing a plurality
of Huffman codes. The class integers c(j) are provided to Huffman encoder 47 to select
the appropriate Huffman table for the identified class.
[0084] In response to an index I(k), limit controller 45 generates a corresponding Huffman
index I
h(k) which identifies an entry in the selected Huffman table. Huffman encoder 47 then
provides the selected Huffman code F
h(k) to bit stream controller 20 for transmission to the far end.
[0085] Typically, limit controller 45 simply forwards the index I(k) for use as the Huffman
index I
h(k). However, the range of possible indices I(k) may exceed the input range of the
corresponding Huffman table. Accordingly, for each class c(j), limit controller 45
includes an index maximum and minimum. The limit controller compares each index I(k)
with the index maximum and minimum. If I(k) exceeds either the index maximum or minimum,
limit controller 45 clips I(k) to equal the respective maximum or minimum and provides
the clipped index to encoder 47 as the corresponding Huffman index.
[0086] For each frame, limit controller 45 also maintains a running tally of the number
of bits required to transmit the Huffman codes. More specifically, the limit controller
includes, for each Huffman table within Huffman encoder 47, a corresponding bit number
table. Each entry in the bit number table indicates the number of bits of a corresponding
Huffman code stored in the Huffman table of encoder 47. Thus, for each Huffman index
I
h(k) generated by limit controller 45, the limit controller internally supplies the
Huffman index to the bit number table to determine the number of bits required to
transmit the corresponding Huffman code F
h(k) identified by the Huffman index I
h(k). The number of bits are then added to the running tally. If the running tally
exceeds a maximum allowable number of bits, limit controller 45 ignores the remaining
indices I(k). Limit controller 45 then prepares a unique Huffman index which identifies
a unique Huffman code for notifying the far end receiver that the allowable number
of bits for the frame has been reached and that the remaining coefficients will not
be coded for transmission.
[0087] To transmit the unique Huffman code, the limit controller must allocate bits for
transmission of the unique code. Accordingly, it first discards the most recent Huffman
code and recomputes the running tally to determine if enough bits are available to
transmit the unique Huffman code. If not, the limit controller repeatedly discards
the most recent Huffman code until enough bits are allocated for the transmission
of the unique Huffman code.
Reconstruction of the Microphone Signal At the Far End
[0088] Referring again to Fig. 1(b), far end audio decoder 28 reconstructs the microphone
signal from the set of encoded signals. More specifically, a Huffman decoder 25 decodes
the Huffman codes S
h(j) to reconstruct the encoded log spectral estimates S
e(j). Decoder 27 (identical to decoder 41 of the audio encoder 16 (Fig. 1(a)) further
decodes the encoded log spectral estimates to reconstruct the quantized spectral estimates
Log
2 S
q 2(j).
[0089] The log spectral estimates Log
2 S
q 2(j) and the received signal index A
i(F) are applied to a bit rate estimator 46 which duplicates the derivation of classes
c(j) and step sizes Q(k) performed by bit rate estimator 42 at the near end. The derived
class information c(j) is provided to a Huffman decoder 47 to decode the Huffman codes
F
h(k). The output of the Huffman decoder 47 is applied to a coefficient reconstruction
module 48, which, based on the derived quantization step sizes Q(k), reconstructs
the original coefficients F
q(k). The bit rate estimator 46 further supplies class information c(j) to a coefficient
fill-in module 50 to notify it of which coefficients were not coded for transmission.
Module 50 then estimates the missing coefficients using the reconstructed log spectral
estimates Log
2 S
q 2(j).
[0090] Finally, the decoded coefficients F
q(k) and the estimated coefficients F
e(k) are supplied to a signal composer 52 which converts the coefficients back to the
time domain and unnormalizes the time domain signal using the reconstructed normalization
gain g
q.
[0091] More specifically, an inverse DCT module 51 merges the decoded and estimated coefficients
F
q(k), F
e(k) and transforms the resultant frequency coefficient values to a time domain signal
m'(n). An inverse normalization module 53 scales the time domain signal m'(n) back
to the original scale. The resultant microphone signal m(n) is applied to an overlap
decoder 55 which removes the redundant samples introduced by windowing module 32 of
the audio encoder 16. (Fig. 1(a)). A signal conditioner 57 filters the resultant microphone
signal and converts it to an analog signal L(t) for driving loudspeaker 29.
Estimation of Uncoded Coefficients
[0092] As explained above, the human ear is less likely to notice distortion of an audio
signal in regions of the signal's spectrum having a relatively large energy level.
Accordingly, bit rate estimator 42 (Fig. 1) adjusts the quantization step size to
finely quantize coefficients having a low energy level and coarsely quantize coefficients
having a large energy level. In apparent contradiction to this approach, the bit rate
estimator simply discards coefficients having a very low energy level.
[0093] The absence of these coefficients would result in noticeable audio artifacts. Accordingly,
a coefficient fill-in module 50 prepares a coefficient estimate of each discarded
coefficient from the spectral estimates. It then inserts the coefficient estimate
in place of the missing coefficient to prevent such audio artifacts. In doing so,
the coefficient module considers the level of the signal index A
i for the frame in which the uncoded coefficient resides. If, for example, the signal
index is low, (indicating that the frame largely consists of background noise), the
fill-in module assumes the missing coefficient represents background noise. Accordingly,
it prepares the coefficient estimate largely from a measure of the background noise
in the frame. However, if the signal index is large, (indicating that the frame consists
largely of a voice signal), the noise fill-in module assumes that the missing coefficient
represents a voice signal. Accordingly, it prepares the coefficient estimate largely
from the value of the spectral estimate corresponding to the band of frequencies which
includes the frequency of the missing coefficient.
[0094] Referring to Fig. 12, coefficient fill-in module 50 (Fig. 1(b)) includes a coefficient
estimator module 82 for each band. Each estimator module 82 includes a noise floor
module 84 for approximating, for each Frame F, the amount of background noise in the
band of frequencies j. The noise estimate S
n(j,F) is derived from a comparison of the log spectral estimate Log
2 S
q 2(j) of the current frame with a noise estimate derived from spectral estimates of
previous frames. An adder 91 adds the log spectral estimate to the log gain, log
2 g
q 2, to unnormalize the log spectral estimate. The unnormalized estimate S
u(j, F) is applied to a comparator 99 which compares S
u(j, F) to the noise estimate S
n(j,F-1) calculated for the previous frame F-1 (S
n is initialized to zero for the first frame). If S
u(j, F) is greater than the previous noise estimate S
n(j, F-1), the noise estimate for the present frame F is computed as follows:

where t
r is a rise time constant provided by table 100. More specifically, table 100 provides
a unique t
r for each value of the signal index A
i(F). (e.g., for A
i(F) = 0, a relatively large time constant t
r is chosen to yield a long rise time. As A
i(F) increases, the selected time constant decreases).
[0095] If S
u(j, F) is less than the previous noise estimate, S
n(j, F - 1), the noise for the current frame is computed as follows:

where t
f is a fall time constant provided by table 102. Table 102, like table 100, provides
a unique constant t
f for each value of the index A
i(F). Adder 93 normalizes the resultant noise estimate S
n(j, F) by subtracting the log gain, log
2 g
2 q. The output of 93 is then applied to a log inverter 94 which computes the inverse
logarithm as show below to provide a normalized noise estimate S
nn(j, F):

[0096] For each frame, the normalized band noise estimates, Snn(j, F), and the spectral
estimate S
q(j, F) are applied to a weighting function 86 which prepares weighted sum of the two
values. Weighting function 86 includes a first table 88 which contains, for each value
of signal index A
i(F), a noise weighting coefficient C
n(F). Similarly, a second table 90 includes for each index A
i(F), a voice weighting index C
a(F). In response to the current value of the audio index A
i(F), table 88 provides a corresponding noise weighting coefficient C
n(F) to multiplier 92. Multiplier 92 computes the product of C
n(F) and S
nn(j, F) to produce a weighted noise value S
nw(j, F). Similarly, table 90 provides a voice weighting coefficient C
a(F) to multiplier 94 to compute a weighted voice value S
vw(j, F) where S
vw = S
q(j, F) C
a. The weighted values are provided to a composer 98 which computes a weight estimate
as follows:

[0097] The weighting coefficients C
a, C
n stored in tables 90, 88 are related as follows C
a = 1 - C
n, where the values of C
n may range between zero (for A
i=7) and one (for A
i=0). Thus, during silence, the noise estimate carries more weight, while the spectral
estimate carries gradually more weight as the audio index increases.
[0098] This weighted estimate is then supplied to signal composer 52 for use in computing
the estimated frequency coefficient F
e(k) for each of the ten uncoded frequency coefficients corresponding to the spectral
estimate.
[0099] More specifically, the weighted estimate, W, is used to control the level of "fill-in"
for each missing frequency coefficient (i.e. those with a class of zero). The fill-in
consists of scaling the output from a random number generator (with a uniform distribution)
and inserting the result, F
e in place of the missing frequency coefficients. The following equation is used to
generate the fill-in for each missing frequency coefficient.

Where "noise" is the output from the random number generator at given instant (or
coef/sample), the range of the random number generator being twice the value n; and
wherein e is a constant, e.g., 3. Note that a new value of noise is generated for
each of the missing frequency coefficients.
1. A method for allocating transmission bits for use in transmitting samples of a digital
signal, the method comprising the steps of:
selecting an aggregate allowable quantization distortion value representing a collective
allowable quantization distortion error for a frame of samples of said digital signal,
selecting from said frame of samples, a set of samples wherein each of a plurality
of samples of said set is greater than a noise threshold,
computing, for each sample of said set, a sample quantization distortion value representing
an allowable quantization distortion error for said sample, wherein the sum of all
sample quantization distortion values for all samples of said set is approximately
equal to said aggregate allowable quantization distortion value,
for each sample of said set, selecting a quantization step size which yields a quantization
distortion error approximately equal to said sample's corresponding quantization distortion
value, and
quantizing said sample using said quantization step size.
2. The method of claim 1 wherein selecting a quantization step size comprises the step
of calculating a quantization step size for each sample to be transmitted based, at
least in part, on the difference between said sample and said sample's corresponding
quantization distortion value.
3. The method of claim 1 wherein said digital signal includes a noise component and a
signal component, and wherein said selection of an aggregate allowable quantization
distortion comprises the steps of:
preparing a signal index representing, for at least one sample of said frame, a signal
magnitude of said signal component relative to a noise magnitude of said noise component,
and
based on said signal index, selecting said aggregate allowable quantization distortion.
4. The method of claim 1 wherein computing said sample quantization distortion value
comprises the step of dividing said aggregate allowable quantization distortion value
by a number of samples in said frame of samples to form a first sample distortion
value.
5. The method of claim 4 wherein selecting a set of samples comprises the step of:
selecting a tentative set of samples of said digital signal wherein each sample of
said tentative set is greater than a noise threshold determined at least in part by
the value of said first sample distortion value, and wherein calculating said sample
quantization distortion value comprises the step of:
adjusting said first sample distortion value by an amount determined by the difference
between said first sample distortion value and at least one sample excluded from said
tentative set.
6. The method of claim 5 wherein selecting said tentative set of samples and adjusting
said first sample distortion value comprise the steps of:
a) identifying any noisy samples of said tentative set, each said noisy sample having
a magnitude below a current value of said first sample distortion value,
b) removing said noisy samples, if any, from said tentative set,
c) if any noisy samples are removed, increasing said first sample distortion value
by an amount determined by the difference between said first sample distortion value
and at least one said noisy sample, and
repeating steps a, b and c until an adjusted first sample distortion value is
reached for which step a identifies no additional noisy samples of said tentative
set which are greater than said adjusted first sample distortion.
7. The method of claim 6 further comprising the step of terminating said adjustment if
said steps a, b and c have been repeated a maximum number of times.
8. The method of claim 7 further comprising the steps of:
after said adjustment is terminated, estimating a bit number for a quantity of bits
required to transmit all samples of said tentative set,
comparing said estimated bit number to a maximum bit number, and
if said estimated bit number is less than or equal to said maximum bit number, selecting
a final noise threshold based on said adjusted first sample distortion value.
9. The method of claim 8 wherein if said estimated bit number exceeds said maximum bit
number, the method further comprises the steps of:
preparing a second sample distortion value,
a) selecting a second tentative set of samples of said digital signal, each of a plurality
of said samples of said second tentative set having a magnitude above said second
sample distortion value,
b) estimating the number of bits required to transmit all samples of said second tentative
set,
c) comparing said estimated bit number to said maximum bit number, and
d) if said estimated bit number is greater than said maximum bit number, increasing
said second sample distortion value by an amount determined by said selection.
10. The method of claim 9 further comprising the step of repeating steps d-g until a second
sample distortion value is reached for which step g determines that said estimated
bit number is less than or equal to said maximum bit number.
11. The method of claim 10 further comprising the steps of:
calculating said sample distortion value from said adjusted first sample distortion
value and said second sample distortion value and,
selecting a final set of samples of said digital signal, each of a plurality of said
samples of said final set having a magnitude above a final threshold determined by
the sample's corresponding sample distortion value.
12. A method for communicating a digital signal which includes a noise component and a
signal component, the method comprising the steps of:
preparing an estimation signal representative of said digital signal, yet having fewer
samples than said digital signal,
preparing a signal index representing, for at least one sample of said estimation
signal, a signal magnitude of said signal component relative to a noise magnitude
of said noise component,
based on said signal index, selecting samples of said digital signal having a sufficiently
large signal component,
transmitting said selected samples of said digital signal,
transmitting said samples of said digital estimation signal, and
reconstructing said digital signal from said transmitted selected samples and estimation
samples.
13. The method of claim 12 wherein said digital signal is a frequency domain speech signal
representative of voice information to be communicated, and each sample of said estimation
signal is a spectral estimate of the frequency domain signal in a corresponding band
of frequencies, and wherein reconstructing said frequency domain speech signal comprises
the steps of:
generating a random number for each nonselected sample of said frequency domain speech
signal,
preparing, from at least one said spectral estimate, a noise estimate of the magnitude
of a noise component of said nonselected sample,
based on said noise estimate, generating a scaling factor, and
scaling said random number according to said scaling factor to produce a reconstructed
sample representative of said nonselected sample.
14. The method of claim 13 wherein said estimation signal and said frequency domain speech
signal each comprise a series of frames, each frame representing said voice information
over a specified window of time, and wherein preparing a first noise estimate for
a current frame comprises the steps of:
for a prior frame of said estimation signal, preparing an initial noise estimate from
at least one spectral estimate representative of a band of frequencies of said prior
frame,
based on the magnitude of said signal index, selecting a rise time constant tr and a fall time constant tf,
adding said rise time constant to said initial noise estimate to form an upper threshold,
subtracting said fall time constant from said initial noise estimate to form a lower
threshold,
comparing a current spectral estimate of said current frame, representative of said
band of frequencies, to said upper and lower thresholds,
if said current spectral estimate is between said thresholds, setting said current
noise estimate equal to said current spectral estimate,
if said current spectral estimate is below said lower threshold, decrementing said
initial noise estimate by said fall time constant to form said current noise estimate,
and
if said current spectral estimate is above said upper threshold, incrementing said
initial noise estimate by said rise time constant to form said current noise estimate.
15. The method of claim 14 wherein generating a scaling factor comprises the step of generating
a weighted sum of said current noise estimate and said current spectral estimate.
16. The method of claim 15 wherein generating said weighted sum comprises the steps of:
selecting a noise coefficient based on the value of said signal index,
selecting a voice coefficient based on the value of said signal index,
multiplying said current noise estimate by said noise coefficient,
multiplying said current spectral estimate by said voice coefficient,
adding the products of said multiplication steps, and
forming said scaling factor from the result of said adding step.
17. An encoding device for allocating transmission bits for use in transmitting samples
of a digital signal, the encoding device comprising:
means for selecting an aggregate allowable quantization distortion value representing
a collective allowable quantization distortion error for a frame of samples of said
digital signal,
means for selecting from said frame of samples, a set of samples wherein each of a
plurality of samples of said set is greater than a noise threshold,
means for computing, for each sample of said set, a sample quantization distortion
value representing an allowable quantization distortion error for said sample, wherein
the sum of all sample quantization distortion values for all samples of said set is
approximately equal to said aggregate allowable quantization distortion value,
means for selecting, for each sample of said set, a quantization step size which yields
a quantization distortion error approximately equal to said sample's corresponding
quantization distortion value, and
means for quantizing said sample using said quantization step size.
18. The encoding device of claim 17 wherein said means for selecting a quantization step
size comprises means for calculating a quantization step size for each sample to be
transmitted based, at least in part, on the difference between said sample and said
sample's corresponding quantization distortion value.
19. The encoding device of claim 17 wherein said digital signal includes a noise component
and a signal component,
and wherein said means for selecting an aggregate allowable quantization distortion
comprises:
means for preparing a signal index representing, for at least one sample of said frame,
a signal magnitude of said signal component relative to a noise magnitude of said
noise component, and
means for selecting said aggregate allowable quantization distortion, based on said
signal index.
20. The encoding device of claim 17 wherein said means for computing said sample quantization
distortion value comprises means for dividing said aggregate allowable quantization
distortion value by a number of samples in said frame of samples to form a first sample
distortion value.
21. The encoding device of claim 20 wherein said means for selecting a set of samples
comprises:
means for selecting a tentative set of samples of said digital signal wherein each
sample of said tentative set is greater than a noise threshold determined at least
in part by the value of said first sample distortion value, and wherein said means
for computing said sample quantization distortion value comprises:
means for adjusting said first sample distortion value by an amount determined by
the difference between said first sample distortion value and at least one sample
excluded from said tentative set.
22. The encoding device of claim 21 wherein said means for selecting said tentative set
of samples and said means for adjusting said first sample distortion value comprise:
means for identifying any noisy samples of said tentative set, each said noisy sample
having a magnitude below a current value of said first sample distortion value,
means for removing said noisy samples, if any, from said tentative set,
means for increasing said first sample distortion value by an amount determined by
the difference between said first sample distortion value and at least one said noisy
sample, if any noisy samples are removed, and
means for repeatedly removing noisy samples and adjusting said first sample distortion
value until an adjusted first sample distortion value is reached for which no additional
noisy samples of said tentative set are greater than said adjusted first sample distortion.
23. The encoding device of claim 22 further comprising means for terminating said adjustment
if said first sample distortion value is adjusted a maximum number of times.
24. The encoding device of claim 23 further comprising:
means for estimating after said adjustment is terminated, a bit number for a quantity
of bits required to transmit all samples of said tentative set,
means for comparing said estimated bit number to a maximum bit number, and
means for selecting a final noise threshold based on said adjusted first sample distortion
value, if said estimated bit number is less than or equal to said maximum bit number.
25. The encoding device of claim 24 wherein the encoding device further comprises:
means for preparing a second sample distortion value, if said estimated bit number
exceeds said maximum bit number,
means for selecting a second tentative set of samples of said digital signal, each
of a plurality of said samples of said second tentative set having a magnitude above
said second sample distortion value,
means for estimating the number of bits required to transmit all samples of said second
tentative set,
means for comparing said estimated bit number to said maximum bit number, and
means for increasing said second sample distortion value by an amount determined by
said selection, if said estimated bit number is greater than said maximum bit number.
26. The encoding device of claim 25 further comprising means for repeatedly adjusting
said second sample distortion value, reselecting said second tentative set of samples,
and estimating the number of bits required to transmit all samples of said second
tentative set until a second sample distortion value is reached for which said estimated
bit number is less than or equal to said maximum bit number.
27. The encoding device of claim 26 further comprising:
means for calculating said sample distortion value from said adjusted first sample
distortion value and said second sample distortion value and,
means for selecting a final set of samples of said digital signal, each of a plurality
of said samples of said final set having a magnitude above a final threshold determined
by the sample's corresponding sample distortion value.
28. An encoding device for communicating a digital signal which includes a noise component
and a signal component, the encoding device comprising:
means for preparing an estimation signal representative of said digital signal, yet
having fewer samples than said digital signal,
means for preparing a signal index representing, for at least one sample of said estimation
signal, the magnitude of said signal component relative to the magnitude of said noise
component,
means for selecting, based on said signal index, samples of said digital signal having
a sufficiently large signal component,
means for transmitting said selected samples of said digital signal,
means for transmitting said samples of said digital estimation signal, and
means for reconstructing said digital signal from said transmitted selected samples
and estimation samples.
29. The encoding device of claim 28 wherein said digital signal is a frequency domain
speech signal representative of voice information to be communicated, and each sample
of said estimation signal is a spectral estimate of the frequency domain signal in
a corresponding band of frequencies, and wherein said means for reconstructing said
frequency domain speech signal comprises:
means for generating a random number for each nonselected sample of said frequency
domain speech signal,
means for preparing, from at least one said spectral estimate, a noise estimate of
the magnitude of a noise component of said nonselected sample,
means for generating a scaling factor based on said noise estimate, and
means for scaling said random number according to said scaling factor to produce a
reconstructed sample representative of said nonselected sample.
30. The encoding device of claim 29 wherein said estimation signal and said frequency
domain speech signal each comprise a series of frames, each frame representing said
voice information over a specified window of time, and wherein said means for preparing
a first noise estimate for a current frame comprises:
means for preparing, for a prior frame of said estimation signal, an initial noise
estimate from at least one spectral estimate representative of a band of frequencies
of said prior frame,
means for selecting, based on the magnitude of said signal index, a rise time constant
tr and a fall time constant tf,
means for adding said rise time constant to said initial noise estimate to form an
upper threshold,
means for subtracting said fall time constant from said initial noise estimate to
form a lower threshold,
means for comparing a current spectral estimate representative of said band of frequencies
in said current frame, to said upper and lower thresholds,
means for setting said current noise estimate equal to said current spectral estimate,
if said current spectral estimate is between said thresholds,
means for decrementing said initial noise estimate by said fall time constant to form
said current noise estimate, if said current spectral estimate is below said lower
threshold, and
means for incrementing said initial noise estimate by said rise time constant to form
said current noise estimate, if said current spectral estimate is above said upper
threshold.
31. The encoding device of claim 30 wherein said means for generating a scaling factor
comprises means for generating a weighted sum of said current noise estimate and said
current spectral estimate.
32. The encoding device of claim 31 wherein said means for generating said weighted sum
comprises:
means for selecting a noise coefficient based on the value of said signal index,
means for selecting a voice coefficient based on the value of said signal index,
means for multiplying said current noise estimate by said noise coefficient,
means for multiplying said current spectral estimate by said voice coefficient,
means for adding the products of said multiplications, and
means for forming said scaling factor from the result of said adding.
1. Verfahren zum Zuteilen von Übertragungsbits zur Verwendung bei der Übertragung von
Abtastwerten eines digitalen Signales, welches Verfahren die Schritte aufweist:
Wählen eines summierten, zulässigen Quantisierungsverzerrungswertes, der einen kollektiven,
zulässigen Quantisierungsverzerrungsfehler für einen Block von Abtastwerten des genannten
digitalen Signales darstellt,
Auswählen einer Gruppe von Abtastwerten aus dem genannten Block von Abtastwerten,
wobei jeder Abtastwert aus einer Mehrzahl von Abtastwerten der genannten Gruppe größer
ist als eine Geräuschschwelle,
für jeden Abtastwert der genannten Gruppe Berechnen eines Abastwert-Quantisierungsverzerrungswertes,
welcher einen zulässigen Quantisierungsverzerrungsfehler für den genannten Abtastwert
darstellt, wobei die Summe sämtlicher Abtastwert-Quantisierungsverzerrungswerte für
sämtliche Abtastwerte der genannten Gruppe etwa gleich dem genannten summierten, zulässigen
Quantisierungsverzerrungswert ist,
für jeden Abtastwert der genannten Gruppe Wählen einer Quantisierungsschrittgröße,
die zu einem Quantisierungsverzerrungsfehler führt, der etwa gleich dem entsprechenden
Quantisierungsverzerrungswert des genannten Abtastwertes ist, und
Quantisieren des genannten Abtastwertes unter Verwendung der genannten Quantisierungsschrittgröße.
2. Verfahren nach Anspruch 1, bei dem das Wählen einer Quantisierungsschrittgröße den
Schritt des Berechnens einer Quantisierungsschrittgröße für jeden zu übertragenden
Abtastwert beinhaltet, basierend zumindest teilweise auf der Differenz zwischen dem
genannten Abtastwert und dem entsprechenden Quantisierungsverzerrungswert des genannten
Abtastwertes.
3. Verfahren nach Anspruch 1, bei dem das genannte digitale Signal eine Geräuschkomponente
und eine Signalkomponente beinhaltet und bei dem die genannte Wahl einer summierten,
zulässigen Quantisierungsverzerrung die Schritte beinhaltet:
Bereiten eines Signalindexes, der für zumindest einen Abtastwert des genannten Blockes
eine Signalgröße der genannten Signalkomponente relativ zu einer Rauschgröße der genannten
Geräuschkomponente darstellt, und
Wählen der genannten summierten, zulässigen Quantisierungsverzerrung auf Grundlage
des genannten Signal indexes.
4. Verfahren nach Anspruch 1, bei dem das Berechnen des genannten Abtastwert-Quantisierungsverzerrungswertes
den Schritt der Division des genannten summierten, zulässigen Quantisierungsverzerrungswertes
durch eine Anzahl von Abtastwerten in dem genannten Block von Abtastwerten beinhaltet,
um einen ersten Abtastwert-Verzerrungswert zu bilden.
5. Verfahren nach Anspruch 4, bei dem das Auswählen einer Gruppe von Abtastwerten den
Schritt beinhaltet:
Auswählen einer provisorischen Gruppe von Abtastwerten des genannten digitalen Signales,
wobei jeder Abtastwert der genannten provisorischen Gruppe größer ist als eine Geräuschschwelle,
die zumindest teilweise durch den Wert des genannten ersten Abtastwert-Verzerrungswertes
bestimmt ist, und wobei das Berechnen des genannten Abtastwert-Quantisierungsverzerrungswertes
den Schritt aufweist:
Verstellen des genannten ersten Abtastwert-Verzerrungswertes um einen Betrag, der
durch die Differenz zwischen dem genannten ersten Abtastwert-Verzerrungswert und zumindest
einem Abtastwert bestimmt ist, der aus der genannten provisorischen Gruppe ausgeschlossen
ist.
6. Verfahren nach Anspruch 5, bei dem das Auswählen der genannten provisorischen Gruppe
von Abtastwerten und das Verstellen des genannten ersten Abtastwert-Verzerrungswertes
die Schritte beinhaltet:
a) Identifizieren jedweder rauschigen Abtastwerte der genannten provisorischen Gruppe,
wobei jeder genannte rauschige Abtastwert eine geringere Größe als ein laufender Wert
des genannten ersten Abtastwert-Verzerrungswertes besitzt,
b) Entfernen der genannten rauschigen Abtastwerte, falls vorhanden, aus der genannten
provisorischen Gruppe,
c) Vergrößern, wenn irgendwelche rauschigen Abtastwerte entfernt werden, des genannten
ersten Abtastwert-Verzerrungswertes um einen Betrag, der durch die Differenz zwischen
dem genannten ersten Abtastwert-Verzerrungswert und zumindest einem genannten rauschigen
Abtastwert bestimmt ist, und
Wiederholen der Schritte a, b und c bis ein eingestellter erster Abtastwert-Verzerrungswert
erreicht ist, für den der Schritt a keine zusätzlichen rauschigen Abtastwerte der
genannten provisorischen Gruppe identifiziert, welche größer sind als die genannte
eingestellte erste Abtastwertverzerrung.
7. Verfahren nach Anspruch 6, das ferner den Schritt des Abbrechens der genannten Verstellung
beinhaltet, wenn die Schritte a, b und c in einer Höchstanzahl von Wiederholungen
durchgeführt sind.
8. Verfahren nach Anspruch 7, außerdem die Schritte beinhaltend:
Abschätzen, nachdem die genannte Verstellung abgebrochen ist, einer Bitanzahl für
eine Menge von Bits, die erforderlich sind, um sämtliche Abtastwerte der genannten
provisorischen Gruppe zu übertragen,
Vergleichen der genannten abgeschätzten Bitanzahl mit einer maximalen Bitanzahl und
Wählen, wenn die genannte abgeschätzte Bitanzahl kleiner oder gleich der genannten
maximalen Bitanzahl ist, einer endgültigen Geräuschschwelle auf Grundlage des genannten
eingestellten ersten Abtastwert-Verzerru ngswertes.
9. Verfahren nach Anspruch 8, bei dem, wenn die genannte abgeschätzte Bitanzahl die genannte
maximale Bitanzahl übersteigt, das Verfahren außerdem die Schritte beinhaltet:
Bereiten eines zweiten Abtastwert-Verzerrungswertes,
a) Auswählen einer zweiten provisorischen Gruppe von Abtastwerten des genannten digitalen
Signales, wobei jeder Abtastwert einer Mehrzahl der genannten Abtastwerte der genannten
zweiten provisorischen Gruppe eine Größe besitzt, die oberhalb des genannten zweiten
Abtastwert-Verzerrungswertes liegt,
b) Abschätzen der Anzahl von Bits, die erforderlich sind, um sämtliche Abtastwerte
der genannten zweiten provisorischen Gruppe zu übertragen,
c) Vergleichen der genannten abgeschätzten Bitanzahl mit der genannten maximalen Bitanzahl
und
d) Vergrößern, wenn die genannte abgeschätzte Bitanzahl größer ist als die genannte
maximale Bitanzahl, des genannten zweiten Abtastwert-Verzerrungswertes um einen Betrag,
der durch die genannte Auswahl bestimmt ist.
10. Verfahren nach Anspruch 9, das ferner den Schritt des Wiederholens der Schritte d
- g beinhaltet, bis ein zweiter Abtastwert-Verzerrungswert erreicht ist, für den der
Schritt g ermittelt, daß die genannte abgeschätzte Bitanzahl kleiner oder gleich der
genannten maximalen Bitanzahl ist.
11. Verfahren nach Anspruch 10, das außerdem die Schritte aufweist:
Berechnen des genannten Abtastwert-Verzerrungswertes aus dem genannten eingestellten,
ersten Abtastwert-Verzerrungswert und dem genannten zweiten Abtastwert-Verzerrungswert
und
Auswählen einer endgültigen Gruppe von Abtastwerten des genannten digitalen Signales,
wobei jeder Abtastwert einer Mehrzahl der genannten Abtastwerte der genannten endgültigen
Gruppe eine Größe besitzt, die oberhalb einer endgültigen Schwelle liegt, die durch
den entsprechenden Abtastwert-Verzerrungswert des Abtastwertes bestimmt ist.
12. Verfahren zur Übertragung eines digitalen Signales, das eine Geräuschkomponente und
eine Signalkomponente beinhaltet, wobei das Verfahren die Schritte beinhaltet:
Bereiten eines abgeschätzten Signales, das für das genannte digitale Signal repräsentativ
ist, jedoch weniger Abtastwerte aufweist als das genannte digitale Signal,
Bereiten eines Signalindexes, der, für zumindest einen Abtastwert des genannten abgeschätzten
Signales, eine Signalgröße der genannten Signalkomponente relativ zu einer Rauschgröße
der genannten Geräuschkomponente darstellt,
Auswählen, basierend auf dem genannten Signalindex, von Abtastwerten des genannten
digitalen Signales, die eine ausreichend große Signalkomponente besitzen,
Übertragen der genannten ausgewählten Abtastwerte des genannten digitalen Signales,
Übertragen der genannten Abtastwerte des genannten digitalen, abgeschätzten Signales
und
Rekonstruieren des genannten digitalen Signales aus den genannten übertragenen ausgewählten
Abtastwerten und den abgeschätzten Abtastwerten.
13. Verfahren nach Anspruch 12, bei dem das genannte digitale Signal ein Frequenzbereich-Sprachsignal
ist, das eine zu übertragende Toninformation darstellt, und jeder Abtastwert des genannten
abgeschätzten Signales ein spektraler Schätzwert des Frequenzbereichsignales in einem
entsprechenden Frequenzband ist und bei dem die Rekonstruktion des genannten Frequenzbereich-Sprachsignales
die Schritte beinhaltet:
Generieren einer Zufallszahl für jeden nicht ausgewählten Abtastwert des genannten
Frequenzbereich-Sprachsignales,
Bereiten, aus zumindest einem genannten spektralen Schätzwert, einen Geräuschschätzwert
für die Größe einer Geräuschkomponente des nicht gewählten Abtastwertes,
Erzeugen, auf Grundlage des genannten Geräuschschätzwertes, eines Skalierungsfaktors
und
Skalieren der genannten Zufallszahl gemäß dem genannten Skalierungsfaktor, um einen
rekonstruierten Abtastwert zu erzeugen, der für den genannten nicht ausgewählten Abtastwert
repräsentativ ist.
14. Verfahren nach Anspruch 13, bei dem das genannte abgeschätzte Signal und das genannte
Frequenzbereich-Sprachsignal je eine Reihe von Blöcken aufweisen, von denen jeder
Block die genannte Toninformation über ein spezifiziertes Zeitfenster darstellt, und
bei dem das Bereiten eines ersten Geräuschschätzwertes für einen laufenden Block die
Schritte beinhaltet:
Bereiten, für einen voraufgehenden Block des genannten abgeschätzten Signales, einen
anfänglichen Geräuschschätzwert aus zumindest einem spektralen Schätzwert, der für
ein Frequenzband des genannten voraufgehenden Blockes repräsentativ ist,
Wählen, basierend auf der Größe des genannten Signalindexes, eine Anstiegszeitkonstante
t, und eine Abfallzeitkonstante tf,
Addieren der genannten Anstiegszeitkonstante zu dem genannten anfänglichen Geräuschschätzwert,
um eine obere Schwelle zu bilden,
Subtrahieren der genannten Abfallzeitkonstante von dem genannten anfänglichen Geräuschschätzwert,
um eine untere Schwelle zu bilden,
Vergleichen eines laufenden spektralen Schätzwertes des genannten laufenden Blockes,
der für das genannte Frequenzband repräsentativ ist, mit der genannten oberen und
der genannten unteren Schwelle,
Setzen, wenn der genannte laufende spektrale Schätzwert sich zwischen den genannten
Schwellen befindet, des genannten laufenden Schätzwertes gleich dem genannten laufenden
spektralen Schätzwert,
Dekrementieren, wenn der genannte spektrale Schätzwert unterhalb der genannten unteren
Schwelle liegt, des genannten anfänglichen Geräuschschätzwertes mit der genannten
Abfallzeitkonstante, um den genannten laufenden Geräuschschätzwert zu bilden, und
Inkrementieren, wenn der genannte laufende spektrale Schätzwert oberhalb der genannten
oberen Schwelle liegt, des genannten anfänglichen Geräuschschätzwertes mit der Anstiegszeitkonstante,
um den genannten laufenden Geräuschschätzwert zu bilden.
15. Verfahren nach Anspruch 14, bei dem das Erzeugen eines Skalierungsfaktors den Schritt
des Erzeugens einer gewichteten Summe aus dem genannten laufenden Geräuschschätzwert
und dem genannten laufenden spektralen Schätzwert beinhaltet.
16. Verfahren nach Anspruch 15, bei dem das Erzeugen der genannten gewichteten Summe die
Schritte beinhaltet:
Wählen eines Geräuschkoeffizienten auf der Basis des Werte des genannten Signalindexes,
Wählen eines Tonkoeffizienten auf Grundlage des Wertes des genannten Signalindexes,
Multiplizieren des genannten laufenden Geräuschschätzwertes mit dem genannten Geräuschkoeffizienten,
Multiplizieren des genannten laufenden spektralen Schätzwertes mit dem genannten Tonkoeffizienten,
Addieren der Produkte der genannten Multiplikationsschritte und
Bilden des genannten Skalierungsfaktors aus dem Ergebnis des genannten Additionsschrittes.
17. Encoder zum Zuteilen von Übertragungsbits zur Verwendung bei der Übertragung von Abtastwerten
eines digitalen Signales, wobei der Encoder aufweist:
ein Mittel zum Wählen eines summierten, zulässigen Quantisierungsverzerrungswertes,
der einen kollektiven, zulässigen Quantisierungsverzerrungsfehler für einen Block
von Abtastwerten des genannten digitalen Signales darstellt,
ein Mittel zum Auswählen einer Gruppe von Abtastwerten aus dem genannten Block von
Abtastwerten, wobei jeder Abtastwert aus einer Mehrzahl von Abtastwerten der genannten
Gruppe größer ist als eine Geräuschschwelle,
ein Mittel, um für jeden Abtastwert der genannten Gruppe einen Abtastwert-Quantisierungsverzerrungswert
zu berechnen, der einen zulässigen Quantisierungsverzerrungsfehler für den genannten
Abtastwert darstellt, wobei die Summe sämtlicher Abtastwert-Quantisierungsverzerrungswerte
für sämtliche Abtastwerte der genannten Gruppe etwa gleich dem genannten summierten,
zulässigen Quantisierungsverzerrungswert ist,
ein Mittel, um für jeden Abtastwert der genannten Gruppe eine Quantisierungsschrittgröße
zu wählen, die zu einem Quantisierungsverzerrungsfehler führt, der etwa gleich dem
entsprechenden Quantisierungsverzerrungswert des genannten Abtastwertes ist, und
ein Mittel zum Quantisieren des genannten Abtastwertes unter Verwendung der genannten
Quantisierungsschrittgröße.
18. Encoder nach Anspruch 17, bei dem das genannte Mittel zum Wählen einer Quantisierungsschrittgröße
ein Mittel zum Berechnen einer Quantisierungsschrittgröße für jeden zu übertragenden
Abtastwert beinhaltet, basierend zumindest teilweise auf der Differenz zwischen dem
genannten Abtastwert und dem entsprechenden Quantisierungswert des genannten Abtastwertes.
19. Encoder nach Anspruch 17, bei dem das genannte digitale Signal eine Geräuschkomponente
und eine Signalkomponente beinhaltet,
und bei dem das genannte Mittel zum Wählen einer summierten, zulässigen Quantisierungsverzerrung
beinhaltet:
ein Mittel zum Bereiten eines Signalindexes, der für zumindest einen Abtastwert des
genannten Blockes eine Signalgröße der genannten Signalkomponente relativ zu einer
Rauschgröße der genannten Geräuschkomponente darstellt, und
ein Mittel zum Wählen der genannten summierten, zulässigen Quantisierungsverzerrung
auf Grundlage des genannten Singalindexes.
20. Encoder nach Anspruch 17, bei dem das genannte Mittel zum Berechnen des genannten
Abtastwert-Quantisierungsverzerrungswertes ein Mittel zum Dividieren des genannten
summierten, zulässigen Quantisierungsverzerrungswertes durch eine Anzahl von Abtastwerten
in dem genannten Block von Abtastwerten beinhaltet, um eine ersten Abtastwert-Verzerrungswert
zu bilden.
21. Encoder nach Anspruch 20, bei dem das genannte Mittel zum Auswählen einer Gruppe von
Abtastwerten aufweist:
ein Mittel zum Auswählen einer provisorischen Gruppe von Abtastwerten des genannten
digitalen Signales, wobei jeder Abtastwert der genannten provisorischen Gruppe größer
ist als eine Geräuschschwelle, die zumindest teilweise durch den Wert des genannten
ersten Abtastwert-Verzerrungswertes bestimmt ist, und wobei das genannte Mittel zum
Berechnen des genannten Abtastwert-Quantisierungsverzerrungswertes aufweist:
ein Mittel zum Verstellen des genannten ersten Abtastwert-Verzerrungswertes um einen
Betrag, der durch die Differenz zwischen dem genannten ersten Abtastwert-Verzerrungswert
und zumindest einem Abtastwert bestimmt ist, der aus der genannten provisorischen
Gruppe ausgeschlossen ist.
22. Encoder nach Anspruch 21, bei dem das genannte Mittel zum Auswählen der genannten
provisorischen Gruppe von Abtastwerten und das genannte Mittel zum Verstellen des
genannten erten Abtastwert-Verzerrungswertes aufweisen:
ein Mittel zum Identifizieren jedweder rauschiger Abtastwerte der genannten provisorischen
Gruppe, wobei jeder genannte rauschige Abtastwert eine geringere Größe als ein laufender
Wert des genannten ersten Abtastwert-Verzerrungswertes besitzt,
ein Mittel zum Entfernen der genannten rauschigen Abtastwerte, falls vorhanden, aus
der genannten provisorischen Gruppe,
ein Mittel zum Vergrößern des genanten ersten Abtastwert-Verzerrungswertes um einen
Betrag, der durch die Differenz zwischen dem genannten ersten Abtastwert-Verzerrungswert
und zumindest einem genannten rauschigen Abtastwert bestimmt ist, wenn irgendwelche
rauschigen Abtastwerte entfernt werden, und
ein Mittel zum wiederholten Entfernen rauschiger Abtastwerte und Verstellen des genannten
ersten Abtastwert-Verzerrungswertes, bis ein eingestellter erster Abtastwert-Verzerrungswert
erreicht ist, bei dem keine zusätzlichen rauschigen Abtastwerte der genannten provisorischen
Gruppe größer sind als die genannte, eingestellte erste Abtastwert-Verzerrung.
23. Encoder nach Anspruch 22, außerdem ein Mittel beinhaltend, um die genannte Verstellung
abzubrechen, wenn der genannte erste Abtastwert-Verzerrungswert eine Größtanzahl von
Malen verstellt ist.
24. Encoder nach Anspruch 23, außerdem aufweisend:
ein Mittel, um, nachdem die genannte Verstellung abgebrochen ist, eine Bitanzahl für
eine Menge von Bits abzuschätzen, die erforderlich sind, um sämtliche Abtastwerte
der genannten provisorischen Gruppe zu übertragen,
ein Mittel zum Vergleichen der genannten abgeschätzten Bitanzahl mit einer maximalen
Bitanzahl und
ein Mittel, um eine endgültige Geräuschschwelle auf Grundlage des genannten eingestellten
ersten Abtastwert-Verzerrungswertes zu wählen, wenn die genannte abgeschätzte Bitanzahl
kleiner oder gleich der genannten maximalen Bitanzahl ist.
25. Encoder nach Anspruch 24, bei dem der Encoder außerdem aufweist:
ein Mittel zum Bereiten eines zweiten Abtastwert-Verzerrungswertes, wenn die genannte
abgeschätzte Bitanzahl die genannte maximale Bitanzahl übersteigt,
ein Mittel zum Auswählen einer zweiten provisorischen Gruppe von Abtastwerten des
genannten digitalen Signales, wobei jeder Abtastwert einer Mehrzahl der genannten
Abtastwerte der genannten zweiten provisorischen Gruppe eine Größe besitzt, die oberhalb
des genannten zweiten Abtastwert-Verzerrungswertes liegt,
ein Mittel zum Abschätzen der Anzahl von Bits, die erforderlich sind, um sämtliche
Abtastwerte der genannten zweiten provisorischen Gruppe zu übertragen,
ein Mittel zum Vergleichen der genannten abgeschätzten Bitanzahl mit der genannten
maximalen Bitanzahl und
ein Mittel, um den genannten zweiten Abtastwert-Verzerrungswert um einen Betrag zu
vergrößern, der durch die genannte Auswahl bestimmt ist, wenn die genannte abgeschätzte
Bitanzahl größer ist als die genannte maximale Bitanzahl.
26. Encoder nach Anspruch 25, außerdem ein Mittel aufweisend, um den genannten zweiten
Abtastwert-Verzerrungswert wiederholt zu verstellen, die genannte zweite provisorische
Gruppe von Abtastwerten neu auszuwählen und die Anzahl von Bits abzuschätzen, die
erforderlich sind, um sämtliche Abtastwerte der genanten zweiten provisorischen Gruppe
zu übertragen, bis ein zweiter Abtastwert-Verzerrungswert erreicht ist, bei dem die
genannte abgeschätzte Bitanzahl kleiner oder gleich der genannten maximalen Bitanzahl
ist.
27. Encoder nach Anspruch 26, außerdem aufweisend:
ein Mittel zum Berechnen des genannten Abtastwert-Verzerrungswertes aus dem genannten
eingestellten, ersten Abtastwert-Verzerrungswert und dem genannten zweiten Abtastwert-Verzerrungswert
und
ein Mittel zum Auswählen einer endgültigen Gruppe von Abtastwerten des genannten digitalen
Signales, wobei jeder Abtastwert einer Mehrzahl der genannten Abtastwerte der genannten
endgültigen Gruppe eine Größe besitzt, die oberhalb einer endgültigen Schwelle liegt,
die durch den entsprechenden Abtastwert-Verzerrungswert des Abtastwertes bestimmt
ist.
28. Encoder zur Übertragung eines digitalen Signales, das eine Geräuschkomponente und
eine Signalkomponente beinhaltet, wobei der Encoder aufweist:
ein Mittel zum Bereiten eines abgeschätzten Signales, das für das genannte digitale
Signal repräsentativ ist, jedoch weniger Abtastwerte aufweist als das genannte digitale
Signal,
ein Mittel zum Bereiten eines Signalindexes, der für zumindest einen Abtastwert des
genannten abgeschätzten Signales die Größe der genannten Signalkomponente relativ
zu der Größe der genannten Geräuschkomponente darstellt,
ein Mittel, um, basierend auf dem genannten Signalindex, Abtastwerte des genannten
digitalen Signales auszuwählen, die eine ausreichend große Signalkomponente besitzen,
ein Mittel zum Übertragen der genannten ausgewählten Abtastwerte des genannten digitalen
Signales,
ein Mittel zum Übertragen der genannten Abtastwerte des genannten digitalen, abgeschätzten
Signales und
ein Mittel zum Rekonstruieren des genannten digitalen Signales aus den genannten übertragenen
ausgewählten Abtastwerten und den abgeschätzten Abtastwerten.
29. Encoder nach Anspruch 28, bei dem das genannte digitale Signal ein Frequenzbereich-Sprachsignal
ist, das eine zu übertragende Toninformation darstellt, und jeder Abtastwert des genannten
abgeschätzten Signales ein spektraler Schätzwert des Frequenzbereichsignales in einem
entsprechenden Frequenzband ist und bei dem das genannte Mittel zur Rekonstruktion
des genannten Frequenzbereich-Sprachsignales aufweist:
ein Mittel zum Generieren einer Zufallszahl für jeden nicht ausgewählten Abtastwert
des genannten Frequenzbereich-Sprachsignales,
ein Mittel, um aus zumindest einem genannten spektralen Schätzwert einen Geräuschschätzwert
für die Größe einer Geräuschkomponente des nicht ausgewählten Abtastwertes zu bereiten,
ein Mittel, um auf der Grundlage des genannten Geräuschschätzwertes einen Skalierungsfaktor
zu erzeugen, und
ein Mittel zum Skalieren der genannten Zufallszahl gemäß dem genannten Skalierungsfaktor,
um einen rekonstruierten Abtastwert zu erzeugen, der für den genannten nicht ausgewählten
Abtastwert repräsentativ ist.
30. Encoder nach Anspruch 29, bei dem das genannte abgeschätzte Signal und das genannte
Frequenzbereich-Sprachsignal je eine Reihe von Blöcken beinhalten, von denen jeder
Block die genannte Toninformation über ein spezifiziertes Zeitfenster darstellt, und
bei der das genannte Mittel zum Bereiten eines ersten Geräuschschätzwertes für einen
laufenden Block aufweist:
ein Mittel, um für einen voraufgehenden Block des genannten abgeschätzten Signales
einen anfänglichen Geräuschschätzwert aus zumindest einem spektralen Schätzwert zu
bereiten, der für ein Frequenzband des genannten voraufgehenden Blockes repräsentativ
ist,
ein Mittel, um auf Grundlage der Größe des genannten Signalindexes eine Anstiegszeitkonstante
tr und eine Abfallzeitkonstante tf zu wählen,
ein Mittel, um die genannte Anstiegszeitkonstante mit dem genannten anfänglichen Geräuschschätzwert
zu addieren, um eine obere Schwelle zu bilden,
ein Mittel, um die genannte Abfallzeitkonstante von dem genannten anfänglichen Geräuschschätzwert
zu subtrahieren, um eine untere Schwelle zu bilden,
ein Mittel zum Vergleichen eines laufenden spektralen Schätzwertes, der für das Frequenzband
in dem genannten laufenden Block repräsentativ ist, mit der genannten oberen und genannten
unteren Schwelle,
ein Mittel, um den genannten laufenden Geräuschschätzwert gleich dem genannten laufenden
spektralen Schätzwert zu setzen, wenn der genannte laufende spektrale Schätzwert zwischen
den genannten Schwellen liegt,
ein Mittel, um den genannten anfänglichen Geräuschschätzwert mit der genannten Abfallzeitkonstante
zu dekrementieren, um den genannten laufenden Geräuschschätzwert zu bilden, wenn der
genannte laufende spektrale Schätzwert unterhalb der genannten unteren Schwelle liegt,
und
ein Mittel, um den genannten anfänglichen Geräuschschätzwert mit der genannten Anstiegszeitkonstante
zu inkrementieren, um den genannten laufenden Geräuschschätzwert zu bilden, wenn der
genannte laufende spektrale Schätzwert oberhalb der genannten oberen Schwelle liegt.
31. Encoder nach Anspruch 30, bei dem das genannte Mittel zum Erzeugen eines Skalierungsfaktors
ein Mittel aufweist, um eine gewichtete Summe des genannten laufenden Geräuschschätzwertes
und des genannten laufenden spektralen Schätzwertes zu erzeugen.
32. Encoder nach Anspruch 31, bei dem das genannte Mittel zum Erzeugen der genannten gewichteten
Summe aufweist:
ein Mittel, um einen Geräuschkoeffizienten auf Grundlage des Wertes des genannten
Signalindexes zu wählen,
ein Mittel, um einen Tonkoeffizienten auf Grundlage des Wertes des genannten Signalindexes
zu wählen,
ein Mittel, um den genannten laufenden Geräuschschätzwert mit dem genannten Geräuschkoeffizienten
zu multiplizieren,
ein Mittel, um den genannten laufenden spektralen Schätzwert mit dem genannten Tonkoeffizienten
zu multiplizieren,
ein Mittel, um die Produkte der genannten Multiplikationen zu addieren, und
ein Mittel, um den genannten Skalierungsfaktor aus dem Ergebnis der genannten Addition
zu bilden.
1. Procédé pour attribuer des éléments binaires de transmission destinés à être utilisés
pour la transmission d'échantillons d'un signal numérique, le procédé comprenant les
étapes consistant à :
sélectionner une valeur de distorsion de quantification acceptable globale représentant
une erreur de distorsion de quantification acceptable collective pour une structure
d'échantillons dudit signal numérique,
sélectionner à partir de ladite structure d'échantillons, un jeu d'échantillons dans
lequel chacune d'une pluralité d'échantillons dudit jeu est supérieure à un seuil
de bruit,
calculer, pour chaque échantillon dudit jeu, une valeur de distorsion de quantification
d'échantillon représentant une erreur de distorsion de quantification acceptable pour
ledit échantillon, et où la somme de toutes les valeurs de distorsion de quantification
d'échantillon pour tous les échantillons dudit jeu est approximativement égale à ladite
valeur de distorsion de quantification acceptable globale,
sélectionner, pour chaque échantillon dudit jeu, une étape de taille de quantification
qui fournit une erreur de distorsion de quantification approximativement égale à ladite
valeur de distorsion de quantification correspondante des échantillons, et
quantifier ledit échantillon en utilisant ladite étape de taille de quantification.
2. Procédé selon la revendication 1 dans lequel la sélection d'une étape de taille de
quantification comprend l'étape consistant à calculer une étape de taille de quantification
pour chaque échantillon qui doit être transmis en se fondant, au moins en partie,
sur la différence entre ledit échantillon et ladite valeur de distorsion de quantification
correspondante des échantillons.
3. Procédé selon la revendication 1 dans lequel ledit signal numérique comprend une composante
de bruit et une composante de signal, et dans lequel ladite sélection d'une distorsion
de quantification acceptable globale comprend les étapes consistant à :
préparer un indice de signal représentant, pour au moins un échantillon de ladite
structure, une grandeur de signal de ladite composante de signal relative à une grandeur
de bruit de ladite composante de bruit, et
en se fondant sur ledit indice de signal, sélectionner ladite distorsion de quantification
acceptable globale.
4. Procédé selon la revendication 1, dans lequel le calcul de ladite valeur de distorsion
de quantification d'échantillon comprend l'étape consistant à diviser ladite valeur
de distorsion de quantification acceptable globale par un nombre d'échantillons dans
ladite structure d'échantillons de façon à former une première valeur de distorsion
d'échantillon.
5. Procédé selon la revendication 4, dans lequel la sélection d'un jeu d'échantillons
comprend l'étape consistant à :
sélectionner un jeu expérimental d'échantillons dudit signal numérique dans lequel
chaque échantillon dudit jeu expérimental est supérieur à un seuil de bruit déterminé
au moins en partie par la valeur de ladite première valeur de distorsion d'échantillon,
et dans lequel le calcul de ladite valeur de distorsion de quantification d'échantillon
comprend l'étape consistant à :
ajuster ladite première valeur de distorsion d'échantillon d'une quantité déterminée
par la différence entre ladite première valeur de distorsion d'échantillon et au moins
un échantillon exclu dudit jeu expérimental.
6. Procédé selon la revendication 5 dans lequel la sélection dudit jeu expérimental d'échantillons
et l'ajustement de ladite première valeur de distorsion d'échantillon comprennent
les étapes consistant à :
a) identifier tous les échantillons bruyants dudit jeu expérimental, chacun desdits
échantillons bruyants ayant une grandeur située en dessous d'une valeur courante de
ladite première valeur de distorsion d'échantillon,
b) retirer lesdits échantillons bruyants, s'il en existe, dudit jeu expérimental,
c) si tous les échantillons bruyants sont retirés, augmenter ladite première valeur
de distorsion d'échantillon d'une quantité déterminée par la différence entre ladite
première valeur de distorsion d'échantillon et au moins un dit échantillon bruyant,
et
répéter les étapes a), b) et c) jusqu'à ce que soit atteinte une première valeur
de distorsion d'échantillon ajustée dans laquelle étape aucun échantillon bruyant
additionnel dudit jeu expérimental n'est identifié qui soit supérieur à ladite première
distorsion d'échantillon ajustée.
7. Procédé selon la revendication 6 comprenant en outre l'étape consistant à terminer
ledit ajustement si les étapes a), b) et c) ont été répétées un nombre maximum de
fois.
8. Procédé selon la revendication 7 comprenant en outre les étapes consistant à :
après terminaison dudit ajustement, estimer un nombre binaire pour une quantité d'éléments
binaires requise de façon à transmettre tous les échantillons dudit jeu expérimental,
comparer ledit nombre binaire estimé à un nombre binaire maximum, et
si ledit nombre binaire estimé est inférieur ou égal audit nombre binaire maximum,
sélectionner un seuil de bruit final en se fondant sur ladite première valeur de distorsion
d'échantillon ajustée.
9. Procédé selon la revendication 8, dans lequel si ledit nombre binaire estimé excède
ledit nombre binaire maximum, le procédé comprend les étapes supplémentaires consistant
à :
préparer une seconde valeur de distorsion d'échantillon,
a) sélectionner un second jeu expérimental d'échantillons dudit signal numérique,
chaque pluralité desdits échantillons dudit second jeu expérimental ayant une grandeur
supérieure à ladite seconde valeur de distorsion d'échantillon,
b) estimer le nombre d'éléments binaires requis pour transmettre tous les échantillons
dudit second jeu expérimental,
c) comparer ledit nombre binaire estimé audit nombre binaire maximum, et
d) si ledit nombre binaire estimé est supérieur audit nombre binaire maximum, augmenter
ladite seconde valeur de distorsion d'échantillon d'une quantité déterminée par ladite
sélection.
10. Procédé selon la revendication 9 comprenant en outre l'étape consistant à répéter
les étapes d à g jusqu'à ce que soit atteinte une seconde valeur de distorsion d'échantillon
pour laquelle l'étape g détermine que ledit nombre binaire estimé est inférieur ou
égal audit nombre binaire maximum.
11. Procédé selon la revendication 10 comprenant en outre les étapes consistant à :
calculer ladite valeur de distorsion d'échantillon à partir de ladite première valeur
de distorsion ajustée d'échantillon et ladite seconde valeur de distorsion d'échantillon
et,
sélectionner un jeu final d'échantillons dudit signal numérique, chaque pluralité
desdits échantillons dudit jeu final ayant une grandeur supérieure à un seuil final
déterminé par la valeur de distorsion d'échantillon correspondante des échantillons.
12. Procédé pour la communication d'un signal numérique qui comprend une composante de
bruit et une composante de signal, le procédé comprenant les étapes consistant à :
préparer un signal d'estimation représentatif dudit signal numérique en ayant déjà
moins d'échantillons que ledit signal numérique,
préparer un indice de signal représentant, pour au moins un échantillon dudit signal
d'estimation, une grandeur de signal de ladite composante de signal relative à une
grandeur de bruit de ladite composante de bruit,
en se fondant sur ledit indice de signal, sélectionner des échantillons dudit signal
numérique ayant une composante de signal suffisamment grande,
transmettre lesdits échantillons sélectionnés dudit signal numérique,
transmettre lesdits échantillons dudit signal d'estimation numérique, et
reconstruire ledit signal numérique à partir desdits échantillons sélectionnés transmis
et des échantillons d'estimation.
13. Procédé selon la revendication 12 dans lequel ledit signal numérique est un signal
à fréquence du domaine de la parole représentatif d'une information vocale destinée
à être communiquée, et chaque échantillon dudit signal d'estimation est une estimation
spectrale du signal à fréquence du domaine dans une bande correspondante de fréquences,
et dans lequel la reconstruction dudit signal à fréquence du domaine de la parole
comprend les étapes consistant à :
engendrer un nombre aléatoire pour chaque échantillon non sélectionné dudit signal
à fréquence du domaine de la parole,
préparer, à partir d'au moins une dite estimation spectrale, une estimation de bruit
de la grandeur d'une composante de bruit dudit échantillon non sélectionné,
en se fondant sur ladite estimation de bruit, engendrer un facteur d'échelle, et
proportionner ledit nombre aléatoire en conformité avec ledit facteur d'échelle de
façon à produire un échantillon reconstruit représentatif dudit échantillon non sélectionné.
14. Procédé selon la revendication 13, dans lequel ledit signal d'estimation et ledit
signal à fréquence du domaine de la parole comprennent chacun une série de structures,
chaque structure représentant ladite information vocale sur une fenêtre spécifiée
de temps, et dans lequel la préparation d'une première estimation de bruit pour une
structure courante comprend les étapes consistant à :
pour une structure préalable dudit signal d'estimation, préparer une estimation initiale
de bruit à partir d'au moins une estimation spectrale représentative d'une bande des
fréquences de ladite structure préalable,
en se fondant sur la grandeur dudit indice de signal, sélectionner une constante de
temps de montée tr et une constante de temps de chute tf,
additionner ladite constante de temps de montée à ladite estimation initiale de bruit
pour former un seuil supérieur,
soustraire ladite constante de temps de chute de ladite estimation initiale de bruit
pour former un seuil inférieur,
comparer une estimation spectrale courante de ladite structure courante, représentative
de ladite bande de fréquences, auxdits seuils supérieur et inférieur,
si ladite estimation spectrale courante est comprise entre lesdits seuils, établir
ladite estimation courante de bruit pour qu'elle soit égale à ladite estimation spectrale
courante,
si ladite estimation spectrale courante est inférieure audit seuil inférieur, décrémenter
ladite estimation initiale de bruit de ladite constante de temps de chute pour former
ladite estimation courante de bruit, et
si ladite estimation spectrale courante est supérieure audit seuil supérieur, incrémenter
ladite estimation initiale de bruit de ladite constante de temps de montée de façon
à former ladite estimation courante de bruit.
15. Procédé selon la revendication 14 dans lequel la génération d'un facteur d'échelle
comprend l'étape consistant à engendrer une somme pondérée de ladite estimation courante
de bruit et de ladite estimation spectrale courante.
16. Procédé selon la revendicaiton 15 dans lequel la génération de ladite somme pondérée
comprend les étapes consistant à :
sélectionner un coefficient de bruit en se fondant sur la valeur dudit indice de signal,
sélectionner un coefficient vocal en se fondant sur la valeur dudit indice de signal,
multiplier ladite estimation courante de bruit par ledit coefficient de bruit,
multiplier ladite estimation spectrale courante par ledit coefficient vocal,
additionner les produits desdites étapes de multiplication, et
former ledit facteur d'échelle à partir du résultat de ladite étape d'addition.
17. Dispositif de codage pour l'attribution d'éléments binaires de transmission destinés
à être utilisés en tant qu'échantillons de transmission d'un signal numérique, le
dispositif de codage comprenant :
des moyens pour sélectionner une valeur de distorsion de quantification acceptable
globale représentant une erreur de distorsion de quantification acceptable collective
pour une structure des échantillons dudit signal numérique,
des moyens pour sélectionner à partir de ladite structure des échantillons, un jeu
d'échantillons dans lequel chaque pluralité d'échantillons dudit jeu est supérieure
à un seuil de bruit,
des moyens pour calculer, pour chaque échantillon dudit jeu, une valeur de distorsion
de quantification d'échantillon représentant une erreur de distorsion de quantification
acceptable pour ledit échantillon, et où la somme de toutes les valeurs de distorsion
de quantification d'échantillon pour tous les échantillons dudit jeu est approximativement
égale à ladite valeur de distorsion de quantification acceptable globale,
des moyens pour sélectionner, pour chaque échantillon dudit jeu une étape de taille
de quantification qui fournit une erreur de distorsion de quantification approximativement
égale à ladite valeur de distorsion de quantification correspondante des échantillons,
et
des moyens pour quantifier ledit échantillon en utilisant ladite étape de taille de
quantification.
18. Dispositif de codage selon la revendication 17 dans lequel lesdits moyens pour sélectionner
une étape de taille de quantification comprennent des moyens pour calculer une étape
de taille de quantification pour chaque échantillon devant être transmis en se fondant,
au moins en partie, sur la différence entre ledit échantillon et ladite valeur de
distorsion de quantification correspondante des échantillons.
19. Dispositif de codage selon la revendication 17 dans lequel ledit signal numérique
comprend une composante de bruit et une composante de signal,
et dans lequel lesdits moyens pour sélectionner une distorsion de quantification acceptable
globale comprennent :
des moyens pour préparer un indice de signal représentant, pour au moins un échantillon
de ladite structure, une grandeur de signal de ladite composante de signal relative
à une grandeur de bruit de ladite composante de bruit, et
des moyens pour sélectionner ladite distorsion de quantification acceptable globale
en se fondant sur ledit indice de signal.
20. Dispositif de codage selon la revendication 17 dans lequel lesdits moyens pour calculer
ladite valeur de distorsion de quantification d'échantillon comprennent des moyens
pour diviser ladite valeur de distorsion de quantification acceptable globale par
un nombre des échantillons de ladite structure des échantillons pour former une première
valeur de distorsion d'échantillon.
21. Dispositif de codage selon la revendication 20 dans lequel lesdits moyens pour sélectionner
un jeu des échantillons comprennent :
des moyens pour sélectionner un jeu expérimental des échantillons dudit signal numérique
dans lequel chaque échantillon dudit jeu expérimental est supérieur à un seuil de
bruit déterminé au moins en partie par la valeur de ladite première valeur de distorsion
d'échantillon, et où lesdits moyens pour calculer ladite valeur de distorsion de quantification
d'échantillon comprennent :
des moyens pour ajuster ladite première valeur de distorsion d'échantillon au moyen
d'une quantité déterminée par la différence entre ladite première valeur de distorsion
d'échantillon et au moins un échantillon exclu dudit jeu expérimental.
22. Dispositif de codage selon la revendication 21 dans lequel lesdits moyens pour sélectionner
ledit jeu expérimental des échantillons et lesdits moyens pour ajuster ladite première
valeur de distorsion d'échantillon comprennent :
des moyens pour identifier tous les échantillons bruyants dudit jeu expérimental,
chaque dit échantillon bruyant ayant une grandeur inférieure à une valeur courante
de ladite première valeur de distorsion d'échantillon,
des moyens pour retirer lesdits échantillons bruyants, s'il en existe, dudit jeu expérimental,
des moyens pour augmenter ladite première valeur de distorsion d'échantillon d'une
quantité déterminée par la différence entre ladite première valeur de distorsion d'échantillon
et au moins un dit échantillon bruyant, si des échantillons bruyants sont retirés,
et
des moyens pour retirer de façon répétitive des échantillons bruyants et ajuster ladite
première valeur de distorsion d'échantillon jusqu'à ce qu'une première valeur de distorsion
d'échantillon ajustée soit atteinte pour laquelle aucun échantillon bruyant additionnel
dudit jeu expérimental ne soit supérieur à ladite première distorsion d'échantillon
ajustée.
23. Dispositif de codage selon la revendication 22 comprenant en outre des moyens pour
terminer ledit ajustement si ladite première valeur de distorsion d'échantillon est
ajustée un nombre maximum de fois.
24. Dispositif de codage selon la revendication 23 comprenant en outre :
des moyens pour estimer après terminaison dudit ajustement, un nombre binaire pour
une quantité d'éléments binaires requis de façon à transmettre tous les échantillons
dudit jeu expérimental,
des moyens pour comparer ledit nombre binaire estimé à un nombre binaire maximum,
et
des moyens pour sélectionner un seuil de bruit final en se fondant sur ladite première
valeur de distorsion d'échantillon ajustée, si ledit nombre binaire estimé est inférieur
ou égal audit nombre binaire maximum.
25. Dispositif de codage selon la revendication 24 dans lequel le dispositif de codage
comprend en outre :
des moyens pour préparer une seconde valeur de distorsion d'échantillon, si ledit
nombre binaire estimé excède ledit nombre binaire maximum,
des moyens pour sélectionner un second jeu expérimental d'échantillons dudit signal
numérique, chaque pluralité desdits échantillons dudit second jeu expérimental ayant
une grandeur supérieure à ladite seconde valeur de distorsion d'échantillon,
des moyens pour estimer le nombre d'éléments binaires requis pour transmettre tous
les échantillons dudit second jeu expérimental,
des moyens pour comparer ledit nombre binaire estimé audit nombre binaire maximum,
et
des moyens pour augmenter ladite seconde valeur de distorsion d'échantillon d'une
quantité déterminée par ladite sélection, si ledit nombre binaire estimé est supérieur
audit nombre binaire maximum.
26. Dispositif de codage selon la revendication 25 comprenant en outre des moyens pour
de façon répétitive ajuster ladite seconde valeur de distorsion d'échantillon, resélectionner
ledit second jeu expérimental des échantillons, et estimer le nombre des éléments
binaires requis pour transmettre tous les échantillons dudit second jeu expérimental
jusqu'à ce que soit atteinte une seconde valeur de distorsion d'échantillon pour laquelle
ledit nombre binaire estimé est inférieur ou égal audit nombre binaire maximum.
27. Dispositif de codage selon la revendication 26 comprenant en outre :
des moyens pour calculer ladite valeur de distorsion d'échantillon à partir de ladite
première valeur de distorsion d'échantillon ajustée et de ladite seconde valeur de
distorsion d'échantillon et,
des moyens pour sélectionner un jeu final des échantillons dudit signal numérique,
chaque pluralité desdits échantillons dudit jeu final ayant une grandeur supérieure
à un seuil final déterminé par ladite valeur de distorsion d'échantillon correspondante
des échantillons.
28. Dispositif de codage pour la communication d'un signal numérique qui comprend une
composante de bruit et une composante de signal, le dispositif de codage comprenant
:
des moyens pour préparer un signal d'estimation représentatif dudit signal numérique
ayant déjà moins d'échantillons que ledit signal numérique,
des moyens pour préparer un indice de signal représentant, pour au moins un échantillon
dudit signal d'estimation, la grandeur de ladite composante de signal relative à la
grandeur de ladite composante de bruit,
des moyens pour sélectionner, en se fondant sur ledit indice de signal, des échantillons
dudit signal numérique ayant une composante de signal suffisamment grande,
des moyens pour transmettre lesdits échantillons sélectionnés dudit signal numérique,
des moyens pour transmettre lesdits échantillons du signal d'estimation numérique,
et
des moyens pour reconstruire ledit signal numérique à partir desdits échantillons
sélectionnés transmis et des échantillons d'estimation.
29. Dispositif de codage selon la revendication 28 dans lequel ledit signal numérique
est un signal à fréquence du domaine de la parole représentant une information vocale
destinée à être communiquée, et chaque échantillon dudit signal d'estimation est une
estimation spectrale du signal à fréquence du domaine dans une bande correspondante
de fréquences, et où lesdits moyens pour reconstruire ledit signal à fréquence du
domaine de la parole comprennent :
des moyens pour engendrer un nombre aléatoire pour chaque échantillon non sélectionné
dudit signal à fréquence du domaine de la parole,
des moyens pour préparer, à partir d'au moins une dite estimation spectrale, une estimation
de bruit de la grandeur d'une composante de bruit dudit échantillon non sélectionné,
des moyens pour engendrer un facteur d'échelle en se fondant sur ladite estimation
de bruit, et
des moyens pour proportionner ledit nombre aléatoire en conformité audit facteur d'échelle
de façon à produire un échantillon reconstruit représentatif dudit échantillon non
sélectionné.
30. Dispositif de codage selon la revendication 29 dans lequel ledit signal d'estimation
et ledit signal à fréquence du domaine de la parole comprennent chacun une série de
structures, chaque structure représentant ladite information vocale sur une fenêtre
spécifiée de temps, et où lesdits moyens pour préparer une première estimation de
bruit pour une structure courante comprennent :
des moyens pour préparer, pour une structure préalable dudit signal d'estimation,
une estimation initiale de bruit à partir d'au moins une estimation spectrale représentative
d'une bande de fréquences de ladite structure préalable,
des moyens pour sélectionner, en se fondant sur la grandeur dudit indice de signal,
une constante de temps de montée tr, et une constante de temps de chute tf,
des moyens pour additionner ladite constante de temps de montée à ladite estimation
initiale de bruit de manière à former un seuil supérieur,
des moyens pour soustraire ladite constante de temps de chute de ladite estimation
initiale de bruit pour former un seuil inférieur,
des moyens pour comparer une estimation spectrale courante représentative de ladite
bande de fréquences dans ladite structure courante, auxdits seuils supérieur et inférieur,
des moyens pour établir ladite estimation courante de bruit pour qu'elle soit égale
à ladite estimation spectrale courante, si ladite estimation spectrale courante est
comprise entre lesdits seuils,
des moyens pour décrémenter ladite estimation initiale de bruit de ladite constante
de temps de chute pour former ladite estimation courante de bruit, si ladite estimation
spectrale courante est inférieure audit seuil inférieur, et
des moyens pour incrémenter ladite estimation initiale de bruit au moyen de ladite
constante de temps de montée de manière à former ladite estimation courante de bruit,
si ladite estimation spectrale courante est supérieure audit seuil supérieur.
31. Dispositif de codage selon la revendication 30 dans lequel lesdits moyens pour engendrer
un facteur d'échelle comprennent des moyens pour engendrer une somme pondérée de ladite
estimation courante de bruit et de ladite estimation spectrale courante.
32. Dispositif de codage selon la revendication 31 dans lequel lesdits moyens pour engendrer
ladite somme pondérée comprennent :
des moyens pour sélectionner un coefficient de bruit en se fondant sur la valeur dudit
indice de signal,
des moyens pour sélectionner un coefficient vocal en se fondant sur la valeur dudit
indice de signal,
des moyens pour multiplier ladite estimation courante de bruit par ledit coefficient
de bruit,
des moyens pour mulitplier ladite estimation spectrale courante par ledit coefficient
vocal,
des moyens pour additionner les produits desdites multiplications, et
des moyens pour former le facteur d'échelle à partir du résultat de ladite addition.