Noise reduction apparatus and method
Granted 18 May 2004 · 1 office action
Current assignee: CLUSTER LLC · originally Ericsson
Law firm: Law firm · Log in to unlock
Attorney: Attorney · Log in to unlock
Inventors: Ali S. Khayrallah, Leonid Krasny · Examiner: Melur Ramakrishnaiah · AU 2643 · TC 2600
Life of the application
11 dated eventsAbstract
A method and noise reduction apparatus comprises a microphone array including a plurality of microphone elements for receiving a training signal including a plurality of training signal samples, and a working signal including a plurality of working signal samples, and at least one frequency domain convertor coupled to the plurality of microphone elements for converting the plurality of training signal samples and the plurality of working signal samples to the frequency domain. A signal spatial correlation matrix estimator is coupled to the at least one frequency domain convertor for estimating a signal spatial correlation matrix using the converted plurality of training signal samples. An inverse noise spatial correlation matrix estimator is coupled to the at least one frequency domain convertor for estimating an inverse noise spatial correlation matrix using the converted plurality of working signal samples. A constrained output generator is coupled to the at least one frequency domain convertor, the signal spatial correlation matrix estimator and the inverse noise spatial correlation matrix estimator for generating a constrained output for the noise reduction apparatus using the converted working signal samples, the estimated signal spatial correlation matrix and the estimated inverse noise spatial correlation matrix.
Description
6 parts›BACKGROUND OF THE INVENTION
This invention is directed to noise reduction, and more particularly, to an apparatus and method for performing noise reduction for a signal received at a microphone array.
A noise reduction apparatus is typically used in conjunction with hands-free mobile terminals (for example, cellular telephones) and speaker phones, or with speech recognition systems, to reduce noise received at a microphone array of the noise reduction apparatus.
The general structure of different array processing algorithms for noise reduction apparatuses utilizing microphone arrays in conjunction with signal processing can be expressed in the frequency domain as U out ( ω ) = ∑ i = 1 N U ( ω , r i ) · H * ( ω , r i )
where U out (ω) and U(ω, r 1 ) are respectively the Fourier transform of the microphone output and the field u(t, r i ) observed at the i-th microphone elements with the spatial coordinates r i , H(ω, r 1 ) is the frequency response of the filter at the i-th element of the microphone array, and N is the number of microphone array elements.
The determination of the functions H(ω, r 1 ) is the major area of concern in array processing. In conventional array processing, the optimization criteria used for the determination of the functions H(ω, r i ) are based on an assumption that the signal field in a limited space, for example an automobile cabin, has a coherent structure. This assumption leads to the following conventional algorithm for the determination of the weighting functions H(ω, r 1 ): H ( ω , r i ) ≡ H 0 ( ω , r i ) = ∑ p = 1 N K N - 1 ( ω ; r i , r p ) G ( ω ; r p , r 0 )
where K N −1 (ω, r 1 , r p ) denotes the elements of the matrix K N −1 (ω) which is the inverse of the noise spatial correlation function matrix K N (ω) with the elements K N (ω; r 1 , r p ). G (ω, r p , r 0 ) is the Green function which describes the propagation channel between the talker with the spatial coordinates r 0 and the p-th array microphone. However, experimental data and theoretical analysis show that the coherent signal field model is unrealistic for many limited or confined spaces such as automobile environments where wall irregularities will scatter the signal waves propogating inside the automobile cabin.
›SUMMARY OF THE INVENTION
A method of reducing noise and a noise reduction apparatus are provided utilizing a microphone array including a plurality of microphone elements for receiving a training signal including a plurality of training signal samples, and a working signal including a plurality of working signal samples. At least one frequency domain convertor is coupled to the plurality of microphone elements for converting the plurality of training signal samples and the plurality of working signal samples to the frequency domain. A signal spatial correlation matrix estimator is coupled to the at least one frequency domain convertor for estimating a signal spatial correlation matrix using the converted plurality of training signal samples, and an inverse noise spatial correlation matrix estimator is coupled to the at least one frequency domain convertor for estimating an inverse noise spatial correlation matrix using the converted plurality of working signal samples. A constrained output generator is coupled to the at least one frequency domain convertor, the signal spatial correlation matrix estimator and the inverse noise spatial correlation matrix estimator for generating a constrained output for the noise reduction apparatus using the converted working signal samples, the estimated signal spatial correlation matrix and the estimated inverse noise spatial correlation matrix.
The noise reduction apparatus may be used in conjunction with or implemented as part of a mobile terminal, a speaker-phone, a speech recognition system, or any other device where noise reduction is desirable.
›BRIEF DESCRIPTION OF THE DRAWINGS
FIG. 1 is a block diagram in accordance with an embodiment of the invention;
FIG. 2 is a flowchart illustrating the training phase in accordance with the embodiment of FIG. 1; and
FIG. 3 is a flowchart illustrating the working phase in accordance with the embodiment of FIG. 1 .
›DETAILED DESCRIPTION OF THE INVENTION · 1 of 3
To avoid the drawbacks of the conventional array processing technique, a new optimization criteria with constraint is not based on the assumption that the signal field in a limited space, for example an automobile cabin, has a coherent structure. The nature of the human auditory system is taken into account in the formulation of the optimization criteria, as significant degradation in the desired signal is unacceptable even if the noise level is greatly reduced. Thus, the optimization problem for the array processing algorithm U out (ω) may be overcome by minimizing the output noise spectral density subject to an equality nonlinear constraint
g S out (ω)= gs (ω)| B (ω)| 2
where g S out ( ω ) = ∑ i = 1 N ∑ p = 1 N K S ( ω ; r i , r p ) H * ( ω , r i ) H ( ω , r p )
is the signal spectral density after array processing, and B(ω) is the constraint function which takes into account the response characteristics of the human auditory system. The constraint function B(ω) may be tailored for greater noise constraint over specific parts of the audible frequency spectrum. For example, the constraint function B(ω) may be selectable to provide greater noise suppression over lower audible frequencies, providing people with hearing difficulties over such lower audible frequencies a clearer (and louder) audible signal from the cellular telephone speaker. The constraint g S out represents the degree of degradation of the desired signal and permits the combination of various frequency bins at the space-time processing output with a priori desired distortion.
According to this optimization criteria, the weighting functions H(ω, r 1 ) are obtained as a solution of the variation problem H ( ω , r i ) = arg { min ∑ i = 1 N ∑ p = 1 N K N ( ω ; r i , r p ) H * ( ω , r i ) H ( ω , r p ) }
subject to the constraint g S out .
The solution of this optimization problem gives the following algorithm for the calculation of weighting functions: H ( ω , r i ) = B ( ω ) ν max ( ω ) E max ( ω , r i )
where E max (ω, r 1 ) are the elements of the eigenvector E max (ω), which corresponds to the largest eigenvalue v max (ω) of the constraint matrix K=K N −1 Ks having elements K ( ω ; r i , r p ) = ∑ m = 1 N K N - 1 ( ω ; r i , r m ) K S ( ω ; r m , r p ) .
The constraint function B(ω) allows the nature of the human auditory system to be taken into account during calculation of the weighting functions.
The working scheme for the proposed array processing algorithm may be divided into two phases, a training phase and a working phase. The training phase provides an estimate of the signal spatial correlation function K S (ω; r 1 , r p ) which is used in the working phase, along with other values, to generate a constrained output for a noise reduction apparatus. A block diagram of a noise reduction apparatus in accordance with an embodiment of the invention is shown in FIG. 1 .
FIG. 1 shows a noise reduction apparatus 100 comprising a microphone array 102 for selectively receiving either a training signal or a working signal and includes a plurality N of microphone elements, for example microphone elements 104 , 106 and 108 . Each microphone element 104 , 106 and 108 of the microphone array 102 is coupled to a corresponding frequency domain convertor 110 , 112 and 114 respectively of frequency domain convertors 115 , the frequency domain convertors 115 for converting the training signal and the working signal to the frequency domain. The frequency domain convertors 115 are coupled to both a signal spatial correlation matrix estimator 120 and an inverse noise spatial correlation matrix estimator 125 . The signal spatial correlation matrix estimator 120 provides an estimate of a signal spatial correlation matrix for the training signal (further discussed below). The inverse noise spatial correlation matrix estimator 125 provides an estimate of the inverse noise spatial correlation matrix using the working signal (further discussed below). The frequency domain convertors 115 , the signal spatial correlation matrix estimator 120 and the inverse noise spatial correlation matrix estimator 125 are further coupled to a constrained output generator 130 .
The constrained output generator includes a first calculator 135 coupled to the signal spatial correlation matrix estimator 120 and the inverse noise spatial correlation matrix estimator 125 for calculating a constraint matrix. The first calculator 135 is coupled to a second calculator 140 which calculates a maximum eigenvalue and a maximum eigenvector of the constraint matrix. The second calculator 140 and the frequence domain convertors 115 are coupled to frequency response filters 145 , which calculate a frequency response of the microphone elements 104 , 106 and 108 . Each of the frequency domain convertors 110 , 112 and 114 is coupled to frequency response filters 146 , 147 and 148 respectively. The frequency response filters 145 are coupled to a summing device 150 which generates the constrained output for the noise reduction apparatus 100 using the frequency response of each of the plurality N microphone elements of the microphone array 102 . A time domain convertor 155 is coupled to the constrained output generator 130 for converting the constrained output from the frequency domain to the time domain. Specifically, the time domain convertor 155 is coupled to the summing device 150 .
In order to estimate the signal spatial correlation function K S (ω; r 1 , r p ) at the aperture of the microphone array 102 , training sequences are recorded through the actual system in the limited or confined space, for example, the automobile environment with all its imperfections. They are recorded during a training phase where little or no ambient automobile noise is present. The training can be done on site in a parked automobile by using the existing hands-free loud speaker in what would be a human speaker's position. The estimate of the signal spatial correlation function then is stored in a memory (not shown) for later use during the working phase. Operation of the noise reduction apparatus 100 of FIG. 1 will be discussed with respect to the flowcharts of FIGS. 2 and 3.
›DETAILED DESCRIPTION OF THE INVENTION · 2 of 3
FIG. 2 is a flowchart illustrating the training phase. In step 200 , sampled training sequences are received as a plurality of training signal samples
{s ( n, r 1 ), . . . , s ( n, r i ), . . . , s ( n, r N )},
which are recorded at the output of the microphone array 102 in the limited space, for example the automobile cabin, when little or no ambient noise is present. Here, s(n, r 1 ) denotes the n-th sample of the training signal which is recorded at the output of the i-th microphone element with spatial coordinates r i .
Once the training signal is received, it is converted to the frequency domain by the plurality of frequency domain converters 115 using, for example, a Fast Fourier Transform (FFT) algorithm. The frequency domain converting technique is running on a frame-block basis. In hands-free mobile telephones each frame contains N 1 =160 samples. To improve the representation of the spectrum, the FFT length is effectively increased by overlapping and windowing, step 210 . Where the FFT with N 0 =256 points (samples), the N 1 samples of the q-th frame are overlapped with the last (N 0 −N 1 ) samples of the previous (q− 1 )th frame. As a result, the q-th frame at the i-th microphone element contains training signal
s q ( n, r 1 )≡ s (q·N 1 −N 0 +n, r 1 ),
where nε[0, N 0 −1] and iε[1, N].
The signals s q (n, r 1 ) are windowed using the smoothed Hanning window w ( n ) = { sin 2 ( π n / ( N 0 - N 1 ) ) 1 sin 2 ( π ( n - N 0 + 1 ) / ( N 0 - N 1 ) )
if n ∈ [ 0 , ( N 0 - N 1 ) / 2 - 1 ]
if n ∈ [ ( N 0 - N 1 ) / 2 , ( N 0 + N 1 ) / 2 - 1 ]
if n ∈ [ ( N 0 + N 1 ) / 2 , ( N 0 - 1 ) ]
Using the windowed, overlapped training signal samples, the FFT is calculated For Kε[0, N 0 −1] and iε[1, N] in step 220 as S q ( k , r i ) = ∑ n = 0 N 0 - 1 w ( n ) · s q ( n , r i ) · exp ( - j2π kn / N 0 ) .
After the training signal samples are converted to the frequency domain, the signal spatial correlation matrix is estimated at the signal spatial correlation matrix estimator 120 , step 230 , for Kε[0, N 0 /2] and iε[1, N], and pε[i, N] as
{circumflex over (K)} Sq ( k, r 1 , r p )= m·{circumflex over (K)} S(q−1) ( k, r 1 , r p )+(1 −m )· S q ( k, r 1 )· S q *( k, r p )
where m is a convergence factor (for example, mε[0.9, 0.95]). {circumflex over (K)} Sq (k, r 1 , r p ) denotes an estimate of the signal spatial correlation matrix at the q-th frame. Initially, {circumflex over (K)} S ( q−1 )(k, r i , r p ) may be set to zero. To minimize the calculations, it may be taken into account that
{circumflex over (K)} Sq ( k, r 1 , r p )=[ {circumflex over (K)} Sq ( k, r p , r i )]*.
After processing of the Q frames, the signal spatial correlation matrix is estimated as
{circumflex over (K)} S ( k, r 1 , r p )≡ {circumflex over (K)} SQ ( k, r i , r p ).
The working phase is illustrated in FIG. 3 . In step 300 , sampled working sequences are received as a plurality of working signal samples
{u ( n, r 1 ), . . . , u ( n, r 1 ), . . . , u ( n, r N )},
which are observed at the microphone elements of the microphone array 102 . For example u(n, r 1 ) is the output signal of the i-th microphone element with the spatial coordinates r 1 . The working sequences are received under normal operating conditions, and thus ambient noise need not be limited.
The working signal samples u q (n, r 1 ) are windowed and overlapped, step 310 , in a similar fashion as for the training phase, described above with respect to step 210 of FIG. 2 . For example, the q-th frame at the i-th microphone element contains the signal
u q ( n, r i )≡ u ( q·N 1 −N 0 +n, r 1 ),
where nε[0, N 0 −1] and iε[1, N].
Using the windowed, overlapped training signal samples, the FFT is calculated by the plurality of frequency domain convertors 115 for kε[0, N 0 −1] and iε[1, N] in step 320 in a similar fashion as in the training phase discussed above with reference to step 220 of FIG. 2, where U q ( k , r i ) = ∑ n = 0 N 0 - 1 w ( n ) · u q ( n , r i ) · exp ( - j2π kn / N 0 ) .
After the working signal has been converted to the frequency domain, the inverse noise spatial correlation matrix estimator 125 estimates the inverse noise spatial correlation matrix K N −1 (ω; r 1 , r p ) using the Recursive Least Square (RLS) algorithm, which has been modified for processing in the frequency domain, step 330 . This algorithm allows direct calculation of the matrix K N −1 (ω; r 1 , r p ). For kε[0, N 0 /2], iε[1, N], and pε[i, N], the inverse noise spatial correlation function is estimated as K ^ Nq - 1 ( k , r i , r p ) = 1 m · { K ^ N ( q - 1 ) - 1 ( k , r i , r p ) - D q ( k , r i ) · D q * ( k , r p ) m + ∑ i = 1 N D q ( k , r i ) · U q * ( k , r i ) }
where K Nq −1 (k, r 1 , r p ) denotes an estimate of the inverse noise spatial correlation matrix at the q-th frame.
The initial matrix for the inverse spatial correlation matrix algorithm can be chosen as K ^ N0 - 1 ( k ; r i , r p ) = a · δ i p
where a is a large constant, and δ 1p is the Kronecker symbol. The functions D q (k, r p ) are calculated using the inverse noise correlation matrix at the previous (q−1)th frame as D q ( k , r p ) = ∑ i = 1 N K ^ N ( q - 1 ) - 1 ( k , r p , r i ) · U q ( k , r i ) .
After the inverse noise spatial correlation matrix is estimated in step 330 , the constraint matrix is calculated by the first calculator 135 , step 340 , using the signal spatial correlation matrix as, for example as calculated in step 230 , and the inverse noise spatial correlation matrix. For kε[0, N 0 /2], iε[1, N], and pε[i, N], the constraint matrix is calculated as K ^ q ( k , r i , r p ) = ∑ m = 1 N K ^ N q - 1 ( k ; r i , r m ) K ^ S ( k ; r m , r p ) .
In step 350 , a maximum eigenvalue v max (k) and a corresponding eigen vector E max (k, r 1 ) of the constraint matrix {circumflex over (K)} q (k, r l , r p ) is calculated by the second calculator 140 for kε[0, N 0 /2], iε[1, N], and pε[i, N]. Calculations may be done using standard matrix computations, similar to that as discussed above with respect to calculation of the constraint matrix {circumflex over (K)} q −{circumflex over (K)} Nq −1 {circumflex over (K)}K s .
›DETAILED DESCRIPTION OF THE INVENTION · 3 of 3
After calculating the maximum eigenvalue v max (k) and the corresponding eigen vector E max (k, r 1 ), the frequency response for the microphone elements 104 , 106 and 108 of the microphone array 102 are calculated by the plurality of frequency response filters 145 for kε[0, N 0 /2], and iε[1, N], step 360 , as H q ( k , r i ) = B ( k ) ν max ( k ) E max ( k , r i ) .
B(k) accounts for the nature of the human auditory system.
In step 370 , the constrained output is generated at the summing device 150 for kε[0, N 0 /2] as U q o u t ( k ) = ∑ i = 1 N U q ( k , r i ) H q * ( k , r i )
and for kε[N 0 /2+1, N 0 −1] as
U q out ( k )=[ U q out (N 0 −k )]*.
The constrained output is then converted to the time domain by time domain convertor 155 in step 380 for nε[0, N 0 −1], by calculating an inverse FFT as u q o u t ( n ) = ∑ k = 0 N 0 - 1 · U q o u t ( k ) exp ( j 2 π k n / N 0 ) .
It would be apparent to one skilled in the art that the noise reduction apparatus may be implemented as discrete components, or as a program operating on a suitable processor. Additionally, the number of microphone elements of the microphone array is not crucial in attaining the advantages of the noise reduction apparatus of the invention. Further, the noise reduction apparatus may be implemented as part of a mobile terminal operating in a communications system utilizing, for example, Code Division Multiple Access or Time Division Multiple Access architecture. The noise reduction apparatus may also be implemented as part of a speaker phone, a speech recognition system or any device where noise reduction is desired. Alternatively, the noise reduction apparatus may be utilized in conjunction with a mobile terminal, speaker phone, speech recognition system or any device where noise reduction is desired. Additionally, although the invention has been described in the context of the limited or confined space being an automobile cabin, the advantages attained would be applicable for any space such as a conference room or other confined or limited area.
Still other aspects, objects and advantages of the invention can be obtained from a study of the specification, the drawings, and the appended claims. It should be understood, however, that the invention could be used in alternate forms where less than all of the advantages of the present invention and preferred embodiments as described above would be obtained.
Claims as granted
19 claimsLog in to read the claims of this application.
Log in to unlockClassifications
6 codes- G10L21/02
Claim changes
SoonSee which claims were amended, added or cancelled during examination, with every added and removed word marked.
The published claims of this application are not paired with the granted ones in what we hold.
File wrapper
See the full prosecution history — every USPTO and applicant action on this file, in order.
Log in to unlockDocuments
Log in to open the documents of this file: the application as filed, every office action and response, the notice of allowance.
Log in to unlockChain of title
See the full assignment history — every owner this patent has passed through, with recordation dates and reel/frame numbers.
Log in to unlock