Internet communication device and method for controlling noise thereof
Granted 17 May 2011 · 2 office actions
Assignee: FORTEMEDIA, INC.
Law firm: Law firm · Log in to unlock
Attorney: Attorney · Log in to unlock
Inventors: Ming Zhang, Xiaoyan Lu · Examiner: Daniel D Abebe · AU 2626 · TC 2600
Life of the application
10 dated eventsAbstract
The invention provides an Internet communication device. The Internet communication device plays a remote audio signal received via a network and transmits an audio signal back to the remote party to complete the communication. The Internet communication device comprises a line-in speech detection module and a line-in channel control module. The line-in speech detection module detects whether the remote audio signal is speech or not to generate a remote speech detection result. The line-in channel control module then attenuates the remote audio signal if the remote speech detection result indicates that the remote audio signal is not speech, thus, all noise including non-stationary noise is removed from the remote audio signal.
Description
6 parts›BACKGROUND OF THE INVENTION
1. Field of the Invention
The invention relates to noise cancellation, and more particularly to noise cancellation in Internet communication devices.
2. Description of the Related Art
Because the cost of traditional circuit-switched telephony is great, Internet phones are frequently used to make domestic long distance and international calls. Consequently, Internet communication devices, such as VoIP devices and Instant Messengers, have become popular. For Instant Messengers such as Skype, MSN Messenger, Yahoo Messenger, Google Talker, and AOL Messenger are examples of software applications for Internet communication. Increased use of Internet communication devices demands increased audio quality of Internet communication devices. One of the greatest obstacles to audio quality of Internet communication devices is noise.
Noise from computer fans, typing, and mouse movement is often received by the microphone of an Internet communication device connected to the computer. Internet communication devices comprising noise suppression modules are typically capable of canceling a majority of the stationary noise with certain level in order not to affect too much on voice quality. In such case, quite some residual noise will be remained, even after noise suppression. In addition, normal noise suppression modules, however, cannot eliminate non-stationary noise.
Because the noise of each party is independent, when multiple parties are VoIP conferencing, the total level of noise is the sum of the noise of each party. Automatic gain control modules connected to Internet communication devices may further amplify and increase noise. Thus, a method for handling noise, particularly on non-stationary noise of Internet communication devices to improve audio quality Internet communication devices is desirable.
›BRIEF SUMMARY OF THE INVENTION
The invention provides an Internet communication devices. An exemplary embodiment of the Internet communication device plays a remote audio signal received through a network and transmits an audio signal to a remote user to complete the communication. The Internet communication device comprises a line-in speech detection module and a line-in channel control module. The line-in speech detection module detects whether or not the remote audio signal is speech to generate a remote speech detection result. The line-in channel control module then attenuates the remote audio signal if the remote speech detection result indicates that the remote audio signal is not speech, thus, noise is removed from the remote audio signal.
A method for controlling noise of an Internet communication device is also provided. The Internet communication device outputs a remote audio signal received from a network and transmits an audio signal to a remote user through the network to complete a conversation. Whether the remote audio signal is speech or not is first detected to generate a remote speech detection result. The remote audio signal is then attenuated if the remote speech detection result indicates that the remote audio signal is not speech, thus, noise is removed from the remote audio signal.
A detailed description is given in the following embodiments with reference to the accompanying drawings.
›BRIEF DESCRIPTION OF THE DRAWINGS
The invention can be more fully understood by reading the subsequent detailed description and examples with references made to the accompanying drawings, wherein:
FIG. 1 is a block diagram of an Internet communication device with noise control according to the invention;
FIG. 2 is a block diagram of a line-in speech detection module according to the invention;
FIG. 3 is a block diagram of a line-in channel control module according to the invention;
FIG. 4 is a block diagram of a microphone speech detection module according to the invention; and
FIG. 5 is a block diagram of an Internet communication device with an array microphone according to the invention.
›DETAILED DESCRIPTION OF THE INVENTION · 1 of 3
The following description is of the best-contemplated mode of carrying out the invention. This description is made for the purpose of illustrating the general principles of the invention and should not be taken in a limiting sense. The scope of the invention is best determined by reference to the appended claims.
FIG. 1 is a block diagram of an Internet communication device 100 with noise control according to the invention. The Internet communication device 100 is connected to a personal computer 108 , which is further connected to a network. The Internet communication device 100 may be a physical IP phone or a software speakerphone module in personal computer 108 . The Internet communication device 100 receives an audio signal from a near-end user and transmits the audio signal to a remote Internet communication device via the network. The Internet communication device 100 also receives a remote audio signal from the remote Internet communication device through the network and then plays the remote audio signal. Thus, communication is conducted between two Internet communication devices. There can be more than one remote Internet communication device communicating with Internet communication device 100 , such as in a multi-party VoIP conference.
The Internet communication device 100 is connected to the personal computer 108 via an interface 110 , such as a USB interface, an analog audio interface, or a software API interface if the Internet communication device 100 is a software speakerphone module. Subsequent to the Internet communication device 100 receiving the remote audio signal through the Interface 110 , the remote audio signal is processed by line-in signal path modules of the Internet communication device 100 before being output by a loudspeaker 122 . The line-in signal path is shown in the lower half of FIG. 1 and includes a line echo cancellation module 112 , a line-in noise suppression module 114 , a line-in speech detection module 102 , a line-in channel control module 104 , a line-in automatic gain control module 116 , a digital to analog converter 118 , and a power amplifier 120 .
The line echo cancellation module 112 removes the echo caused by the network or line from the remote audio signal. The line-in noise suppression module 114 then removes some stationary noise from the remote audio signal. Only part of the stationary noise, however, can be eliminated because the remote audio is attenuated in conjunction with the elimination of the stationary noise. In addition, non-stationary noise cannot be removed by the line-in noise suppression module 114 . Thus, two modules, the line-in speech detection module 102 and the line-in channel control module 104 , are added to the Internet communication device 100 to cancel the residual noise and non-stationary noise carried by the remote audio signal.
The line-in speech detection module 102 first detects whether or not the remote audio signal is real speech. If the remote audio signal is real speech, a remote speech detection result with a value of 1 is generated. Otherwise, a remote speech detection result with a value of 0 is generated. The remote speech detection result is delivered to the line-in channel control module 104 . If the remote speech detection result indicates that the remote audio signal is not speech, the line-in channel control module 104 attenuates the remote audio signal. For example, the line-in channel control module 104 mutes a non-speech remote audio signal. Thus, all noise including non-stationary noise is removed from the remote audio signal. The line-in automatic gain control module 116 then adjusts the signal level of the remote audio signal to an appropriate level. After being further converted to an analog signal and amplified by power amplifier 120 , the remote audio signal is output by loudspeaker 122 , allowing the user to hear the remote audio signal with no noise.
The microphone 130 receives an audio signal from a user. The audio signal is then processed by line-out signal path modules of Internet communication device 100 before transmission via interface 110 to a network. The line-out signal path is shown in the upper half of FIG. 1 and includes an analog to digital converter 132 , an acoustic echo cancellation module 134 , a noise suppression module 136 , a microphone speech detection module 106 , and an automatic gain control module 138 . The microphone speech detection module 106 is added to the Internet communication device 100 to cancel all noise including non-stationary noise carried by the audio signal. Similar to the line-in speech detection module 102 , the microphone speech detection module 106 detects whether or not the audio signal is speech to generate a speech detection result. If the speech detection result indicates that the audio signal is not speech, the automatic gain control module 138 does not amplify the audio signal. Thus, the residual noise and non-stationary noise carried by the audio signal are prevented from being amplified before transmission.
FIG. 2 is a block diagram of a line-in speech detection module 200 according to the invention. The line-in speech detection module 200 includes a short-term power calculation module 202 , a long-term power calculation module 204 , a noise estimation module 206 , two comparators 208 and 210 , a detector module 212 , and a harmonic detection module 214 . The short-term power calculation module 202 measures a short-term power Ps(n) of the remote audio signal L(n) with a faster update speed. The long-term power calculation module 204 measures a long-term power P l (n) of the remote audio signal L(n) with a slower update speed. The short-term power Ps(n) and the long-term power P l (n) are determined according to the following algorithm:
P s ( n )=α s ·P s ( n− 1)+(1−α s )· L ( n )· L ( n ); and (1)
P l ( n )=α l ·P l ( n− 1)+(1−α l )· L ( n )· L ( n ); (2)
wherein the L(n) is the remote audio signal, the α s is a predetermined short-term smoothing parameter, the α l is a predetermined long-term smoothing parameter and the n is a sample index. The short-term smoothing parameter α s and the long-term smoothing parameter α l are chosen that (1−α l ) is at least one order less than (1−α s ), such that the short-term power Ps(n) is updated faster than the long-term power P l (n).
›DETAILED DESCRIPTION OF THE INVENTION · 2 of 3
The noise estimation module 206 derives a noise power estimate P n (n) from a noise estimate N(m) of the remote audio signal. The frequency domain noise estimate N(m) is obtained from the line-in noise suppression module 114 of FIG. 1 . The time domain noise power estimate P n (n) is determined according to the following algorithms:
Q ( k ) = 1 M ∑ m = 1 M N ( m ) · N ( m ) ; and ( 3 ) P n ( n )= Q ([2 n/M ]); (4)
wherein the k is a frame index, M is a frame size for frequency domain processing, and the function [x] denotes an integer closest to x.
After the short-term power Ps(n), the long-term power P l (n), and the noise power estimate P n (n) are obtained, they are delivered to the comparators 208 and 210 . The comparator 208 compares the difference between the short-term and the long-term powers Ps(n) and P l (n) with a first threshold T 1 (n) to generate a first comparison result C 1 (n). The comparator 210 compares the difference between the long-term power P l (n) and the noise power estimate P n (n) with a second threshold T 2 (n) to generate a second comparison result C 2 (n). The first comparison result C 1 (n) and the second comparison result C 2 (n) are determined according to the following algorithms:
wherein the function |x| denotes the absolute value of x, and log(x) denotes basis-10 logarithm of x.
If the first comparison result C 1 (n) indicates that the short-term power Ps(n) is much greater than the long-term power P l (n), and the second comparison result C 2 (n) indicates that the long-term power P l (n) is much greater than the long-term power P n (n), both the first comparison result C 1 (n) and the second comparison result C 2 (n) are true, and the detector module 212 enables a detector output D(n) to trigger the harmonic detection module 214 . Thus, the detector output D(n) is determined according to the following algorithm:
When triggered by the detector output D(n), the harmonic detection module 214 perform harmonic analysis on the remote audio signal L(n) to detect whether the remote audio signal L(n) consists of real speech or not. If the remote audio signal L(n) comprises speech, the harmonic detection module 214 generates a remote speech detection result S(n) with the value “1”, indicating the existence of speech. Thus, the line-in channel control module 104 of FIG. 1 can mutes the remote audio signal L(n) according to the remote speech detection result S(n). In one embodiment, the harmonic detection module 214 may perform harmonic analysis based on the method provided by E. Fisher, etc. in the “Generalized likelihood ratio test for voiced-unvoiced decision in noisy speech using the harmonic model”, IEEE Trans. On Audio, Speech and Language Processing, Vol. 14, No. 2, March 2006, or the method provided by J. Tabrikian, etc. in the “Tracking speech in a noisy environment using the harmonic model”, IEEE Trans. Speech and Audio Processing, Vol. 12, No. 1, January 2004.
FIG. 3 is a block diagram of a line-in channel control module 300 according to the invention. The line-in channel control module 300 includes a detection frequency module 302 , a speech period control module 304 , and an attenuation control module 306 . The detection frequency module 302 counts a frequency that the remote speech detection result S(n) is true during a speech period of a speech period signal G(n) to determine a detection frequency V(n), wherein the speech period is a period during which the speech period signal G(n) is true. The detection frequency V(n) is determined according to the following algorithm:
The speech period control module 304 then generates the speech period signal G(n) to control the attenuation of the remote audio signal L(n) according to the detection frequency V(n) and the remote speech detection result S(n). If the detection frequency V(n) is greater than a frequency threshold B, the speech period is extended by the speech period control module 304 . Otherwise, the speech period is shortened if the detection frequency is less than the frequency threshold B. Thus, during a conversation between two Internet communication devices, the remote audio signal L(n) is not repeatedly muted for short periods with high frequency, thus eliminating harsh, potentially ear damaging sound in remote audio signal L(n). The attenuation control module 306 then mutes the remote audio signal L(n) according to the speech period signal G(n) to obtain the remote audio signal L′(n). The speech period signal G(n) is determined according to the following algorithms:
FIG. 4 is a block diagram of a microphone speech detection module 400 according to the invention. The microphone speech detection module 400 includes a comparator 402 , a pitch detection module 404 , a transformation module 406 , and a detector module 408 . The transformation module 406 converts a time-domain remote detection signal V f (n) indicating the existence of speech of the remote audio signal to a frequency-domain remote detection signal V f (m). Thus, if the remote detection signal V f (m) is positive, a conversation is underway and the probability that the audio signal comprises speech is greater. The frequency-domain remote detection signal V f (m) is determined according to the following algorithm:
wherein m is a frame index, and M is a frame size for frequency domain processing.
The comparator 402 determines whether a difference between a power P x (m) of the audio signal and a stationary noise estimate power P n (m) of the audio signal is greater than a third threshold T x (m) to obtain a third comparison result C f (m). If the third comparison result C f (m) is true, it means that the power P x (m) of the audio signal is much larger than the stationary noise estimate power P n (m), and the audio signal may comprise speech. Thus, the pitch detection module 404 is triggered to perform pitch detection on the audio signal X(m) to generate a pitch detection signal D x (m). If the pitch detection is positive, the audio signal is confirmed to comprise speech. In one embodiment, the pitch detection module 404 performs pitch detection based on the method provided by D. Huang, etc. in “Speech pitch detection in noisy environment using multi-rate adaptive lossless FIR filters”, ISCAS'04, 22-26 May 2004, or the method provided by L. Hui, etc. in “A Pitch Detection Algorithm Based on AMDF and ACF”, ICASSP'06, 14-19 May 2006.
›DETAILED DESCRIPTION OF THE INVENTION · 3 of 3
If both the pitch detection signal D x (m) and the remote detection signal V f (m) are true, a conversation between Internet communication devices is underway, and the detector module 408 enables the speech detection result S x (n). Thus, the automatic gain control module 138 of FIG. 1 can then amplify audio signal X(m) according to speech detection result S x (n). The speech detection result S x (n) is determined according to the following algorithms:
wherein S x (m) is the speech detection result of frequency domain, the S x (n) is the speech detection result of time domain, and the function [x] denotes an integer closest to x.
FIG. 5 is a block diagram of a Internet communication device 500 with an array microphone according to the invention. The Internet communication device 500 is roughly similar to the Internet communication device 100 of FIG. 1 , except for an array microphone and the beam-forming module 535 . The array microphone includes two microphones 530 and 531 to receive two audio signals at different locations, and the beam-forming module 535 can suppress noise from the beam. The beam-forming module 535 can also provide in-beam and out-of-beam information I for the microphone speech detection module 506 . Thus, the microphone speech detection module 506 generates the speech detection result with better precision.
The invention provides a method for controlling noise of an Internet communication device. A line-in speech detection module is added to detect the speech of a remote audio signal sent by a far-end talker, and the remote audio signal is muted by a line-in channel control module if the remote audio signal is not speech. A microphone speech detection module is added to detect the speech of an audio signal received from a near-end talker, and the audio signal is not amplified if the audio signal is not speech. Thus, the noise including non-stationary noise is eliminated from the remote audio signal and the audio signal, and the audio quality of the Internet communication device is improved.
While the invention has been described by way of example and in terms of preferred embodiment, it is to be understood that the invention is not limited thereto. To the contrary, it is intended to cover various modifications and similar arrangements (as would be apparent to those skilled in the art). Therefore, the scope of the appended claims should be accorded the broadest interpretation so as to encompass all such modifications and similar arrangements.
Claims as granted
22 claimsLog in to read the claims of this application.
Log in to unlockClassifications
5 codes- G10L21/00
Claim changes
SoonSee which claims were amended, added or cancelled during examination, with every added and removed word marked.
The published claims of this application are not paired with the granted ones in what we hold.
File wrapper
See the full prosecution history — every USPTO and applicant action on this file, in order.
Log in to unlockDocuments
Log in to open the documents of this file: the application as filed, every office action and response, the notice of allowance.
Log in to unlockChain of title
See the full assignment history — every owner this patent has passed through, with recordation dates and reel/frame numbers.
Log in to unlock