USPatent applicationPatented

Method for recognizing multi-dimensional anomalous urban traffic event based on ternary gaussian mixture model

Granted 12 Apr 2022 · no office action yet

Life of the application

7 dated events
⤢ drag to zoom20202022202420262028203020322034203620382040ProsecutionOwnershipTerm & fees
ProsecutionOwnershipTerm & feeshover for detail · click to open

Abstract

A method for recognizing multi-dimensional anomalous urban traffic events based on a ternary Gaussian mixture model includes: reading a data sample of urban road traffic events; randomly dividing the data sample into a first subsample and a second subsample; performing modeling based on the first subsample by using the ternary Gaussian mixture model to obtain a second ternary Gaussian mixture model to calculate a distribution probability p of any sample point; clustering the second subsample, recognizing an outlier in the second subsample, and labeling the outlier and a normal point to obtain a labeled subsample; calculating the labeled subsample to obtain the distribution probability p corresponding to each sample point in the labeled subsample; when a new traffic event occurs, obtaining features of three dimensions of the new traffic event, calculating a distribution probability p by using the second model, and recognizing the new traffic event as anomalous if p<t-score.

Description

6 parts
›CROSS REFERENCE TO THE RELATED APPLICATIONS

This application is the national phase entry of International Application No. PCT/CN2020/084556, filed on Apr. 13, 2020, which is based upon and claims priority to Chinese Patent Application No. 201910820821.2, filed on Aug. 30, 2019, the entire contents of which are incorporated herein by reference.

›TECHNICAL FIELD

The present invention belongs to the technical field of intelligent traffic applications, and more particularly, relates to a method for intelligent recognition of multi-dimensional anomalous urban traffic events based on a ternary Gaussian mixture model and clustering.

›BACKGROUND

The comprehensive perception of urban road traffic conditions, especially the recognition and warning of anomalous urban traffic events, provides data support and a theoretical basis for alleviating traffic congestion and increasing traffic safety, and thus has important implications for improving the urban traffic management and decision-making capacity. At present, the main research focus is on the recognition of anomalous events on expressways, but there is a lack of research on the recognition of anomalous urban traffic events.

›SUMMARY

An objective of the present invention is to study, recognize, and determine anomalous urban traffic events based on traffic big data by using an artificial intelligence algorithm.

To achieve the above-mentioned objective, the technical solutions of the present invention provide a method for recognizing multi-dimensional anomalous urban traffic events based on a ternary Gaussian mixture model, including the following steps:

step 1: reading a data sample S of urban road traffic events, wherein an input X of the data sample S includes features of three dimensions: a traffic event quantity based on an event sequence, a weather condition, and a traffic congestion index;

step 2: randomly dividing the data sample S into a subsample S 1 and a subsample S 2 ;

step 3: performing modeling based on the subsample S 1 by using the ternary Gaussian mixture model to obtain a ternary Gaussian mixture model M, wherein the ternary Gaussian mixture model M is configured to calculate a distribution probability p of any sample point;

step 4: clustering the subsample S 2 by using a density-based spatial clustering of applications with noise (DBSCAN) algorithm, recognizing an outlier in the subsample S 2 , and labeling the outlier and a normal point to change the subsample S 2 to a labeled subsample S 3 ;

step 5: calculating the subsample S 3 by using the ternary Gaussian mixture model M obtained in step 3 to obtain a distribution probability p corresponding to each sample point x in the subsample S 3 , wherein a distribution probability p allowing F1score to reach a maximum is a threshold t-score, and F1score is calculated by the following formula:

wherein

tp represents a quantity of true-positive sample points, fp represents a quantity of false-positive sample points, fn represents a quantity of false-negative sample points, the true-positive sample point is defined as a sample point with both an anomalous model prediction result and an anomalous actual result, the false-positive sample point is defined as a sample point with an anomalous model prediction result but a normal actual result, and the false-negative sample point is defined as a sample point with a normal model prediction result but an anomalous actual result, wherein a method of selecting the threshold t-score includes the following steps:

step 501 : initializing an initial value of p′ and a highest value best_f1 of F1score, as 0, and selecting a step, wherein step=(max(P 3 )−min(P 3 ))/1000, wherein P 3 represents a set of the distribution probability p corresponding to each sample point x in the S 3 ;

step 502 : setting the value of p 1 ′ to a sum of a minimum value in the P 3 and one step, namely, p 1 ′=min(P 3 )+step;

step 503 : extracting a sample point whose distribution probability p is less than p 1 ′ from the subsample S 3 , determining, by using the ternary Gaussian mixture model M, that the sample point is an outlier, calculating the F1score, and denoting the calculated value as f1;

step 504: comparing f1 and best_f1; if f1 is greater than best_f1, setting the value of best_f1 to f1, and assigning the value of p 1 ′ to p′, namely, p′=p 1 ; and if f1 is not greater than best_f1, keeping the value of best_f1 and the value of p′ unchanged; and

step 505 : repeating steps 502 to 504 cyclically, and increasing p 1 ′ by one step each time until p 1 ′=max (P 3 ), wherein

the final value of p′ is the threshold t-score of an anomalous event on the urban road section; and

step 6: when a new traffic event occurs, obtaining features of three dimensions of the new traffic event, calculating a distribution probability p by using the ternary Gaussian mixture model M, and recognizing the new traffic event as anomalous if p<t-score.

Preferably, in step 2, the ratio of the subsample S 1 to the subsample S 2 is 9:1.

Preferably, in step 5, the distribution probability p is calculated by the following formula:

wherein p (x; μ, Σ) represents a distribution probability of a sample point x in the subsample S 3 , μ represents a mean vector of each dimension in the subsample S 3 , μ=[μ 1 , μ 2 , μ 3 ], wherein μ 1 represents a mean value of the traffic event quantity, μ 2 represents a mean value of the weather condition, μ 3 represents a mean value of the traffic congestion index, Σ represents a covariance matrix of each dimension in the subsample S 3 ,

∑ = σ 1 2 0 0 0 σ 2 2 0 0 0 σ 3 2 ,

σ 1 represents a standard deviation of the traffic event quantity, σ 2 represents a standard deviation of the weather condition, and σ 3 represents a standard deviation of the traffic congestion index.

In the present invention, anomalous urban traffic events are automatically recognized and determined by using an artificial intelligence algorithm. The recognition of the anomalous events is not limited to a single alert, but involves the comprehensive consideration of event data such as alerts, accidents, and construction. The method is thus applicable to a whole city at a macro level, a region at a meso level, and a road section at a micro level.

›BRIEF DESCRIPTION OF THE DRAWINGS

FIG. 1 shows the overall process of recognizing anomalous urban events based on a ternary Gaussian mixture model and clustering; and

FIGS. 2A-2C show the effect and threshold of anomaly recognition based on the ternary Gaussian mixture model. The ternary features are Traffic congestion index, Traffic event quantity and Weather condition. The original image is 3D, but for the convenience of display, it is represented by 2D profile: FIG. 2A is a two-dimensional (2D) cross-sectional view of the anomaly recognition and threshold effect of the two features which are Traffic event quantity and Traffic congestion index. FIG. 2B is a two-dimensional (2D) cross-sectional view of the anomaly recognition and threshold effect of the two features which are Traffic event quantity and Weather condition. FIG. 2C is a two-dimensional (2D) cross-sectional view of the anomaly recognition and threshold effect of the two features which are Traffic congestion index and Weather condition.

›DETAILED DESCRIPTION OF THE EMBODIMENTS

The present invention will be described in detail below with reference to the specific embodiments. It should be understood that these embodiments are only used to describe the present invention rather than to limit the scope of the present invention. In addition, those skilled in the art may make various changes and modifications to the present invention after reading the content of the present invention, and these equivalent forms shall also fall within the scope defined by the appended claims of the present invention.

According to the present invention, anomalous urban traffic events are recognized automatically. Not only a warning point with a high incidence of anomalies, but also a point with problematic data quality and a point with missing data are detected and then defined as anomalous events. The method includes the following steps:

Step 1: a data sample S of urban road traffic events is read, wherein an input X of the data sample S includes features of three dimensions: a traffic event quantity based on an event sequence, a weather condition, and a traffic congestion index, namely, S=[x 1 (1 . . . n) , x 2 (1 . . . n) , x 3 (1 . . . n) ], wherein x 1 n represents a traffic event quantity of the n th sample in the data sample S, x 2 n represents a weather condition of the n th sample in the data sample S, and x 3 n represents a traffic congestion index of the n th sample in the data sample S.

Step 2: the data sample S is randomly divided into a subsample S 1 and a subsample S 2 , wherein a ratio of the subsample S 1 to the subsample S 2 is 9:1, namely, S 1 =[x 1 (1 . . . m) , x 2 (1 . . . m) , x 3 (1 . . . m) ] and S 2 =[x 1 (1 . . . n-m) , x 2 (1 . . . n-m) , x 3 (1 . . . n-m) ], wherein m represents a quantity of samples in the subsample S 1 .

Step 3: modeling is performed based on the subsample S 1 =[x 1 (1 . . . m) , x 2 (1 . . . m) , x 3 (1 . . . m) ] by using the ternary Gaussian mixture model to obtain a ternary Gaussian mixture model M, wherein the ternary Gaussian mixture model M is configured to calculate a distribution probability p of any sample point, and a ternary Gaussian distribution is calculated by the following formula:

wherein, p represents a probability distribution; x represents a single sample point, and there are a total of m sample points in the subsample S 1 ; μ represents a mean vector of each dimension in the subsample S 1 , to be specific, μ 1 =Mean(x 1 (1 . . . m) ), representing a mean value of the traffic event quantity, μ 2 =Mean(x 2 (1 . . . m) ), representing a mean value of the weather condition, and μ 3 =Mean(x 3 (1 . . . m) ), representing a mean value of the traffic congestion index, and μ=[μ 1 , μ 2 , μ 3 ]; Σ represents a covariance matrix of each dimension in the subsample S 3 ,

∑ = σ 1 2 0 0 0 σ 2 2 0 0 0 σ 3 2 ,

σ 1 represents a standard deviation of the traffic event quantity, σ 2 represents a standard deviation of the weather condition, and σ 3 represents a standard deviation of the traffic congestion index; and T represents transposition of a matrix.

Step 4: the subsample S 2 is clustered by using a density-based spatial clustering of applications with noise (DBSCAN) algorithm, an outlier in the subsample S 2 is recognized, and the outlier and a normal point (0 represents the normal point, and 1 represents the outlier) are labeled to obtain a labeled subsample S 3 , namely S 3 =[x 1 (1 . . . n-m) , x 2 (1 . . . n-m) , x 3 (1 . . . n-m) , y (1 . . . n-m) ].

Step 5: the subsample S 3 is calculated by using the ternary Gaussian mixture model M obtained in step 3 to obtain a p-value corresponding to each sample point x in the subsample S 3 as P 3 =[p 3 1 , p 3 2 , p 3 3 , . . . , p 3 n-m ], wherein a p′ value allowing F1score to reach a maximum is a threshold t-score, and F1score is calculated by the following formula:

F ⁢ ⁢ 1 ⁢ score = 2 · precision · recall precision + recall ,

wherein

tp represents a quantity of true-positive sample points, fp represents a quantity of false-positive sample points, fn represents a quantity of false-negative sample points, the true-positive sample point is defined as a sample point with both an anomalous model prediction result and an anomalous actual result, the false-positive sample point is defined as a sample point with an anomalous model prediction result but a normal actual result, and the false-negative sample point is defined as a sample point with a normal model prediction result but an anomalous actual result.

In step 5, a method of selecting the threshold t-score includes the following steps:

Step 501 : an initial value of p′ and a highest value best_f1 of F1score are initialized as 0, and a step is selected, wherein step=(max(P 3 )−min(P 3 ))/1000, wherein P 3 represents a set of the distribution probability p corresponding to each sample point x in the S 3 .

Step 502 : the value of p 1 ′ is set to a sum of a minimum value in the P 3 and one step, namely, p 1 ′=min(P 3 )+step.

Step 503 : a sample point whose distribution probability p is less than p 1 ′ is extracted from the subsample S 3 , the ternary Gaussian mixture model M determines that the sample point is an outlier, the F1score is calculated, and the calculated value is denoted as f1.

Step 504 : f1 is compared with best_f1; if f1 is greater than best_f1, the value of best_f1 is set to f1, and the value of p 1 ′ is assigned to p′, namely, p′=p 1 ; and if f1 is not greater than best_f1, the value of best_f1 and the value of p′ are kept unchanged.

Step 505 : steps 502 to 504 are repeated cyclically, and p 1 ′ is increased by one step each time until p 1 ′=max (P 3 ).

The final value of p′ is the threshold t-score of an anomalous event on the urban road section.

Step 6: When a new traffic event occurs, features of three dimensions of the new traffic event are obtained, a distribution probability p is calculated by using the ternary Gaussian mixture model M, and the new traffic event is recognized as anomalous if p<t-score.

Claims as granted

3 claims

Log in to read the claims of this application.

Log in to unlock

Classifications

2 codes
IPC · International Patent Classification
Section G — Physics
  • G08G1/01
  • G06N7/00

Claim changes

Soon
Coming soonHow the claims changed between publication and grant

See which claims were amended, added or cancelled during examination, with every added and removed word marked.

AmendedAddedCancelledUnchanged

The published claims of this application are not paired with the granted ones in what we hold.

File wrapper

⤢ drag to zoomApr 2020Jul 2020Oct 2020Jan 2021Apr 2021Jul 2021Oct 2021Jan 2022Apr 2022USPTOApplicantNotice of allowance
USPTOApplicanthover for detail · click to open
Pendency
2.0 y
729 days filing → grant
Office actions
0
none on record
Examiner
Ian Jen
art unit 3664 · TC 3600
Citations: 7 back · 2 forward

See the full prosecution history — every USPTO and applicant action on this file, in order.

Log in to unlock

Documents

Log in to open the documents of this file: the application as filed, every office action and response, the notice of allowance.

Log in to unlock

Chain of title

⤢ drag to zoom2022202420262028203020322034203620382040Owner 1
Titlehover for detail · click to open

See the full assignment history — every owner this patent has passed through, with recordation dates and reel/frame numbers.

Log in to unlock