www.icgst.com
home
Password
Community
Styles
Feedback
Sign Up
Sign in
Paper Details:
Downloads:
1403
Serial Number:
P1121317270
Title:
A Model For Improving Classifier Accuracy using Outlier Analysis
Authors:
Lakshmi Sreenivasa Reddy.D and D Ramchander. M
Abstract:
Outlier analysis is an important task for data mining to find outliers in datasets. Anomalies are objects; they have different behavior and do not follow with the remaining objects in the datasets. Outliers do not follow the rules formed by other data objects in the dataset. Detecting outliers efficiently is an important issue in many fields like science, medicine and Engineering. There are many applications in real life like fraud detection in electronic commerce, network intrusion detection. So many methods are available to detect outliers in numerical datasets. But limited methods are available for categorical datasets. We propose a method to detect outliers in categorical data based on disturbance made by the object. This method finds outliers based on each record disturbance score from our Algorithm that has great intuitive appeal. In this paper we call these scores as BAD scores. This algorithm utilizes the frequency of each value in the dataset .Greedy method needs k- scans of dataset to find ‘k’ outliers. Our method needs only one scan of dataset and it calculates BAD score of each record directly. It avoids the problem of giving ‘k’ as an input.AVF method shows less time complexity and accuracy comparing with other methods like Greedy and FPOF, FDOD and Greedy has good accuracy comparing with other methods like AVF and FPOF, FDOD (which are based on frequency patterns of all combinations of values in each record) and time complexity is multiple of AVF. Our algorithm shows better results in accuracy than AVF algorithm and Greedy. But this method has reached nearest to AVF in time complexity. This algorithm has been applied on Nursery dataset and Bank dataset taken from “UCI Machine Learning Repository”. In this paper we have extended Normal distribution [11], and Fuzzy concept [12] to BAD score [13] .We have excluded numerical attributes from Datasets for our analysis. The experimental results show that it is efficient for outlier detection in categorical dataset.
Keywords:
Data Mining, Outlier detection, BAD Score, NAVF, FuzzyAVF
Journal/Conference:
International Journal of Artificial Intelligence and Machine Learning
Volume:
15
Issue:
1
Submission Date:
4/22/2013 12:00:00 AM
Review Date:
5/30/2013 1:11:51 PM
Publishing Date:
9/14/2015 10:04:43 AM
Article Downloads:
1403
Download:
Facebook