Showing posts with label Confusion Matrix. Show all posts
Showing posts with label Confusion Matrix. Show all posts

Friday, November 23, 2018

Understanding Misclassifier


A common issue that has to be addressed during model building  is the understanding the output and churning it . One important understanding is required in the term 'misclassifier'.

It is also required to understand the following terms
True positive , false positive, true negative , false negative.

Depending upon the business / technical problem the important misclassifier  need to be identified and churned .

For instance,
False positive means predicting a negative result as a positive result.

In the below example  a sample fictitious data is used for understanding.
It has 32 variables and 1000 records. The data is scaled suitably to use it for analysis, which is a typical data preprocessing activity.
It is partitioned for testing and training. 80% -20 %

Two models were build- ANN and CRT.

How should we understand the results?







                     Figure1 : Model built- ANN and CRT using SPSS Modeler.






Figure2: Results from CART algorithm 


Figure3: Response from the Artificial and Neural Network. 


Based on the given output by the models the question that arises is which should be chosen ?
Considering the figure 2, the false positive 64 % and the figure 3 gives a result of 55.4 % which mean that the the former model is saying the result as positive which is higher than the ANN model.
What is the conclusion ?
If we are bothered on a model which should say the facts as such , then both of those are dangerous. Say for instance we are working on a hospital data for which treatment is given for  a symptom and the model predicts that the person has a disease. Isn't this dangerous ? Hence work on the  misclassifier and refine the model . 

Friday, August 24, 2018

Confusion Matrix -- Is it a confusion all the time ?

Part -2 of confusion matrix
How to read a confusion matrix

Supposing the machine predicts that a person gets a particular disease.

 the matrix displays as

Confirmed
Not confirmed
Confirmed
35
20
Not confirmed
25
15

What do you understand? This has to be understood as the rows denote the actual values and the columns form the predicted values. Let us go deep now.

 I have given summation of the rows and columns

Predicted values

Confirmed
Not confirmed
total
Actual values
Confirmed
35
20
55
Not confirmed
25
15
40
total
60
35
95


Here in the above matrix , the important class is 'confirmed.' We don't bother about the other part. 

Now comes the terminologies. Specificity


Again what do we mean? The machine predicts that 35 people has a disease against the total count of 55.
Secondly the machine predicts that 15 people do not have a disease against the total of 40 people without a disease.

What should we take from this ?
35/55  =.6363 or  the specificity is 63 .63 %  We expect the machine to be accurate about 90-95 or still higher as the case may be to predict people with a disease.
 Sensitivity= True positive / (true positive + False negative ) -----> The sensitive part of our problem


Secondly the machine predicts 15/40 don't have a disease. We are not too much bothered about it. 
ie 37.5 %  is the specificity.

Specificity= true negative / ( false positive + true negative)


Monday, May 7, 2018

Reading a confusion Matrix... A misclassifier


How to read a confusion matrix
suppose you have 95 records available for prediction. The record has a target to predict whether the person has a particular disease or not.
 the matrix displays as

Confirmed
Not confirmed
Confirmed
35
20
Not confirmed
25
15

What do you understand?

 I have given summation of the rows and columns

Predicted values

Confirmed
Not confirmed
total
Actual values
Confirmed
35
20
55
Not confirmed
25
15
40
total
60
35
95

Can you get some information?
The green cells are the actuals, while the yellows are the predicted values and the pink are the errors.
Actual values are the one which are tested physically and the results are available on hand. Which means in our case 55 people had the disease and 40 don’t have.
What did the algorithm do ? It has predicted as 60 people who have disease and 35 who do not have.
How to interpret? Out of the 55 people, the algorithm has made a wrong prediction of classifying 20 people as they don’t have a disease. 
Secondly, out of 45 people who do not have disease it has predicted 25 people that they may have a disease, which is a very serious issue.