Firstly, how can you get such a big volume of data? What kind of file ? What is the source ?
Any industry which is associated with sensor data, e commerce sites, online portals , telecom industries have such a huge data. These can't be stored in the normal csv format or in excel files.
These could be in server , if not in the files as json format. The volume of the files could be in GBs .
coming to the point of 5.5 million records [ in Indian context- More than Half a crore data], size of the files could be in the range of 3.5 GBs in the json format.
How could one read those file and get the insights ?
While there are several tools available , open source tools such as R and Python , does the magic.
Python has several features , reading the file could be easier with just two commands. One need to be more cautious as it involves ' Date - time ' as a variable, which will can often put us in trouble.
Getting insights:
Before starting the analysis part, what is important is to understand the type of conversion of data that has happened. There are possibilities that the same type of data is not understood by the package.
'int' can also be read as 'char'.
Since it is in the json format, chances are that the variables get stored as 'dict' format.
The major work will be to clean up the data and make it workable.
'Groupby' can used in cases to draw conclusions and form tables
Once things are set in order, it is the usual dataframe and analysis part is just a usual piece of cake.
Happy Analysis !
Tuesday, May 21, 2019
Tuesday, April 16, 2019
How Excel can put you in trouble...
I have a Large data , the head of it will look like the one as given below
My file did not have headers. I have 10 such files, need to perform analysis and get some insights.
The very first step I did was to insert a row and gave the names of the headers .
Saved the file in the csv format. Used R to read the file . It gave the summary . Every thing went fine.
I was about to compare the column of mobile_num of one file with another, it gave unexpected output which put me in lot of trouble.
What went wrong ?
After several attempts to check my code, test my logic, I was not convinced that my R prog has malfunctioned. Finally decided to see the data what R has read. The mobile numbers were read as 6 numbers followed by zeros as given in the exponential value. Truly unimaginable to the core!! how did it happen. I don't know. But I have figured out what should have been done to avoid this.
1. While saving a column with exponential number / a 12 digit number it has to be formatted as number with no decimal places.
2. Insert a row above
3. paste the header names
4. save it and re open for confirmation.
You could save atleast your own time if you know the above.
2.33546E+11
|
ee
|
21-04-2017
13:04
|
ACTIVE
|
Pop
|
4
|
2017
|
2.33246E+11
|
gg
|
21-04-2017
15:43
|
ACTIVE
|
SMS
|
4
|
2017
|
2.33245E+11
|
hh
|
22-04-2017
21:52
|
ACTIVE
|
Intg
|
4
|
2017
|
2.33244E+11
|
jj
|
22-04-2017
22:09
|
ACTIVE
|
Intg
|
4
|
2017
|
My file did not have headers. I have 10 such files, need to perform analysis and get some insights.
The very first step I did was to insert a row and gave the names of the headers .
mobile_num
|
keyword
|
Date and time
|
status_name
|
mode_name
|
month
|
year
|
2.33546E+11
|
ee
|
21-04-2017
13:04
|
ACTIVE
|
Pop
|
4
|
2017
|
Saved the file in the csv format. Used R to read the file . It gave the summary . Every thing went fine.
I was about to compare the column of mobile_num of one file with another, it gave unexpected output which put me in lot of trouble.
What went wrong ?
After several attempts to check my code, test my logic, I was not convinced that my R prog has malfunctioned. Finally decided to see the data what R has read. The mobile numbers were read as 6 numbers followed by zeros as given in the exponential value. Truly unimaginable to the core!! how did it happen. I don't know. But I have figured out what should have been done to avoid this.
1. While saving a column with exponential number / a 12 digit number it has to be formatted as number with no decimal places.
2. Insert a row above
3. paste the header names
4. save it and re open for confirmation.
You could save atleast your own time if you know the above.
Thursday, April 11, 2019
Few of the problems we face...
While we know the syntax, algorithms, context and have the ability to anlayse the given data which may also be in a weird format, there are always challenges which consumes tons of our time by throwing error.
The error and the challenges I faced- quite a few,
1. A csv file which contained Date column and quite other variables. While R was reading this file, it interpreted as character, the character was something like ' 44326'. I was struggling with all my known skills/ strategies, I was not able to succeed.
Finally when gone through the data, I found the String ' Total' at the end of the date column, which was almost impossible for me to identify.
2. The second challenge was the date conversion, while the package may be R or Python, there are lot more chances that the date is not interpreted properly and the program might run.While I was cross checking the data I found the error, by the time I found it hours have gone !.
3.Use of reserved words / Standard words for the package - Example- count.
4. Space between the variables in the header
5. Blank lines in the first row
Lot more... will add then and there.
The error and the challenges I faced- quite a few,
1. A csv file which contained Date column and quite other variables. While R was reading this file, it interpreted as character, the character was something like ' 44326'. I was struggling with all my known skills/ strategies, I was not able to succeed.
Finally when gone through the data, I found the String ' Total' at the end of the date column, which was almost impossible for me to identify.
2. The second challenge was the date conversion, while the package may be R or Python, there are lot more chances that the date is not interpreted properly and the program might run.While I was cross checking the data I found the error, by the time I found it hours have gone !.
3.Use of reserved words / Standard words for the package - Example- count.
4. Space between the variables in the header
5. Blank lines in the first row
Lot more... will add then and there.
Thursday, March 28, 2019
Capabilities of python
What special things can
be done using python?
Recently I happen to
analyse a set of csv files atleast 100.
Each file contains several fields- variables, among them the
most is the purchaser Id, purchase date .
The task is,
List all the purchaser’s id and find them in the successive
date files, to identify whether the
purchaser has visited again.
If available list the dates and the value of all purchases
he has made.
The illustration
Day1 – 01-12-2018
Ids value of
purchase
123 10
234 15
456 25
Get the summary – 3 customers , Rs 50 /-
Find these ids in the next day – 02-12-2018
If available get the summary- only one purchaser id is
available – value is Rs 20
Extend this for the next 99 days. This is cycle 1.
Pick up the list of purchasers in the day 2 and continue
this process for the next 98 days.
WoW , what a wonderful tool is this ?
It is calculating the number of files available, running the iterations – say 100 factorial,
tabulating summarizing the results
as output in excel format, and what else, the capability is still more.
Friday, February 8, 2019
Thursday, February 7, 2019
Market basket Analysis
Few terminologies:
Transaction is a set of items (Itemset).
Confidence : It is the measure of uncertainty or trust worthiness associated with each discovered
pattern.
Support : It is the measure of how often the collection of items in an association occur together as percentage of all transactions
Frequent itemset : If an itemset satisfies minimum support,then it is a frequent itemset.
Strong Association rules: Rules that satisfy both a minimum support threshold and a minimum
confidence threshold
In Association rule mining, we first find all frequent itemsets and then generate strong association rules from the frequent itemsets
Apriori algorithm
is the most established algorithm for finding frequent item sets mining.
The basic
principle of Apriori is “Any subset of a frequent itemset must be frequent”.
We use these
frequent itemsets to generate association rules.
Association rule
mining
Finding frequent patterns, associations, correlations, or
causal structures among sets of items in transaction databases
Understand customer buying habits by finding associations
and correlations between the different items that customers place in their
“shopping basket”
Applications: Basket data analysis, cross‐marketing,
catalog design, loss‐leader analysis, web log analysis, fraud detection (supervisor‐>examiner)
Rule form
Antecedent →Consequent [support, confidence]
(support and confidence are user defined measures of interestingness)
Let the rule discovered be {Jamun,...} → {Potato Chips}
Potato chips as consequent => Can be used to determine what should be done to boost its sales
Jamun in the antecedent => Can be used to see which products would be affected if the store discontinues selling Jamun
Jamun in antecedent and Potato chips in the consequent => Can be used to see what products should be sold with Jamun to promote sale of Potato Chips
Find all itemsets that have high support,These are known as frequent itemsets. Generate association rules from frequent itemsets
Let us see an example
:
Finding the support, confidence and lift of i)
shirts and ties, ii) trousers and ties.
|
Transaction Id
|
Shirts
|
Trousers
|
Ties
|
|
001
|
1
|
1
|
1
|
|
002
|
0
|
1
|
0
|
|
003
|
1
|
0
|
1
|
|
004
|
1
|
0
|
1
|
|
005
|
1
|
1
|
0
|
For our data there are 3 transactions with both shirts and ties (shirts ∩
ties) out of total 5 transactions.
Support =3/5 =0.6 or 60%
Confidence for association is calculated using the following formula:
In our example, there are 3 transaction for both shirts and ties
together out of 4 transactions for shirts. The calculation for confidence
for our dataset is:
Confidence =3/4 =0.75 or 75 %
A third useful metric for association analysis is lift; it is defined
as:
Expected confidence in the above formula is presence of ties in the
overall dataset i.e. there are 4 instances of ties purchase out of 5.
The value for lift, 125%, shows that purchases of the ties improve when
the customers buy shirts. The point to note here is
that if the customer buys a shirt, does his chance of buying ties go up i.e.
value of lift above 100
Lift= confidence / expected confidence = P( ties | shirts) / P(ties)
Expected confidence in the above formula is
presence of ties in the overall dataset i.e. there are 4 instances of ties
purchase out of 5.
Lift = (3/4) / ( 3/5) = 15/12 =1.25 or 125%
Similarly for the trousers and ties
Support: 1/5
Confidence: 1/3
Lift= ( 1/3 )/ (3/5)= 5/9 or 55.6%
Customer Life Time Value
Just like we use Net Present Value (NPV) to evaluate
investments and companies, we use CLV to evaluate customer relationships CLV is
the expected NPV of the cash flows from a customer relationship CLV is defined
as the discounted sum of all future customer revenue streams minus product and
servicing costs and re marketing costs
Let us assume we are analyzing the customer life time value of a company which offers services of a kind . The following are obtained.
Where ,
This is a simple model of calculation. The formula differs when the service is offered before the contribution from the customer. [ Eg- Credit cards ]
Segmentation of data -K- Means
- Demographic segmentation
In business, demographic segmentation is when
an organization uses data about the demographic characteristics of its
customers to better target and enhance its marketing efforts.
By segmenting a market by demographic variables
such as the age of the customer, gender, income, education, religion, and
family life cycle, business professionals are able to create groups of
customers that display similar wants and needs. Examples
Age
Gender
Income
Education
Family life cycle
other segmentationss
- Behavioral segmentation
- Psychographic segmentation
- Geographic segmentation
Illustration of the algorithm- Choose two random centeroids.
Step:3
Step:4
What we expect, but what we get will be different.
This can
be accomplished by making the WCSS -within-cluster sums of squares as minimum as possible. This metric determines the optimum number
of clusters .
Friday, November 23, 2018
Understanding Misclassifier
A common issue that has to be addressed during model building is the understanding the output and churning it . One important understanding is required in the term 'misclassifier'.
It is also required to understand the following terms
True positive , false positive, true negative , false negative.
Depending upon the business / technical problem the important misclassifier need to be identified and churned .
For instance,
False positive means predicting a negative result as a positive result.
In the below example a sample fictitious data is used for understanding.
It has 32 variables and 1000 records. The data is scaled suitably to use it for analysis, which is a typical data preprocessing activity.
It is partitioned for testing and training. 80% -20 %
Two models were build- ANN and CRT.
How should we understand the results?
Figure1 : Model built- ANN and CRT using SPSS Modeler.
Figure2: Results from CART algorithm
Figure3: Response from the Artificial and Neural Network.
Based on the given output by the models the question that arises is which should be chosen ?
Considering the figure 2, the false positive 64 % and the figure 3 gives a result of 55.4 % which mean that the the former model is saying the result as positive which is higher than the ANN model.
What is the conclusion ?
If we are bothered on a model which should say the facts as such , then both of those are dangerous. Say for instance we are working on a hospital data for which treatment is given for a symptom and the model predicts that the person has a disease. Isn't this dangerous ? Hence work on the misclassifier and refine the model .
How to start python / learning programming ?
So, you have researched , googled, and finally came to a conclusion that need python for enhancement in career, business, automation and so on.....But where and how to start.
With my experience in automation,predictive analytics and interest in software, I have touched few points,if you think useful pick those. I walked through all the steps mentioned here.
1.How to start learning python ? Which are the best resources- websites / e- books / institutes/ tutors are the questions that are frequently asked by people ..
My one line answer is your passion, your interest ,money and the google.
Start from any mode, but start today.
All the text books/ contents/ courses of online can give you information or make you walk around.
In my opinion swimming cannot be taught online. Likewise is the programming.
You will not have the same set of theoretical programming that you have learnt in labs/ classes in real applications.Therefore the basics can be learnt in 15- 20 hours of simple training- it can even be self taught. I repeat- the word ' PASSION'.
2.You start a simple program yourself,with a problem statement- addition of two number/ multiplication - logical questions or of any kind - be a self starter- later you can magnify the problem level and improve your programming skill.
It needs lot of guts to say this ! The reason being most of the resources give problems and also solution which hinders our way of thinking. Solve it yourself--- YOU WILL GAIN CONFIDENCE.
3. Time matters- More the time you devote for programming, more will be your ability.
4. There is no substitute for your creativity and problem solving ability- use that. If you have not got any programming exposure- that is the best thing needed- you can think in a different mode.
All you need is the logic and not the syntax. I know the experts with lot of experience still google for the syntax.
5. The software's capability is infinite- it means the problem you think today might have been experienced by most others years ago. The other skill needed is ' GOOGLING' the solution for your syntax.
While we study in primary school, we were exposed to the languages,alphabets, words and sentences. Can mugging up an answer for a question help us anywhere ? Similarly is the programming
Know the data types, syntax - for few and start writing programs. That's all the one you need.
Joining in groups, answering questions asked in stackoverflow, attending contests can definitely make you pro!!!
With my experience in automation,predictive analytics and interest in software, I have touched few points,if you think useful pick those. I walked through all the steps mentioned here.
1.How to start learning python ? Which are the best resources- websites / e- books / institutes/ tutors are the questions that are frequently asked by people ..
My one line answer is your passion, your interest ,money and the google.
Start from any mode, but start today.
All the text books/ contents/ courses of online can give you information or make you walk around.
In my opinion swimming cannot be taught online. Likewise is the programming.
You will not have the same set of theoretical programming that you have learnt in labs/ classes in real applications.Therefore the basics can be learnt in 15- 20 hours of simple training- it can even be self taught. I repeat- the word ' PASSION'.
2.You start a simple program yourself,with a problem statement- addition of two number/ multiplication - logical questions or of any kind - be a self starter- later you can magnify the problem level and improve your programming skill.
It needs lot of guts to say this ! The reason being most of the resources give problems and also solution which hinders our way of thinking. Solve it yourself--- YOU WILL GAIN CONFIDENCE.
3. Time matters- More the time you devote for programming, more will be your ability.
4. There is no substitute for your creativity and problem solving ability- use that. If you have not got any programming exposure- that is the best thing needed- you can think in a different mode.
All you need is the logic and not the syntax. I know the experts with lot of experience still google for the syntax.
5. The software's capability is infinite- it means the problem you think today might have been experienced by most others years ago. The other skill needed is ' GOOGLING' the solution for your syntax.
While we study in primary school, we were exposed to the languages,alphabets, words and sentences. Can mugging up an answer for a question help us anywhere ? Similarly is the programming
Know the data types, syntax - for few and start writing programs. That's all the one you need.
Joining in groups, answering questions asked in stackoverflow, attending contests can definitely make you pro!!!
Wednesday, September 26, 2018
Python coding
Most often Data science professionals especially who use open source tools like python encounter issues related to date. Conversion to different format, short forms etc. This is an attempt to bring few of the issues related to date in one window .
How to get the last 6 months from the current date ->( format May 2018 )?
datelist=[ ]
datem = datetime.date.today().replace(day=1)
for i in range (7):
x = datem + relativedelta(months=-i)
x=datetime.datetime.strftime(x,'%b %Y')
datelist.append(x)
del datelist[0]
How to get the last 6 months from the current date ->( format May 2018 )?
datelist=[ ]
datem = datetime.date.today().replace(day=1)
for i in range (7):
x = datem + relativedelta(months=-i)
x=datetime.datetime.strftime(x,'%b %Y')
datelist.append(x)
del datelist[0]
Friday, August 24, 2018
Confusion Matrix -- Is it a confusion all the time ?
Part -2 of confusion matrix
How to read a confusion matrix
Supposing the machine predicts that a person gets a particular disease.
Supposing the machine predicts that a person gets a particular disease.
the matrix displays as
Confirmed
|
Not confirmed
| |
Confirmed
|
35
|
20
|
Not confirmed
|
25
|
15
|
What do you understand? This has to be understood as the rows denote the actual values and the columns form the predicted values. Let us go deep now.
I have given summation of the rows and columns
Predicted values
| ||||
Confirmed
|
Not confirmed
|
total
| ||
Actual values
|
Confirmed
|
35
|
20
|
55
|
Not confirmed
|
25
|
15
|
40
| |
total
|
60
|
35
|
95
| |
Here in the above matrix , the important class is 'confirmed.' We don't bother about the other part.
Now comes the terminologies. Specificity
Again what do we mean? The machine predicts that 35 people has a disease against the total count of 55.
Secondly the machine predicts that 15 people do not have a disease against the total of 40 people without a disease.
What should we take from this ?
35/55 =.6363 or the specificity is 63 .63 % We expect the machine to be accurate about 90-95 or still higher as the case may be to predict people with a disease.
Sensitivity= True positive / (true positive + False negative ) -----> The sensitive part of our problem
Secondly the machine predicts 15/40 don't have a disease. We are not too much bothered about it.
ie 37.5 % is the specificity.
Specificity= true negative / ( false positive + true negative)
Subscribe to:
Posts (Atom)



















