Showing posts with label python. Show all posts
Showing posts with label python. Show all posts

Thursday, July 25, 2019

Problem solving using R and Python a caution



Recently, I had a task of identifying repeated customers. For instance customer A purchases a product on 17th May and later comes on 22nd May , 24th May and so on. Likewise customer B purchases a product on 18th May and purchases similar products in  later dates. I had  a data of  3000+ customers and the repeated visit list of 21000. I need to identify  day wise  list of repetition something like this. The expected outcome is “ on which date of purchase yielded Maximum repeated customer’. Later we could identify the revenue.
Date of
 First purchase/Repeat
18 May
19 May
20 May
17 May
3
4
7
18 May
12
6
5
19 May
10
4
2

I used python for reading both the lists , coded  and computed the daywise list.  Out of curiosity , used merge command and tabulated daywise using R to cross check. To my shock there was a difference of 25 %. It spoiled my entire enthusiasm of doing  something useful as there is a substantial difference which otherwise should not be.  Compared the  tabulated values daywise obtained from python and the values obtained from R.
This simple piece of work has cost me few hours , but the learning is forever, which I thought of sharing .
Findings:   There is a marginal difference in every row. This is due to the fact that the machine/ data source has  duplicated the data while capturing the repeated customers. Cross checking is always required , may be with a different tool / strategy / approach. This might have resulted in wrong computations and projections, but I have avoided them in full.

Tuesday, May 21, 2019

Is 55 Lakh data Difficult to analyse

Firstly, how can you get such a big volume of data? What kind of file ? What is the source ?

Any industry which is associated with sensor data,  e commerce sites, online portals ,  telecom industries  have such a huge data. These can't be stored in the normal csv format or in excel files.
These could be in server ,  if not in the files as json format. The volume of the files could be in GBs .

coming to the point of 5.5 million records [ in Indian context- More than Half a crore data], size of the files could be in the range of  3.5 GBs in the json format.


How could one read those file and get the insights ?

While there are several tools available , open source tools such as R and Python , does the magic.

 Python has several features , reading the file could be easier with just two commands. One need to be more cautious as it involves  ' Date - time '  as a variable, which will  can often put us in trouble.

Getting insights: 
Before starting the analysis part, what is important is to understand the type of conversion of data that  has happened. There are possibilities that the same type of data is not  understood by the package.
'int' can also be read as 'char'.
Since it is in the json format, chances are that the variables get stored as 'dict' format. 
The major work will be to clean up the data and make it workable.


'Groupby'  can used in cases to draw conclusions and form tables

Once things are set in order, it is the usual dataframe and analysis part is just a usual piece of cake.
Happy Analysis !

Thursday, March 28, 2019

Capabilities of python


What  special things can be done using python?
Recently I happen to  analyse a set of csv files atleast 100.
Each file contains several fields- variables, among them the most is the purchaser Id, purchase date .
The task is,
List all the purchaser’s id and find them in the successive date  files, to identify whether the purchaser has visited again.
If available list the dates and the value of all purchases he has made.
The illustration
Day1 – 01-12-2018
Ids          value of purchase
123         10
234         15
456         25
Get the summary – 3 customers , Rs 50 /-

Find these ids in the next day – 02-12-2018
If available get the summary- only one purchaser id is available – value is  Rs 20
Extend this for the next 99 days. This is cycle 1.
Pick up the list of purchasers in the day 2 and continue this process for the next 98 days.

WoW , what a wonderful tool is this ?
It is calculating the number of files available, running  the iterations – say 100 factorial, tabulating   summarizing  the results  as output in excel format, and what else, the capability is still more.




Saturday, April 28, 2018

What's there in wine ? Part 2 PCA- validating with SPSS Modeler

Are you curious?
Will  the data set which was analysed in  Python, when tested with spss modeler give the same result?
The data of Wine and its components were given as input to the spss modeler. The process which were done in the python was done in spss modeler.( Scaling,and partition).

The same input conditions were given, keeping the customer segment as the target variable.
Feature scaling :( x- min(x)/ range (x)



The choice of the factors was done on based on the  results of python 55 % variance.



Partition : 80-20  --> Training to testing

Number of components: 2 ( the factors which were 14 were shrunk to 2)





A filter node is connected to the nugget to filter only the factors and the customer segment to perform the logistic regression.
Finally the analysis node give the results in the form of a confusion matrix.
Let us see in detail.




Here are the results .

The confusion matrix,


So what is to be  noted ?


  • The variables which were 14 in number were reduced to two factors
  • The equation of two factors were given above.
  • The logistic regression models gives the results with an accuracy of 97.23 %
  • The prediction of spss modeler for the testing set is perfect with a misclassifier of just 1 which is the same as the python.( the previous post)
Post your comments and views.

Thursday, March 29, 2018

IMPUTING DATA IN PYTHON

A very basic and important thing in the data analysis and in machine learning is imputing. We cannot delete the record as it has some missing values. It may contain some valuable information. One of the strategies is to impute. That is we could put the mean values of the row/ column depending on the need and fill the cells. Python has some in built packages which does this function. I am giving the steps involved in doing this.