Automated EDA with Python

Pandas, Python
In this post, we will investigate the pandas_profiling and sweetviz packages, which can be used to speed up EDA (exploratory data analysis) with Python. In a previous article, we talked about an analagous package in R (see this link). Getting started with pandas_profiling pandas_profiling can be installed using pip, like this: Next, let's read in our dataset. The data we'll be using is a heart attack-related dataset, which can be found here. Now, let's import ProfileReport from pandas_profiling. If you're running this code in Jupyter Notebook, you should see the report generated within your notebook file. The report shows several pieces of analysis. First, it gives a summary glimpse of the data, giving the number of variables, observations, missing values and percentages, data type information, and number of duplicate rows…
Read More