Exploratory Data Analytics
Exploratory Data Analytics (EDA) is used to understand the characteristics, patterns, relationships and distributions present in a dataset before applying advanced models.
1. Introduction to Exploratory Data Analytics
Exploratory Data Analytics, commonly called EDA, is the process of examining and summarizing a dataset using statistical methods and graphical techniques.
Objectives of EDA
- Understand the basic characteristics of data.
- Identify patterns and trends.
- Detect missing values and outliers.
- Understand distributions of variables.
- Find relationships between variables.
- Generate hypotheses for further analysis.
- Support selection of suitable statistical or machine learning methods.
2. Descriptive Statistics
Descriptive statistics summarizes and describes the important features of a dataset using numerical measures and tables.
Main Categories
Measures of Central Tendency
Mean, median and mode describe the central or typical value of a dataset.
Measures of Dispersion
Range, variance and standard deviation describe how widely the observations are spread.
Measures of Shape
Skewness and kurtosis describe the shape of a distribution.
Graphical Measures
Box plots, histograms and other visualizations help understand the distribution of data.
3. Mean
The arithmetic mean is the average value of all observations in a dataset.
For observations x₁, x₂, ..., xₙ:
Example
Consider the data: 10, 20, 30, 40, 50
Sum = 150 and number of observations = 5. Therefore, mean = 30.
Advantages
- Easy to calculate.
- Uses every observation.
- Useful for many statistical calculations.
Limitation
Mean can be strongly affected by extreme values or outliers.
4. Standard Deviation
Standard deviation measures how much the observations deviate or spread around the mean.
For a sample, the denominator is commonly replaced by n − 1.
Interpretation
- Small standard deviation → values are relatively close to the mean.
- Large standard deviation → values are more widely spread.
- Standard deviation has the same units as the original variable.
5. Skewness
Skewness measures the degree of asymmetry of a probability distribution around its mean.
Types of Skewness
Positive Skewness
The distribution has a longer tail toward the right. The mean is generally greater than the median.
Negative Skewness
The distribution has a longer tail toward the left. The mean is generally less than the median.
Zero Skewness
The distribution is approximately symmetric around its center.
Skew
Skew
Negative skew → longer left tail.
6. Kurtosis
Kurtosis describes the shape of a distribution, particularly the degree of tail heaviness and concentration around the center relative to a reference distribution.
Common Categories
Mesokurtic
A distribution with kurtosis characteristics similar to the normal distribution.
Leptokurtic
A distribution with relatively heavier tails and a sharper central concentration.
Platykurtic
A distribution with relatively lighter tails and a flatter central shape.
7. Box Plot
A Box Plot, also called a box-and-whisker plot, is a graphical method for displaying the distribution of numerical data using quartiles.
Five-Number Summary
Important Terms
- Q1: First quartile or 25th percentile.
- Median: Middle value or 50th percentile.
- Q3: Third quartile or 75th percentile.
- IQR: Interquartile Range = Q3 − Q1.
- Whiskers: Lines extending from the box toward the lower and upper values according to the plotting convention.
- Outliers: Observations unusually far from the central portion of the data.
- Compare distributions.
- Identify spread.
- Identify median.
- Detect possible outliers.
- Understand skewness.
8. Pivot Table
A Pivot Table is a data summarization technique that organizes and aggregates data according to selected categories.
Example
Suppose a sales dataset contains:
- Product
- Region
- Sales
- Quantity
A pivot table can summarize total sales by region and product.
| Region | Product A | Product B | Total |
|---|---|---|---|
| North | ₹20,000 | ₹15,000 | ₹35,000 |
| South | ₹18,000 | ₹22,000 | ₹40,000 |
| West | ₹25,000 | ₹17,000 | ₹42,000 |
Advantages
- Quickly summarizes large datasets.
- Makes comparison easier.
- Supports aggregation.
- Helps identify trends and patterns.
9. Heat Map
A Heat Map is a graphical representation in which different values are represented using variations in color intensity.
Example
| Maths | OS | DBMS | CN | |
|---|---|---|---|---|
| Student A | 90 | 72 | 85 | 78 |
| Student B | 65 | 88 | 74 | 92 |
| Student C | 80 | 70 | 95 | 83 |
Applications
- Correlation matrices.
- Student performance analysis.
- Business dashboards.
- Website activity analysis.
- Financial data analysis.
- Geographical data visualization.
10. Correlation Statistics
Correlation measures the degree and direction of association between two variables.
Correlation Coefficient
The Pearson correlation coefficient is commonly represented by r and ranges from −1 to +1.
r = +1
Perfect positive linear correlation.
r = 0
No linear correlation.
r = −1
Perfect negative linear correlation.
0 < r < 1
Positive linear association.
Positive Correlation
When one variable tends to increase as the other increases, the variables have positive correlation.
Negative Correlation
When one variable tends to increase as the other decreases, the variables have negative correlation.
11. ANOVA
ANOVA stands for Analysis of Variance. It is a statistical method used to test whether the means of multiple groups differ significantly.
Basic Idea
Variation
Variation
Hypotheses
Null Hypothesis (H₀): The group means are equal.
Alternative Hypothesis (H₁): At least one group mean differs from the others.
F-Statistic
Applications of ANOVA
- Comparing average performance of different groups.
- Comparing different teaching methods.
- Comparing treatment groups in experiments.
- Comparing average sales across regions.
- Analyzing the effect of different factors.
12. EDA Techniques – Quick Comparison
| Technique | Purpose |
|---|---|
| Mean | Measures the average value of observations. |
| Standard Deviation | Measures the spread or variability of data. |
| Skewness | Measures asymmetry of a distribution. |
| Kurtosis | Describes tail behavior and distribution shape. |
| Box Plot | Shows quartiles, median, spread and possible outliers. |
| Pivot Table | Summarizes and aggregates data by categories. |
| Heat Map | Represents values using color intensity. |
| Correlation | Measures association between variables. |
| ANOVA | Tests differences among multiple group means. |
13. Unit III Quick Revision
EDA: Exploration and understanding of data using statistics and visualization.
Mean: Average value of observations.
Standard Deviation: Measure of data dispersion around the mean.
Skewness: Measure of distribution asymmetry.
Kurtosis: Describes tail behavior and distribution shape.
Box Plot: Shows five-number summary, spread and possible outliers.
Pivot Table: Summarizes data through grouping and aggregation.
Heat Map: Uses color intensity to represent numerical values.
Correlation: Measures strength and direction of association between variables.
ANOVA: Tests whether means of multiple groups differ significantly.
🎯 RGPV Exam-Oriented Important Questions
Important Questions – Unit III
💡 RGPV Exam Tip
Start with a definition, write the formula where applicable, draw a simple diagram/table and explain the concept with an example.
For 14 Marks:Write introduction + definition + formula + diagram + detailed explanation + example + applications/advantages + short conclusion.