CD404 – Introduction to Data Science

CSE – Data Science / Data Science | IV Semester

UNIT – III

Exploratory Data Analytics

Exploratory Data Analytics (EDA) is used to understand the characteristics, patterns, relationships and distributions present in a dataset before applying advanced models.

1. Introduction to Exploratory Data Analytics

Exploratory Data Analytics, commonly called EDA, is the process of examining and summarizing a dataset using statistical methods and graphical techniques.

Definition:
Exploratory Data Analytics is the systematic analysis of data using statistical summaries and visualizations to discover patterns, relationships, trends and unusual observations.

Objectives of EDA

Raw Dataset
→
Statistics
→
Visualization
→
Patterns
→
Insights

2. Descriptive Statistics

Descriptive statistics summarizes and describes the important features of a dataset using numerical measures and tables.

Main Categories

Measures of Central Tendency

Mean, median and mode describe the central or typical value of a dataset.

Measures of Dispersion

Range, variance and standard deviation describe how widely the observations are spread.

Measures of Shape

Skewness and kurtosis describe the shape of a distribution.

Graphical Measures

Box plots, histograms and other visualizations help understand the distribution of data.

3. Mean

The arithmetic mean is the average value of all observations in a dataset.

Mean = Sum of all observations / Number of observations

For observations x₁, x₂, ..., xₙ:

x̄ = (x₁ + x₂ + ... + xₙ) / n

Example

Consider the data: 10, 20, 30, 40, 50

Sum = 150 and number of observations = 5. Therefore, mean = 30.

Advantages

Limitation

Mean can be strongly affected by extreme values or outliers.

4. Standard Deviation

Standard deviation measures how much the observations deviate or spread around the mean.

Definition:
Standard deviation is a measure of dispersion that indicates the typical amount by which observations differ from their mean.
σ = √[ Σ(x − μ)² / N ]

For a sample, the denominator is commonly replaced by n − 1.

Interpretation

Exam Point: Standard deviation is one of the most important measures of dispersion used to understand variability in data.

5. Skewness

Skewness measures the degree of asymmetry of a probability distribution around its mean.

Types of Skewness

Positive Skewness

The distribution has a longer tail toward the right. The mean is generally greater than the median.

Negative Skewness

The distribution has a longer tail toward the left. The mean is generally less than the median.

Zero Skewness

The distribution is approximately symmetric around its center.

Negative
Skew
→
Symmetric
→
Positive
Skew
Remember: Positive skew → longer right tail.
Negative skew → longer left tail.

6. Kurtosis

Kurtosis describes the shape of a distribution, particularly the degree of tail heaviness and concentration around the center relative to a reference distribution.

Common Categories

Mesokurtic

A distribution with kurtosis characteristics similar to the normal distribution.

Leptokurtic

A distribution with relatively heavier tails and a sharper central concentration.

Platykurtic

A distribution with relatively lighter tails and a flatter central shape.

Exam Point: Skewness mainly describes asymmetry, while kurtosis describes tail behavior and distribution shape.

7. Box Plot

A Box Plot, also called a box-and-whisker plot, is a graphical method for displaying the distribution of numerical data using quartiles.

Five-Number Summary

Minimum
→
Q1
→
Median
→
Q3
→
Maximum

Important Terms

Uses of Box Plot:
  • Compare distributions.
  • Identify spread.
  • Identify median.
  • Detect possible outliers.
  • Understand skewness.

8. Pivot Table

A Pivot Table is a data summarization technique that organizes and aggregates data according to selected categories.

Definition:
A Pivot Table summarizes large datasets by grouping records according to selected fields and applying aggregation functions such as sum, count, average, minimum and maximum.

Example

Suppose a sales dataset contains:

A pivot table can summarize total sales by region and product.

Region Product A Product B Total
North ₹20,000 ₹15,000 ₹35,000
South ₹18,000 ₹22,000 ₹40,000
West ₹25,000 ₹17,000 ₹42,000

Advantages

9. Heat Map

A Heat Map is a graphical representation in which different values are represented using variations in color intensity.

Definition:
A heat map visualizes numerical values in a matrix or grid using color intensity, making high and low values easy to identify.

Example

Maths OS DBMS CN
Student A 90 72 85 78
Student B 65 88 74 92
Student C 80 70 95 83

Applications

10. Correlation Statistics

Correlation measures the degree and direction of association between two variables.

Definition:
Correlation is a statistical measure that indicates how strongly two variables are associated and whether they move in the same or opposite directions.

Correlation Coefficient

The Pearson correlation coefficient is commonly represented by r and ranges from −1 to +1.

−1 ≤ r ≤ +1

r = +1

Perfect positive linear correlation.

r = 0

No linear correlation.

r = −1

Perfect negative linear correlation.

0 < r < 1

Positive linear association.

Positive Correlation

When one variable tends to increase as the other increases, the variables have positive correlation.

Negative Correlation

When one variable tends to increase as the other decreases, the variables have negative correlation.

Important: Correlation indicates association, but correlation alone does not prove that one variable causes another.

11. ANOVA

ANOVA stands for Analysis of Variance. It is a statistical method used to test whether the means of multiple groups differ significantly.

Definition:
ANOVA is a hypothesis-testing technique used to compare the means of two or more groups by analyzing variation within groups and variation between groups.

Basic Idea

Total Variation
→
Between-Group
Variation
+
Within-Group
Variation

Hypotheses

Null Hypothesis (H₀): The group means are equal.

Alternative Hypothesis (H₁): At least one group mean differs from the others.

F-Statistic

F = Between-Group Variability / Within-Group Variability

Applications of ANOVA

Exam Point: ANOVA determines whether observed differences among group means are statistically significant rather than simply due to random variation.

12. EDA Techniques – Quick Comparison

Technique Purpose
Mean Measures the average value of observations.
Standard Deviation Measures the spread or variability of data.
Skewness Measures asymmetry of a distribution.
Kurtosis Describes tail behavior and distribution shape.
Box Plot Shows quartiles, median, spread and possible outliers.
Pivot Table Summarizes and aggregates data by categories.
Heat Map Represents values using color intensity.
Correlation Measures association between variables.
ANOVA Tests differences among multiple group means.
📥 Download Handwritten Notes

13. Unit III Quick Revision

EDA: Exploration and understanding of data using statistics and visualization.

Mean: Average value of observations.

Standard Deviation: Measure of data dispersion around the mean.

Skewness: Measure of distribution asymmetry.

Kurtosis: Describes tail behavior and distribution shape.

Box Plot: Shows five-number summary, spread and possible outliers.

Pivot Table: Summarizes data through grouping and aggregation.

Heat Map: Uses color intensity to represent numerical values.

Correlation: Measures strength and direction of association between variables.

ANOVA: Tests whether means of multiple groups differ significantly.

🎯 RGPV Exam-Oriented Important Questions

Important Questions – Unit III

1. What is Exploratory Data Analytics? Explain its objectives.
2. Explain descriptive statistics with suitable examples.
3. Define Mean and Standard Deviation. Explain their significance.
4. Explain Skewness and its different types with diagrams.
5. What is Kurtosis? Explain its different types.
6. Explain Box Plot and its components with a suitable diagram.
7. What is a Pivot Table? Explain its uses with an example.
8. Explain Heat Map and its applications in Data Analytics.
9. Define Correlation. Explain positive, negative and zero correlation.
10. Explain Pearson's correlation coefficient and its interpretation.
11. What is ANOVA? Explain its basic concept and hypotheses.
12. Differentiate between correlation and ANOVA.
13. Explain different EDA techniques used for understanding a dataset.

💡 RGPV Exam Tip

For 7 Marks:

Start with a definition, write the formula where applicable, draw a simple diagram/table and explain the concept with an example.

For 14 Marks:

Write introduction + definition + formula + diagram + detailed explanation + example + applications/advantages + short conclusion.