CD404 – Introduction to Data Science

CSE – Data Science / Data Science | IV Semester

UNIT – II

Data Collection and Data Pre-Processing

This unit explains how data is collected from different sources and prepared for analysis. It covers data collection strategies, data cleaning, integration, transformation, reduction and discretization.

1. Data Collection Strategies

Data collection is the process of gathering relevant information from different sources for analysis and decision making. The quality of collected data directly affects the quality of Data Science results.

Definition:
Data Collection is the systematic process of obtaining data from relevant sources according to the objectives of a Data Science project.

Common Data Sources

📋 Surveys

Data is collected by asking questions to individuals or groups. Surveys may be conducted online or offline.

🗄️ Databases

Existing organizational databases can provide structured transactional and operational data.

🌐 Web Data

Information can be obtained from websites and publicly available online sources.

🔌 APIs

Application Programming Interfaces allow applications to retrieve data from other software systems.

📡 Sensors

Sensors generate data continuously from physical devices, machines and IoT systems.

📁 Files

CSV, Excel, JSON and other files can be used as sources of structured or semi-structured data.

Primary and Secondary Data

Primary Data Secondary Data
Collected directly by the researcher or organization for a specific purpose. Already collected by another person, organization or system and reused for analysis.
Examples: surveys, interviews, experiments and observations. Examples: government reports, existing databases, published datasets and research reports.

2. Data Pre-Processing Overview

Raw data is usually incomplete, inconsistent, noisy or stored in different formats. Therefore, it must be processed before being used for analysis or machine learning.

Definition:
Data Pre-Processing is the process of converting raw data into a clean, consistent and suitable form for analysis.

Need for Data Pre-Processing

Major Steps

Raw Data
→
Cleaning
→
Integration
→
Transformation
→
Reduction
→
Discretization
→
Prepared Data

3. Data Cleaning

Data cleaning is the process of detecting and correcting inaccurate, incomplete, inconsistent or duplicate data. It is one of the most important stages of data pre-processing.

Problems Addressed by Data Cleaning

Missing Values

Some records may not contain values for one or more attributes.

Duplicate Data

The same record may appear multiple times in a dataset.

Inconsistent Data

Different representations may be used for the same value.

Outliers

Extremely unusual observations may affect analysis and models.

Noise

Random errors or unwanted variations may be present in the data.

Invalid Values

Some values may violate the expected range or data type.

Handling Missing Values

Common methods for handling missing data include:

Example: If the age values are 20, 22, 24, missing, 26, the missing value may be replaced using an appropriate statistical method depending on the dataset.

Duplicate Removal

Duplicate records are identified and unnecessary copies are removed so that the same observation does not receive extra importance.

4. Data Integration

Data integration combines data from multiple sources into a single consistent dataset.

Definition:
Data Integration is the process of combining data from different sources to provide a unified view of information.

Example

Consider an organization having:

These sources can be integrated to create a unified customer dataset for analysis.

Challenges in Data Integration

5. Data Transformation

Data transformation converts data into a suitable format or scale for analysis and modeling.

Common Transformation Techniques

Normalization

Scales numerical values into a common range so that features with different scales can be compared.

Aggregation

Combines detailed data into summarized information. For example, daily sales can be aggregated into monthly sales.

Encoding

Converts categorical information into numerical representation when required by an algorithm.

Generalization

Replaces low-level data with higher-level concepts.

Normalization

Normalization is commonly used when numerical attributes have very different ranges.

Example: Suppose one feature represents age from 0–100 while another represents annual income from 0–1,000,000. Transformation can bring them to comparable scales.

6. Data Reduction

Data reduction reduces the size or complexity of a dataset while attempting to preserve its important information.

Definition:
Data Reduction is the process of obtaining a smaller representation of data while maintaining its essential analytical characteristics.

Need for Data Reduction

Major Data Reduction Techniques

Dimensionality Reduction

Reduces the number of attributes or features in a dataset.

Numerosity Reduction

Uses a smaller representation of data instead of storing every original observation.

Data Cube Aggregation

Aggregates detailed data into summarized levels.

Feature Selection

Selects the most useful features and removes irrelevant or redundant attributes.

7. Data Discretization

Data discretization converts continuous numerical values into a limited number of intervals or categories.

Definition:
Data Discretization is the process of transforming continuous data into discrete intervals or categories.

Example

Suppose age is represented as:

18, 21, 25, 32, 41, 56, 67, 72

It can be converted into categories such as:

Young
18–30
→
Adult
31–50
→
Senior
51+

Common Discretization Methods

Advantages

8. Data Pre-Processing Techniques – Quick Comparison

Technique Main Purpose Example
Data Cleaning Remove or correct inaccurate, incomplete and inconsistent data. Handling missing values.
Data Integration Combine data from multiple sources. Combining customer and sales databases.
Data Transformation Convert data into a suitable format or scale. Normalization and encoding.
Data Reduction Reduce data size or dimensionality. Feature selection.
Data Discretization Convert continuous values into intervals. Age → Young, Adult, Senior.
📥 Download Handwritten Notes

9. Unit II Quick Revision

Data Collection: Gathering relevant data from suitable sources.

Pre-Processing: Preparing raw data for analysis.

Cleaning: Handling missing, duplicate, noisy and inconsistent data.

Integration: Combining data from multiple sources.

Transformation: Changing data into a suitable format or scale.

Reduction: Reducing the size or dimensionality of data while preserving important information.

Discretization: Converting continuous values into discrete intervals.

🎯 RGPV Exam-Oriented Important Questions

Important Questions – Unit II

1. Explain different Data Collection Strategies.
2. What is Data Pre-Processing? Explain its need and major steps.
3. Explain Data Cleaning and various techniques for handling missing data.
4. Explain Data Integration with its challenges.
5. What is Data Transformation? Explain different transformation techniques.
6. Explain Data Reduction and its major techniques.
7. What is Data Discretization? Explain its methods with examples.
8. Differentiate between Data Cleaning, Integration, Transformation, Reduction and Discretization.
9. Explain the complete Data Pre-Processing pipeline with a diagram.

💡 RGPV Exam Tip

For 7 Marks:

Write the definition, explain the technique using 5–7 clear points, and add a simple example or flow diagram.

For 14 Marks:

Start with an introduction and definition, draw the pre-processing flowchart, explain each stage in detail, give examples, mention advantages/challenges and conclude with its importance.