Data Collection and Data Pre-Processing
This unit explains how data is collected from different sources and prepared for analysis. It covers data collection strategies, data cleaning, integration, transformation, reduction and discretization.
1. Data Collection Strategies
Data collection is the process of gathering relevant information from different sources for analysis and decision making. The quality of collected data directly affects the quality of Data Science results.
Common Data Sources
📋 Surveys
Data is collected by asking questions to individuals or groups. Surveys may be conducted online or offline.
🗄️ Databases
Existing organizational databases can provide structured transactional and operational data.
🌐 Web Data
Information can be obtained from websites and publicly available online sources.
🔌 APIs
Application Programming Interfaces allow applications to retrieve data from other software systems.
📡 Sensors
Sensors generate data continuously from physical devices, machines and IoT systems.
📁 Files
CSV, Excel, JSON and other files can be used as sources of structured or semi-structured data.
Primary and Secondary Data
| Primary Data | Secondary Data |
|---|---|
| Collected directly by the researcher or organization for a specific purpose. | Already collected by another person, organization or system and reused for analysis. |
| Examples: surveys, interviews, experiments and observations. | Examples: government reports, existing databases, published datasets and research reports. |
2. Data Pre-Processing Overview
Raw data is usually incomplete, inconsistent, noisy or stored in different formats. Therefore, it must be processed before being used for analysis or machine learning.
Need for Data Pre-Processing
- Raw data may contain missing values.
- Data may contain duplicate records.
- Different sources may use different formats.
- Data may contain errors or inconsistent values.
- Some features may contain unnecessary information.
- Different numerical scales may affect analysis.
Major Steps
3. Data Cleaning
Data cleaning is the process of detecting and correcting inaccurate, incomplete, inconsistent or duplicate data. It is one of the most important stages of data pre-processing.
Problems Addressed by Data Cleaning
Missing Values
Some records may not contain values for one or more attributes.
Duplicate Data
The same record may appear multiple times in a dataset.
Inconsistent Data
Different representations may be used for the same value.
Outliers
Extremely unusual observations may affect analysis and models.
Noise
Random errors or unwanted variations may be present in the data.
Invalid Values
Some values may violate the expected range or data type.
Handling Missing Values
Common methods for handling missing data include:
- Deleting records containing missing values.
- Replacing missing values with the mean.
- Replacing missing values with the median.
- Replacing categorical missing values with the mode.
- Using a suitable prediction or imputation technique.
Duplicate Removal
Duplicate records are identified and unnecessary copies are removed so that the same observation does not receive extra importance.
4. Data Integration
Data integration combines data from multiple sources into a single consistent dataset.
Example
Consider an organization having:
- Customer information in one database.
- Sales information in another database.
- Website activity stored in log files.
These sources can be integrated to create a unified customer dataset for analysis.
Challenges in Data Integration
- Different data formats.
- Different naming conventions.
- Duplicate records.
- Different units of measurement.
- Conflicting values.
- Different database structures.
5. Data Transformation
Data transformation converts data into a suitable format or scale for analysis and modeling.
Common Transformation Techniques
Normalization
Scales numerical values into a common range so that features with different scales can be compared.
Aggregation
Combines detailed data into summarized information. For example, daily sales can be aggregated into monthly sales.
Encoding
Converts categorical information into numerical representation when required by an algorithm.
Generalization
Replaces low-level data with higher-level concepts.
Normalization
Normalization is commonly used when numerical attributes have very different ranges.
6. Data Reduction
Data reduction reduces the size or complexity of a dataset while attempting to preserve its important information.
Need for Data Reduction
- Reduces storage requirements.
- Reduces processing time.
- Improves computational efficiency.
- Makes visualization easier.
- Can improve model training efficiency.
Major Data Reduction Techniques
Dimensionality Reduction
Reduces the number of attributes or features in a dataset.
Numerosity Reduction
Uses a smaller representation of data instead of storing every original observation.
Data Cube Aggregation
Aggregates detailed data into summarized levels.
Feature Selection
Selects the most useful features and removes irrelevant or redundant attributes.
7. Data Discretization
Data discretization converts continuous numerical values into a limited number of intervals or categories.
Example
Suppose age is represented as:
It can be converted into categories such as:
18–30
31–50
51+
Common Discretization Methods
- Equal-Width Binning: The complete range is divided into intervals of equal width.
- Equal-Frequency Binning: Data is divided so that each interval contains approximately the same number of observations.
- Clustering-Based Discretization: Similar values are grouped using clustering techniques.
- Domain-Based Discretization: Intervals are created according to domain knowledge.
Advantages
- Simplifies continuous data.
- Can make patterns easier to understand.
- May reduce the effect of small variations.
- Can be useful for some classification techniques.
8. Data Pre-Processing Techniques – Quick Comparison
| Technique | Main Purpose | Example |
|---|---|---|
| Data Cleaning | Remove or correct inaccurate, incomplete and inconsistent data. | Handling missing values. |
| Data Integration | Combine data from multiple sources. | Combining customer and sales databases. |
| Data Transformation | Convert data into a suitable format or scale. | Normalization and encoding. |
| Data Reduction | Reduce data size or dimensionality. | Feature selection. |
| Data Discretization | Convert continuous values into intervals. | Age → Young, Adult, Senior. |
9. Unit II Quick Revision
Data Collection: Gathering relevant data from suitable sources.
Pre-Processing: Preparing raw data for analysis.
Cleaning: Handling missing, duplicate, noisy and inconsistent data.
Integration: Combining data from multiple sources.
Transformation: Changing data into a suitable format or scale.
Reduction: Reducing the size or dimensionality of data while preserving important information.
Discretization: Converting continuous values into discrete intervals.
🎯 RGPV Exam-Oriented Important Questions
Important Questions – Unit II
💡 RGPV Exam Tip
Write the definition, explain the technique using 5–7 clear points, and add a simple example or flow diagram.
For 14 Marks:Start with an introduction and definition, draw the pre-processing flowchart, explain each stage in detail, give examples, mention advantages/challenges and conclude with its importance.