CD404 – Introduction to Data Science

CSE – Data Science / Data Science | IV Semester

UNIT – IV

Model Development

This unit introduces regression-based model development, visualization techniques for model evaluation, polynomial regression, pipelines, in-sample evaluation measures, prediction and decision making.

1. Introduction to Model Development

Model development is the process of selecting a suitable mathematical or statistical model, training it using available data, evaluating its performance and using it to make predictions.

Definition:
Model Development is the process of building a mathematical or statistical representation of a real-world relationship using data for prediction, analysis or decision making.

General Model Development Process

Collect Data
→
Prepare Data
→
Select Model
→
Train Model
→
Evaluate
→
Predict

2. Simple Linear Regression

Simple Linear Regression is used to model the relationship between one independent variable and one dependent variable.

Definition:
Simple Linear Regression predicts the value of a dependent variable using a single independent variable by fitting a straight line through the data.

Regression Equation

y = b₀ + b₁x

Where:

Example

Suppose we want to predict a student's marks based on study hours. Study hours are the independent variable and marks are the dependent variable.

Study Hours
(X)
→
Regression
Model
→
Predicted Marks
(Y)

Applications

3. Multiple Linear Regression

Multiple Linear Regression is an extension of simple linear regression in which two or more independent variables are used to predict a dependent variable.

y = b₀ + b₁x₁ + b₂x₂ + ... + bₙxₙ

Here, x₁, x₂, ..., xₙ are independent variables and y is the dependent variable.

Example

House price may depend on:

Area
+
Bedrooms
+
Location
+
Age
→
House Price

Simple vs Multiple Regression

Simple Regression Multiple Regression
Uses one independent variable. Uses two or more independent variables.
y = b₀ + b₁x y = b₀ + b₁x₁ + b₂x₂ + ... + bₙxₙ
Simpler model. Can represent more complex relationships.

4. Model Evaluation Using Visualization

Visualization helps understand whether a model represents the data appropriately and whether errors or unusual patterns exist.

Common Visualization Techniques

📈 Regression Plot

Shows the observed data points along with the fitted regression line.

📊 Residual Plot

Shows the difference between observed and predicted values.

📉 Distribution Plot

Compares the distribution of actual and predicted values.

📋 Prediction Plot

Helps compare predicted values with actual observations.

Why visualization?

A numerical evaluation metric alone may not reveal unusual patterns. Visualization can expose systematic errors, outliers, non-linearity and unequal variance.

5. Residual Plot

A residual is the difference between an observed value and the corresponding predicted value.

Residual = Actual Value − Predicted Value
Residual Plot:
A residual plot displays residuals against predicted values or an explanatory variable to help identify patterns in model errors.

Good Residual Plot

For a suitable linear model, residuals should generally appear randomly scattered around zero without a clear systematic pattern.

Actual Value
−
Predicted Value
=
Residual

What Can a Residual Plot Reveal?

6. Distribution Plot

A distribution plot shows how values are distributed across different ranges.

It can be used to compare the distribution of actual values with predicted values.

Uses

Exam Point: Distribution plots are useful for visually comparing the behavior of actual and predicted data and identifying major differences between them.

7. Polynomial Regression

Polynomial Regression is used when the relationship between the independent and dependent variables is better represented by a curved relationship rather than a straight line.

y = b₀ + b₁x + b₂x² + b₃x³ + ... + bₙxⁿ

Example

A simple linear model may not adequately represent a curved relationship between variables. Polynomial regression introduces higher-order terms such as x² and x³.

Input X
→
X, X², X³
→
Polynomial
Model
→
Prediction

Advantages

Limitation

Using a polynomial degree that is too high can make the model unnecessarily complex and may cause poor generalization.

8. Polynomial Regression and Pipelines

A pipeline is a sequence of data processing and modeling steps that are executed in an organized manner.

Definition:
A pipeline combines multiple processing steps and a model into a single workflow so that the same sequence can be applied consistently.

Typical Pipeline

Raw Data
→
Feature
Transformation
→
Polynomial
Features
→
Regression
→
Prediction

Benefits of Pipelines

9. Measures for In-Sample Evaluation

In-sample evaluation measures how well a model fits the same data that was used to train the model.

Important Measures

Mean Squared Error (MSE)

Measures the average squared difference between actual and predicted values.

Root Mean Squared Error (RMSE)

Square root of MSE. It is expressed in the same units as the target variable.

Mean Absolute Error (MAE)

Measures the average absolute difference between actual and predicted values.

R² Score

Indicates the proportion of variation in the dependent variable explained by the model.

Mean Squared Error

MSE = (1/n) Σ(yᵢ − ŷᵢ)²

Root Mean Squared Error

RMSE = √MSE

Mean Absolute Error

MAE = (1/n) Σ|yᵢ − ŷᵢ|

R² Score

R² = 1 − (SSres / SStot)
General Interpretation: For error measures such as MAE, MSE and RMSE, lower values generally indicate smaller prediction errors. For R², a value closer to 1 indicates that the model explains a larger proportion of the variation in the target for the evaluated data.

10. Prediction

Prediction is the process of using a trained model to estimate the value of an unknown or future observation.

Prediction Process

Historical
Data
→
Train Model
→
New Input
→
Model
→
Predicted
Output

Examples

11. Decision Making

Data-driven decision making uses predictions, statistical analysis and available evidence to support decisions.

Definition:
Decision making is the process of selecting an appropriate action based on information, analysis, predictions and relevant constraints.

Data Science Decision Process

Data
→
Analysis
→
Prediction
→
Interpretation
→
Decision
→
Action

Important Factors

12. Regression Techniques – Quick Comparison

Technique Independent Variables Relationship
Simple Linear Regression One Linear
Multiple Linear Regression Two or More Linear
Polynomial Regression One or More Curved / Polynomial
📥 Download Handwritten Notes

13. Unit IV Quick Revision

Simple Regression: Uses one independent variable to predict a dependent variable.

Multiple Regression: Uses two or more independent variables.

Residual: Actual value − Predicted value.

Residual Plot: Used to inspect patterns in prediction errors.

Distribution Plot: Used to understand and compare data distributions.

Polynomial Regression: Models curved relationships using polynomial terms.

Pipeline: Combines multiple preprocessing and modeling steps into one organized workflow.

In-Sample Evaluation: Measures model performance on the data used for fitting.

Prediction: Estimation of an output for new or unknown input.

Decision Making: Using data, analysis and predictions to support appropriate actions.

🎯 RGPV Exam-Oriented Important Questions

Important Questions – Unit IV

1. Explain Simple Linear Regression with its equation and example.
2. Explain Multiple Linear Regression with a suitable example.
3. Differentiate between Simple and Multiple Regression.
4. Explain model evaluation using visualization techniques.
5. What is a Residual Plot? Explain its importance.
6. Explain Distribution Plot and its applications in model evaluation.
7. What is Polynomial Regression? Explain with an equation.
8. Explain Polynomial Regression using Pipelines.
9. Explain different measures used for In-Sample Evaluation.
10. Explain MSE, RMSE, MAE and R² score.
11. Explain Prediction and Decision Making in Data Science.
12. Explain the complete Model Development process with a diagram.

💡 RGPV Exam Tip

For 7 Marks:

Write the definition, equation, diagram/flowchart and explain the major points with one suitable example.

For 14 Marks:

Write introduction + definition + mathematical equation + workflow diagram + detailed explanation + example + advantages/limitations + evaluation measures + conclusion.