Model Development
This unit introduces regression-based model development, visualization techniques for model evaluation, polynomial regression, pipelines, in-sample evaluation measures, prediction and decision making.
1. Introduction to Model Development
Model development is the process of selecting a suitable mathematical or statistical model, training it using available data, evaluating its performance and using it to make predictions.
General Model Development Process
2. Simple Linear Regression
Simple Linear Regression is used to model the relationship between one independent variable and one dependent variable.
Regression Equation
Where:
- y = predicted dependent variable
- x = independent variable
- b₀ = intercept
- b₁ = slope or regression coefficient
Example
Suppose we want to predict a student's marks based on study hours. Study hours are the independent variable and marks are the dependent variable.
(X)
Model
(Y)
Applications
- Sales prediction.
- Price prediction.
- Demand forecasting.
- Performance prediction.
- Trend analysis.
3. Multiple Linear Regression
Multiple Linear Regression is an extension of simple linear regression in which two or more independent variables are used to predict a dependent variable.
Here, x₁, x₂, ..., xₙ are independent variables and y is the dependent variable.
Example
House price may depend on:
- Area of house
- Number of bedrooms
- Location
- Age of building
Simple vs Multiple Regression
| Simple Regression | Multiple Regression |
|---|---|
| Uses one independent variable. | Uses two or more independent variables. |
| y = b₀ + b₁x | y = b₀ + b₁x₁ + b₂x₂ + ... + bₙxₙ |
| Simpler model. | Can represent more complex relationships. |
4. Model Evaluation Using Visualization
Visualization helps understand whether a model represents the data appropriately and whether errors or unusual patterns exist.
Common Visualization Techniques
📈 Regression Plot
Shows the observed data points along with the fitted regression line.
📊 Residual Plot
Shows the difference between observed and predicted values.
📉 Distribution Plot
Compares the distribution of actual and predicted values.
📋 Prediction Plot
Helps compare predicted values with actual observations.
A numerical evaluation metric alone may not reveal unusual patterns. Visualization can expose systematic errors, outliers, non-linearity and unequal variance.
5. Residual Plot
A residual is the difference between an observed value and the corresponding predicted value.
Good Residual Plot
For a suitable linear model, residuals should generally appear randomly scattered around zero without a clear systematic pattern.
What Can a Residual Plot Reveal?
- Non-linearity.
- Outliers.
- Unequal variance.
- Systematic prediction errors.
- Potential model inadequacy.
6. Distribution Plot
A distribution plot shows how values are distributed across different ranges.
It can be used to compare the distribution of actual values with predicted values.
Uses
- Understand the shape of data.
- Compare actual and predicted distributions.
- Identify concentration of observations.
- Identify possible outliers.
- Understand spread and central tendency.
7. Polynomial Regression
Polynomial Regression is used when the relationship between the independent and dependent variables is better represented by a curved relationship rather than a straight line.
Example
A simple linear model may not adequately represent a curved relationship between variables. Polynomial regression introduces higher-order terms such as x² and x³.
Model
Advantages
- Can model curved relationships.
- More flexible than simple linear regression.
- Useful when the relationship is non-linear.
Limitation
Using a polynomial degree that is too high can make the model unnecessarily complex and may cause poor generalization.
8. Polynomial Regression and Pipelines
A pipeline is a sequence of data processing and modeling steps that are executed in an organized manner.
Typical Pipeline
Transformation
Features
Benefits of Pipelines
- Organizes the modeling workflow.
- Reduces repetitive code.
- Ensures consistent preprocessing.
- Helps avoid mistakes in the modeling process.
- Makes experiments easier to reproduce.
9. Measures for In-Sample Evaluation
In-sample evaluation measures how well a model fits the same data that was used to train the model.
Important Measures
Mean Squared Error (MSE)
Measures the average squared difference between actual and predicted values.
Root Mean Squared Error (RMSE)
Square root of MSE. It is expressed in the same units as the target variable.
Mean Absolute Error (MAE)
Measures the average absolute difference between actual and predicted values.
R² Score
Indicates the proportion of variation in the dependent variable explained by the model.
Mean Squared Error
Root Mean Squared Error
Mean Absolute Error
R² Score
10. Prediction
Prediction is the process of using a trained model to estimate the value of an unknown or future observation.
Prediction Process
Data
Output
Examples
- Predicting house prices.
- Predicting sales.
- Predicting product demand.
- Predicting student performance.
- Predicting energy consumption.
11. Decision Making
Data-driven decision making uses predictions, statistical analysis and available evidence to support decisions.
Data Science Decision Process
Important Factors
- Accuracy and reliability of the model.
- Quality of the available data.
- Business or application requirements.
- Cost of incorrect decisions.
- Interpretability of the results.
- Relevant constraints and risks.
12. Regression Techniques – Quick Comparison
| Technique | Independent Variables | Relationship |
|---|---|---|
| Simple Linear Regression | One | Linear |
| Multiple Linear Regression | Two or More | Linear |
| Polynomial Regression | One or More | Curved / Polynomial |
13. Unit IV Quick Revision
Simple Regression: Uses one independent variable to predict a dependent variable.
Multiple Regression: Uses two or more independent variables.
Residual: Actual value − Predicted value.
Residual Plot: Used to inspect patterns in prediction errors.
Distribution Plot: Used to understand and compare data distributions.
Polynomial Regression: Models curved relationships using polynomial terms.
Pipeline: Combines multiple preprocessing and modeling steps into one organized workflow.
In-Sample Evaluation: Measures model performance on the data used for fitting.
Prediction: Estimation of an output for new or unknown input.
Decision Making: Using data, analysis and predictions to support appropriate actions.
🎯 RGPV Exam-Oriented Important Questions
Important Questions – Unit IV
💡 RGPV Exam Tip
Write the definition, equation, diagram/flowchart and explain the major points with one suitable example.
For 14 Marks:Write introduction + definition + mathematical equation + workflow diagram + detailed explanation + example + advantages/limitations + evaluation measures + conclusion.