CD404 – Introduction to Data Science

CSE – Data Science / Data Science | IV Semester

UNIT – V

Model Evaluation

This unit explains how to evaluate the performance of a Data Science model on unseen data. It covers generalization error, out-of-sample metrics, cross validation, overfitting, underfitting, model selection, Ridge Regression and Grid Search.

1. Generalization Error

A machine learning or statistical model should perform well not only on the data used for training but also on new, unseen data. The difference between the model's expected performance on unseen data and its performance on training data is related to generalization error.

Definition:
Generalization error is the error made by a trained model when it is applied to previously unseen data from the same underlying problem distribution.

Training Error vs Generalization Error

Training Error Generalization Error
Error measured using the data used to train the model. Error measured on new or unseen data.
May be very small for a highly complex model. Shows how well the model is expected to perform on new data.

Importance

2. Out-of-Sample Evaluation Metrics

Out-of-sample evaluation measures model performance on data that was not used to fit the model.

Definition:
Out-of-sample evaluation is the process of evaluating a model using observations that were not used during model training.

Common Regression Metrics

Mean Absolute Error (MAE)

Measures the average absolute difference between actual and predicted values.

Mean Squared Error (MSE)

Measures the average squared prediction error.

Root Mean Squared Error (RMSE)

Square root of MSE and expressed in the target variable's units.

R² Score

Measures the proportion of variation explained by the model on the evaluated data.

MAE = (1/n) Σ |yᵢ − ŷᵢ|
MSE = (1/n) Σ (yᵢ − ŷᵢ)²
RMSE = √MSE
General Rule: For MAE, MSE and RMSE, lower values generally indicate smaller prediction errors. For R², a value closer to 1 generally indicates that more variation is explained by the model on the evaluated dataset.

3. Cross Validation

Cross Validation is a resampling technique used to estimate how well a model will perform on unseen data.

Definition:
Cross Validation divides the available data into multiple parts and repeatedly trains and evaluates the model using different parts for training and validation.

K-Fold Cross Validation

In K-Fold Cross Validation, the dataset is divided into K approximately equal-sized folds. The model is trained using K−1 folds and evaluated on the remaining fold. This process is repeated K times.

Dataset
→
Fold 1
+
Fold 2
+
...
+
Fold K

Basic Process

Divide Data
→
Train on K−1
Folds
→
Validate on
Remaining Fold
→
Repeat K Times
→
Average Score

Advantages

4. Overfitting

Overfitting occurs when a model learns the training data too closely, including noise or random variations, and therefore performs poorly on unseen data.

Definition:
Overfitting is a situation in which a model performs very well on training data but has poor performance on unseen or validation data.

Characteristics

Causes

Ways to Reduce Overfitting

Remember: Overfitting generally means: Excellent training performance + Poor unseen-data performance.

5. Underfitting

Underfitting occurs when a model is too simple to capture the important patterns in the data.

Definition:
Underfitting is a situation in which a model performs poorly on both training data and unseen data because it fails to capture the underlying relationship adequately.

Causes

Ways to Reduce Underfitting

6. Overfitting vs Underfitting

Feature Overfitting Underfitting
Model Complexity Too complex Too simple
Training Error Usually very low Usually high
Test Error Usually high Usually high
Learning Learns noise and details Fails to learn important patterns
Solution Regularization, simpler model, better validation More suitable/complex model, better features

7. Model Selection

Model selection is the process of choosing the most suitable model from multiple candidate models based on appropriate evaluation criteria.

Definition:
Model Selection is the process of comparing candidate models and choosing one that provides an appropriate balance between performance, complexity and generalization.

Important Factors

Model Selection Process

Candidate
Models
→
Cross
Validation
→
Compare
Metrics
→
Select
Model
→
Final
Evaluation

8. Prediction Using Ridge Regression

Ridge Regression is a regularized version of linear regression. It adds a penalty based on the squared magnitude of model coefficients.

Definition:
Ridge Regression is a linear regression technique that adds L2 regularization to the objective function to discourage excessively large coefficients.

Ridge Objective

Minimize: Σ(yᵢ − ŷᵢ)² + α Σβⱼ²

Here, α controls the strength of regularization and the coefficient penalty is applied to the model parameters (excluding the intercept in the usual formulation).

Why Ridge Regression?

Ridge Prediction Process

Training
Data
→
Choose α
→
Train Ridge
Model
→
Evaluate
→
Predict
Important: The regularization parameter α should be selected using a suitable validation strategy rather than simply choosing it because it produces the lowest training error.

9. Testing Multiple Parameters Using Grid Search

Grid Search is a systematic method for testing combinations of selected hyperparameter values and identifying a combination that performs well according to a chosen validation metric.

Definition:
Grid Search evaluates a predefined set of hyperparameter combinations, usually using cross validation, and selects the combination with the best validation performance according to the chosen scoring criterion.

Example

Suppose a Ridge Regression model is tested with the following values of α:

α = 0.01
α = 0.1
α = 1
α = 10
α = 100

Grid Search evaluates the specified values and identifies a suitable value based on cross-validation performance.

Grid Search Process

Define
Parameters
→
Create
Parameter Grid
→
Cross
Validation
→
Evaluate
Combinations
→
Select Best
Parameters

Advantages

Limitation

If many parameters and many possible values are included, the number of combinations can become large and computationally expensive.

10. Complete Model Evaluation Pipeline

Dataset
→
Train / Test
Split
→
Model
Training
→
Cross
Validation
→
Hyperparameter
Tuning
→
Final
Evaluation

A good evaluation workflow separates model development from final testing and uses validation procedures to make informed model-selection decisions.

11. Model Evaluation Metrics – Quick Comparison

Metric / Technique Purpose General Interpretation
MAE Average absolute prediction error. Lower is generally better.
MSE Average squared prediction error. Lower is generally better.
RMSE Square root of MSE. Lower is generally better.
R² Proportion of target variation explained by the model. Higher is generally better for comparable evaluation settings.
Cross Validation Estimates performance across multiple train/validation splits. Helps assess generalization.
Grid Search Tests predefined hyperparameter combinations. Selects the best combination according to the chosen validation score.
📥 Download Handwritten Notes

12. Unit V Quick Revision

Generalization Error: Error of a model on unseen data.

Out-of-Sample Evaluation: Evaluating a model using data not used during model fitting.

Cross Validation: Repeatedly trains and validates a model using different portions of the available data.

Overfitting: Model learns training data too closely and performs poorly on unseen data.

Underfitting: Model is too simple to capture important patterns.

Model Selection: Choosing a suitable model based on performance and generalization.

Ridge Regression: Linear regression with L2 regularization.

Grid Search: Systematic testing of predefined hyperparameter combinations.

🎯 RGPV Exam-Oriented Important Questions

Important Questions – Unit V

1. What is Generalization Error? Explain its importance.
2. Explain Out-of-Sample Evaluation Metrics.
3. Explain Cross Validation and K-Fold Cross Validation with a diagram.
4. What is Overfitting? Explain its causes and methods to reduce it.
5. What is Underfitting? Explain its causes and remedies.
6. Differentiate between Overfitting and Underfitting.
7. Explain Model Selection in Data Science.
8. Explain Ridge Regression and its importance in model evaluation.
9. Explain prediction using Ridge Regression.
10. What is Grid Search? Explain testing multiple parameters.
11. Explain the relationship between Cross Validation and Grid Search.
12. Explain the complete Model Evaluation process with a suitable diagram.

💡 RGPV Exam Tip

For 7 Marks:

Write the definition, explain the concept in clear points, include a suitable flowchart and give one example.

For 14 Marks:

Write introduction + definition + diagram + detailed explanation + formula/metrics + advantages/limitations + example + conclusion.

🎓 CD404 Complete Unit Revision

Unit I

Introduction to Data Science

Unit II

Data Collection and Data Pre-Processing

Unit III

Exploratory Data Analytics

Unit IV

Model Development

Unit V

Model Evaluation