Model Evaluation
This unit explains how to evaluate the performance of a Data Science model on unseen data. It covers generalization error, out-of-sample metrics, cross validation, overfitting, underfitting, model selection, Ridge Regression and Grid Search.
1. Generalization Error
A machine learning or statistical model should perform well not only on the data used for training but also on new, unseen data. The difference between the model's expected performance on unseen data and its performance on training data is related to generalization error.
Training Error vs Generalization Error
| Training Error | Generalization Error |
|---|---|
| Error measured using the data used to train the model. | Error measured on new or unseen data. |
| May be very small for a highly complex model. | Shows how well the model is expected to perform on new data. |
Importance
- Helps measure real-world model performance.
- Helps identify overfitting.
- Supports model selection.
- Encourages evaluation on data not used for fitting.
2. Out-of-Sample Evaluation Metrics
Out-of-sample evaluation measures model performance on data that was not used to fit the model.
Common Regression Metrics
Mean Absolute Error (MAE)
Measures the average absolute difference between actual and predicted values.
Mean Squared Error (MSE)
Measures the average squared prediction error.
Root Mean Squared Error (RMSE)
Square root of MSE and expressed in the target variable's units.
R² Score
Measures the proportion of variation explained by the model on the evaluated data.
3. Cross Validation
Cross Validation is a resampling technique used to estimate how well a model will perform on unseen data.
K-Fold Cross Validation
In K-Fold Cross Validation, the dataset is divided into K approximately equal-sized folds. The model is trained using K−1 folds and evaluated on the remaining fold. This process is repeated K times.
Basic Process
Folds
Remaining Fold
Advantages
- Uses available data efficiently.
- Provides a more reliable estimate than a single split in many cases.
- Helps compare different models.
- Can help identify overfitting.
4. Overfitting
Overfitting occurs when a model learns the training data too closely, including noise or random variations, and therefore performs poorly on unseen data.
Characteristics
- Very low training error.
- Relatively high validation or test error.
- Model is unnecessarily complex.
- Model may capture noise in the training dataset.
Causes
- Excessively complex model.
- Too many features relative to available data.
- Insufficient training data.
- Training for too long in some learning algorithms.
- Noisy training data.
Ways to Reduce Overfitting
- Use cross validation.
- Reduce model complexity.
- Remove irrelevant features.
- Use regularization.
- Collect more representative training data where feasible.
- Use suitable model selection techniques.
5. Underfitting
Underfitting occurs when a model is too simple to capture the important patterns in the data.
Causes
- Model is too simple.
- Important features may be missing.
- Insufficient training.
- Incorrect model assumptions.
Ways to Reduce Underfitting
- Use a more suitable model.
- Add useful features.
- Increase model complexity when justified.
- Train the model adequately.
6. Overfitting vs Underfitting
| Feature | Overfitting | Underfitting |
|---|---|---|
| Model Complexity | Too complex | Too simple |
| Training Error | Usually very low | Usually high |
| Test Error | Usually high | Usually high |
| Learning | Learns noise and details | Fails to learn important patterns |
| Solution | Regularization, simpler model, better validation | More suitable/complex model, better features |
7. Model Selection
Model selection is the process of choosing the most suitable model from multiple candidate models based on appropriate evaluation criteria.
Important Factors
- Predictive performance.
- Generalization ability.
- Model complexity.
- Computational requirements.
- Interpretability.
- Application requirements.
Model Selection Process
Models
Validation
Metrics
Model
Evaluation
8. Prediction Using Ridge Regression
Ridge Regression is a regularized version of linear regression. It adds a penalty based on the squared magnitude of model coefficients.
Ridge Objective
Here, α controls the strength of regularization and the coefficient penalty is applied to the model parameters (excluding the intercept in the usual formulation).
Why Ridge Regression?
- Helps control model complexity.
- Can reduce the effect of large coefficients.
- Useful when predictors are correlated.
- Can improve generalization in appropriate situations.
Ridge Prediction Process
Data
Model
9. Testing Multiple Parameters Using Grid Search
Grid Search is a systematic method for testing combinations of selected hyperparameter values and identifying a combination that performs well according to a chosen validation metric.
Example
Suppose a Ridge Regression model is tested with the following values of α:
Grid Search evaluates the specified values and identifies a suitable value based on cross-validation performance.
Grid Search Process
Parameters
Parameter Grid
Validation
Combinations
Parameters
Advantages
- Systematic hyperparameter testing.
- Easy to understand and implement.
- Can be combined with cross validation.
- Useful for model tuning.
Limitation
If many parameters and many possible values are included, the number of combinations can become large and computationally expensive.
10. Complete Model Evaluation Pipeline
Split
Training
Validation
Tuning
Evaluation
A good evaluation workflow separates model development from final testing and uses validation procedures to make informed model-selection decisions.
11. Model Evaluation Metrics – Quick Comparison
| Metric / Technique | Purpose | General Interpretation |
|---|---|---|
| MAE | Average absolute prediction error. | Lower is generally better. |
| MSE | Average squared prediction error. | Lower is generally better. |
| RMSE | Square root of MSE. | Lower is generally better. |
| R² | Proportion of target variation explained by the model. | Higher is generally better for comparable evaluation settings. |
| Cross Validation | Estimates performance across multiple train/validation splits. | Helps assess generalization. |
| Grid Search | Tests predefined hyperparameter combinations. | Selects the best combination according to the chosen validation score. |
12. Unit V Quick Revision
Generalization Error: Error of a model on unseen data.
Out-of-Sample Evaluation: Evaluating a model using data not used during model fitting.
Cross Validation: Repeatedly trains and validates a model using different portions of the available data.
Overfitting: Model learns training data too closely and performs poorly on unseen data.
Underfitting: Model is too simple to capture important patterns.
Model Selection: Choosing a suitable model based on performance and generalization.
Ridge Regression: Linear regression with L2 regularization.
Grid Search: Systematic testing of predefined hyperparameter combinations.
🎯 RGPV Exam-Oriented Important Questions
Important Questions – Unit V
💡 RGPV Exam Tip
Write the definition, explain the concept in clear points, include a suitable flowchart and give one example.
For 14 Marks:Write introduction + definition + diagram + detailed explanation + formula/metrics + advantages/limitations + example + conclusion.
🎓 CD404 Complete Unit Revision
Unit I
Introduction to Data Science
Unit II
Data Collection and Data Pre-Processing
Unit III
Exploratory Data Analytics
Unit IV
Model Development
Unit V
Model Evaluation