Lesson 53 of 60 · python
Model Selection – Cross‑Validation
Duration: 20 minutes
Cross‑Validation (CV)
Cross‑validation provides reliable estimates of model performance on unseen data.
k‑fold cross‑validation
- Split data into
kequal folds. - Train on
k‑1folds, validate on the held‑out fold. - Repeat
ktimes, each fold becomes validation once. - Aggregate scores (mean, std).
Using scikit‑learn cross_val_score
from sklearn.model_selection import cross_val_score
from sklearn.ensemble import GradientBoostingClassifier
model = GradientBoostingClassifier(random_state=42)
cv_scores = cross_val_score(model, X, y, cv=5, scoring='accuracy')
print('CV Accuracy scores:', cv_scores)
print('Mean accuracy:', cv_scores.mean())
Stratified k‑fold (preserves class distribution)
from sklearn.model_selection import StratifiedKFold
skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
cv_scores = cross_val_score(model, X, y, cv=skf, scoring='f1')
Nested cross‑validation (hyperparameter tuning + performance estimate)
from sklearn.model_selection import GridSearchCV, cross_val_score
param_grid = {'n_estimators': [100, 200], 'learning_rate': [0.01, 0.1]}
inner_cv = StratifiedKFold(n_splits=3, shuffle=True, random_state=42)
outer_cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
clf = GridSearchCV(GradientBoostingClassifier(random_state=42), param_grid, cv=inner_cv, scoring='roc_auc')
nested_score = cross_val_score(clf, X, y, cv=outer_cv, scoring='roc_auc')
print('Nested CV ROC‑AUC:', nested_score.mean())
TimeSeriesSplit (for temporal data)
from sklearn.model_selection import TimeSeriesSplit
tscv = TimeSeriesSplit(n_splits=5)
cv_scores = cross_val_score(model, X, y, cv=tscv, scoring='neg_mean_squared_error')
Best practices
- Shuffle data before splitting unless time series.
- Use stratification for classification.
- Keep a final hold‑out test set untouched for the last evaluation.