Skip to main content
Brave Programmer Logo

BraveProgrammer

BraveProgrammer

HomeProjectsBlogsCoursesLessonsAbout

Site footer

BraveProgrammer

Free coding courses, practical tutorials, and real projects from BraveProgrammer. Learn web development with React, Next.js, and TypeScript.

Navigation

  • Home
  • Projects
  • Blogs
  • Courses

Resources

  • About
  • Lessons

© 2026 BraveProgrammer. All rights reserved.

  1. Courses
  2. /
  3. Master Data Science with Python

Lesson 53 of 60 · python

Model Selection – Cross‑Validation

Duration: 20 minutes

Cross‑Validation (CV)

Cross‑validation provides reliable estimates of model performance on unseen data.

k‑fold cross‑validation

  • Split data into k equal folds.
  • Train on k‑1 folds, validate on the held‑out fold.
  • Repeat k times, each fold becomes validation once.
  • Aggregate scores (mean, std).

Using scikit‑learn cross_val_score

from sklearn.model_selection import cross_val_score
from sklearn.ensemble import GradientBoostingClassifier

model = GradientBoostingClassifier(random_state=42)
cv_scores = cross_val_score(model, X, y, cv=5, scoring='accuracy')
print('CV Accuracy scores:', cv_scores)
print('Mean accuracy:', cv_scores.mean())

Stratified k‑fold (preserves class distribution)

from sklearn.model_selection import StratifiedKFold
skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
cv_scores = cross_val_score(model, X, y, cv=skf, scoring='f1')

Nested cross‑validation (hyperparameter tuning + performance estimate)

from sklearn.model_selection import GridSearchCV, cross_val_score

param_grid = {'n_estimators': [100, 200], 'learning_rate': [0.01, 0.1]}
inner_cv = StratifiedKFold(n_splits=3, shuffle=True, random_state=42)
outer_cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)

clf = GridSearchCV(GradientBoostingClassifier(random_state=42), param_grid, cv=inner_cv, scoring='roc_auc')
nested_score = cross_val_score(clf, X, y, cv=outer_cv, scoring='roc_auc')
print('Nested CV ROC‑AUC:', nested_score.mean())

TimeSeriesSplit (for temporal data)

from sklearn.model_selection import TimeSeriesSplit
tscv = TimeSeriesSplit(n_splits=5)
cv_scores = cross_val_score(model, X, y, cv=tscv, scoring='neg_mean_squared_error')

Best practices

  • Shuffle data before splitting unless time series.
  • Use stratification for classification.
  • Keep a final hold‑out test set untouched for the last evaluation.

Info

Cross‑validation can be computationally expensive; consider using fewer folds or smaller subsets for quick prototyping.

Previous: Dimensionality Reduction – Principal Component Analysis (PCA)Next: Hyperparameter Tuning – Grid Search & Randomized Search