Skip to main content
Brave Programmer Logo

BraveProgrammer

BraveProgrammer

HomeProjectsBlogsCoursesLessonsAbout

Site footer

BraveProgrammer

Free coding courses, practical tutorials, and real projects from BraveProgrammer. Learn web development with React, Next.js, and TypeScript.

Navigation

  • Home
  • Projects
  • Blogs
  • Courses

Resources

  • About
  • Lessons

© 2026 BraveProgrammer. All rights reserved.

  1. Courses
  2. /
  3. Master Data Science with Python

Lesson 55 of 60 · python

Ensemble Methods – Gradient Boosting (XGBoost)

Duration: 30 minutes

Gradient Boosting & XGBoost

Gradient boosting builds models sequentially, each new model correcting errors of its predecessor. XGBoost is a highly optimized implementation.

Installing XGBoost

pip install xgboost

Basic usage

import xgboost as xgb
from sklearn.metrics import accuracy_score

# Convert data to DMatrix (efficient internal format)
dtrain = xgb.DMatrix(X_train, label=y_train)
dtest = xgb.DMatrix(X_test, label=y_test)

params = {
    'objective': 'binary:logistic',
    'eval_metric': 'logloss',
    'eta': 0.1,
    'max_depth': 5,
    'subsample': 0.8,
    'colsample_bytree': 0.8,
    'seed': 42
}

bst = xgb.train(params, dtrain, num_boost_round=200, evals=[(dtest, 'test')], early_stopping_rounds=20)

# Predict probabilities and threshold at 0.5
y_pred_proba = bst.predict(dtest)
y_pred = (y_pred_proba > 0.5).astype(int)
print('Test accuracy:', accuracy_score(y_test, y_pred))

Feature importance

xgb.plot_importance(bst, max_num_features=10, importance_type='gain')
plt.show()

Hyperparameter tuning (example with RandomizedSearchCV)

from sklearn.model_selection import RandomizedSearchCV
from xgboost import XGBClassifier

param_dist = {
    'n_estimators': randint(100, 500),
    'max_depth': randint(3, 10),
    'learning_rate': uniform(0.01, 0.3),
    'subsample': uniform(0.6, 0.4),
    'colsample_bytree': uniform(0.6, 0.4)
}

xgb_clf = XGBClassifier(objective='binary:logistic', eval_metric='logloss', use_label_encoder=False)
rand_search = RandomizedSearchCV(xgb_clf, param_distributions=param_dist, n_iter=30, cv=5, scoring='roc_auc', n_jobs=-1, random_state=42)
rand_search.fit(X_train, y_train)
print('Best XGB params:', rand_search.best_params_)

Advantages of XGBoost

  • Handles missing values natively.
  • Regularization (lambda, alpha).
  • Parallelized tree construction.
  • Supports custom loss functions.

When to use?

  • Tabular data with complex interactions.
  • Kaggle competitions – often a top performer.

Info

Beware of over‑fitting; use early stopping and keep a separate test set.

Previous: Hyperparameter Tuning – Grid Search & Randomized SearchNext: Deep Learning Foundations – TensorFlow Basics