Lesson 55 of 60 · python
Ensemble Methods – Gradient Boosting (XGBoost)
Duration: 30 minutes
Gradient Boosting & XGBoost
Gradient boosting builds models sequentially, each new model correcting errors of its predecessor. XGBoost is a highly optimized implementation.
Installing XGBoost
pip install xgboost
Basic usage
import xgboost as xgb
from sklearn.metrics import accuracy_score
# Convert data to DMatrix (efficient internal format)
dtrain = xgb.DMatrix(X_train, label=y_train)
dtest = xgb.DMatrix(X_test, label=y_test)
params = {
'objective': 'binary:logistic',
'eval_metric': 'logloss',
'eta': 0.1,
'max_depth': 5,
'subsample': 0.8,
'colsample_bytree': 0.8,
'seed': 42
}
bst = xgb.train(params, dtrain, num_boost_round=200, evals=[(dtest, 'test')], early_stopping_rounds=20)
# Predict probabilities and threshold at 0.5
y_pred_proba = bst.predict(dtest)
y_pred = (y_pred_proba > 0.5).astype(int)
print('Test accuracy:', accuracy_score(y_test, y_pred))
Feature importance
xgb.plot_importance(bst, max_num_features=10, importance_type='gain')
plt.show()
Hyperparameter tuning (example with RandomizedSearchCV)
from sklearn.model_selection import RandomizedSearchCV
from xgboost import XGBClassifier
param_dist = {
'n_estimators': randint(100, 500),
'max_depth': randint(3, 10),
'learning_rate': uniform(0.01, 0.3),
'subsample': uniform(0.6, 0.4),
'colsample_bytree': uniform(0.6, 0.4)
}
xgb_clf = XGBClassifier(objective='binary:logistic', eval_metric='logloss', use_label_encoder=False)
rand_search = RandomizedSearchCV(xgb_clf, param_distributions=param_dist, n_iter=30, cv=5, scoring='roc_auc', n_jobs=-1, random_state=42)
rand_search.fit(X_train, y_train)
print('Best XGB params:', rand_search.best_params_)
Advantages of XGBoost
- Handles missing values natively.
- Regularization (
lambda,alpha). - Parallelized tree construction.
- Supports custom loss functions.
When to use?
- Tabular data with complex interactions.
- Kaggle competitions – often a top performer.