Skip to main content
Brave Programmer Logo

BraveProgrammer

BraveProgrammer

HomeProjectsBlogsCoursesLessonsAbout

Site footer

BraveProgrammer

Free coding courses, practical tutorials, and real projects from BraveProgrammer. Learn web development with React, Next.js, and TypeScript.

Navigation

  • Home
  • Projects
  • Blogs
  • Courses

Resources

  • About
  • Lessons

© 2026 BraveProgrammer. All rights reserved.

  1. Courses
  2. /
  3. Master Data Science with Python

Lesson 49 of 60 · python

Decision Trees and Random Forests

Duration: 30 minutes

Decision Trees & Random Forests

Tree‑based models are intuitive, handle mixed data types, and require little preprocessing.

Decision Tree basics

  • Splits data based on feature thresholds.
  • Uses impurity measures: Gini, entropy (information gain).

Training a Decision Tree with scikit‑learn

from sklearn.tree import DecisionTreeClassifier
from sklearn import tree

clf = DecisionTreeClassifier(max_depth=5, random_state=42)
clf.fit(X_train, y_train)

# Visualize the tree
plt.figure(figsize=(12,8))
tree.plot_tree(clf, filled=True, feature_names=X.columns, class_names=['No','Yes'])
plt.show()

Random Forests (ensemble of trees)

from sklearn.ensemble import RandomForestClassifier
rf = RandomForestClassifier(n_estimators=200, max_features='sqrt', random_state=42)
rf.fit(X_train, y_train)

# Feature importance
importances = pd.Series(rf.feature_importances_, index=X.columns)
importances.sort_values(ascending=False).plot(kind='bar')
plt.title('Feature Importances from Random Forest')
plt.show()

Hyperparameters to tune

  • n_estimators (number of trees)
  • max_depth
  • min_samples_split / min_samples_leaf
  • max_features

Advantages & disadvantages

Decision TreeRandom Forest
InterpretabilityHigh (visual)Medium (feature importance)
OverfittingProneLess prone (averaging)
Training speedFastSlower (many trees)

Out‑of‑Bag (OOB) error estimate

rf.oob_score_   # gives OOB accuracy for classification

Info

Random Forests handle missing values internally (by using surrogate splits) but it's better to clean data first.

Previous: Supervised Learning – Logistic RegressionNext: Model Evaluation Metrics for Classification