Skip to main content
Brave Programmer Logo

BraveProgrammer

BraveProgrammer

HomeProjectsBlogsCoursesLessonsAbout

Site footer

BraveProgrammer

Free coding courses, practical tutorials, and real projects from BraveProgrammer. Learn web development with React, Next.js, and TypeScript.

Navigation

  • Home
  • Projects
  • Blogs
  • Courses

Resources

  • About
  • Lessons

© 2026 BraveProgrammer. All rights reserved.

  1. Courses
  2. /
  3. Master Data Science with Python

Lesson 42 of 60 · python

Handling Missing Values: Imputation Strategies

Duration: 20 minutes

Handling Missing Values

Missing data can bias models. Choose an imputation strategy based on the variable type and missingness pattern.

Simple strategies

  • Drop rows (df.dropna())
  • Fill with constant (df.fillna(0))
  • Mean/median imputation
  • Forward/backward fill (ffill, bfill)

Example: Mean imputation for numeric column

median_price = df['price'].median()
df['price'].fillna(median_price, inplace=True)

Categorical imputation

mode_category = df['category'].mode()[0]
df['category'].fillna(mode_category, inplace=True)

Advanced: Using scikit‑learn SimpleImputer

from sklearn.impute import SimpleImputer

imputer = SimpleImputer(strategy='mean')
numeric_data = df.select_dtypes(include='number')
df[numeric_data.columns] = imputer.fit_transform(numeric_data)

Visualizing missingness with missingno

pip install missingno
import missingno as msno
msno.matrix(df)
plt.show()

Choosing the right method

  • MCAR (Missing Completely at Random): dropping may be OK.
  • MAR (Missing at Random): use predictive imputation.
  • MNAR (Missing Not at Random): consider domain‑specific reasoning.

Info

Never impute the target variable – keep it as missing if you need to predict it.

Previous: Real‑World Data Challenges: Messy DatasetsNext: Removing Duplicates and Detecting Outliers