Skip to main content
Brave Programmer Logo

BraveProgrammer

BraveProgrammer

HomeProjectsBlogsCoursesLessonsAbout

Site footer

BraveProgrammer

Free coding courses, practical tutorials, and real projects from BraveProgrammer. Learn web development with React, Next.js, and TypeScript.

Navigation

  • Home
  • Projects
  • Blogs
  • Courses

Resources

  • About
  • Lessons

© 2026 BraveProgrammer. All rights reserved.

  1. Courses
  2. /
  3. Master Data Science with Python

Lesson 22 of 60 · python

Handling Missing Data in Pandas

Duration: 20 minutes

Handling Missing Data

Real‑world datasets often contain missing values (NaN). Pandas offers tools to detect, drop, or impute them.

Detecting missing values

df.isna().sum()   # count per column

Dropping rows or columns

# Drop rows with any missing value
clean_rows = df.dropna()

# Drop columns with more than 50% missing
threshold = len(df) * 0.5
clean_cols = df.dropna(axis=1, thresh=threshold)

Filling missing values

# Fill with a constant
filled = df.fillna(0)

# Forward fill (propagate last valid observation)
ffill = df.fillna(method='ffill')

Imputation with statistics

# Fill numeric columns with median
numeric_cols = df.select_dtypes(include='number').columns
for col in numeric_cols:
    median = df[col].median()
    df[col].fillna(median, inplace=True)

Using scikit‑learn SimpleImputer

from sklearn.impute import SimpleImputer

imputer = SimpleImputer(strategy='mean')
numeric_data = df[numeric_cols].values
imputed = imputer.fit_transform(numeric_data)
df[numeric_cols] = imputed

Info

Always examine why data is missing; imputation can bias models.

Previous: Selecting & Filtering Data in PandasNext: Data Cleaning: Duplicates, Renaming, and Type Conversion