Skip to main content
Brave Programmer Logo

BraveProgrammer

BraveProgrammer

HomeProjectsBlogsCoursesLessonsAbout

Site footer

BraveProgrammer

Free coding courses, practical tutorials, and real projects from BraveProgrammer. Learn web development with React, Next.js, and TypeScript.

Navigation

  • Home
  • Projects
  • Blogs
  • Courses

Resources

  • About
  • Lessons

© 2026 BraveProgrammer. All rights reserved.

  1. Courses
  2. /
  3. Master Data Science with Python

Lesson 28 of 60 · python

Categorical Data and Encoding Techniques

Duration: 20 minutes

Categorical Data & Encoding

Machine learning models require numeric input, so we must convert categorical variables.

One‑Hot Encoding with get_dummies

df_onehot = pd.get_dummies(df, columns=['city', 'gender'], drop_first=True)
print(df_onehot.head())

Label Encoding with scikit‑learn

from sklearn.preprocessing import LabelEncoder
le = LabelEncoder()
df['dept_code'] = le.fit_transform(df['department'])

Frequency Encoding

freq = df['category'].value_counts() / len(df)
df['category_freq'] = df['category'].map(freq)

Target Encoding (mean encoding)

target_means = df.groupby('city')['target'].mean()
df['city_target_enc'] = df['city'].map(target_means)

When to use which?

  • Low cardinality (<10) → One‑Hot.
  • Medium cardinality (10‑100) → Frequency / Target.
  • High cardinality (>100) → Hashing or embeddings.

Info

Always fit encoders on the training set only and apply them to the test set to avoid data leakage.

Previous: Time Series Data with PandasNext: Applying Functions: `apply`, `map`, and Vectorized Operations