Lesson 44 of 60 · python
Feature Scaling: Normalization and Standardization
Duration: 20 minutes
Feature Scaling
Many ML algorithms (e.g., K‑means, SVM) are sensitive to feature magnitude. Scaling brings features onto a comparable scale.
Normalization (min‑max scaling)
from sklearn.preprocessing import MinMaxScaler
scaler = MinMaxScaler()
scaled = scaler.fit_transform(df[['price', 'quantity']])
scaled_df = pd.DataFrame(scaled, columns=['price_norm', 'quantity_norm'])
Standardization (zero‑mean, unit‑variance)
from sklearn.preprocessing import StandardScaler
std_scaler = StandardScaler()
std_scaled = std_scaler.fit_transform(df[['price', 'quantity']])
std_df = pd.DataFrame(std_scaled, columns=['price_std', 'quantity_std'])
When to use which?
- Normalization: when you need bounded values (0‑1), e.g., neural networks with sigmoid.
- Standardization: works well for algorithms assuming Gaussian distribution.
Scaling pipelines with Pipeline
from sklearn.pipeline import Pipeline
pipe = Pipeline([
('imputer', SimpleImputer(strategy='median')),
('scaler', StandardScaler())
])
X_scaled = pipe.fit_transform(df[numeric_features])
Visualizing before/after scaling
fig, ax = plt.subplots(1,2, figsize=(12,5))
ax[0].hist(df['price'], bins=30, color='skyblue')
ax[0].set_title('Original')
ax[1].hist(scaled_df['price_norm'], bins=30, color='orange')
ax[1].set_title('Normalized')
plt.show()