Lesson 48 of 60 · python
Supervised Learning – Logistic Regression
Duration: 25 minutes
Logistic Regression
Logistic regression is used for binary classification. It models the log‑odds of the probability of the positive class.
Logistic function
p = 1 / (1 + exp(-(β0 + β1·x1 + …)))
Using scikit‑learn
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, classification_report, confusion_matrix
log_reg = LogisticRegression(max_iter=1000, solver='lbfgs')
log_reg.fit(X_train, y_train)
y_pred = log_reg.predict(X_test)
print('Accuracy:', accuracy_score(y_test, y_pred))
print(classification_report(y_test, y_pred))
print('Confusion Matrix:\n', confusion_matrix(y_test, y_pred))
Multiclass extension
- One‑vs‑Rest (OvR) – default.
- Multinomial – set
multi_class='multinomial'.
Regularization
- L2 (default) controls overfitting.
- Adjust via
C(inverse of regularization strength).
Class imbalance handling
- Class weighting (
class_weight='balanced'). - Resampling (SMOTE, undersampling).
log_reg_bal = LogisticRegression(class_weight='balanced')
log_reg_bal.fit(X_train, y_train)
Visualizing decision boundary (2‑D example)
import numpy as np
import matplotlib.pyplot as plt
# Create a grid
x_min, x_max = X[:, 0].min() - 1, X[:, 0].max() + 1
y_min, y_max = X[:, 1].min() - 1, X[:, 1].max() + 1
xx, yy = np.meshgrid(np.arange(x_min, x_max, 0.1), np.arange(y_min, y_max, 0.1))
Z = log_reg.predict(np.c_[xx.ravel(), yy.ravel()])
Z = Z.reshape(xx.shape)
plt.contourf(xx, yy, Z, alpha=0.3)
plt.scatter(X[:,0], X[:,1], c=y, edgecolor='k')
plt.title('Logistic Regression Decision Boundary')
plt.show()
Interpreting coefficients
- Positive coefficient → increase odds of class 1.
- Magnitude indicates strength of impact.