Lesson 30 of 60 · python
Pandas Project: Titanic Data Analysis
Duration: 30 minutes
Pandas Project – Titanic
The Titanic dataset is a classic beginner data‑science project. We'll load, clean, explore, and visualize the data.
Step 1: Load the data
import pandas as pd
import seaborn as sns
titanic_url = "https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv"
df = pd.read_csv(titanic_url)
print(df.head())
Step 2: Quick overview
print(df.info())
print(df.describe(include='all'))
Step 3: Handle missing values
# Age has missing values – fill with median
median_age = df['Age'].median()
df['Age'].fillna(median_age, inplace=True)
# Cabin has many missing – drop column
df.drop(columns=['Cabin'], inplace=True)
Step 4: Feature engineering
# Extract title from name (Mr, Mrs, etc.)
import re
def get_title(name):
match = re.search(r',\s*([^\.]+)\.', name)
return match.group(1).strip() if match else ''
df['Title'] = df['Name'].apply(get_title)
# Encode Sex as binary
df['Sex'] = df['Sex'].map({'male': 0, 'female': 1})
Step 5: Exploratory analysis
# Survival rate by gender
surv_by_sex = df.groupby('Sex')['Survived'].mean()
print(surv_by_sex)
# Visualize with seaborn
import matplotlib.pyplot as plt
sns.barplot(x='Sex', y='Survived', data=df)
plt.title('Survival Rate by Gender')
plt.show()
Step 6: Grouped statistics
# Average fare by passenger class
fare_by_class = df.groupby('Pclass')['Fare'].mean()
print(fare_by_class)
Step 7: Save cleaned data
df.to_csv('titanic_cleaned.csv', index=False)
What you learned
- Loading CSV from a remote URL.
- Data cleaning: handling missing values and dropping columns.
- Feature engineering using regex and mapping.
- Simple visualizations with seaborn.
- Saving the processed data for later modeling.