Lesson 24 of 60 · python
GroupBy and Aggregations
Duration: 20 minutes
GroupBy for Summarizing Data
groupby splits data into groups, applies a function, then combines the results.
Simple groupby
total_sales = df.groupby('store_id')['revenue'].sum()
print(total_sales)
Multiple aggregations
agg = df.groupby(['store_id', 'product_id']).agg({
'units': ['sum', 'mean'],
'revenue': 'sum'
})
print(agg.head())
Using named aggregation (pandas ≥ 0.25)
result = df.groupby('department').agg(
total_employees=('employee_id', 'count'),
avg_salary=('salary', 'mean')
)
print(result)
Transform vs. Apply
# Center units per store (subtract store mean)
centered = df.groupby('store_id')['units'].transform(lambda x: x - x.mean())
df['units_centered'] = centered
Filtering groups with filter
# Keep stores with total revenue > $1M
big_stores = df.groupby('store_id').filter(lambda x: x['revenue'].sum() > 1_000_000)