Lesson 15 of 60 · python
Hands‑On NumPy Project: Synthetic Sales Dataset
Duration: 30 minutes
Hands‑On NumPy Project
In this project you will create a synthetic sales dataset, clean it, and compute key metrics using NumPy.
Step 1: Generate synthetic data
import numpy as np
rng = np.random.default_rng(2024)
n = 1000
# Columns: [store_id, product_id, units_sold, price]
store_ids = rng.integers(1, 11, size=n) # 10 stores
product_ids = rng.integers(100, 200, size=n) # 100 products
units = rng.poisson(lam=20, size=n) # realistic sales volume
price = rng.uniform(5, 100, size=n).round(2)
sales = np.column_stack((store_ids, product_ids, units, price))
print(sales[:5])
Step 2: Compute total revenue per store
revenue = sales[:,2] * sales[:,3]
# Aggregate revenue per store using np.bincount (store IDs start at 1)
rev_per_store = np.bincount(store_ids, weights=revenue)
print('Revenue per store (index = store_id):', rev_per_store[1:])
Step 3: Identify top‑selling product
units_per_product = np.bincount(product_ids, weights=units)
top_product = product_ids[np.argmax(units_per_product)]
print('Top selling product ID:', top_product)
Step 4: Visual sanity check (using Matplotlib later)
import matplotlib.pyplot as plt
plt.bar(range(1, len(rev_per_store)), rev_per_store[1:])
plt.xlabel('Store ID')
plt.ylabel('Revenue ($)')
plt.title('Revenue per Store')
plt.show()
What you learned
- Creating synthetic data with NumPy's random API.
- Vectorized arithmetic for revenue.
- Aggregating using
np.bincount. - Preparing data for downstream visualizations.