Lesson 11 of 60 · python
Random Sampling and Statistics with NumPy
Duration: 20 minutes
Random Sampling and Statistics
NumPy’s random module (numpy.random) is a powerful tool for generating synthetic data.
Simple RNG (Generator API)
rng = np.random.default_rng(2024)
# Uniform distribution
uniform = rng.uniform(low=0.0, high=1.0, size=5)
print("Uniform:", uniform)
# Normal distribution
normal = rng.normal(loc=0.0, scale=1.0, size=5)
print("Normal:", normal)
Sampling without replacement
population = np.arange(10)
sample = rng.choice(population, size=4, replace=False)
print(sample)
Basic statistics
data = rng.normal(5, 2, size=1000)
print("Mean:", data.mean())
print("Median:", np.median(data))
print("25th percentile:", np.percentile(data, 25))
Correlation coefficient
x = rng.normal(size=100)
y = 2*x + rng.normal(scale=0.5, size=100)
print("Pearson r:", np.corrcoef(x, y)[0,1])
Choosing the right method
- MCAR (Missing Completely at Random): dropping may be OK.
- MAR (Missing at Random): use predictive imputation.
- MNAR (Missing Not at Random): consider domain‑specific reasoning.