Skip to main content
Brave Programmer Logo

BraveProgrammer

BraveProgrammer

HomeProjectsBlogsCoursesLessonsAbout

Site footer

BraveProgrammer

Free coding courses, practical tutorials, and real projects from BraveProgrammer. Learn web development with React, Next.js, and TypeScript.

Navigation

  • Home
  • Projects
  • Blogs
  • Courses

Resources

  • About
  • Lessons

© 2026 BraveProgrammer. All rights reserved.

  1. Courses
  2. /
  3. Master Data Science with Python

Lesson 11 of 60 · python

Random Sampling and Statistics with NumPy

Duration: 20 minutes

Random Sampling and Statistics

NumPy’s random module (numpy.random) is a powerful tool for generating synthetic data.

Simple RNG (Generator API)

rng = np.random.default_rng(2024)

# Uniform distribution
uniform = rng.uniform(low=0.0, high=1.0, size=5)
print("Uniform:", uniform)

# Normal distribution
normal = rng.normal(loc=0.0, scale=1.0, size=5)
print("Normal:", normal)

Sampling without replacement

population = np.arange(10)
sample = rng.choice(population, size=4, replace=False)
print(sample)

Basic statistics

data = rng.normal(5, 2, size=1000)
print("Mean:", data.mean())
print("Median:", np.median(data))
print("25th percentile:", np.percentile(data, 25))

Correlation coefficient

x = rng.normal(size=100)
y = 2*x + rng.normal(scale=0.5, size=100)
print("Pearson r:", np.corrcoef(x, y)[0,1])

Choosing the right method

  • MCAR (Missing Completely at Random): dropping may be OK.
  • MAR (Missing at Random): use predictive imputation.
  • MNAR (Missing Not at Random): consider domain‑specific reasoning.

Info

Never impute the target variable – keep it as missing if you need to predict it.

Previous: Mathematical Operations with NumPyNext: Linear Algebra Basics with NumPy