Lesson 14 of 60 · python
Performance Tips: Vectorization, Broadcasting, and Memory Layout
Duration: 20 minutes
Performance Tips for NumPy
Writing fast NumPy code is about thinking in whole‑array operations.
Vectorize instead of looping
# Slow Python loop
s = 0
for i in range(1_000_000):
s += i*i
# Fast NumPy vectorized version
arr = np.arange(1_000_000)
s_fast = np.sum(arr**2)
Use in‑place operations to reduce temporary arrays
arr = np.arange(10, dtype=np.float64)
arr *= 2 # modifies arr in place, no new array created
Control memory layout using order
C_contig = np.arange(12).reshape(3,4, order='C') # row‑major
F_contig = np.arange(12).reshape(3,4, order='F') # column‑major
Avoid unnecessary copies with np.r_ and np.c_
# Concatenation without extra copy
a = np.arange(3)
b = np.arange(3,6)
combined = np.r_[a, b]
Use np.einsum for complex summations
# Compute trace of a matrix efficiently
M = np.random.rand(5,5)
trace = np.einsum('ii', M)