NumPy Essentials: Arrays and Math Operations
Learners will create and manipulate NumPy arrays, perform vectorised arithmetic, and understand broadcasting so they can handle numerical data efficiently.
NumPy Essentials: Arrays and Math Operations is a free Machine Learning Academy lesson on CoddyKit — lesson 2 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the Machine Learning Academy learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.
Why NumPy Is the Foundation of ML
NumPy is the foundation under every ML library. Its superpower is vectorized math: it works on whole arrays at once in fast C code, leaving Python loops in the dust.
import numpy as np
import time
# Speed comparison: NumPy vs Python list
n = 1_000_000
python_list = list(range(n))
np_array = np.arange(n)
# Python loop (slow)
start = time.time()
result = [x * 2 for x in python_list]
print(f'Python loop: {time.time() - start:.3f}s')
# NumPy vectorised (fast)
start = time.time()
result = np_array * 2
print(f'NumPy vectorised: {time.time() - start:.4f}s')Creating NumPy Arrays
The core NumPy object is the array. Its most important trait is shape — a tuple giving the size of each dimension, like (rows, cols). The code shows how to make them.
import numpy as np
# From Python list
a = np.array([1, 2, 3, 4, 5])
print('1D array:', a, '| shape:', a.shape) # (5,)
# 2D matrix from nested list
M = np.array([[1, 2, 3], [4, 5, 6]])
print('2D matrix shape:', M.shape) # (2, 3)
# Built-in generators
zeros = np.zeros((3, 4)) # 3x4 matrix of zeros
ones = np.ones((2, 5)) # 2x5 matrix of ones
range_arr = np.arange(0, 10, 2) # [0 2 4 6 8]
linspace = np.linspace(0, 1, 5) # 5 evenly spaced points
random = np.random.randn(3, 3) # 3x3 standard normalArray Indexing and Slicing
You index arrays with [row, col] and slice with start:stop:step. One gotcha: a slice is a view, not a copy — change it and you change the original. Use .copy() to be safe.
import numpy as np
M = np.array([[1, 2, 3, 4],
[5, 6, 7, 8],
[9, 10, 11, 12]])
# Single element
print(M[1, 2]) # 7 (row 1, column 2)
# Slice rows and columns
print(M[0:2, 1:3]) # rows 0-1, columns 1-2 -> [[2,3],[6,7]]
# All rows, last column
print(M[:, -1]) # [4, 8, 12]
# Boolean indexing
print(M[M > 6]) # [7, 8, 9, 10, 11, 12]Vectorised Arithmetic Operations
NumPy math is element-wise by default: a + b adds matching elements with no loop. It runs in compiled C, which is why it flies on millions of numbers. See the code.
import numpy as np
a = np.array([1.0, 2.0, 3.0, 4.0])
b = np.array([10.0, 20.0, 30.0, 40.0])
print('Addition:', a + b) # [11. 22. 33. 44.]
print('Multiply:', a * b) # [10. 40. 90. 160.]
print('Power:', a ** 2) # [ 1. 4. 9. 16.]
print('Divide:', b / a) # [10. 10. 10. 10.]
# Scalar operations apply to all elements
print('Add scalar:', a + 100) # [101. 102. 103. 104.]
print('Square root:', np.sqrt(a)) # [1. 1.41 1.73 2.]Broadcasting: Operating on Different Shapes
Broadcasting lets NumPy combine different shapes by stretching the smaller one to fit. It's how you add a bias vector to a whole batch without writing any loops.
import numpy as np
# Matrix + vector (broadcasting)
matrix = np.array([[1, 2, 3],
[4, 5, 6],
[7, 8, 9]]) # shape (3, 3)
bias = np.array([10, 20, 30]) # shape (3,) -> broadcasts to (3, 3)
result = matrix + bias
print(result)
# [[11 22 33]
# [14 25 36]
# [17 28 39]]
# Normalize each column to zero mean (ML preprocessing)
mean = matrix.mean(axis=0) # shape (3,)
centered = matrix - mean # broadcasts across rowsAggregation Functions
Aggregations like mean and sum collapse an array down. The axis picks the direction: axis=0 goes down columns, axis=1 goes across rows. The code shows both.
import numpy as np
X = np.array([[1, 2, 3],
[4, 5, 6],
[7, 8, 9]], dtype=float)
print('Global mean:', X.mean()) # 5.0
print('Column means:', X.mean(axis=0)) # [4. 5. 6.]
print('Row means:', X.mean(axis=1)) # [2. 5. 8.]
print('Global std:', X.std()) # ~2.58
print('Column max:', X.max(axis=0)) # [7. 8. 9.]
print('Row sum:', X.sum(axis=1)) # [ 6. 15. 24.]Matrix Multiplication: The Heart of ML
Matrix multiplication is the heart of ML — every neural net layer and prediction is one. Use @ (not *, which is element-wise). Inner shapes must match: (m,k) by (k,n).
import numpy as np
# Linear regression prediction: y_hat = X @ weights + bias
X = np.random.randn(100, 5) # 100 samples, 5 features
weights = np.random.randn(5) # one weight per feature
bias = 0.5
y_hat = X @ weights + bias # shape: (100,)
print('Predictions shape:', y_hat.shape)
# Matrix-matrix multiplication (e.g., two weight layers)
A = np.random.randn(4, 3) # (4, 3)
B = np.random.randn(3, 5) # (3, 5)
C = A @ B # (4, 5)
print('A @ B shape:', C.shape)Reshaping and Stacking Arrays
Reshaping changes an array's shape without touching its data — like flattening a 28x28 image into 784 numbers. Stacking with vstack or hstack joins arrays together.
import numpy as np
# Reshape: 1D to 2D
a = np.arange(12) # [0, 1, ..., 11]
matrix = a.reshape(3, 4) # shape (3, 4)
print('Reshaped:', matrix.shape)
# Use -1 to infer one dimension automatically
flat = matrix.reshape(-1) # back to 1D (12,)
row = matrix.reshape(1, -1) # shape (1, 12)
col = matrix.reshape(-1, 1) # shape (12, 1)
# Stack two arrays row-wise
A = np.array([[1, 2], [3, 4]])
B = np.array([[5, 6], [7, 8]])
stacked = np.vstack([A, B]) # shape (4, 2)Random Number Generation for ML
Random numbers drive weight init, shuffling, and train-test splits. Always set a seed so results repeat — without it, every run differs and debugging gets painful.
import numpy as np
# Set seed for reproducibility
rng = np.random.default_rng(seed=42)
# Common distributions used in ML
uniform = rng.uniform(0, 1, size=(3, 3)) # Uniform [0, 1)
normal = rng.normal(0, 1, size=(3, 3)) # Standard normal
integers = rng.integers(0, 10, size=5) # Random integers
# Shuffle an array
data = np.arange(10)
rng.shuffle(data)
print('Shuffled:', data)
# Random sampling without replacement
idxs = rng.choice(100, size=20, replace=False) # 20 unique indicesBoolean Masks and Fancy Indexing
Boolean masking filters arrays in one clean line: compare to a condition, then select the matches. Fancy indexing picks elements by an explicit list of positions.
import numpy as np
scores = np.array([85, 42, 91, 67, 55, 78, 33, 95])
# Boolean mask: select scores above 70
mask = scores > 70
print('Mask:', mask) # [T F T F F T F T]
print('High scores:', scores[mask]) # [85 91 78 95]
# Count how many passed
print('Passed:', mask.sum()) # 4
# Fancy indexing: select by explicit indices
indices = np.array([0, 2, 7])
print('Selected:', scores[indices]) # [85 91 95]
# Replace outliers
scores[scores < 50] = 50 # clip low scores to 50NumPy in scikit-learn Workflows
scikit-learn wants shapes right: X as 2D (samples, features), y as 1D (samples). A shape like (100,) instead of (100, 1) is a classic error — reshape(-1, 1) fixes it.
import numpy as np
from sklearn.linear_model import LinearRegression
# X must be 2D: (n_samples, n_features)
X = np.array([1, 2, 3, 4, 5]) # shape (5,) -- WRONG
# Fix:
X = X.reshape(-1, 1) # shape (5, 1) -- CORRECT
y = np.array([2.1, 4.0, 5.9, 8.1, 10.0]) # shape (5,) -- correct
model = LinearRegression()
model.fit(X, y)
print('Learned slope:', round(model.coef_[0], 2)) # ~2.0Quick Check
Test your understanding of Machine Learning with Python concepts from this lesson.
Lesson Recap
You learned the core of NumPy: arrays underpin all ML, vectorized math and broadcasting kill slow loops, and the @ operator runs every model. Next up: Pandas. 🐼
Frequently asked questions
Is the “NumPy Essentials: Arrays and Math Operations” lesson free?
Yes — the full text of “NumPy Essentials: Arrays and Math Operations” is free to read here on the web, and the Machine Learning Academy course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the Machine Learning Academy course, upgrade to CoddyKit PRO.
What will I learn in “NumPy Essentials: Arrays and Math Operations”?
Learners will create and manipulate NumPy arrays, perform vectorised arithmetic, and understand broadcasting so they can handle numerical data efficiently. You practise Machine Learning Academy with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.
Do I need any experience to start Machine Learning Academy?
No prior experience is required. Machine Learning Academy on CoddyKit is structured for beginners through advanced learners; this is — lesson 2 of 4, so you can start here or from the beginning and move at your own pace.
How long does the “NumPy Essentials: Arrays and Math Operations” lesson take?
Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.
Can I write and run code in this Machine Learning Academy lesson?
Yes. Every Machine Learning Academy lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.
All lessons in this course
- Installing Anaconda and Jupyter Notebook
- NumPy Essentials: Arrays and Math Operations
- Pandas for Data Manipulation
- Visualising Data with Matplotlib and Seaborn