0Pricing
Machine Learning Academy · 课时

NumPy 基础:数组与数学运算

您将创建和操作 NumPy 数组,执行向量化算术运算,并理解广播机制,从而高效处理数值数据

NumPy 基础:数组与数学运算 是 CoddyKit 上的免费 Machine Learning Academy 课时。 这是第 2 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Machine Learning Academy 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Machine Learning Academy 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

Why NumPy Is the Foundation of ML

NumPy is the foundation under every ML library. Its superpower is vectorized math: it works on whole arrays at once in fast C code, leaving Python loops in the dust.

import numpy as np
import time

# Speed comparison: NumPy vs Python list
n = 1_000_000
python_list = list(range(n))
np_array = np.arange(n)

# Python loop (slow)
start = time.time()
result = [x * 2 for x in python_list]
print(f'Python loop: {time.time() - start:.3f}s')

# NumPy vectorised (fast)
start = time.time()
result = np_array * 2
print(f'NumPy vectorised: {time.time() - start:.4f}s')

Creating NumPy Arrays

The core NumPy object is the array. Its most important trait is shape — a tuple giving the size of each dimension, like (rows, cols). The code shows how to make them.

import numpy as np

# From Python list
a = np.array([1, 2, 3, 4, 5])
print('1D array:', a, '| shape:', a.shape)  # (5,)

# 2D matrix from nested list
M = np.array([[1, 2, 3], [4, 5, 6]])
print('2D matrix shape:', M.shape)  # (2, 3)

# Built-in generators
zeros = np.zeros((3, 4))   # 3x4 matrix of zeros
ones = np.ones((2, 5))     # 2x5 matrix of ones
range_arr = np.arange(0, 10, 2)   # [0 2 4 6 8]
linspace = np.linspace(0, 1, 5)   # 5 evenly spaced points
random = np.random.randn(3, 3)    # 3x3 standard normal

Array Indexing and Slicing

You index arrays with [row, col] and slice with start:stop:step. One gotcha: a slice is a view, not a copy — change it and you change the original. Use .copy() to be safe.

import numpy as np

M = np.array([[1, 2, 3, 4],
              [5, 6, 7, 8],
              [9, 10, 11, 12]])

# Single element
print(M[1, 2])   # 7 (row 1, column 2)

# Slice rows and columns
print(M[0:2, 1:3])  # rows 0-1, columns 1-2 -> [[2,3],[6,7]]

# All rows, last column
print(M[:, -1])  # [4, 8, 12]

# Boolean indexing
print(M[M > 6])  # [7, 8, 9, 10, 11, 12]

Vectorised Arithmetic Operations

NumPy math is element-wise by default: a + b adds matching elements with no loop. It runs in compiled C, which is why it flies on millions of numbers. See the code.

import numpy as np

a = np.array([1.0, 2.0, 3.0, 4.0])
b = np.array([10.0, 20.0, 30.0, 40.0])

print('Addition:', a + b)        # [11. 22. 33. 44.]
print('Multiply:', a * b)        # [10. 40. 90. 160.]
print('Power:', a ** 2)          # [ 1.  4.  9. 16.]
print('Divide:', b / a)          # [10. 10. 10. 10.]

# Scalar operations apply to all elements
print('Add scalar:', a + 100)    # [101. 102. 103. 104.]
print('Square root:', np.sqrt(a)) # [1. 1.41 1.73 2.]

Broadcasting: Operating on Different Shapes

Broadcasting lets NumPy combine different shapes by stretching the smaller one to fit. It's how you add a bias vector to a whole batch without writing any loops.

import numpy as np

# Matrix + vector (broadcasting)
matrix = np.array([[1, 2, 3],
                   [4, 5, 6],
                   [7, 8, 9]])  # shape (3, 3)

bias = np.array([10, 20, 30])   # shape (3,) -> broadcasts to (3, 3)

result = matrix + bias
print(result)
# [[11 22 33]
#  [14 25 36]
#  [17 28 39]]

# Normalize each column to zero mean (ML preprocessing)
mean = matrix.mean(axis=0)  # shape (3,)
centered = matrix - mean    # broadcasts across rows

Aggregation Functions

Aggregations like mean and sum collapse an array down. The axis picks the direction: axis=0 goes down columns, axis=1 goes across rows. The code shows both.

import numpy as np

X = np.array([[1, 2, 3],
              [4, 5, 6],
              [7, 8, 9]], dtype=float)

print('Global mean:', X.mean())           # 5.0
print('Column means:', X.mean(axis=0))   # [4. 5. 6.]
print('Row means:', X.mean(axis=1))      # [2. 5. 8.]
print('Global std:', X.std())            # ~2.58
print('Column max:', X.max(axis=0))      # [7. 8. 9.]
print('Row sum:', X.sum(axis=1))         # [ 6. 15. 24.]

Matrix Multiplication: The Heart of ML

Matrix multiplication is the heart of ML — every neural net layer and prediction is one. Use @ (not *, which is element-wise). Inner shapes must match: (m,k) by (k,n).

import numpy as np

# Linear regression prediction: y_hat = X @ weights + bias
X = np.random.randn(100, 5)   # 100 samples, 5 features
weights = np.random.randn(5)  # one weight per feature
bias = 0.5

y_hat = X @ weights + bias    # shape: (100,)
print('Predictions shape:', y_hat.shape)

# Matrix-matrix multiplication (e.g., two weight layers)
A = np.random.randn(4, 3)   # (4, 3)
B = np.random.randn(3, 5)   # (3, 5)
C = A @ B                    # (4, 5)
print('A @ B shape:', C.shape)

Reshaping and Stacking Arrays

Reshaping changes an array's shape without touching its data — like flattening a 28x28 image into 784 numbers. Stacking with vstack or hstack joins arrays together.

import numpy as np

# Reshape: 1D to 2D
a = np.arange(12)          # [0, 1, ..., 11]
matrix = a.reshape(3, 4)   # shape (3, 4)
print('Reshaped:', matrix.shape)

# Use -1 to infer one dimension automatically
flat = matrix.reshape(-1)  # back to 1D (12,)
row = matrix.reshape(1, -1)  # shape (1, 12)
col = matrix.reshape(-1, 1)  # shape (12, 1)

# Stack two arrays row-wise
A = np.array([[1, 2], [3, 4]])
B = np.array([[5, 6], [7, 8]])
stacked = np.vstack([A, B])  # shape (4, 2)

Random Number Generation for ML

Random numbers drive weight init, shuffling, and train-test splits. Always set a seed so results repeat — without it, every run differs and debugging gets painful.

import numpy as np

# Set seed for reproducibility
rng = np.random.default_rng(seed=42)

# Common distributions used in ML
uniform = rng.uniform(0, 1, size=(3, 3))   # Uniform [0, 1)
normal = rng.normal(0, 1, size=(3, 3))     # Standard normal
integers = rng.integers(0, 10, size=5)     # Random integers

# Shuffle an array
data = np.arange(10)
rng.shuffle(data)
print('Shuffled:', data)

# Random sampling without replacement
idxs = rng.choice(100, size=20, replace=False)  # 20 unique indices

Boolean Masks and Fancy Indexing

Boolean masking filters arrays in one clean line: compare to a condition, then select the matches. Fancy indexing picks elements by an explicit list of positions.

import numpy as np

scores = np.array([85, 42, 91, 67, 55, 78, 33, 95])

# Boolean mask: select scores above 70
mask = scores > 70
print('Mask:', mask)  # [T F T F F T F T]
print('High scores:', scores[mask])  # [85 91 78 95]

# Count how many passed
print('Passed:', mask.sum())  # 4

# Fancy indexing: select by explicit indices
indices = np.array([0, 2, 7])
print('Selected:', scores[indices])  # [85 91 95]

# Replace outliers
scores[scores < 50] = 50  # clip low scores to 50

NumPy in scikit-learn Workflows

scikit-learn wants shapes right: X as 2D (samples, features), y as 1D (samples). A shape like (100,) instead of (100, 1) is a classic error — reshape(-1, 1) fixes it.

import numpy as np
from sklearn.linear_model import LinearRegression

# X must be 2D: (n_samples, n_features)
X = np.array([1, 2, 3, 4, 5])   # shape (5,) -- WRONG
# Fix:
X = X.reshape(-1, 1)            # shape (5, 1) -- CORRECT

y = np.array([2.1, 4.0, 5.9, 8.1, 10.0])  # shape (5,) -- correct

model = LinearRegression()
model.fit(X, y)
print('Learned slope:', round(model.coef_[0], 2))  # ~2.0

Quick Check

Test your understanding of Machine Learning with Python concepts from this lesson.

Lesson Recap

You learned the core of NumPy: arrays underpin all ML, vectorized math and broadcasting kill slow loops, and the @ operator runs every model. Next up: Pandas. 🐼

常见问题解答

「NumPy 基础:数组与数学运算」课时是免费的吗?

是的 — 「NumPy 基础:数组与数学运算」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Machine Learning Academy 课程的其余内容,请升级到 CoddyKit PRO。 Machine Learning Academy 课程共包含 4 节课。

「NumPy 基础:数组与数学运算」这节课中我会学到什么?

您将创建和操作 NumPy 数组,执行向量化算术运算,并理解广播机制,从而高效处理数值数据 你通过在浏览器中直接运行的动手代码来练习 Machine Learning Academy,全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 Machine Learning Academy 需要有经验吗?

无需任何先前经验。CoddyKit 上的 Machine Learning Academy 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 2 节课,共 4 节。

「NumPy 基础:数组与数学运算」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 Machine Learning Academy 课中编写并运行代码吗?

能。每节 Machine Learning Academy 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. 安装 Anaconda 与 Jupyter Notebook
  2. NumPy 基础:数组与数学运算
  3. 使用 Pandas 操作数据
  4. 使用 Matplotlib 和 Seaborn 可视化数据
← 返回 Machine Learning Academy