PyTorchテンソル:作成、演算、GPUへの転送
PythonリストとNumPy配列からテンソルを作成し、要素単位の演算と行列演算を行い、.to('cuda')でテンソルをGPUへ移動します。
「PyTorchテンソル:作成、演算、GPUへの転送」はCoddyKit上の無料Machine Learning Academyレッスンです。 これはレッスン1/4です。 下記で完全なレッスンを無料で読むことができます。その後、ブラウザ内の組み込みコードエディタと24時間対応のAIチューターでハンズオン演習できます。 これはMachine Learning Academy学習パスの一部であり、ウェブとCoddyKitアプリ全体で進捗が同期されます。 Machine Learning Academyコースには全4レッスンが含まれています。
このレッスンの一部はまだ翻訳されておらず、英語で表示されています。
What Is a PyTorch Tensor?
A tensor is the fundamental data structure in PyTorch — essentially an n-dimensional array similar to a NumPy array but with built-in GPU support and automatic differentiation. Tensors can be scalars (0D), vectors (1D), matrices (2D), or higher-dimensional structures. PyTorch tensors track computation history, enabling automatic gradient computation for training neural networks.
import torch
# Scalar (0-dimensional tensor)
scalar = torch.tensor(3.14)
print(scalar.shape) # torch.Size([])
# Vector (1D)
vector = torch.tensor([1.0, 2.0, 3.0])
print(vector.shape) # torch.Size([3])
# Matrix (2D)
matrix = torch.tensor([[1, 2], [3, 4]])
print(matrix.shape) # torch.Size([2, 2])Creating Tensors: Common Methods
PyTorch provides many factory functions to create tensors with specific values or shapes. torch.zeros and torch.ones fill tensors with constants; torch.rand samples from a uniform distribution; torch.randn samples from a standard normal distribution. These are the building blocks for weight initialisation and synthetic data generation.
import torch
zeros = torch.zeros(3, 4) # 3x4 matrix of zeros
ones = torch.ones(2, 3) # 2x3 matrix of ones
rand_uniform = torch.rand(3, 3) # uniform in [0, 1)
rand_normal = torch.randn(3, 3) # standard normal N(0,1)
arange = torch.arange(0, 10, 2) # [0, 2, 4, 6, 8]
print(zeros.dtype) # torch.float32 (default)
print(arange) # tensor([0, 2, 4, 6, 8])Creating Tensors from NumPy Arrays
You will often start with a NumPy array (from Pandas, scikit-learn, etc.) and need to convert it to a PyTorch tensor. torch.from_numpy shares memory with the NumPy array — modifying one modifies the other. Alternatively, torch.tensor makes a copy. Knowing which to use prevents surprising bugs in data pipelines.
import torch
import numpy as np
arr = np.array([1.0, 2.0, 3.0])
# Shares memory with arr
t_shared = torch.from_numpy(arr)
# Makes an independent copy
t_copy = torch.tensor(arr)
arr[0] = 99.0
print(t_shared) # tensor([99., 2., 3.]) <- changed
print(t_copy) # tensor([1., 2., 3.]) <- unchangedTensor Data Types (dtypes)
Tensors have a dtype that controls numeric precision and memory usage. The default is torch.float32, which is the standard for neural network weights. torch.float64 offers higher precision at double the memory; torch.int64 is used for integer labels. Mismatched dtypes cause runtime errors, so always check and cast explicitly with .float() or .to(dtype).
import torch
f32 = torch.tensor([1.0, 2.0]) # float32 by default
f64 = torch.tensor([1.0, 2.0], dtype=torch.float64)
i64 = torch.tensor([1, 2, 3]) # int64 by default
print(f32.dtype) # torch.float32
print(i64.dtype) # torch.int64
# Cast to float32
labels = i64.float()
print(labels.dtype) # torch.float32Element-Wise Arithmetic Operations
PyTorch overloads the standard Python arithmetic operators for element-wise operations on tensors of the same shape. Addition, subtraction, multiplication, and division all work element-wise. These operations are highly optimised and will run on the GPU when tensors are on a CUDA device. In-place operations (e.g., add_) modify the tensor directly without allocating new memory.
import torch
a = torch.tensor([1.0, 2.0, 3.0])
b = torch.tensor([4.0, 5.0, 6.0])
print(a + b) # tensor([5., 7., 9.])
print(a * b) # tensor([ 4., 10., 18.])
print(b / a) # tensor([4., 2.5, 2.])
print(a ** 2) # tensor([1., 4., 9.])
# In-place add
a.add_(1.0)
print(a) # tensor([2., 3., 4.])Matrix Multiplication with torch.matmul
Matrix multiplication is the core operation inside every neural network layer. torch.matmul (or the @ operator) performs matrix multiplication and handles batches of matrices automatically with broadcasting. For 2D inputs it computes standard matrix product; for 3D or higher it performs batched matrix multiply. This is how a linear layer computes output = input @ weight.T + bias.
import torch
A = torch.randn(3, 4) # 3x4
B = torch.randn(4, 5) # 4x5
C = torch.matmul(A, B) # 3x5
print(C.shape) # torch.Size([3, 5])
# Equivalent using @ operator
C2 = A @ B
print(torch.allclose(C, C2)) # True
# Batched matmul
batch_A = torch.randn(8, 3, 4) # batch of 8 matrices
batch_B = torch.randn(8, 4, 5)
result = batch_A @ batch_B # torch.Size([8, 3, 5])Reshaping Tensors: view and reshape
Changing the shape of a tensor without changing its data is one of the most common operations in deep learning. view requires the tensor to be contiguous in memory and returns a view (shared data); reshape works on non-contiguous tensors by copying if needed. Use -1 as a wildcard dimension and PyTorch infers the correct size. Flattening a 2D feature map to a 1D vector before a linear layer is a typical use case.
import torch
t = torch.arange(12).float() # tensor of 12 elements
print(t.shape) # torch.Size([12])
m = t.view(3, 4) # reshape to 3x4
print(m.shape) # torch.Size([3, 4])
m2 = t.view(2, -1) # PyTorch infers 6 columns
print(m2.shape) # torch.Size([2, 6])
# Flatten to 1D
flat = m.reshape(-1)
print(flat.shape) # torch.Size([12])Broadcasting: Operating on Different Shapes
Broadcasting allows PyTorch to perform operations on tensors with different shapes by implicitly expanding dimensions. The rules are borrowed from NumPy: dimensions are aligned from the right, and a size-1 dimension can be stretched to match the other tensor. Broadcasting avoids explicit tiling of data, saving memory. It is used constantly in neural network layers to add bias vectors to batched output matrices.
import torch
# Matrix (3x4) + vector (4,) -- vector broadcast over rows
matrix = torch.ones(3, 4)
bias = torch.tensor([1.0, 2.0, 3.0, 4.0]) # shape (4,)
result = matrix + bias
print(result.shape) # torch.Size([3, 4])
print(result[0]) # tensor([2., 3., 4., 5.])
# Column vector (3,1) * row vector (1,4) -> (3,4)
col = torch.arange(1, 4).float().unsqueeze(1) # (3,1)
row = torch.arange(1, 5).float().unsqueeze(0) # (1,4)
print((col * row).shape) # torch.Size([3, 4])Checking Devices: CPU vs CUDA
Each tensor lives on a device: either cpu or a CUDA GPU such as cuda:0. Checking device availability with torch.cuda.is_available() lets you write device-agnostic code. All tensors involved in a computation must be on the same device — attempting to add a CPU tensor and a GPU tensor raises a runtime error. The standard pattern is to create a device variable and move everything to it.
import torch
device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
print('Using device:', device)
# Create tensor directly on the chosen device
t = torch.randn(3, 3, device=device)
print(t.device)
# Move an existing CPU tensor to the device
cpu_tensor = torch.tensor([1.0, 2.0, 3.0])
gpu_tensor = cpu_tensor.to(device)
print(gpu_tensor.device)Moving Tensors to GPU with .to('cuda')
Training neural networks on a GPU can be 10-100x faster than on CPU for large models. After confirming CUDA availability, move tensors to the GPU with .to('cuda') or .cuda(). Move them back to CPU for NumPy conversion with .cpu() (NumPy cannot access GPU memory directly). The .detach() call removes a tensor from the computation graph before converting to NumPy.
import torch
# Simulate GPU workflow (falls back to CPU gracefully)
device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
# Move model weights and data to GPU
weights = torch.randn(256, 128).to(device)
inputs = torch.randn(32, 128).to(device) # batch of 32
outputs = inputs @ weights.T # matmul on device
print(outputs.shape) # torch.Size([32, 256])
# Convert back to NumPy for visualization
np_out = outputs.detach().cpu().numpy()
print(type(np_out)) # <class 'numpy.ndarray'>Useful Tensor Attributes and Utilities
Several tensor attributes and utility functions are essential for debugging and shaping data. .shape gives the dimensions; .numel() returns total element count; .dtype shows the data type; .requires_grad indicates whether gradients will be tracked. Functions like torch.cat, torch.stack, and torch.squeeze are used constantly to assemble batches and remove size-1 dimensions.
import torch
t = torch.randn(4, 3, 2)
print(t.shape) # torch.Size([4, 3, 2])
print(t.numel()) # 24
print(t.dtype) # torch.float32
# Concatenate along dim 0
a = torch.ones(2, 3)
b = torch.zeros(3, 3)
cat = torch.cat([a, b], dim=0) # shape (5, 3)
print(cat.shape)
# Remove size-1 dimensions
x = torch.randn(1, 5, 1)
print(x.squeeze().shape) # torch.Size([5])Quick Check
Test your understanding of Machine Learning with Python concepts from this lesson.
Lesson Recap
In this lesson you learned: tensors are PyTorch's core data structure supporting n-dimensional arrays on CPU or GPU, creation functions like torch.zeros, torch.randn, and torch.from_numpy give flexible ways to initialise data, and moving tensors to GPU with .to(device) is the key step to accelerate deep learning training. Next up we explore automatic differentiation with Autograd.
よくある質問
「PyTorchテンソル:作成、演算、GPUへの転送」レッスンは無料ですか?
はい。「PyTorchテンソル:作成、演算、GPUへの転送」の完全なテキストはこのウェブで無料で読めます。インタラクティブに演習し(組み込みコードエディタと24時間対応のAIチューター)、Machine Learning Academyコースの残りをアンロックするには、CoddyKit PROにアップグレードしてください。 Machine Learning Academyコースには全4レッスンが含まれています。
「PyTorchテンソル:作成、演算、GPUへの転送」で何を学びますか?
PythonリストとNumPy配列からテンソルを作成し、要素単位の演算と行列演算を行い、.to('cuda')でテンソルをGPUへ移動します。 ブラウザで直接実行するハンズオンコードでMachine Learning Academyを演習し、24時間対応のAIチューターがレッスンを進める中での質問に答えます。
Machine Learning Academyを始めるのに経験は必要ですか?
事前経験は必要ありません。CoddyKitのMachine Learning Academyは初級者から上級者向けに構成されているため、ここから始めるか最初から始めて、自分のペースで進むことができます。 これはレッスン1/4です。
「PyTorchテンソル:作成、演算、GPUへの転送」レッスンにはどのくらい時間がかかりますか?
ほとんどのCoddyKitレッスンは約5~10分かかります。各レッスンはコンパクトでインタラクティブなので、着実に進歩し、ウェブとアプリ全体で正確に前回の場所から再開できます。
このMachine Learning Academyレッスンでコードを書いて実行できますか?
はい。すべてのMachine Learning Academyレッスンに組み込みコードエディタが含まれているため、ブラウザでリアルコードを書いて実行し、即座のAIフィードバックを取得できます。ローカル設定は不要です。
このコースのすべてのレッスン
- PyTorchテンソル:作成、演算、GPUへの転送
- Autograd:逆伝播のための自動微分
- nn.Moduleによるフィードフォワードネットワークの構築
- 学習ループ:損失、オプティマイザー、エポック