0Pricing
Pandas & NumPy Academy · Урок

Создание DataFrames

Создавайте DataFrames из словарей списков, списков словарей и массивов NumPy, а затем изучайте их форму, столбцы и типы данных

«Создание DataFrames» — бесплатный урок Pandas & NumPy Academy на CoddyKit. Это урок 1 из 4. Ты можешь прочитать весь урок бесплатно ниже — а потом практиковать его прямо в браузере с встроенным редактором кода и ИИ-репетитором 24/7. Это часть пути обучения Pandas & NumPy Academy, и твой прогресс синхронизируется между веб-версией и приложением CoddyKit. Курс Pandas & NumPy Academy содержит 4 уроков всего.

Части этого урока еще не переведены и отображаются на английском.

What Is a DataFrame?

A DataFrame is a two-dimensional, size-mutable, labelled data structure — essentially a table with rows and columns, similar to a spreadsheet or SQL table. Each column is a Pandas Series sharing the same row index. DataFrames can hold columns of different dtypes, making them perfect for real-world datasets that mix numbers, text, and dates.

import pandas as pd

df = pd.DataFrame({'name': ['Alice', 'Bob', 'Carol'],
                   'age': [30, 25, 35],
                   'score': [88.5, 72.0, 95.3]})
print(df)

Creating from a Dict of Lists

The most common way to construct a DataFrame is from a dict of lists: the keys become column names and the lists become column values. All lists must be the same length. The row index defaults to 0, 1, 2, ... unless you specify one with index=. This pattern mirrors how CSV data is often represented in Python.

import pandas as pd

df = pd.DataFrame({
    'city': ['Berlin', 'Paris', 'Rome'],
    'pop': [3.6, 2.2, 2.8],
    'country': ['DE', 'FR', 'IT']
})
print(df.shape)    # (3, 3)
print(df.dtypes)

Creating from a List of Dicts

You can also pass a list of dicts, where each dict represents one row. Missing keys in any dict produce NaN for that column. This format is common when consuming JSON APIs that return a list of record objects. Pandas aligns all keys across dicts to form the column set automatically.

import pandas as pd

rows = [
    {'name': 'Alice', 'age': 30, 'dept': 'Eng'},
    {'name': 'Bob',   'age': 25},              # 'dept' missing -> NaN
    {'name': 'Carol', 'age': 35, 'dept': 'HR'}
]
df = pd.DataFrame(rows)
print(df)

Creating from a NumPy Array

Passing a 2-D NumPy array creates a DataFrame with integer column names (0, 1, 2, ...) and a RangeIndex by default. Supply columns= and index= to add meaningful labels. This is the bridge between NumPy numerical computation and Pandas labelled data analysis.

import pandas as pd
import numpy as np

arr = np.array([[1, 2, 3], [4, 5, 6], [7, 8, 9]])
df = pd.DataFrame(arr,
                  columns=['x', 'y', 'z'],
                  index=['r1', 'r2', 'r3'])
print(df)

Setting a Custom Index at Creation

The index= parameter sets the row labels at creation time. Common choices include date strings, IDs, or category names that make row lookups meaningful. A well-chosen index makes .loc access intuitive and enables powerful time-series operations when the index is a DatetimeIndex.

import pandas as pd

df = pd.DataFrame(
    {'revenue': [100, 200, 150], 'costs': [80, 160, 90]},
    index=['Jan', 'Feb', 'Mar']
)
print(df)
print(df.loc['Feb'])  # select row 'Feb'

Inspecting columns, dtypes, and index

Three attributes immediately tell you the structure of a DataFrame: df.columns gives an Index of column names, df.dtypes returns a Series mapping each column to its dtype, and df.index describes the row labels. These are the first checks to run on any new DataFrame to understand what data you have before transforming it.

import pandas as pd

df = pd.DataFrame({'a': [1, 2], 'b': [3.0, 4.0], 'c': ['x', 'y']})
print(df.columns)  # Index(['a', 'b', 'c'], dtype='object')
print(df.dtypes)
# a      int64
# b    float64
# c     object
print(df.index)    # RangeIndex(start=0, stop=2, step=1)

Specifying Column Order

When constructing from a dict, column order follows dict insertion order (Python 3.7+). If you need a specific order, pass the columns= parameter with a list of column names. Any name not in the data will produce a NaN column; any name in the data but not in the list will be excluded. This lets you select and order columns at construction time.

import pandas as pd

data = {'c': [3, 6], 'a': [1, 4], 'b': [2, 5]}
df = pd.DataFrame(data, columns=['a', 'b', 'c'])
print(df.columns.tolist())  # ['a', 'b', 'c']

Creating from a Dict of Series

Passing a dict of Pandas Series aligns on the union of all indices. Where a Series is missing a label that another has, the result cell is NaN. This is the most index-aware construction method and is used when combining separately computed columns that may have different row counts or labels.

import pandas as pd

s1 = pd.Series([1, 2, 3], index=['a', 'b', 'c'])
s2 = pd.Series([10, 20], index=['b', 'c'])
df = pd.DataFrame({'col1': s1, 'col2': s2})
print(df)
#    col1  col2
# a   1.0   NaN
# b   2.0  10.0
# c   3.0  20.0

shape, len, and size

df.shape returns a tuple (nrows, ncols). len(df) returns the number of rows. df.size returns the total number of cells (rows × columns). These are the quickest sanity checks after loading data to confirm you received the expected number of records and that no columns were silently dropped.

import pandas as pd

df = pd.DataFrame({'a': range(5), 'b': range(5), 'c': range(5)})
print(df.shape)  # (5, 3)
print(len(df))   # 5
print(df.size)   # 15  (5 rows x 3 cols)

Creating an Empty DataFrame

You can create an empty DataFrame with specific column names and dtypes for use as a template or accumulator. Append rows with pd.concat([df, new_row])m. Specifying dtypes at creation avoids expensive type inference when rows are added later. An empty DataFrame is also useful as a seed in iterative data collection loops.

import pandas as pd

df = pd.DataFrame(columns=['name', 'score'])
new_row = pd.DataFrame([{'name': 'Alice', 'score': 90}])
df = pd.concat([df, new_row], ignore_index=True)
print(df)

Copying a DataFrame

Assigning a DataFrame to a new variable creates a reference, not a copy — modifying one modifies the other. Use df.copy() to create an independent copy. This is important before making transformations you do not want to apply to the original, or when you want to keep a pre-cleaning backup while building a cleaned version.

import pandas as pd

original = pd.DataFrame({'a': [1, 2, 3]})
ref = original          # same object!
copy = original.copy()  # independent copy

ref['a'] = 99
print(original['a'].tolist())  # [99 99 99] -- ref modified original
print(copy['a'].tolist())      # [1, 2, 3]  -- copy unchanged

Quick Check

Test your understanding of creating DataFrames from this lesson.

Lesson Recap

In this lesson you learned: DataFrames can be created from dicts of lists, lists of dicts, or NumPy arrays, the columns, dtypes, and index attributes describe the table structure, and df.copy() is required to get an independent copy rather than a reference. Next up we select specific columns and rows using bracket notation, .loc, and .iloc.

Часто задаваемые вопросы

Урок «Создание DataFrames» бесплатный?

Да — полный текст урока «Создание DataFrames» бесплатно доступен здесь в веб-версии. Чтобы практиковать его интерактивно (встроенный редактор кода и ИИ-репетитор 24/7) и разблокировать остальной курс Pandas & NumPy Academy, подпишись на CoddyKit PRO. Курс Pandas & NumPy Academy содержит 4 уроков всего.

Чему я научусь в уроке «Создание DataFrames»?

Создавайте DataFrames из словарей списков, списков словарей и массивов NumPy, а затем изучайте их форму, столбцы и типы данных Ты практикуешь Pandas & NumPy Academy с помощью реального кода, который запускаешь прямо в браузере, и ИИ-репетитор 24/7 отвечает на твои вопросы во время урока.

Нужен ли мне опыт, чтобы начать Pandas & NumPy Academy?

Предыдущий опыт не требуется. Pandas & NumPy Academy на CoddyKit структурирован для всех уровней — от новичков до продвинутых, поэтому ты можешь начать отсюда или с самого начала и учиться в своем темпе. Это урок 1 из 4.

Сколько времени занимает урок «Создание DataFrames»?

Большинство уроков CoddyKit занимают около 5–10 минут. Каждый из них компактный и интерактивный, поэтому ты постоянно делаешь прогресс и продолжаешь с того же места в веб-версии и приложении.

Можно ли писать и запускать код в этом уроке Pandas & NumPy Academy?

Да. Каждый урок Pandas & NumPy Academy включает встроенный редактор кода, поэтому ты пишешь и запускаешь реальный код прямо в браузере и получаешь моментальную обратную связь от AI — локальная установка не требуется.

Все уроки этого курса

  1. Создание DataFrames
  2. Выбор столбцов и строк
  3. Добавление и удаление столбцов
  4. Базовая проверка DataFrame
← Назад к Pandas & NumPy Academy