0Pricing
Pandas & NumPy Academy · Lesson

Inspecting Column Data Types

Read the dtypes attribute, distinguish int64 from float64 and object, and spot columns that need conversion.

Inspecting Column Data Types is a free Pandas & NumPy Academy lesson on CoddyKit — lesson 1 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the Pandas & NumPy Academy learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.

Why Column Data Types Matter

Every column in a Pandas DataFrame has a dtype (data type) that determines how values are stored in memory and which operations are valid on them. An integer column with dtype object (string) cannot be summed. A date stored as a string is treated as text, not a time series. Incorrect dtypes are one of the most common causes of silent bugs and unexpected results in data analysis pipelines.

Always inspect dtypes as the very first step after loading a new dataset.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'age': ['25', '30', '35'],      # looks like int, stored as object
    'salary': [50000, 60000, 70000], # int64
    'hired': ['2022-01-01', '2023-06-15', '2024-03-10']  # looks like date
})

print(df.dtypes)
# age       object
# salary     int64
# hired     object
# dtype: object

The dtypes Attribute

DataFrame.dtypes returns a Series whose index is the column names and whose values are the Pandas dtype of each column. Common dtypes you will encounter are int64, float64, object (strings and mixed), bool, datetime64[ns], and category. The dtype name tells you both the kind of data and how many bytes each value uses.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'int_col': [1, 2, 3],
    'float_col': [1.5, 2.5, 3.5],
    'str_col': ['a', 'b', 'c'],
    'bool_col': [True, False, True]
})

print(df.dtypes)
# int_col      int64
# float_col  float64
# str_col     object
# bool_col      bool
# dtype: object

Identifying the object Dtype

The object dtype is Pandas' catch-all for columns that cannot be stored as a numeric or boolean type. It typically means string data, but it can also indicate a mixed-type column (e.g., integers mixed with strings). Object columns consume more memory than specialised dtypes and are slower to operate on. When you see object, ask: should this be a string, a number, a category, or a date?

import pandas as pd

df = pd.DataFrame({'mixed': [1, 'two', 3.0, None]})
print(df.dtypes)
# mixed    object

# Check actual Python types inside the column
print(df['mixed'].apply(type).value_counts())
# <class 'int'>      1
# <class 'str'>      1
# <class 'float'>    2  (includes NaN which is float)

info() for a Complete Type Overview

df.info() prints a concise table showing every column's name, non-null count, and dtype, plus total memory usage. This one command gives you a complete snapshot of dataset quality: you can see which columns have missing data and which have wrong dtypes at the same time, making it the standard first command after loading a dataset.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'id': [1, 2, 3],
    'score': [88.5, 92.0, np.nan],
    'grade': ['A', 'A', 'B']
})

df.info()
# <class 'pandas.core.frame.DataFrame'>
# RangeIndex: 3 entries, 0 to 2
# Data columns (total 3 columns):
#  #   Column  Non-Null Count  Dtype
# ---  ------  --------------  -----
#  0   id      3 non-null      int64
#  1   score   2 non-null      float64
#  2   grade   3 non-null      object
# dtypes: float64(1), int64(1), object(1)
# memory usage: 200.0+ bytes

Common dtype Problems After read_csv

When Pandas reads a CSV file, it infers dtypes by scanning values. Several common problems arise: numeric columns read as object if there are comma-thousand-separators or currency symbols; boolean columns read as object if stored as 'Yes'/'No' strings; and date columns read as object because Pandas doesn't parse dates unless told to. Recognising these patterns lets you fix them quickly with type conversion.

import pandas as pd
import io

csv = '''price,date,is_active
'1,200',2024-01-15,Yes
'800',2024-02-20,No
'''

df = pd.read_csv(io.StringIO(csv))
print(df.dtypes)
# price       object    <- should be int
# date        object    <- should be datetime64
# is_active   object    <- should be bool

Memory Usage of Different dtypes

Different dtypes consume vastly different amounts of memory. An int64 column uses 8 bytes per value, while an object column storing short strings may use 50-200 bytes per value. For DataFrames with millions of rows, choosing the right dtype can reduce memory from gigabytes to megabytes. Call df.memory_usage(deep=True) to see the byte cost per column.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'int64_col': np.arange(1_000_000, dtype='int64'),
    'float32_col': np.arange(1_000_000, dtype='float32'),
    'obj_col': ['a'] * 1_000_000
})

mem = df.memory_usage(deep=True) / 1024**2  # MB
print(mem.round(2))
# Index         0.00
# int64_col     7.63
# float32_col   3.81
# obj_col      55.26  <- much more for object!

Spotting Wrong-Typed Numeric Columns

A common issue is a numeric column stored as object because of stray formatting characters. You can detect this by checking if pd.to_numeric(df['col'], errors='coerce') produces a different result than the original column. Any values that cannot be parsed become NaN, revealing exactly which rows caused the type inference to fail.

import pandas as pd

df = pd.DataFrame({'revenue': ['1000', '2000', '$3000', '4,000']})

# Check which values can't be parsed as numbers
parsed = pd.to_numeric(df['revenue'], errors='coerce')
print(parsed)
# 0    1000.0
# 1    2000.0
# 2       NaN   <- '$3000' failed
# 3       NaN   <- '4,000' failed

Checking for Integer vs Float Ambiguity

When a column has any NaN values, Pandas stores integer columns as float64 because the standard NumPy integer types cannot represent NaN. The resulting floats like 25.0 look like integers but are stored as floats. Pandas introduced nullable integer types (Int64 with capital I) that support NaN while maintaining integer semantics.

import pandas as pd
import numpy as np

df = pd.DataFrame({'count': [1, 2, np.nan, 4]})
print(df['count'].dtype)  # float64 — because NaN is float

# Nullable integer type supports NaN
df['count'] = df['count'].astype('Int64')  # capital I!
print(df)
#    count
# 0      1
# 1      2
# 2   <NA>
# 3      4
print(df['count'].dtype)  # Int64

select_dtypes() for Type-Based Filtering

df.select_dtypes(include=) lets you select all columns of a particular type family. Common arguments: 'number' (all numeric), 'object' (strings), 'bool', 'datetime', 'category'. This is very useful at the start of a pipeline to apply numeric cleaning only to numeric columns and string cleaning only to text columns.

import pandas as pd

df = pd.DataFrame({
    'name': ['Alice', 'Bob'],
    'age': [25, 30],
    'salary': [50000.0, 70000.0],
    'dept': ['Eng', 'HR']
})

numeric = df.select_dtypes(include='number')
print('Numeric columns:', numeric.columns.tolist())
# ['age', 'salary']

text = df.select_dtypes(include='object')
print('Text columns:', text.columns.tolist())
# ['name', 'dept']

Building a dtype Report

A practical EDA habit is building a dtype report that lists each column's dtype, unique value count, and sample values. This helps you quickly identify columns that need conversion before analysis begins. You can generate such a report with a simple DataFrame constructed from column-wise aggregations.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'age': [25, 30, 35],
    'salary': [50000, 60000, 70000],
    'dept': ['Eng', 'HR', 'Eng'],
    'hired': ['2022-01-01', '2023-06-15', '2024-03-10']
})

report = pd.DataFrame({
    'dtype': df.dtypes,
    'unique_count': df.nunique(),
    'sample': df.iloc[0]
})
print(report)

Dtype Consistency in Production Pipelines

In production pipelines that process new data every day, dtypes can drift — a column that was always int64 might arrive as float64 if a single NaN appears. Add dtype assertions at the start of the pipeline to catch schema changes early. Failing loudly with a clear error is far better than silently producing wrong results downstream.

import pandas as pd

df = pd.DataFrame({'user_id': [1, 2, 3], 'score': [88.0, 92.0, 76.0]})

expected_dtypes = {'user_id': 'int64', 'score': 'float64'}

for col, expected in expected_dtypes.items():
    actual = str(df[col].dtype)
    assert actual == expected, (
        f'Column {col}: expected {expected}, got {actual}'
    )
print('All dtypes are correct')

Quick Check

Test your understanding of inspecting column data types.

Lesson Recap

In this lesson you learned: df.dtypes shows each column's type, object dtype is the catch-all for strings and mixed types, and df.info() gives a complete type+nullability overview. Use select_dtypes() to filter columns by type family and add dtype assertions to detect schema drift in production. Next up we convert columns to the correct dtypes using astype().

Frequently asked questions

Is the “Inspecting Column Data Types” lesson free?

Yes — the full text of “Inspecting Column Data Types” is free to read here on the web, and the Pandas & NumPy Academy course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the Pandas & NumPy Academy course, upgrade to CoddyKit PRO.

What will I learn in “Inspecting Column Data Types”?

Read the dtypes attribute, distinguish int64 from float64 and object, and spot columns that need conversion. You practise Pandas & NumPy Academy with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.

Do I need any experience to start Pandas & NumPy Academy?

No prior experience is required. Pandas & NumPy Academy on CoddyKit is structured for beginners through advanced learners; this is — lesson 1 of 4, so you can start here or from the beginning and move at your own pace.

How long does the “Inspecting Column Data Types” lesson take?

Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.

Can I write and run code in this Pandas & NumPy Academy lesson?

Yes. Every Pandas & NumPy Academy lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.

All lessons in this course

  1. Inspecting Column Data Types
  2. Casting with astype()
  3. Categorical Data Type
  4. Parsing Dates Correctly
← Back to Pandas & NumPy Academy