設定用dictによるパイプラインのパラメーター化
ハードコードされたファイルパスや列名を実行時に渡す設定用dictに置き換え、パイプラインを再利用可能にします。
「設定用dictによるパイプラインのパラメーター化」はCoddyKit上の無料Pandas & NumPy Academyレッスンです。 これはレッスン2/4です。 下記で完全なレッスンを無料で読むことができます。その後、ブラウザ内の組み込みコードエディタと24時間対応のAIチューターでハンズオン演習できます。 これはPandas & NumPy Academy学習パスの一部であり、ウェブとCoddyKitアプリ全体で進捗が同期されます。 Pandas & NumPy Academyコースには全4レッスンが含まれています。
このレッスンの一部はまだ翻訳されておらず、英語で表示されています。
The Problem with Hardcoded Values
A pipeline with hardcoded file paths, column names, and threshold values breaks whenever the environment changes — a different server, a renamed column, or a changed business rule. Every change requires editing the pipeline code itself, creating risk of introducing bugs. The solution is to externalise all variable values into a configuration dictionary that is loaded at runtime and passed to the pipeline functions.
import pandas as pd
# BAD: hardcoded values scattered through code
df = pd.read_csv('/data/orders_2024.csv')
df = df.dropna(subset=['revenue', 'quantity'])
df = df[df['revenue'] < 5000]
df.to_parquet('/output/orders_clean.parquet')
print('Hardcoded paths and thresholds are fragile')Defining a Config Dictionary
Replace every hardcoded value with an entry in a configuration dictionary. Group related settings logically: input/output paths together, cleaning thresholds together, column name mappings together. The config dict becomes the single source of truth for all pipeline parameters. Changing one value in the config updates every function that uses it without touching the function bodies.
CONFIG = {
'input_path': '/data/orders_2024.csv',
'output_path': '/output/orders_clean.parquet',
'required_cols': ['order_id', 'order_date', 'revenue', 'quantity'],
'date_cols': ['order_date'],
'revenue_cap': 5000,
'min_quantity': 1,
'categorical_cols': ['region', 'category']
}
print('Config loaded:', list(CONFIG.keys()))Passing Config to Extract Functions
The extract function reads all its parameters from the config: the input path, the date columns to parse, and any encoding or delimiter settings. This means running the same pipeline against a test dataset or a different month's file requires only a config change — no code change. You can maintain separate configs for development, staging, and production environments.
def extract(config):
return pd.read_csv(
config['input_path'],
parse_dates=config.get('date_cols', [])
)
df = extract(CONFIG)
print('Extracted:', df.shape)Passing Config to Transform Functions
Each transformation function receives the full config and extracts the values it needs. Functions should use config.get('key', default) with sensible defaults so the pipeline is robust against incomplete configs. A function that requires a threshold of 5000 by default but can be overridden via config is both safe and flexible.
def transform(df, config):
required = config.get('required_cols', [])
cap = config.get('revenue_cap', float('inf'))
min_qty = config.get('min_quantity', 1)
return (
df
.dropna(subset=required)
.query(f'quantity >= {min_qty}')
.assign(revenue=lambda d: d['quantity'] * d['unit_price'])
.assign(revenue_capped=lambda d: d['revenue'].clip(upper=cap))
)
df_clean = transform(df, CONFIG)
print(df_clean.shape)Loading Config from a JSON File
For production pipelines, store the config in a JSON file rather than a Python dictionary hardcoded in the script. Load it with json.load() at pipeline start. This allows operations teams to change thresholds without access to the Python code, and enables config versioning through Git — every config change is a tracked commit with a description of the business reason for the change.
import json
# config.json would contain the same keys as CONFIG above
# with open('config.json') as f:
# config = json.load(f)
# Example: write and read back
with open('/tmp/pipeline_config.json', 'w') as f:
json.dump(CONFIG, f, indent=2)
with open('/tmp/pipeline_config.json') as f:
loaded_config = json.load(f)
print('Loaded config from JSON:', loaded_config['revenue_cap'])Environment-Specific Configs
Maintain separate config files for each environment: config_dev.json, config_staging.json, and config_prod.json. Determine which to load based on an environment variable. This pattern prevents accidental use of production file paths during development and keeps environment-specific secrets (like database credentials) out of the shared code repository.
import os
ENV = os.environ.get('PIPELINE_ENV', 'dev')
CONFIG_PATH = f'config_{ENV}.json'
# In practice:
# with open(CONFIG_PATH) as f:
# config = json.load(f)
print(f'Using config for environment: {ENV}')
print(f'Config file: {CONFIG_PATH}')Column Name Remapping via Config
Source data often has column names that differ from your internal naming convention. Rather than hard-coding df.rename(columns={'OrderDate': 'order_date', 'Qty': 'quantity'}) in the pipeline body, store the rename mapping in the config. This makes the pipeline agnostic to the source column names and easy to adapt when the upstream data provider changes their export format.
CONFIG['column_rename'] = {
'OrderDate': 'order_date',
'Qty': 'quantity',
'UnitPrice': 'unit_price',
'OrderID': 'order_id'
}
def rename_columns(df, config):
return df.rename(columns=config.get('column_rename', {}))
print('Column rename mapping stored in config.')Aggregation Config: Dynamic groupby Keys
The aggregation phase often groups by different columns for different use cases. Store the group keys and aggregation specs in the config rather than hardcoding them. This lets analysts produce different summary tables (by region, by category, by month) by changing the config, without touching the aggregation function. The function becomes a general-purpose aggregator driven entirely by configuration.
CONFIG['agg_spec'] = {
'group_by': ['region', 'category'],
'agg_cols': {
'revenue': ['sum', 'mean'],
'quantity': ['sum', 'count']
}
}
def aggregate(df, config):
spec = config['agg_spec']
return df.groupby(spec['group_by']).agg(spec['agg_cols'])
result = aggregate(df_clean, CONFIG)
print(result.head())Validating the Config at Startup
Validate the config at pipeline start to catch missing or invalid keys before any data is loaded. A pipeline that runs for 10 minutes and then fails because revenue_cap was a string instead of a float wastes time. Validate all required keys exist, values are the correct type, and paths are accessible with a short config-check function that runs before any expensive I/O.
def validate_config(config):
required_keys = ['input_path', 'output_path', 'required_cols']
for key in required_keys:
assert key in config, f'Config missing key: {key}'
assert isinstance(config['required_cols'], list), 'required_cols must be a list'
assert isinstance(config.get('revenue_cap', 1), (int, float)), 'revenue_cap must be numeric'
print('Config validation passed.')
validate_config(CONFIG)Merging Default Config with User Config
Allow users to provide a partial config that overrides only the values they care about. Merge the user config on top of a default config using {**defaults, **user_config}. This pattern provides sensible defaults while remaining fully configurable. It is the same pattern used by popular Python libraries that accept configuration as a dict or kwargs.
DEFAULT_CONFIG = {
'revenue_cap': 10000,
'min_quantity': 1,
'date_cols': ['order_date'],
'required_cols': ['order_id', 'revenue']
}
user_config = {'revenue_cap': 5000, 'input_path': '/data/q1.csv'}
final_config = {**DEFAULT_CONFIG, **user_config}
print('Final config revenue_cap:', final_config['revenue_cap']) # 5000
print('Final config min_quantity:', final_config['min_quantity']) # 1 (from default)Storing the Config with the Output
Save the config alongside the output file so anyone inspecting the output can immediately reproduce the pipeline run that created it. Store it as a JSON sidecar file with the same name as the output but a .config.json extension. Include the pipeline run timestamp in the saved config for full traceability of every output file.
import json
from datetime import datetime
def save_with_config(df, config):
output_path = config['output_path']
config_path = output_path.replace('.parquet', '.config.json')
run_metadata = {**config, 'run_at': datetime.now().isoformat()}
with open(config_path, 'w') as f:
json.dump(run_metadata, f, indent=2, default=str)
df.to_parquet(output_path, index=False)
print(f'Saved data to {output_path}')
print(f'Saved config to {config_path}')Quick Check
Test your understanding of Data Analysis concepts from this lesson.
Lesson Recap
In this lesson you learned: replacing hardcoded values with a config dictionary loaded from JSON, passing config to extract, transform, and aggregate functions for full parameterisation, and validating configs at startup and saving them alongside output files for reproducibility. Next up we explore testing pipeline steps with row-count checks and assertion guards.
よくある質問
「設定用dictによるパイプラインのパラメーター化」レッスンは無料ですか?
はい。「設定用dictによるパイプラインのパラメーター化」の完全なテキストはこのウェブで無料で読めます。インタラクティブに演習し(組み込みコードエディタと24時間対応のAIチューター)、Pandas & NumPy Academyコースの残りをアンロックするには、CoddyKit PROにアップグレードしてください。 Pandas & NumPy Academyコースには全4レッスンが含まれています。
「設定用dictによるパイプラインのパラメーター化」で何を学びますか?
ハードコードされたファイルパスや列名を実行時に渡す設定用dictに置き換え、パイプラインを再利用可能にします。 ブラウザで直接実行するハンズオンコードでPandas & NumPy Academyを演習し、24時間対応のAIチューターがレッスンを進める中での質問に答えます。
Pandas & NumPy Academyを始めるのに経験は必要ですか?
事前経験は必要ありません。CoddyKitのPandas & NumPy Academyは初級者から上級者向けに構成されているため、ここから始めるか最初から始めて、自分のペースで進むことができます。 これはレッスン2/4です。
「設定用dictによるパイプラインのパラメーター化」レッスンにはどのくらい時間がかかりますか?
ほとんどのCoddyKitレッスンは約5~10分かかります。各レッスンはコンパクトでインタラクティブなので、着実に進歩し、ウェブとアプリ全体で正確に前回の場所から再開できます。
このPandas & NumPy Academyレッスンでコードを書いて実行できますか?
はい。すべてのPandas & NumPy Academyレッスンに組み込みコードエディタが含まれているため、ブラウザでリアルコードを書いて実行し、即座のAIフィードバックを取得できます。ローカル設定は不要です。
このコースのすべてのレッスン
- 変換手順の関数化
- 設定用dictによるパイプラインのパラメーター化
- アサーションによるパイプライン手順のテスト
- パイプライン実行のスケジューリングとログ記録