0Pricing
Pandas & NumPy Academy · Lesson

Parameterising Pipelines with Config Dicts

Replace hardcoded file paths and column names with a config dict passed at runtime to make the pipeline reusable.

Parameterising Pipelines with Config Dicts is a free Pandas & NumPy Academy lesson on CoddyKit — lesson 2 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the Pandas & NumPy Academy learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.

The Problem with Hardcoded Values

A pipeline with hardcoded file paths, column names, and threshold values breaks whenever the environment changes — a different server, a renamed column, or a changed business rule. Every change requires editing the pipeline code itself, creating risk of introducing bugs. The solution is to externalise all variable values into a configuration dictionary that is loaded at runtime and passed to the pipeline functions.

import pandas as pd

# BAD: hardcoded values scattered through code
df = pd.read_csv('/data/orders_2024.csv')
df = df.dropna(subset=['revenue', 'quantity'])
df = df[df['revenue'] < 5000]
df.to_parquet('/output/orders_clean.parquet')
print('Hardcoded paths and thresholds are fragile')

Defining a Config Dictionary

Replace every hardcoded value with an entry in a configuration dictionary. Group related settings logically: input/output paths together, cleaning thresholds together, column name mappings together. The config dict becomes the single source of truth for all pipeline parameters. Changing one value in the config updates every function that uses it without touching the function bodies.

CONFIG = {
    'input_path': '/data/orders_2024.csv',
    'output_path': '/output/orders_clean.parquet',
    'required_cols': ['order_id', 'order_date', 'revenue', 'quantity'],
    'date_cols': ['order_date'],
    'revenue_cap': 5000,
    'min_quantity': 1,
    'categorical_cols': ['region', 'category']
}
print('Config loaded:', list(CONFIG.keys()))

Passing Config to Extract Functions

The extract function reads all its parameters from the config: the input path, the date columns to parse, and any encoding or delimiter settings. This means running the same pipeline against a test dataset or a different month's file requires only a config change — no code change. You can maintain separate configs for development, staging, and production environments.

def extract(config):
    return pd.read_csv(
        config['input_path'],
        parse_dates=config.get('date_cols', [])
    )

df = extract(CONFIG)
print('Extracted:', df.shape)

Passing Config to Transform Functions

Each transformation function receives the full config and extracts the values it needs. Functions should use config.get('key', default) with sensible defaults so the pipeline is robust against incomplete configs. A function that requires a threshold of 5000 by default but can be overridden via config is both safe and flexible.

def transform(df, config):
    required = config.get('required_cols', [])
    cap = config.get('revenue_cap', float('inf'))
    min_qty = config.get('min_quantity', 1)

    return (
        df
        .dropna(subset=required)
        .query(f'quantity >= {min_qty}')
        .assign(revenue=lambda d: d['quantity'] * d['unit_price'])
        .assign(revenue_capped=lambda d: d['revenue'].clip(upper=cap))
    )

df_clean = transform(df, CONFIG)
print(df_clean.shape)

Loading Config from a JSON File

For production pipelines, store the config in a JSON file rather than a Python dictionary hardcoded in the script. Load it with json.load() at pipeline start. This allows operations teams to change thresholds without access to the Python code, and enables config versioning through Git — every config change is a tracked commit with a description of the business reason for the change.

import json

# config.json would contain the same keys as CONFIG above
# with open('config.json') as f:
#     config = json.load(f)

# Example: write and read back
with open('/tmp/pipeline_config.json', 'w') as f:
    json.dump(CONFIG, f, indent=2)

with open('/tmp/pipeline_config.json') as f:
    loaded_config = json.load(f)

print('Loaded config from JSON:', loaded_config['revenue_cap'])

Environment-Specific Configs

Maintain separate config files for each environment: config_dev.json, config_staging.json, and config_prod.json. Determine which to load based on an environment variable. This pattern prevents accidental use of production file paths during development and keeps environment-specific secrets (like database credentials) out of the shared code repository.

import os

ENV = os.environ.get('PIPELINE_ENV', 'dev')
CONFIG_PATH = f'config_{ENV}.json'

# In practice:
# with open(CONFIG_PATH) as f:
#     config = json.load(f)

print(f'Using config for environment: {ENV}')
print(f'Config file: {CONFIG_PATH}')

Column Name Remapping via Config

Source data often has column names that differ from your internal naming convention. Rather than hard-coding df.rename(columns={'OrderDate': 'order_date', 'Qty': 'quantity'}) in the pipeline body, store the rename mapping in the config. This makes the pipeline agnostic to the source column names and easy to adapt when the upstream data provider changes their export format.

CONFIG['column_rename'] = {
    'OrderDate': 'order_date',
    'Qty': 'quantity',
    'UnitPrice': 'unit_price',
    'OrderID': 'order_id'
}

def rename_columns(df, config):
    return df.rename(columns=config.get('column_rename', {}))

print('Column rename mapping stored in config.')

Aggregation Config: Dynamic groupby Keys

The aggregation phase often groups by different columns for different use cases. Store the group keys and aggregation specs in the config rather than hardcoding them. This lets analysts produce different summary tables (by region, by category, by month) by changing the config, without touching the aggregation function. The function becomes a general-purpose aggregator driven entirely by configuration.

CONFIG['agg_spec'] = {
    'group_by': ['region', 'category'],
    'agg_cols': {
        'revenue': ['sum', 'mean'],
        'quantity': ['sum', 'count']
    }
}

def aggregate(df, config):
    spec = config['agg_spec']
    return df.groupby(spec['group_by']).agg(spec['agg_cols'])

result = aggregate(df_clean, CONFIG)
print(result.head())

Validating the Config at Startup

Validate the config at pipeline start to catch missing or invalid keys before any data is loaded. A pipeline that runs for 10 minutes and then fails because revenue_cap was a string instead of a float wastes time. Validate all required keys exist, values are the correct type, and paths are accessible with a short config-check function that runs before any expensive I/O.

def validate_config(config):
    required_keys = ['input_path', 'output_path', 'required_cols']
    for key in required_keys:
        assert key in config, f'Config missing key: {key}'
    assert isinstance(config['required_cols'], list), 'required_cols must be a list'
    assert isinstance(config.get('revenue_cap', 1), (int, float)), 'revenue_cap must be numeric'
    print('Config validation passed.')

validate_config(CONFIG)

Merging Default Config with User Config

Allow users to provide a partial config that overrides only the values they care about. Merge the user config on top of a default config using {**defaults, **user_config}. This pattern provides sensible defaults while remaining fully configurable. It is the same pattern used by popular Python libraries that accept configuration as a dict or kwargs.

DEFAULT_CONFIG = {
    'revenue_cap': 10000,
    'min_quantity': 1,
    'date_cols': ['order_date'],
    'required_cols': ['order_id', 'revenue']
}

user_config = {'revenue_cap': 5000, 'input_path': '/data/q1.csv'}

final_config = {**DEFAULT_CONFIG, **user_config}
print('Final config revenue_cap:', final_config['revenue_cap'])  # 5000
print('Final config min_quantity:', final_config['min_quantity'])  # 1 (from default)

Storing the Config with the Output

Save the config alongside the output file so anyone inspecting the output can immediately reproduce the pipeline run that created it. Store it as a JSON sidecar file with the same name as the output but a .config.json extension. Include the pipeline run timestamp in the saved config for full traceability of every output file.

import json
from datetime import datetime

def save_with_config(df, config):
    output_path = config['output_path']
    config_path = output_path.replace('.parquet', '.config.json')

    run_metadata = {**config, 'run_at': datetime.now().isoformat()}
    with open(config_path, 'w') as f:
        json.dump(run_metadata, f, indent=2, default=str)

    df.to_parquet(output_path, index=False)
    print(f'Saved data to {output_path}')
    print(f'Saved config to {config_path}')

Quick Check

Test your understanding of Data Analysis concepts from this lesson.

Lesson Recap

In this lesson you learned: replacing hardcoded values with a config dictionary loaded from JSON, passing config to extract, transform, and aggregate functions for full parameterisation, and validating configs at startup and saving them alongside output files for reproducibility. Next up we explore testing pipeline steps with row-count checks and assertion guards.

Frequently asked questions

Is the “Parameterising Pipelines with Config Dicts” lesson free?

Yes — the full text of “Parameterising Pipelines with Config Dicts” is free to read here on the web, and the Pandas & NumPy Academy course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the Pandas & NumPy Academy course, upgrade to CoddyKit PRO.

What will I learn in “Parameterising Pipelines with Config Dicts”?

Replace hardcoded file paths and column names with a config dict passed at runtime to make the pipeline reusable. You practise Pandas & NumPy Academy with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.

Do I need any experience to start Pandas & NumPy Academy?

No prior experience is required. Pandas & NumPy Academy on CoddyKit is structured for beginners through advanced learners; this is — lesson 2 of 4, so you can start here or from the beginning and move at your own pace.

How long does the “Parameterising Pipelines with Config Dicts” lesson take?

Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.

Can I write and run code in this Pandas & NumPy Academy lesson?

Yes. Every Pandas & NumPy Academy lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.

All lessons in this course

  1. Structuring Transformation Steps as Functions
  2. Parameterising Pipelines with Config Dicts
  3. Testing Pipeline Steps with Assertions
  4. Scheduling and Logging Pipeline Runs
← Back to Pandas & NumPy Academy