Paramétrer les pipelines avec des dictionnaires de configuration
Remplacez les chemins de fichiers et les noms de colonnes codés en dur par un dictionnaire de configuration transmis à l’exécution afin de rendre le pipeline réutilisable.
Paramétrer les pipelines avec des dictionnaires de configuration est une leçon Pandas & NumPy Academy gratuite sur CoddyKit. Ceci est la leçon 2 sur 4. Tu peux lire la leçon complète ci-dessous gratuitement — puis la pratiquer en direct dans le navigateur avec un éditeur de code intégré et un tuteur IA 24/7. Elle fait partie du parcours d'apprentissage Pandas & NumPy Academy, et ta progression se synchronise sur le web et l'application CoddyKit. Le cours Pandas & NumPy Academy comprend 4 leçons au total.
Certaines parties de cette leçon n'ont pas encore été traduites et s'affichent en anglais.
The Problem with Hardcoded Values
A pipeline with hardcoded file paths, column names, and threshold values breaks whenever the environment changes — a different server, a renamed column, or a changed business rule. Every change requires editing the pipeline code itself, creating risk of introducing bugs. The solution is to externalise all variable values into a configuration dictionary that is loaded at runtime and passed to the pipeline functions.
import pandas as pd
# BAD: hardcoded values scattered through code
df = pd.read_csv('/data/orders_2024.csv')
df = df.dropna(subset=['revenue', 'quantity'])
df = df[df['revenue'] < 5000]
df.to_parquet('/output/orders_clean.parquet')
print('Hardcoded paths and thresholds are fragile')Defining a Config Dictionary
Replace every hardcoded value with an entry in a configuration dictionary. Group related settings logically: input/output paths together, cleaning thresholds together, column name mappings together. The config dict becomes the single source of truth for all pipeline parameters. Changing one value in the config updates every function that uses it without touching the function bodies.
CONFIG = {
'input_path': '/data/orders_2024.csv',
'output_path': '/output/orders_clean.parquet',
'required_cols': ['order_id', 'order_date', 'revenue', 'quantity'],
'date_cols': ['order_date'],
'revenue_cap': 5000,
'min_quantity': 1,
'categorical_cols': ['region', 'category']
}
print('Config loaded:', list(CONFIG.keys()))Passing Config to Extract Functions
The extract function reads all its parameters from the config: the input path, the date columns to parse, and any encoding or delimiter settings. This means running the same pipeline against a test dataset or a different month's file requires only a config change — no code change. You can maintain separate configs for development, staging, and production environments.
def extract(config):
return pd.read_csv(
config['input_path'],
parse_dates=config.get('date_cols', [])
)
df = extract(CONFIG)
print('Extracted:', df.shape)Passing Config to Transform Functions
Each transformation function receives the full config and extracts the values it needs. Functions should use config.get('key', default) with sensible defaults so the pipeline is robust against incomplete configs. A function that requires a threshold of 5000 by default but can be overridden via config is both safe and flexible.
def transform(df, config):
required = config.get('required_cols', [])
cap = config.get('revenue_cap', float('inf'))
min_qty = config.get('min_quantity', 1)
return (
df
.dropna(subset=required)
.query(f'quantity >= {min_qty}')
.assign(revenue=lambda d: d['quantity'] * d['unit_price'])
.assign(revenue_capped=lambda d: d['revenue'].clip(upper=cap))
)
df_clean = transform(df, CONFIG)
print(df_clean.shape)Loading Config from a JSON File
For production pipelines, store the config in a JSON file rather than a Python dictionary hardcoded in the script. Load it with json.load() at pipeline start. This allows operations teams to change thresholds without access to the Python code, and enables config versioning through Git — every config change is a tracked commit with a description of the business reason for the change.
import json
# config.json would contain the same keys as CONFIG above
# with open('config.json') as f:
# config = json.load(f)
# Example: write and read back
with open('/tmp/pipeline_config.json', 'w') as f:
json.dump(CONFIG, f, indent=2)
with open('/tmp/pipeline_config.json') as f:
loaded_config = json.load(f)
print('Loaded config from JSON:', loaded_config['revenue_cap'])Environment-Specific Configs
Maintain separate config files for each environment: config_dev.json, config_staging.json, and config_prod.json. Determine which to load based on an environment variable. This pattern prevents accidental use of production file paths during development and keeps environment-specific secrets (like database credentials) out of the shared code repository.
import os
ENV = os.environ.get('PIPELINE_ENV', 'dev')
CONFIG_PATH = f'config_{ENV}.json'
# In practice:
# with open(CONFIG_PATH) as f:
# config = json.load(f)
print(f'Using config for environment: {ENV}')
print(f'Config file: {CONFIG_PATH}')Column Name Remapping via Config
Source data often has column names that differ from your internal naming convention. Rather than hard-coding df.rename(columns={'OrderDate': 'order_date', 'Qty': 'quantity'}) in the pipeline body, store the rename mapping in the config. This makes the pipeline agnostic to the source column names and easy to adapt when the upstream data provider changes their export format.
CONFIG['column_rename'] = {
'OrderDate': 'order_date',
'Qty': 'quantity',
'UnitPrice': 'unit_price',
'OrderID': 'order_id'
}
def rename_columns(df, config):
return df.rename(columns=config.get('column_rename', {}))
print('Column rename mapping stored in config.')Aggregation Config: Dynamic groupby Keys
The aggregation phase often groups by different columns for different use cases. Store the group keys and aggregation specs in the config rather than hardcoding them. This lets analysts produce different summary tables (by region, by category, by month) by changing the config, without touching the aggregation function. The function becomes a general-purpose aggregator driven entirely by configuration.
CONFIG['agg_spec'] = {
'group_by': ['region', 'category'],
'agg_cols': {
'revenue': ['sum', 'mean'],
'quantity': ['sum', 'count']
}
}
def aggregate(df, config):
spec = config['agg_spec']
return df.groupby(spec['group_by']).agg(spec['agg_cols'])
result = aggregate(df_clean, CONFIG)
print(result.head())Validating the Config at Startup
Validate the config at pipeline start to catch missing or invalid keys before any data is loaded. A pipeline that runs for 10 minutes and then fails because revenue_cap was a string instead of a float wastes time. Validate all required keys exist, values are the correct type, and paths are accessible with a short config-check function that runs before any expensive I/O.
def validate_config(config):
required_keys = ['input_path', 'output_path', 'required_cols']
for key in required_keys:
assert key in config, f'Config missing key: {key}'
assert isinstance(config['required_cols'], list), 'required_cols must be a list'
assert isinstance(config.get('revenue_cap', 1), (int, float)), 'revenue_cap must be numeric'
print('Config validation passed.')
validate_config(CONFIG)Merging Default Config with User Config
Allow users to provide a partial config that overrides only the values they care about. Merge the user config on top of a default config using {**defaults, **user_config}. This pattern provides sensible defaults while remaining fully configurable. It is the same pattern used by popular Python libraries that accept configuration as a dict or kwargs.
DEFAULT_CONFIG = {
'revenue_cap': 10000,
'min_quantity': 1,
'date_cols': ['order_date'],
'required_cols': ['order_id', 'revenue']
}
user_config = {'revenue_cap': 5000, 'input_path': '/data/q1.csv'}
final_config = {**DEFAULT_CONFIG, **user_config}
print('Final config revenue_cap:', final_config['revenue_cap']) # 5000
print('Final config min_quantity:', final_config['min_quantity']) # 1 (from default)Storing the Config with the Output
Save the config alongside the output file so anyone inspecting the output can immediately reproduce the pipeline run that created it. Store it as a JSON sidecar file with the same name as the output but a .config.json extension. Include the pipeline run timestamp in the saved config for full traceability of every output file.
import json
from datetime import datetime
def save_with_config(df, config):
output_path = config['output_path']
config_path = output_path.replace('.parquet', '.config.json')
run_metadata = {**config, 'run_at': datetime.now().isoformat()}
with open(config_path, 'w') as f:
json.dump(run_metadata, f, indent=2, default=str)
df.to_parquet(output_path, index=False)
print(f'Saved data to {output_path}')
print(f'Saved config to {config_path}')Quick Check
Test your understanding of Data Analysis concepts from this lesson.
Lesson Recap
In this lesson you learned: replacing hardcoded values with a config dictionary loaded from JSON, passing config to extract, transform, and aggregate functions for full parameterisation, and validating configs at startup and saving them alongside output files for reproducibility. Next up we explore testing pipeline steps with row-count checks and assertion guards.
Questions Fréquemment Posées
La leçon « Paramétrer les pipelines avec des dictionnaires de configuration » est-elle gratuite ?
Oui — le texte complet de « Paramétrer les pipelines avec des dictionnaires de configuration » est gratuit à lire ici sur le web. Pour la pratiquer de manière interactive (un éditeur de code intégré et un tuteur IA 24/7) et déverrouiller le reste du cours Pandas & NumPy Academy, passe à CoddyKit PRO. Le cours Pandas & NumPy Academy comprend 4 leçons au total.
Qu'est-ce que j'apprendrai dans « Paramétrer les pipelines avec des dictionnaires de configuration » ?
Remplacez les chemins de fichiers et les noms de colonnes codés en dur par un dictionnaire de configuration transmis à l’exécution afin de rendre le pipeline réutilisable. Tu pratiques Pandas & NumPy Academy avec du code pratique que tu exécutes directement dans le navigateur, et un tuteur IA 24/7 répond à tes questions au fur et à mesure que tu avances dans la leçon.
Dois-je avoir de l'expérience pour commencer Pandas & NumPy Academy ?
Aucune expérience préalable n'est requise. Pandas & NumPy Academy sur CoddyKit est structuré pour les débutants jusqu'aux apprenants avancés, donc tu peux commencer ici ou depuis le début et avancer à ton rythme. Ceci est la leçon 2 sur 4.
Combien de temps prend la leçon « Paramétrer les pipelines avec des dictionnaires de configuration » ?
La plupart des leçons CoddyKit prennent environ 5–10 minutes. Chacune est courte et interactive, tu progresses régulièrement et tu repiques exactement où tu t'es arrêté sur le web et l'app.
Peux-tu écrire et exécuter du code dans cette leçon Pandas & NumPy Academy ?
Oui. Chaque leçon Pandas & NumPy Academy inclut un éditeur de code intégré, tu écris et exécutes du vrai code directement dans ton navigateur et tu reçois des retours IA instantanés — aucune configuration locale requise.
Toutes les leçons de ce cours
- Structurer les étapes de transformation en fonctions
- Paramétrer les pipelines avec des dictionnaires de configuration
- Tester les étapes du pipeline avec des assertions
- Planifier et journaliser les exécutions du pipeline