Avoiding Data Leakage with Scikit-learn Pipelines

A model can score exceptionally well in a notebook and still fail on genuinely new data. One common reason is data leakage: information that would not be available at prediction time has influenced training or evaluation.

Leakage is not always an obvious target column accidentally left in a feature matrix. It can enter through preprocessing fitted on the full dataset, features created after the outcome occurred, duplicated entities across a split, or tuning decisions repeatedly made against a test set.

Scikit-learn pipelines help prevent an important class of leakage by keeping transformations and estimators together. They do not automatically fix a bad split or a feature that encodes the answer, so the evaluation design still matters.

Split observations before fitting preprocessing; each training fold learns its own transformation, which is then applied to held-out observations.

What information is leaking?

It helps to distinguish two related problems:

  • Preprocessing leakage: a transformation learns from validation or test examples. For example, a scaler estimates its mean and variance using all rows before the train-test split.
  • Feature or target leakage: a feature contains information unavailable at the moment a real prediction would be made. For example, using a claim-resolution code to predict whether a claim will later be denied.

A pipeline can address the first problem when it is fitted correctly. It cannot determine whether a feature is legitimate for the intended prediction time; that requires understanding how the data was generated.

Split first, fit later

For a conventional binary classification dataset with independent rows, create the holdout split before learning any data-dependent preprocessing. Put imputers, encoders, feature selectors, and the estimator in a single pipeline:

from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler

numeric_features = ["age", "account_value"]
categorical_features = ["region", "plan"]

X = data[numeric_features + categorical_features]
y = data["renewed"]

X_train, X_test, y_train, y_test = train_test_split(
    X,
    y,
    test_size=0.2,
    random_state=42,
    stratify=y,
)

numeric_pipeline = Pipeline([
    ("imputer", SimpleImputer(strategy="median")),
    ("scaler", StandardScaler()),
])

categorical_pipeline = Pipeline([
    ("imputer", SimpleImputer(strategy="most_frequent")),
    ("encoder", OneHotEncoder(handle_unknown="ignore")),
])

preprocessor = ColumnTransformer([
    ("numeric", numeric_pipeline, numeric_features),
    ("categorical", categorical_pipeline, categorical_features),
])

model = Pipeline([
    ("preprocessor", preprocessor),
    ("classifier", LogisticRegression(max_iter=1000)),
])

model.fit(X_train, y_train)
test_score = model.score(X_test, y_test)

The imputer statistics, scaler values, and category vocabulary are learned during model.fit from X_train. At evaluation time, model.score applies those already-fitted transformations to X_test; it does not fit them again.

stratify=y preserves approximate class proportions in this split; it does not make observations independent or make random splitting appropriate for every dataset. If rows are related by person, device, household, or another entity, use a group-aware split. If prediction is about future periods, split chronologically. Choose the split to match how the model will be used.

Validate the whole pipeline during cross-validation

During cross-validation, pass the pipeline itself to the validation function. Each training fold then fits its own preprocessing steps, and the corresponding validation fold is transformed without contributing to those fitted statistics.

from sklearn.model_selection import StratifiedKFold, cross_validate

cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
results = cross_validate(
    model,
    X_train,
    y_train,
    cv=cv,
    scoring=["accuracy", "f1"],
    return_train_score=False,
)

print(results["test_f1"].mean(), results["test_f1"].std())

Cross-validation estimates performance under the split strategy you chose. It does not eliminate leakage from duplicate entities, time-dependent data, post-outcome fields, or preprocessing performed before the pipeline. If you tune hyperparameters, keep the final test set out of that loop; use nested cross-validation when an unbiased model-selection estimate is required for a small dataset.

Audit the feature timeline

For every feature, ask: Would this value be known at the exact time the model is expected to make its prediction? Document when the prediction happens and when each input becomes available.

Common red flags include:

  • Status fields updated after the target event.
  • Aggregates calculated using future records or the full observation period.
  • IDs or timestamps that encode a collection process correlated with the label.
  • The same customer, patient, or device appearing on both sides of a split when deployment targets unseen entities.
  • Duplicate or near-duplicate rows crossing the train and test boundary.
  • Feature engineering that groups or normalizes using statistics computed over all data.

Some features are valid in one prediction scenario and leakage in another. A status code may be legitimate for predicting a later event if it is available beforehand; it is leakage if it is written after the outcome. The data dictionary alone may not establish that chronology, so confirm it with the team that creates the field.

Keep the test set honest

Use training data for fitting and cross-validation, and reserve the test set for a final evaluation after modeling choices are settled. Repeatedly checking test results and changing features or thresholds based on them turns the test set into part of the selection process.

For time series, random shuffling can let future patterns influence evaluation of earlier periods. Use a time-aware split and ensure every transformation respects the same temporal boundary. For grouped data, use group-aware validation so the same entity cannot appear in both a training fold and its validation fold when deployment targets new entities.

A practical leakage review

Before trusting a score, verify:

  1. The split reflects the real prediction setting: independent rows, groups, time, or another constraint.
  2. Data-dependent transformations are inside the pipeline and are fitted only on training folds.
  3. Feature values existed before the prediction point and do not reveal the target indirectly.
  4. Duplicate records and shared entities are handled consistently with the intended deployment scenario.
  5. Hyperparameter selection does not use the final test set.
  6. The full preprocessing and estimator pipeline is saved and reused for inference.
  7. Baselines and error slices are reviewed; a suspiciously high score deserves investigation, not just celebration.

The takeaway

Use scikit-learn Pipeline and ColumnTransformer to make preprocessing part of model fitting, and split before fitting any transformation that learns from data. Then examine the feature timeline and split strategy: those are domain questions a library cannot answer for you.

A trustworthy evaluation is designed to imitate the data the model will actually see. The goal is not the highest score on a convenient split; it is a credible estimate of performance on the next real prediction.

Sources

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.