Cross-Validation: Trust Issues, But Make It Statistics

I have trust issues with single numbers. Someone says "my model gets 88% accuracy," and my first thought is: on which split? Because if you shuffle the data differently, you might get 79%, and the story changes completely.
Cross-validation exists for people like me.
The idea
Instead of one train/test split, split the data into k parts (usually 5 or 10). Train on four parts, test on the fifth. Repeat until every part has been the test set once.

You end up with five scores instead of one, and the spread between them is as informative as the average.

from sklearn.model_selection import cross_val_score
from sklearn.ensemble import RandomForestClassifier
scores = cross_val_score(RandomForestClassifier(random_state=0), X, y, cv=5, scoring="accuracy")
print(f"{scores.mean():.3f} ± {scores.std():.3f}")
If the spread is wide, your model is sensitive to which data it sees, which usually means you need more data, a simpler model, or both.
Variants for real life
Stratified k-fold keeps the class proportions the same in every fold. Essential when one class is rare; otherwise one fold might have no positive examples at all. (scikit-learn uses stratified folds by default for classifiers when you pass an integer cv.)
Group k-fold keeps related rows together. If you have several photos of the same patient, or several transactions from the same customer, they must all be in training or all in testing. Otherwise the model "recognises" the patient instead of learning the disease, and your score is inflated.
from sklearn.model_selection import GroupKFold
cv = GroupKFold(n_splits=5)
scores = cross_val_score(model, X, y, cv=cv, groups=patient_ids)
Time series split never trains on the future to predict the past. Each fold trains on earlier data and tests on the period right after.
from sklearn.model_selection import TimeSeriesSplit
cv = TimeSeriesSplit(n_splits=5)
Shuffling time series data is one of the most common ways to build a model that looks brilliant and fails on day one.
The leakage trap
Any preprocessing that learns from data (scaling, imputing missing values, selecting features) must happen inside each fold, using only that fold's training data. The clean way is a pipeline:
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
model = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000))
scores = cross_val_score(model, X, y, cv=5)
When not to bother
With millions of rows, a single large hold-out set is already stable, and five full trainings may be too expensive. Cross-validation matters most when data is small to medium, which, outside big tech, is most of the time.
Trust issues, it turns out, are a sign of a good data scientist. Ask the question five times. Believe the pattern, not the one lucky answer.