Hyperparameter Tuning, or: Expensive Guessing

· 2 min read · Syed Omar Faruk Towaha
Hyperparameter Tuning, or: Expensive Guessing

Every machine learning model comes with knobs: learning rate, tree depth, number of layers, regularisation strength. These are hyperparameters: settings you choose before training, as opposed to the parameters the model learns.

Tuning them is a bit like seasoning a curry when you can't taste it directly. You add some salt, wait 40 minutes for the result, and then decide whether to add more salt.

Grid search: try everything

Pick a few values for each knob and try every combination.

Grid of results
A grid of validation scores. The best spot is usually a broad plateau, not a single magic cell.

It's simple and easy to understand. It's also expensive: 5 values for each of 4 hyperparameters is 625 training runs. And it wastes effort on knobs that don't matter much.

Random search: surprisingly better

Pick random combinations within sensible ranges. With the same budget, random search often finds better settings than grid search, because it tries more distinct values of the hyperparameters that actually matter, instead of stepping through a rigid grid.

from scipy.stats import loguniform, randint
from sklearn.model_selection import RandomizedSearchCV
from sklearn.ensemble import GradientBoostingClassifier

search = RandomizedSearchCV(
    GradientBoostingClassifier(),
    {
        "learning_rate": loguniform(1e-3, 0.3),
        "max_depth": randint(2, 10),
        "n_estimators": randint(100, 600),
    },
    n_iter=40, cv=5, scoring="roc_auc", n_jobs=-1, random_state=42,
)
search.fit(X_train, y_train)
print(search.best_params_)

Notice the log-uniform distribution for learning rate. The difference between 0.001 and 0.01 matters as much as between 0.01 and 0.1, so you search on a log scale.

Smarter search: learn from previous tries

Bayesian optimisation tools such as Optuna look at the results so far and choose the next combination where improvement seems likely. They also stop unpromising runs early. For expensive models this saves a lot of compute.

import optuna

def objective(trial):
    lr = trial.suggest_float("lr", 1e-4, 1e-1, log=True)
    depth = trial.suggest_int("depth", 2, 10)
    return train_and_validate(lr, depth)   # returns validation score

study = optuna.create_study(direction="maximize")
study.optimize(objective, n_trials=50)

The rules that keep you honest

  1. Tune on validation data, report on test data. If you tune against the test set, your reported score is a fantasy.
  2. Use cross-validation when data is small, so a lucky split doesn't fool you.
  3. Fix a budget. Tuning expands to fill all available GPUs and weekends.
  4. Prefer plateaus over peaks. A setting that's good across a broad region is more robust than one sharp best cell.

The uncomfortable truth

After the first few sensible choices, tuning usually gives modest gains, often a point or two. Better data, better features or fixing a label bug often give ten. If you've spent three days tuning and haven't looked at the misclassified examples, close the tuning script and open the data. The curry needs better ingredients, not more salt.

// related

// prefer the terminal?

Open the terminal blog and type read hyperparameter-tuning-expensive-guessing.