Overfitting: When Your Model Memorised the Exam Answers

· 3 min read · Syed Omar Faruk Towaha
Overfitting: When Your Model Memorised the Exam Answers

In every class there was a student who got full marks without understanding anything. They had last year's question paper and memorised every answer. Then the teacher changed one number in one question, and the whole strategy collapsed.

That student is an overfitted model.

What overfitting is

A model overfits when it learns the training data too well, including its noise and accidents, instead of the general pattern. It performs brilliantly on data it has seen and poorly on data it hasn't. And new data is the only data that matters, because that's what the model meets in real life.

How to see it

Split your data. Train on one part, and measure on another part the model never sees during training (the validation set). Then watch both numbers.

Training vs validation loss
Training loss keeps falling. Validation loss turns around. That turn is overfitting starting.

Early on, both losses fall together; the model is learning real patterns. Then validation loss bottoms out and starts rising while training loss keeps falling. From that point, every extra epoch is memorisation, not learning.

Why it happens

How to fix it

Get more data. The most boring answer and often the best one.

Use a simpler model, or fewer features. A linear model that generalises beats a deep network that memorises.

Regularisation. Penalise complexity so the model prefers simpler explanations:

from sklearn.linear_model import Ridge

model = Ridge(alpha=1.0)   # larger alpha = stronger penalty on big weights
model.fit(X_train, y_train)

Early stopping. Stop training when validation loss stops improving:

from tensorflow import keras

stop = keras.callbacks.EarlyStopping(monitor="val_loss", patience=3, restore_best_weights=True)
model.fit(X_train, y_train, validation_data=(X_val, y_val), epochs=100, callbacks=[stop])

Dropout (for neural networks): randomly switch off neurons during training so the network can't rely on any single path. It's like making the student study with random pages of the textbook missing; they're forced to understand the subject.

Data augmentation (for images, audio, text): flip, crop, rotate, add noise. You create "new" examples that teach the model what doesn't matter.

The opposite problem

Underfitting is when the model is too simple to learn the pattern at all. Both training and validation scores are bad. The fix is the reverse: more features, a more flexible model, longer training. Machine learning is mostly the art of standing between these two cliffs.

The exam-day test

Before trusting any model, ask: has it ever been tested on data it truly hasn't seen, collected the way real future data will be collected? If not, you don't know whether you've built the student who understands or the one who memorised. And you'll find out on exam day.

// related

// prefer the terminal?

Open the terminal blog and type read overfitting-memorised-the-exam.