Overfitting: When Your Model Memorised the Exam Answers

In every class there was a student who got full marks without understanding anything. They had last year's question paper and memorised every answer. Then the teacher changed one number in one question, and the whole strategy collapsed.
That student is an overfitted model.
What overfitting is
A model overfits when it learns the training data too well, including its noise and accidents, instead of the general pattern. It performs brilliantly on data it has seen and poorly on data it hasn't. And new data is the only data that matters, because that's what the model meets in real life.
How to see it
Split your data. Train on one part, and measure on another part the model never sees during training (the validation set). Then watch both numbers.

Early on, both losses fall together; the model is learning real patterns. Then validation loss bottoms out and starts rising while training loss keeps falling. From that point, every extra epoch is memorisation, not learning.
Why it happens
- Model too flexible for the data. A huge neural network trained on 500 rows can memorise all of them.
- Too little data. Fewer examples make it easier to memorise individual cases.
- Training too long. As in the chart above.
- Leaky features. A column that accidentally contains the answer (an "approved_date" column in a loan-approval model) makes the model look like a genius until it's deployed.
How to fix it
Get more data. The most boring answer and often the best one.
Use a simpler model, or fewer features. A linear model that generalises beats a deep network that memorises.
Regularisation. Penalise complexity so the model prefers simpler explanations:
from sklearn.linear_model import Ridge
model = Ridge(alpha=1.0) # larger alpha = stronger penalty on big weights
model.fit(X_train, y_train)
Early stopping. Stop training when validation loss stops improving:
from tensorflow import keras
stop = keras.callbacks.EarlyStopping(monitor="val_loss", patience=3, restore_best_weights=True)
model.fit(X_train, y_train, validation_data=(X_val, y_val), epochs=100, callbacks=[stop])
Dropout (for neural networks): randomly switch off neurons during training so the network can't rely on any single path. It's like making the student study with random pages of the textbook missing; they're forced to understand the subject.
Data augmentation (for images, audio, text): flip, crop, rotate, add noise. You create "new" examples that teach the model what doesn't matter.
The opposite problem
Underfitting is when the model is too simple to learn the pattern at all. Both training and validation scores are bad. The fix is the reverse: more features, a more flexible model, longer training. Machine learning is mostly the art of standing between these two cliffs.
The exam-day test
Before trusting any model, ask: has it ever been tested on data it truly hasn't seen, collected the way real future data will be collected? If not, you don't know whether you've built the student who understands or the one who memorised. And you'll find out on exam day.