Feature Engineering Is Where the Magic (and the Pain) Lives

When people imagine machine learning, they imagine the algorithm: neural networks, gradient boosting, transformers. When people do machine learning, they spend most of their time on something less glamorous: deciding what information to feed the algorithm.
That's feature engineering, and it's where a lot of the real performance comes from.

A model only knows what you show it
Say you're predicting whether a customer will cancel a subscription. Your raw table has signup_date, last_login, plan, monthly_price and support_tickets.
A model can technically use last_login as a raw timestamp. But what matters is how long ago that was. You know that instantly. The model has to discover it from scratch, if it can at all. So you tell it.
Useful feature recipes
Time since things happened
import pandas as pd
today = pd.Timestamp("2025-06-01")
df["days_since_login"] = (today - df["last_login"]).dt.days
df["tenure_days"] = (today - df["signup_date"]).dt.days
Ratios instead of raw counts
Ten support tickets means something very different for a customer of five years than for a customer of five days.
df["tickets_per_month"] = df["support_tickets"] / (df["tenure_days"] / 30).clip(lower=1)
Calendar features
Day of week, month, end of month, holidays. Sales, traffic and electricity usage all care deeply about whether it's a Friday.
df["signup_weekday"] = df["signup_date"].dt.dayofweek
df["is_month_end"] = df["last_login"].dt.is_month_end
Recent behaviour vs overall behaviour
Logins in the last 7 days compared to the average week. A drop is often the strongest signal of all: people rarely cancel suddenly; they drift away first.
Grouping rare categories
A city column with 4,000 values, most appearing once, is noise. Keep the top cities and lump the rest into "other."
The traps
- Leakage. Never build features using information you wouldn't have at prediction time. "Number of days until cancellation" is a fantastic feature and a complete cheat.
- Computing statistics on the full dataset. If you fill missing values with the mean of all data, information from the test set sneaks into training. Fit transformations on training data only (scikit-learn pipelines help with this).
- Feature explosion. A thousand features you don't understand is not insight. It's a haystack.
Where domain knowledge wins
The best features usually come from talking to people who know the business. A sales manager will tell you that customers who never set up a second user almost always cancel. That's a feature no algorithm would have invented from raw columns, and it's worth more than a week of hyperparameter tuning.
Deep learning reduces the need for hand-crafted features on images and text. For the tabular data that runs most businesses, though, thoughtful features are still where the magic, and much of the pain, lives.