Why Features Matter
A model is only as good as the information you give it. Raw data — timestamps, category strings, pixel values — often contains the right signal but in the wrong form. Feature engineering is the process of transforming raw data into representations that make the patterns explicit and learnable.
Consider predicting taxi fares. The raw data might include pickup time "2024-01-15 08:32:00". A model cannot easily learn that Monday rush-hour trips cost more unless you extract: hour of day, day of week, and a binary is-rush-hour flag. Those derived features expose the signal. Feature engineering is often the single highest-leverage activity in an ML project.
The best features come from understanding the problem, not from automated search. A meteorologist knows that dew point interacts with temperature to predict rain probability. A banker knows that the ratio of debt-to-income is more predictive than either value alone. Talk to domain experts before touching the data.
Numerical Features
Numerical features are the most common type, but they often need transformation before a model can use them effectively. Three transformations cover the majority of real-world cases:
Scaling
Many models — linear regression, SVMs, neural networks, k-NN — are sensitive to the scale of features. A feature ranging from 0 to 1,000,000 will dominate features ranging 0 to 1 in gradient-based optimization. Two standard approaches:
Rule of thumb: use standardization for most situations. Use min-max when you need bounded outputs (e.g., for pixel values fed to CNNs). Tree-based models (random forests, gradient boosting) are scale-invariant and don't need either.
Log Transform
Many real-world quantities have right-skewed distributions — income, city population, website traffic, transaction amounts. Most values are small but a few are enormous. Log-transforming compresses the upper tail and makes the distribution more symmetric, which many models handle far better.
Use log(x + 1) to handle zero values safely. If the feature has negative values, consider a Yeo-Johnson or Box-Cox transform instead.
Binning (Discretization)
Sometimes a continuous feature has a non-monotonic relationship with the target — for example, age might relate to income in a curved, multi-modal way. Binning converts a continuous variable into discrete categories (e.g., age → "under 25", "25–40", "40–60", "over 60"), allowing the model to learn a different weight for each bin. This is especially powerful for linear models and when you have strong domain intuition about breakpoints.
Categorical Features
Models require numeric inputs. Categorical features (country, product category, color) must be encoded. The encoding strategy can have a large impact on model performance:
| Encoding | When to Use | Caution |
|---|---|---|
| One-Hot | Low-cardinality categories (under ~20 unique values). Creates a binary column per category. | Explodes dimensionality for high-cardinality features. Sparse representations. |
| Label / Ordinal | Categories with a natural order (e.g., small/medium/large → 1/2/3). Also works for tree-based models where ordinality doesn't matter. | Implies false ordering for nominal categories in distance-based models. |
| Target Encoding | High-cardinality categories (e.g., city with 10,000 values). Replace each category with the mean target value for that category. | Can cause severe data leakage if applied before train/val split. Always use cross-fold target encoding. |
| Frequency Encoding | Replace each category with its count or frequency in the training set. Captures rarity without target leakage risk. | Different categories with the same frequency get the same encoding. |
Target encoding computes the mean of the target per category — which means it uses label information. If done naively on the full training set before splitting, the encoded feature for each row "knows" its own label. Always encode within cross-validation folds: compute encodings from the other folds, apply to the current fold. Scikit-learn's TargetEncoder handles this correctly.
Text Features
When your data includes raw text — reviews, descriptions, tweets — you need a way to convert words into numbers that capture meaning.
TF-IDF
Term Frequency–Inverse Document Frequency (TF-IDF) is the workhorse for classical text ML. It weights each word by how often it appears in a document (TF) divided by how many documents contain it (IDF). Common words like "the" get low weights; rare but informative words get high weights.
Word Embeddings
TF-IDF treats words as independent. Word embeddings (Word2Vec, GloVe, FastText) map each word to a dense vector — typically 50 to 300 dimensions — where semantically similar words are close in vector space. The classic example: king − man + woman ≈ queen. Document-level embeddings can be obtained by averaging word vectors or using a pretrained sentence encoder (e.g., Sentence-BERT).
For most practical NLP tasks today, fine-tuning a pretrained transformer (BERT, RoBERTa) gives the best results — but TF-IDF remains competitive for simple classification with small datasets.
Time Features
Timestamps are information-rich but need careful extraction. A raw Unix timestamp is nearly meaningless to most models. The goal is to extract the underlying signals:
Cyclical Encoding in Detail
Treating hour of day as a plain integer (0–23) implies that midnight (0) and 11 PM (23) are far apart — but they're one hour apart in reality. The sin/cos encoding wraps the feature around a circle:
Interaction Features
Sometimes two features together are more predictive than either alone. A credit score of 700 means something different for a borrower with $10K income versus $200K income. The interaction feature (credit_score × income) captures this joint effect that linear models would otherwise miss.
Common interaction strategies:
- Polynomial features — all products up to degree N: x₁², x₁x₂, x₂²,... Scikit-learn's
PolynomialFeaturesautomates this. - Ratio features — debt-to-income, price-per-square-foot, clicks-per-impression. Ratios often capture the meaningful quantity.
- Product features — manual domain-informed products: temperature × humidity, ad_spend × seasonality.
- Group statistics — mean of target per user, user's historical avg spend. Aggregations from the same entity across rows.
Missing Values as Features
Missing data is rarely random. A patient not having a lab result recorded may mean the test was not ordered — which is itself informative. Before imputing missing values, always create a binary missingness indicator feature: 1 if the value was missing, 0 otherwise. Then impute the missing values (mean/median for numerical, mode for categorical).
This way the model can learn both "what the value is" and "whether it was missing" — capturing both the imputed signal and the missingness pattern.
Automated Feature Engineering
Libraries like Featuretools automate the creation of interaction features from relational data (multiple tables). They perform "deep feature synthesis" — systematically applying aggregation and transformation primitives across entity relationships. Tools like AutoFeat perform symbolic regression to discover non-linear feature combinations. These tools are most valuable when you have rich relational structure and limited domain knowledge, or when you want to stress-test whether manual features are leaving signal on the table.
Feature engineering creates new features from existing ones. Feature selection decides which features to keep. Both are critical. Too many features — especially correlated or noisy ones — can hurt performance through the curse of dimensionality and overfitting. Techniques like recursive feature elimination (RFE), permutation importance, and LASSO regularization help identify which engineered features actually matter.
The Feature Engineering Workflow
In practice, feature engineering is iterative. A reliable process:
- Start with domain knowledge — list features a human expert would find informative.
- Explore the data — histograms, scatter plots, correlation matrices. Understand distributions and relationships.
- Create a baseline — train a simple model on raw features to establish a benchmark.
- Engineer and evaluate — add new features one group at a time, measure impact on validation score.
- Prune — remove features that don't improve validation performance. Less is often more.
- Repeat — features that don't help with a linear model might help with a tree model, and vice versa.
Feature engineering transforms raw data into learnable representations. Numerical features benefit from scaling, log transforms, and binning. Categorical features need one-hot, label, or target encoding depending on cardinality. Text becomes TF-IDF vectors or word embeddings. Time features need cyclical encoding and lag construction. Always create missingness indicators before imputing. The best features come from domain knowledge — automated tools help but don't replace understanding the problem.