Feature Engineering
A model is only as good as the information you give it. Feature engineering transforms raw data — timestamps, category strings, text — into representations that make patterns explicit and learnable.
The Right Representation
Timestamp "08:32 Mon" can predict taxi fares — but only if you extract rush hour, day of week, and hour of day. Raw data contains the signal; engineering makes it visible.
Scale, Transform, Bin
- Standardization — zero mean, unit variance. Best for gradient-based models
- Min-Max — scale to [0,1]. Sensitive to outliers
- Log transform — tames right-skewed distributions (income, traffic)
- Binning — discretize to capture non-monotonic relationships
- Tree models are scale-invariant — no scaling needed
Standardization vs. Normalization
Two approaches, two use-cases:
Encoding Categories
- One-Hot — binary column per category. Good for low cardinality (<20)
- Label / Ordinal — integer mapping. Use when order exists or with trees
- Target Encoding — replace with mean target per category. Powerful for high cardinality — but beware data leakage
- Frequency Encoding — replace with count/frequency. Leakage-free
Target Encoding Leakage
Target encoding uses label information — each row "knows" its own target. If encoded naively before splitting, this causes data leakage: inflated training scores that collapse on real data.
TF-IDF & Embeddings
TF-IDF weights words by frequency in a document × rarity across all documents. Common words get low scores; informative rare words get high scores.
Word embeddings (Word2Vec, GloVe) put semantically similar words close in vector space. king − man + woman ≈ queen.
Timestamps Need Unpacking
- Extract — hour, day of week, month, year, week of year
- Cyclical encoding — sin/cos so midnight ≈ 11 PM numerically
- Lag features — value at t−1, t−7, t−30 (autocorrelation)
- Rolling stats — 7-day mean, 30-day std (trend & seasonality)
- Time-since — days since last purchase, hours since login
Wrapping Around a Circle
Hour 23 and hour 0 are one hour apart — but numerically far. Encode with sin and cosine to wrap the feature around a circle of period P:
Two Features, One Signal
A credit score of 700 means different things for $10K vs $200K income. The interaction (credit × income) captures what either feature alone cannot.
Missingness Is a Signal
Missing data is rarely random. A patient without a lab result may mean the test was never ordered — which is itself informative.
Key Takeaways
- Feature engineering often has more impact than model choice
- Scale numerical features for gradient-based models; trees don't need it
- Log-transform right-skewed distributions (income, traffic, counts)
- Use target encoding for high-cardinality categoricals — inside CV folds only
- Encode time cyclically with sin/cos; build lag and rolling features for forecasting
- Always create a missingness indicator before imputing