ML 101
M10 · L01
Practical ML

Feature Engineering

A model is only as good as the information you give it. Feature engineering transforms raw data — timestamps, category strings, text — into representations that make patterns explicit and learnable.

01 / 13
ML 101
M10 · L01
Why It Matters

The Right Representation

Timestamp "08:32 Mon" can predict taxi fares — but only if you extract rush hour, day of week, and hour of day. Raw data contains the signal; engineering makes it visible.

Domain Knowledge
The best features come from understanding the problem, not from automated search. Talk to domain experts first.
02 / 13
ML 101
M10 · L01
Numerical Features

Scale, Transform, Bin

  • Standardization — zero mean, unit variance. Best for gradient-based models
  • Min-Max — scale to [0,1]. Sensitive to outliers
  • Log transform — tames right-skewed distributions (income, traffic)
  • Binning — discretize to capture non-monotonic relationships
  • Tree models are scale-invariant — no scaling needed
03 / 13
ML 101
M10 · L01
Scaling Formulas

Standardization vs. Normalization

Two approaches, two use-cases:

Z-score (Standardization)
x'=\frac{x-\mu}{\sigma}
Min-Max (Normalization)
x'=\frac{x-x_{\min}}{x_{\max}-x_{\min}}
04 / 13
ML 101
M10 · L01
Categorical Features

Encoding Categories

  • One-Hot — binary column per category. Good for low cardinality (<20)
  • Label / Ordinal — integer mapping. Use when order exists or with trees
  • Target Encoding — replace with mean target per category. Powerful for high cardinality — but beware data leakage
  • Frequency Encoding — replace with count/frequency. Leakage-free
05 / 13
ML 101
M10 · L01
Danger Zone

Target Encoding Leakage

Target encoding uses label information — each row "knows" its own target. If encoded naively before splitting, this causes data leakage: inflated training scores that collapse on real data.

Solution
Always encode inside cross-validation folds: compute encodings from other folds, apply to current fold. Scikit-learn's TargetEncoder does this automatically.
06 / 13
ML 101
M10 · L01
Text Features

TF-IDF & Embeddings

TF-IDF weights words by frequency in a document × rarity across all documents. Common words get low scores; informative rare words get high scores.

Word embeddings (Word2Vec, GloVe) put semantically similar words close in vector space. king − man + woman ≈ queen.

07 / 13
ML 101
M10 · L01
Time Features

Timestamps Need Unpacking

  • Extract — hour, day of week, month, year, week of year
  • Cyclical encoding — sin/cos so midnight ≈ 11 PM numerically
  • Lag features — value at t−1, t−7, t−30 (autocorrelation)
  • Rolling stats — 7-day mean, 30-day std (trend & seasonality)
  • Time-since — days since last purchase, hours since login
08 / 13
ML 101
M10 · L01
Cyclical Encoding

Wrapping Around a Circle

Hour 23 and hour 0 are one hour apart — but numerically far. Encode with sin and cosine to wrap the feature around a circle of period P:

Cyclical Encoding
x_{\sin}=\sin\!\left(\tfrac{2\pi x}{P}\right),\;x_{\cos}=\cos\!\left(\tfrac{2\pi x}{P}\right)
09 / 13
ML 101
M10 · L01
Interaction Features

Two Features, One Signal

A credit score of 700 means different things for $10K vs $200K income. The interaction (credit × income) captures what either feature alone cannot.

Ratio
Debt / Income
Product
Temp × Humidity
Group
User Avg Spend
10 / 13
ML 101
M10 · L01
Missing Values

Missingness Is a Signal

Missing data is rarely random. A patient without a lab result may mean the test was never ordered — which is itself informative.

Best Practice
Before imputing: create a binary indicator feature (1 = was missing, 0 = observed). Then impute. The model can learn both the imputed value and the missingness pattern.
11 / 13
ML 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

12 / 13
ML 101
Key Takeaways
Summary

Key Takeaways

  • Feature engineering often has more impact than model choice
  • Scale numerical features for gradient-based models; trees don't need it
  • Log-transform right-skewed distributions (income, traffic, counts)
  • Use target encoding for high-cardinality categoricals — inside CV folds only
  • Encode time cyclically with sin/cos; build lag and rolling features for forecasting
  • Always create a missingness indicator before imputing
13 / 13