Encoding Categorical Data
Turning categories into numbers without inventing an order — how much label encoding really costs, why it depends on the model, what cardinality does to memory, and the leak hiding inside target encoding.
Getting real data ready for a model — missing values, encoding, scaling, outliers, and the train/test split.
View all tagsTurning categories into numbers without inventing an order — how much label encoding really costs, why it depends on the model, what cardinality does to memory, and the leak hiding inside target encoding.
Which models need scaling and which are provably indifferent, why one outlier ruins MinMaxScaler, what log transforms actually fix, and how large the scaling leak really is.
Defensive data loading in pandas — handling type inference traps, hidden missingness sentinels, automated pre-flight audits, and memory-efficient array conversions.
Why a value is missing decides what you may do about it — MCAR, MAR and MNAR, what mean imputation costs a distribution, when a smarter imputer pays, and the fit/transform contract that keeps it honest.
What the IQR rule really flags, why outliers hide each other, the extreme point no single column can see, which treatment actually works, and the features a model cannot invent for itself.
Why a held-out set is mandatory, how much the score depends purely on random_state, and the four situations where a random split reports a number that is simply false.