Skip to main content

Descriptive Statistics & Data Distributions

Topic - Descriptive statistics summarizes high-dimensional collections of observations into concise mathematical indicators of location, spread, and distribution shape. Mastering these metrics enables robust exploratory data analysis, outlier identification, and the diagnostic evaluation of normality assumptions.


1. Intuition & Architectural Flow​

When analyzing an unfamiliar dataset, raw tables of thousands of values conceal structural patterns. We condense this complexity into four foundational mathematical dimensions:


2. Measures of Central Tendency​

Central tendency identifies the single central value around which numerical observations concentrate.

Core Mathematical Formulations​


  • Arithmetic Mean (xˉ\bar{x}): The sum of all observations divided by sample count nn:

xˉ=1n∑i=1nxi\bar{x} = \frac{1}{n} \sum_{i=1}^n x_i

  • Merits: Utilizes all observations; mathematically tractable for downstream algebraic derivations.
  • Demerits: Severe sensitivity to outliers; inapplicable to nominal categorical features.
  • Trimmed Mean (α\alpha-Trimmed): Arithmetic mean computed after discarding the lowest and highest α%\alpha\% of ordered observations:

xˉα=1n−2k∑i=k+1n−kx(i)where k=⌊n⋅α⌋\bar{x}_{\alpha} = \frac{1}{n - 2k} \sum_{i=k+1}^{n-k} x_{(i)} \quad \text{where } k = \lfloor n \cdot \alpha \rfloor

  • Engineering Application: Used in Olympic judging and robust data pipelines to prevent extreme values from distorting baseline trends (scipy.stats.trim_mean).

3. Measures of Dispersion & Outlier Fences​

Dispersion quantifies the spread, variability, and reliability of the central tendency measure. High dispersion signifies that the central tendency provides a weaker representation of individual observations.

Dispersion Metrics Decomposition​


  • Range: R=xmax⁡−xmin⁡R = x_{\max} - x_{\min}. Quickest to compute, but highly unreliable because it depends exclusively on two extreme observations.
  • Quartiles (Q1,Q2,Q3Q_1, Q_2, Q_3): Partition values dividing sorted data into 4 equal quarters (25%25\%, 50%50\%, 75%75\%).
  • Interquartile Range (IQR): The spread of the middle 50%50\% of the observations:

IQR=Q3−Q1\text{IQR} = Q_3 - Q_1

  • Tukey's Outlier Fences: Outliers are mathematically flagged if they breach the inner fences:

Lower Bound=Q1−1.5×IQR,Upper Bound=Q3+1.5×IQR\text{Lower Bound} = Q_1 - 1.5 \times \text{IQR}, \quad \text{Upper Bound} = Q_3 + 1.5 \times \text{IQR}


4. Distribution Shape: Skewness & Kurtosis​

Beyond center and spread, distribution shapes determine whether parametric statistical procedures (such as Z-tests or Linear Regression) are valid.

Skewness Formulations​

Skewness quantifies the asymmetry and direction of a distribution's tail:

  1. Karl Pearson's Coefficient of Skewness (SkS_k): Based on the divergence between mean, median, and mode:

    Sk=xˉ−Modes≈3(xˉ−Median)sS_k = \frac{\bar{x} - \text{Mode}}{s} \approx \frac{3(\bar{x} - \text{Median})}{s}

    Values typically fall within [−3,+3][-3, +3]. If Sk=0S_k = 0, the distribution is symmetric.

  2. Bowley's Coefficient of Skewness (SbS_b): Based purely on quartile partition values:

    Sb=(Q3−Q2)−(Q2−Q1)Q3−Q1=Q3−2Q2+Q1Q3−Q1S_b = \frac{(Q_3 - Q_2) - (Q_2 - Q1)}{Q_3 - Q_1} = \frac{Q_3 - 2Q_2 + Q_1}{Q_3 - Q_1}

    Values are bounded strictly within [−1,+1][-1, +1]. Ideal for open-ended or heavily outlier-prone datasets.

  3. Adjusted Fisher-Pearson Standardized 3rd Moment (g1g_1): Standard algorithm used by Python (scipy.stats.skew and pandas.DataFrame.skew):

    g1=n(n−1)(n−2)∑i=1n(xi−xˉs)3g_1 = \frac{n}{(n-1)(n-2)} \sum_{i=1}^n \left( \frac{x_i - \bar{x}}{s} \right)^3

Kurtosis Formulations​

Kurtosis measures the "tailedness" and propensity of a distribution to generate extreme outlier observations relative to a Gaussian distribution.

Fisher’s Excess Kurtosis (K)=[1n∑i=1n(xi−xˉs)4]−3\text{Fisher's Excess Kurtosis } (K) = \left[ \frac{1}{n} \sum_{i=1}^n \left( \frac{x_i - \bar{x}}{s} \right)^4 \right] - 3

Kurtosis TypeExcess Kurtosis (KK)Tail & Peak BehaviorOutlier Propensity
MesokurticK=0K = 0Standard normal Gaussian bell curveBaseline normal
LeptokurticK>0K > 0Fat, heavy tails; sharp, tall central peakHigh frequency of extreme outliers
PlatykurticK<0K < 0Light, thin tails; flat, broad central peakLow frequency of extreme outliers

5. Step-by-Step Numerical Walkthrough​

Academic exams frequently provide a small raw vector and require calculating the complete descriptive parameter battery by hand.

Problem​

Given the sample dataset representing manufacturing component diameters in millimeters:

X=[7,8,8,9,10,10,10,11,12,14,30]X = [7, 8, 8, 9, 10, 10, 10, 11, 12, 14, 30]

Compute:

  1. Mean (xˉ\bar{x}), Median (x~\tilde{x}), Mode
  2. Quartiles (Q1,Q3Q_1, Q_3), IQR\text{IQR}, and Tukey's Outlier Bounds
  3. Sample Variance (s2s^2) and Standard Deviation (ss)
  4. Karl Pearson's and Bowley's Skewness coefficients

Step 1: Central Tendency​

Sort the data (n=11n = 11, odd): Xsorted=[7,8,8,9,10,10,10,11,12,14,30]X_{\text{sorted}} = [7, 8, 8, 9, 10, 10, 10, 11, 12, 14, 30]

  • Mean:

    xˉ=7+8+8+9+10+10+10+11+12+14+3011=12911≈11.727 mm\bar{x} = \frac{7 + 8 + 8 + 9 + 10 + 10 + 10 + 11 + 12 + 14 + 30}{11} = \frac{129}{11} \approx 11.727 \text{ mm}

  • Median: The 11+12=6\frac{11 + 1}{2} = 6-th element: x~=10 mm\tilde{x} = 10 \text{ mm}.

  • Mode: 10 mm10 \text{ mm} occurs three times (unimodal).

Step 2: Quartiles and Outlier Fences​

  • Lower half (elements 1 to 5): [7,8,8,9,10]  ⟹  Q1=8.0[7, 8, 8, 9, 10] \implies Q_1 = 8.0.

  • Upper half (elements 7 to 11): [10,11,12,14,30]  ⟹  Q3=12.0[10, 11, 12, 14, 30] \implies Q_3 = 12.0.

  • Interquartile Range:

    IQR=Q3−Q1=12.0−8.0=4.0 mm\text{IQR} = Q_3 - Q_1 = 12.0 - 8.0 = 4.0 \text{ mm}

  • Outlier Bounds:

    • Lower Fence=Q1−1.5×IQR=8.0−(1.5×4.0)=8.0−6.0=2.0 mm\text{Lower Fence} = Q_1 - 1.5 \times \text{IQR} = 8.0 - (1.5 \times 4.0) = 8.0 - 6.0 = 2.0 \text{ mm}
    • Upper Fence=Q3+1.5×IQR=12.0+(1.5×4.0)=12.0+6.0=18.0 mm\text{Upper Fence} = Q_3 + 1.5 \times \text{IQR} = 12.0 + (1.5 \times 4.0) = 12.0 + 6.0 = 18.0 \text{ mm}
    • Conclusion: The value 30 mm>18.0 mm30 \text{ mm} > 18.0 \text{ mm} is mathematically flagged as an extreme outlier.

Step 3: Sample Variance and Standard Deviation​

Compute deviations from sample mean xˉ≈11.727\bar{x} \approx 11.727:

∑i=111(xi−xˉ)2≈(7−11.73)2+2(8−11.73)2+⋯+(30−11.73)2≈392.18\sum_{i=1}^{11} (x_i - \bar{x})^2 \approx (7-11.73)^2 + 2(8-11.73)^2 + \dots + (30-11.73)^2 \approx 392.18

s2=392.18n−1=392.1810=39.218 mm2s^2 = \frac{392.18}{n - 1} = \frac{392.18}{10} = 39.218 \text{ mm}^2

s=39.218≈6.262 mms = \sqrt{39.218} \approx 6.262 \text{ mm}

  • Coefficient of Variation:

    CV=6.26211.727×100%≈53.4%\text{CV} = \frac{6.262}{11.727} \times 100\% \approx 53.4\%

Step 4: Skewness Measures​

  • Karl Pearson's Skewness:

    Sk=3(xˉ−x~)s=3(11.727−10)6.262=3(1.727)6.262≈+0.827(Right-Skewed)S_k = \frac{3(\bar{x} - \tilde{x})}{s} = \frac{3(11.727 - 10)}{6.262} = \frac{3(1.727)}{6.262} \approx +0.827 \quad (\text{Right-Skewed})

  • Bowley's Skewness:

    Sb=Q3−2Q2+Q1Q3−Q1=12.0−2(10)+8.012.0−8.0=12−20+84=04=0.0S_b = \frac{Q_3 - 2Q_2 + Q_1}{Q_3 - Q_1} = \frac{12.0 - 2(10) + 8.0}{12.0 - 8.0} = \frac{12 - 20 + 8}{4} = \frac{0}{4} = 0.0

  • Key Diagnostic Insight: Bowley's metric evaluates to 0.00.0 (indicating median symmetry in the core quartiles), whereas Pearson's metric evaluates to +0.827+0.827 due to the powerful pull of the outlier 3030 on the arithmetic mean!


6. Implementation Lab​

tip

Implementation Lab: Descriptive Statistics & Data Distributions in Python

Execute the complete descriptive statistics pipeline across synthetic and real-world tabular data.

Open In Colab

Key Experiments to Run:

  • Experiment 1 (Outlier Resistance): Progressively increase a single outlier value from 100100 to 10,00010,000 and record the divergence between mean() and median().
  • Experiment 2 (Kurtosis & Tail Risk): Compare a normal Gaussian sample against a Student's t-distribution (df=3df = 3) to visualize excess kurtosis in histogram tail densities.

Comparative Implementation​


import numpy as np
import pandas as pd
from scipy import stats

data = [7, 8, 8, 9, 10, 10, 10, 11, 12, 14, 30]
s = pd.Series(data)

# Central Tendency & Dispersion
mean_val = s.mean()
median_val = s.median()
mode_val = s.mode()[0]
std_val = s.std(ddof=1) # Bessel's corrected sample std
var_val = s.var(ddof=1)
cv_val = stats.variation(s, ddof=1) * 100

# Robust Quartiles & IQR
q1 = s.quantile(0.25)
q3 = s.quantile(0.75)
iqr = q3 - q1
lower_fence = q1 - 1.5 * iqr
upper_fence = q3 + 1.5 * iqr
outliers = s[(s < lower_fence) | (s > upper_fence)].tolist()

# Distribution Shape
skew_val = s.skew() # Fisher-Pearson adjusted skewness
kurt_val = s.kurt() # Fisher's excess kurtosis (Normal = 0)

print(f"Mean: {mean_val:.2f}, Median: {median_val:.2f}, Mode: {mode_val}")
print(f"Sample Std: {std_val:.2f}, CV: {cv_val:.2f}%")
print(f"IQR: {iqr:.2f}, Outliers: {outliers}")
print(f"Skewness: {skew_val:.3f}, Excess Kurtosis: {kurt_val:.3f}")

7. Interactive Exploration: Outlier Impact Sandbox​

Use the slider below to introduce an extreme outlier into an otherwise symmetric dataset and observe how the mean, median, standard deviation, and skewness react:

Dynamic Outlier Sensitivity Sandbox

Adjust the value of the 6th data point (baseline set: [10, 12, 14, 16, 18]):

Outlier Value
10
Mean (Shifted)
13.33
Median (Robust)
13.00
Std Dev
3.27
Pearson Skew
0.31

8. Exam Traps & Operational Nuances​


  • The Trap: Imputing missing values using the mean on right-skewed features (such as household income or web session durations).
  • The Consequence: The imputed values are pulled artificially high by extreme values, injecting artificial variance and distorting regression coefficients.
  • The Solution: Use median imputation or IQR trimming for skewed features; reserve mean imputation strictly for verified symmetric distributions.

9. Summary & Cheatsheet​

Formulas & Definitions

  • AM vs GM vs HM: AM for sums; GM for growth rates/CAGR; HM for rates/speeds.
  • Sample Variance: s2=1n−1∑(xi−xˉ)2s^2 = \frac{1}{n-1}\sum (x_i - \bar{x})^2 (Bessel's correction).
  • Coefficient of Variation: CV=(s/xˉ)×100%\text{CV} = (s / \bar{x}) \times 100\% (unitless relative spread).
  • Tukey Fences: Q1−1.5IQRQ_1 - 1.5\text{IQR} and Q3+1.5IQRQ_3 + 1.5\text{IQR}.

Distribution Profiles

  • Right-Skewed: \text{Mode} &lt; \text{Median} &lt; \text{Mean}; Skewness &gt; 0.
  • Left-Skewed: \text{Mean} &lt; \text{Median} &lt; \text{Mode}; Skewness &lt; 0.
  • Leptokurtic: Excess K &gt; 0, heavy tails, fat outlier risks.
  • Platykurtic: Excess K &lt; 0, light tails, fewer outliers.

info

Key Takeaways

  • Takeaway 1: The median and IQR provide robust, outlier-resistant representations of center and spread when data is skewed.
  • Takeaway 2: Standard deviation cannot be used to compare dispersion across different units; use the dimensionless Coefficient of Variation (CV) instead.
  • Takeaway 3: Positive skewness elongates the right tail, dragging the mean toward high values while leaving the median resistant.

Next Section: Covariance & Correlation Analysis - Explore directional association and standardized linear strength.


10. Active Recall & Practice​

Test your understanding by answering first, then clicking to reveal the underlying principles.

Checkpoint Quiz​

Interactive Checkpoint: Self-Test

A vehicle travels 120 km at 60 km/h, and returns the same 120 km distance at 40 km/h. What is the correct average speed over the entire trip?

Review Flashcards​

1. [FORMULA] Why do we divide by n−1n - 1 instead of nn when computing sample variance?

  • To apply Bessel's correction, which corrects the negative bias introduced because sample deviations are calculated relative to xˉ\bar{x} rather than the true population mean μ\mu.
  • Dividing by n−1n - 1 makes s2s^2 an unbiased estimator of σ2\sigma^2.
2. [DISTRIBUTION] If Mean = 45, Median = 50, and Mode = 60, what is the shape of the distribution?

  • Since Mean(45)<Median(50)<Mode(60)\text{Mean} (45) < \text{Median} (50) < \text{Mode} (60), the distribution is negatively skewed (left-skewed) with an elongated tail extending toward the left.
3. [ENGINEERING] A financial analyst compares the volatility of Bitcoin (price: 60,000 USD, SD: 3,000 USD) and a penny stock (price: 2 USD, SD: 0.5 USD). Which is more volatile?

  • Compare their Coefficients of Variation:
    • Bitcoin: CV=(3000/60000)×100%=5%\text{CV} = (3000 / 60000) \times 100\% = 5\%.
    • Penny Stock: CV=(0.5/2)×100%=25%\text{CV} = (0.5 / 2) \times 100\% = 25\%.
  • The penny stock has five times higher relative volatility despite its much smaller absolute standard deviation.