Descriptive Statistics & Data Distributions
Topic - Descriptive statistics summarizes high-dimensional collections of observations into concise mathematical indicators of location, spread, and distribution shape. Mastering these metrics enables robust exploratory data analysis, outlier identification, and the diagnostic evaluation of normality assumptions.
1. Intuition & Architectural Flow
When analyzing an unfamiliar dataset, raw tables of thousands of values conceal structural patterns. We condense this complexity into four foundational mathematical dimensions:
2. Measures of Central Tendency
Central tendency identifies the single central value around which numerical observations concentrate.
Core Mathematical Formulations
- 1. Arithmetic & Trimmed Mean
- 2. Geometric & Harmonic Mean
- 3. Median & Mode
- Arithmetic Mean (): The sum of all observations divided by sample count :
- Merits: Utilizes all observations; mathematically tractable for downstream algebraic derivations.
- Demerits: Severe sensitivity to outliers; inapplicable to nominal categorical features.
- Trimmed Mean (-Trimmed): Arithmetic mean computed after discarding the lowest and highest of ordered observations:
- Engineering Application: Used in Olympic judging and robust data pipelines to prevent extreme values from distorting baseline trends (
scipy.stats.trim_mean).
- Geometric Mean (GM): The -th root of the product of non-negative observations:
- Primary Application: Calculating average compound growth rates (e.g., CAGR, investment multipliers, financial portfolio indexation).
- Harmonic Mean (HM): The reciprocal of the arithmetic mean of reciprocals:
- Primary Application: Averaging rates and ratios across equal units of work (e.g., average vehicle speed over equal distance segments, -score in machine learning).
- Fundamental Inequality: For any positive dataset: , with equality holding if and only if all values are identical.
- Median ( or ): The positional midpoint dividing sorted data into two equal halves ( above, below):
- Merits: Robust to extreme outliers; valid for ordinal rankings.
- Mode: The value occurring with greatest empirical frequency.
- Can be unimodal (1 peak), bimodal (2 distinct peaks), or multimodal ( peaks).
- Applicable to nominal categorical data (e.g., modal operating system).
3. Measures of Dispersion & Outlier Fences
Dispersion quantifies the spread, variability, and reliability of the central tendency measure. High dispersion signifies that the central tendency provides a weaker representation of individual observations.
Dispersion Metrics Decomposition
- 1. Range & Interquartile Range (IQR)
- 2. Variance & Standard Deviation
- 3. Coefficient of Variation (CV)
- Range: . Quickest to compute, but highly unreliable because it depends exclusively on two extreme observations.
- Quartiles (): Partition values dividing sorted data into 4 equal quarters (, , ).
- Interquartile Range (IQR): The spread of the middle of the observations:
- Tukey's Outlier Fences: Outliers are mathematically flagged if they breach the inner fences:
- Population Variance () vs Sample Variance ():
- Bessel's Correction (): Dividing sample variance by underestimates true population variance because sample deviations are measured around the sample mean rather than the true parameter . Dividing by produces an unbiased estimator ().
- Standard Deviation (): The square root of variance: . Unlike variance (which is expressed in squared units), standard deviation is expressed in the identical units of the original dataset. Range: .
- Definition: The ratio of the standard deviation to the mean, expressed as a unitless percentage:
- Purpose: Standard deviation cannot compare variability across features with differing units (e.g., comparing volatility of stock prices in dollars versus trading volume in shares) or differing means.
- Interpretation: A dataset with lower CV possesses higher relative consistency and stability.
4. Distribution Shape: Skewness & Kurtosis
Beyond center and spread, distribution shapes determine whether parametric statistical procedures (such as Z-tests or Linear Regression) are valid.
Skewness Formulations
Skewness quantifies the asymmetry and direction of a distribution's tail:
-
Karl Pearson's Coefficient of Skewness (): Based on the divergence between mean, median, and mode:
Values typically fall within . If , the distribution is symmetric.
-
Bowley's Coefficient of Skewness (): Based purely on quartile partition values:
Values are bounded strictly within . Ideal for open-ended or heavily outlier-prone datasets.
-
Adjusted Fisher-Pearson Standardized 3rd Moment (): Standard algorithm used by Python (
scipy.stats.skewandpandas.DataFrame.skew):
Kurtosis Formulations
Kurtosis measures the "tailedness" and propensity of a distribution to generate extreme outlier observations relative to a Gaussian distribution.
| Kurtosis Type | Excess Kurtosis () | Tail & Peak Behavior | Outlier Propensity |
|---|---|---|---|
| Mesokurtic | Standard normal Gaussian bell curve | Baseline normal | |
| Leptokurtic | Fat, heavy tails; sharp, tall central peak | High frequency of extreme outliers | |
| Platykurtic | Light, thin tails; flat, broad central peak | Low frequency of extreme outliers |
5. Step-by-Step Numerical Walkthrough
Academic exams frequently provide a small raw vector and require calculating the complete descriptive parameter battery by hand.
Problem
Given the sample dataset representing manufacturing component diameters in millimeters:
Compute:
- Mean (), Median (), Mode
- Quartiles (), , and Tukey's Outlier Bounds
- Sample Variance () and Standard Deviation ()
- Karl Pearson's and Bowley's Skewness coefficients
Step 1: Central Tendency
Sort the data (, odd):
-
Mean:
-
Median: The -th element: .
-
Mode: occurs three times (unimodal).
Step 2: Quartiles and Outlier Fences
-
Lower half (elements 1 to 5): .
-
Upper half (elements 7 to 11): .
-
Interquartile Range:
-
Outlier Bounds:
- Conclusion: The value is mathematically flagged as an extreme outlier.
Step 3: Sample Variance and Standard Deviation
Compute deviations from sample mean :
-
Coefficient of Variation:
Step 4: Skewness Measures
-
Karl Pearson's Skewness:
-
Bowley's Skewness:
-
Key Diagnostic Insight: Bowley's metric evaluates to (indicating median symmetry in the core quartiles), whereas Pearson's metric evaluates to due to the powerful pull of the outlier on the arithmetic mean!
6. Implementation Lab
Implementation Lab: Descriptive Statistics & Data Distributions in Python
Execute the complete descriptive statistics pipeline across synthetic and real-world tabular data.
Key Experiments to Run:
- Experiment 1 (Outlier Resistance): Progressively increase a single outlier value from to and record the divergence between
mean()andmedian(). - Experiment 2 (Kurtosis & Tail Risk): Compare a normal Gaussian sample against a Student's t-distribution () to visualize excess kurtosis in histogram tail densities.
Comparative Implementation
- 1. Production (Pandas & SciPy)
- 2. Pure Python / NumPy from Scratch
import numpy as np
import pandas as pd
from scipy import stats
data = [7, 8, 8, 9, 10, 10, 10, 11, 12, 14, 30]
s = pd.Series(data)
# Central Tendency & Dispersion
mean_val = s.mean()
median_val = s.median()
mode_val = s.mode()[0]
std_val = s.std(ddof=1) # Bessel's corrected sample std
var_val = s.var(ddof=1)
cv_val = stats.variation(s, ddof=1) * 100
# Robust Quartiles & IQR
q1 = s.quantile(0.25)
q3 = s.quantile(0.75)
iqr = q3 - q1
lower_fence = q1 - 1.5 * iqr
upper_fence = q3 + 1.5 * iqr
outliers = s[(s < lower_fence) | (s > upper_fence)].tolist()
# Distribution Shape
skew_val = s.skew() # Fisher-Pearson adjusted skewness
kurt_val = s.kurt() # Fisher's excess kurtosis (Normal = 0)
print(f"Mean: {mean_val:.2f}, Median: {median_val:.2f}, Mode: {mode_val}")
print(f"Sample Std: {std_val:.2f}, CV: {cv_val:.2f}%")
print(f"IQR: {iqr:.2f}, Outliers: {outliers}")
print(f"Skewness: {skew_val:.3f}, Excess Kurtosis: {kurt_val:.3f}")
import numpy as np
from typing import Dict, Any
def compute_distribution_profile(arr: list) -> Dict[str, Any]:
n = len(arr)
sorted_arr = sorted(arr)
# Mean and Median
mean = sum(arr) / n
mid = n // 2
median = sorted_arr[mid] if n % 2 != 0 else (sorted_arr[mid - 1] + sorted_arr[mid]) / 2.0
# Sample Variance & Standard Deviation (ddof=1)
variance = sum((x - mean) ** 2 for x in arr) / (n - 1)
std_dev = variance ** 0.5
# Bowley Skewness via index approximation
q1 = sorted_arr[int(0.25 * (n - 1))]
q3 = sorted_arr[int(0.75 * (n - 1))]
bowley_skew = (q3 - 2 * median + q1) / (q3 - q1) if (q3 - q1) != 0 else 0.0
# Excess Kurtosis (4th Central Moment)
m4 = sum((x - mean) ** 4 for x in arr) / n
m2 = sum((x - mean) ** 2 for x in arr) / n
excess_kurtosis = (m4 / (m2 ** 2)) - 3.0
return {
"mean": mean,
"median": median,
"std_dev": std_dev,
"bowley_skewness": bowley_skew,
"excess_kurtosis": excess_kurtosis
}
metrics = compute_distribution_profile([7, 8, 8, 9, 10, 10, 10, 11, 12, 14, 30])
print(metrics)
7. Interactive Exploration: Outlier Impact Sandbox
Use the slider below to introduce an extreme outlier into an otherwise symmetric dataset and observe how the mean, median, standard deviation, and skewness react:
Dynamic Outlier Sensitivity Sandbox
Adjust the value of the 6th data point (baseline set: [10, 12, 14, 16, 18]):
8. Exam Traps & Operational Nuances
- 1. Mean Imputation Fallacy
- 2. Kurtosis Offset Confusion
- The Trap: Imputing missing values using the mean on right-skewed features (such as household income or web session durations).
- The Consequence: The imputed values are pulled artificially high by extreme values, injecting artificial variance and distorting regression coefficients.
- The Solution: Use median imputation or IQR trimming for skewed features; reserve mean imputation strictly for verified symmetric distributions.
- The Trap: Confusing Pearson's Kurtosis (, where Normal ) with Fisher's Excess Kurtosis (, where Normal ).
- The Exam Strategy: Always verify whether the question asks for "Kurtosis" or "Excess Kurtosis". If a software package outputs negative values (e.g., ), it is reporting excess kurtosis, indicating a platykurtic distribution.
9. Summary & Cheatsheet
Formulas & Definitions
- AM vs GM vs HM: AM for sums; GM for growth rates/CAGR; HM for rates/speeds.
- Sample Variance: (Bessel's correction).
- Coefficient of Variation: (unitless relative spread).
- Tukey Fences: and .
Distribution Profiles
- Right-Skewed: \text{Mode} < \text{Median} < \text{Mean}; Skewness > 0.
- Left-Skewed: \text{Mean} < \text{Median} < \text{Mode}; Skewness < 0.
- Leptokurtic: Excess K > 0, heavy tails, fat outlier risks.
- Platykurtic: Excess K < 0, light tails, fewer outliers.
Key Takeaways
- Takeaway 1: The median and IQR provide robust, outlier-resistant representations of center and spread when data is skewed.
- Takeaway 2: Standard deviation cannot be used to compare dispersion across different units; use the dimensionless Coefficient of Variation (CV) instead.
- Takeaway 3: Positive skewness elongates the right tail, dragging the mean toward high values while leaving the median resistant.
Next Section: Covariance & Correlation Analysis - Explore directional association and standardized linear strength.
10. Active Recall & Practice
Test your understanding by answering first, then clicking to reveal the underlying principles.
Checkpoint Quiz
Interactive Checkpoint: Self-Test
A vehicle travels 120 km at 60 km/h, and returns the same 120 km distance at 40 km/h. What is the correct average speed over the entire trip?
Review Flashcards
1. [FORMULA] Why do we divide by instead of when computing sample variance?
- To apply Bessel's correction, which corrects the negative bias introduced because sample deviations are calculated relative to rather than the true population mean .
- Dividing by makes an unbiased estimator of .
2. [DISTRIBUTION] If Mean = 45, Median = 50, and Mode = 60, what is the shape of the distribution?
- Since , the distribution is negatively skewed (left-skewed) with an elongated tail extending toward the left.
3. [ENGINEERING] A financial analyst compares the volatility of Bitcoin (price: 60,000 USD, SD: 3,000 USD) and a penny stock (price: 2 USD, SD: 0.5 USD). Which is more volatile?
- Compare their Coefficients of Variation:
- Bitcoin: .
- Penny Stock: .
- The penny stock has five times higher relative volatility despite its much smaller absolute standard deviation.