Skip to main content

Data & Sampling Methods

Topic - Statistics transforms raw, unstructured observations into actionable inferences about populations. Before computing parameters, one must categorize variable types, understand data structures across time and space, and select an unbiased sampling technique that guarantees statistical validity.


1. Intuition & Architectural Flow​

Statistical inference bridges the gap between what we can feasibly observe (a sample) and the complete universe of entities we wish to understand (a population).

The field is bifurcated into two primary branches:

  1. Descriptive Statistics: Summarizing, organizing, and visualizing the concrete features of an observed collection of data points without extending conclusions beyond the sample.
  2. Inferential Statistics: Using mathematical probability theory, hypothesis testing, and estimation models to generalize conclusions from a sample to an unobserved population.

Characteristics and Functions of Statistics​

Statistics operates under distinct empirical characteristics:

  • Aggregation of Facts: Isolated, single figures (e.g., "John earns 50,000 USD") are not statistics; only aggregates of figures across subjects reveal distributions.
  • Multifactorial Causation: Real-world data points are shaped by numerous interacting variables rather than a single isolated cause.
  • Systematic Collection: Data must be gathered systematically for a pre-planned purpose; haphazard observation invalidates statistical inference.
  • Comparability: Data points must be homogeneous enough with respect to their measurement scales to permit valid relative comparisons.

2. Taxonomy of Data and Variables​

Statistical models process numbers, but the numerical representation depends strictly on the underlying variable scale.

Variable Scales Decomposition​


  • Discrete Variables: Variables whose values are distinct, separate integers resulting from counting processes.
    • Examples: Number of defective parts per batch, daily website visits, number of hospital beds occupied.
    • Condition: Countably finite or countably infinite (x∈N∪{0}x \in \mathbb{N} \cup \{0\}).
  • Continuous Variables: Variables that can take an uncountably infinite number of real values within an interval.
    • Examples: Temperature in Celsius, transaction processing latency, stock prices, aircraft flight velocity.
    • Condition: Uncountably infinite (x∈Rx \in \mathbb{R}).
warning

The Categorical Encoding Fallacy

Assigning arbitrary integer IDs to categorical levels (e.g., mapping Male = 0, Female = 1 or Linux = 1, Windows = 2, macOS = 3) does not convert qualitative categories into numerical variables.

Computing arithmetic means or standard deviations on integer encodings of nominal categories yields mathematically meaningless results (e.g., an "average OS" of 2.32.3 has no physical interpretation).


3. Sampling Methodologies​

Conducting a census across an entire population (NN) is frequently cost-prohibitive, computationally impractical, or physically destructive (e.g., testing the lifespan of every light bulb produced). Therefore, sampling selects a representative subset (n≪Nn \ll N).

Probability vs Non-Probability Sampling​

Probability Sampling Designs​

Every individual in the population has a known, non-zero probability of selection, enabling rigorous mathematical estimation of sampling error:

  1. Simple Random Sampling (SRS): Every member of the population has an identical probability of inclusion (P=1/NP = 1/N), and all combinations of nn elements are equally likely. Requires an exhaustive sampling frame.
  2. Systematic Sampling: Elements are selected at a constant periodic interval k=N/nk = N/n after a randomly selected starting point r∈[1,k]r \in [1, k].
  3. Stratified Sampling: The population is segmented into mutually exclusive, homogeneous subgroups (strata based on attributes like age or department). A random sample is drawn from each stratum. Guarantees proportional representation of minority subpopulations.
  4. Cluster Sampling: The population is naturally grouped into heterogeneous mini-populations (clusters, such as school districts or regional branches). A random sample of entire clusters is selected, and all entities within selected clusters are surveyed.
  5. Multistage Sampling: Hierarchical sampling combining multiple methods sequentially (e.g., randomly picking cities →\rightarrow picking schools within cities →\rightarrow randomly selecting students within schools).

Non-Probability Sampling Designs​

Selection relies on convenience, availability, or researcher judgment. These methods cannot produce unbiased population parameter estimates or calculate confidence intervals:

  • Convenience Sampling: Surveying units that are easiest to reach (e.g., interviewing visitors entering a shopping mall).
  • Purposive / Judgmental Sampling: Selecting specific units based on expert judgment of what constitutes a "typical" case.
  • Snowball Sampling: Initial participants recruit subsequent contacts from their social or professional networks (essential for hard-to-reach populations, such as rare disease patients).
  • Quota Sampling: Specifying fixed proportions of demographic traits (e.g., 50% women, 50% men), but selecting the individuals within those categories via convenience rather than randomization.

4. Sampling Error Mechanics​

When calculating sample statistics (xˉ,s\bar{x}, s) to estimate population parameters (μ,σ\mu, \sigma), errors naturally emerge:

Total Error=Sampling (Random) Error+Non-Sampling (Systematic) Error\text{Total Error} = \text{Sampling (Random) Error} + \text{Non-Sampling (Systematic) Error}


  • Nature: Unsystematic, chance variations between the sample statistic and the true population parameter.
  • Directionality: Symmetric; just as likely to overestimate as underestimate the true parameter.
  • Mitigation: Increasing sample size (nn) strictly decreases the standard error of the estimate inversely proportional to the square root of nn:

SE(xˉ)=σn\text{SE}(\bar{x}) = \frac{\sigma}{\sqrt{n}}


5. Step-by-Step Numerical Walkthrough​

Academic examinations frequently test the execution of Systematic Sampling intervals and Stratified Allocation.

Problem: Stratified Proportional Allocation​

A university has N=2,000N = 2,000 students distributed across three schools:

  • School of Engineering (N1=1,000N_1 = 1,000)
  • School of Business (N2=600N_2 = 600)
  • School of Humanities (N3=400N_3 = 400)

Design a representative stratified sample of total size n=200n = 200 using proportional allocation. Calculate the sampling interval kk if systematic sampling were applied to the entire population.

Step 1: Compute Proportional Sampling Fraction​

The uniform sampling fraction ff across all strata is:

f=nN=2002000=0.10(10%)f = \frac{n}{N} = \frac{200}{2000} = 0.10 \quad (10\%)

Step 2: Calculate Strata Sample Sizes (nhn_h)​

nh=Nh×nNn_h = N_h \times \frac{n}{N}

  1. Engineering: n1=1000×0.10=100n_1 = 1000 \times 0.10 = 100 students.
  2. Business: n2=600×0.10=60n_2 = 600 \times 0.10 = 60 students.
  3. Humanities: n3=400×0.10=40n_3 = 400 \times 0.10 = 40 students.

∑h=13nh=100+60+40=200=n\sum_{h=1}^3 n_h = 100 + 60 + 40 = 200 = n

Step 3: Compute Systematic Sampling Interval (kk)​

If the entire list of 2,0002,000 students is ordered alphabetically:

k=⌊Nn⌋=2000200=10k = \left\lfloor \frac{N}{n} \right\rfloor = \frac{2000}{200} = 10

Select a random integer rr uniformly distributed between 11 and 1010. If r=7r = 7, sample student IDs: 7,17,27,37,…,19977, 17, 27, 37, \dots, 1997.


6. Implementation Lab​

tip

Implementation Lab: Sampling Distributions & Bias in Python

Experiment with probability vs convenience sampling on simulated synthetic populations.

Open In Colab

Key Experiments to Run:

  • Experiment 1 (Sample Size vs Standard Error): Scale nn from 1010 to 10,00010,000 and observe the convergence of xˉ→μ\bar{x} \to \mu.
  • Experiment 2 (Selection Bias Demonstration): Sample exclusively from an unrepresentative tail of a distribution to verify that larger nn does not correct systematic bias.

Comparative Implementation​


import numpy as np
import pandas as pd
from sklearn.model_selection import train_test_split

# Generate a synthetic population of 10,000 records
np.random.seed(42)
departments = np.random.choice(['Engineering', 'Business', 'Humanities'], size=10000, p=[0.5, 0.3, 0.2])
salaries = np.random.normal(loc=75000, scale=15000, size=10000)

pop_df = pd.DataFrame({'Department': departments, 'Salary': salaries})

# 1. Simple Random Sampling (n = 200)
srs_sample = pop_df.sample(n=200, random_state=42)

# 2. Stratified Sampling preserving exact department ratios (n = 200)
strat_sample, _ = train_test_split(pop_df, train_size=200, stratify=pop_df['Department'], random_state=42)

print("Population Department Distribution:")
print(pop_df['Department'].value_counts(normalize=True))
print("\nStratified Sample Distribution:")
print(strat_sample['Department'].value_counts(normalize=True))

7. Interactive Exploration: Sampling Architecture​

The widget below demonstrates how changing the sampling technique impacts parameter estimation on a population with distinct subgroups:

Sampling Design Evaluator

Select a sampling methodology to evaluate its structural trade-offs:

Simple Random Sampling (SRS)
Population Coverage: Equal chance for all individuals. Minorities may be underrepresented in small samples by random chance.
Primary Design Risk: High sample variance if the population has distinct heterogeneous sub-clusters.
Variance Property: Standard Error = σ / √n

8. Exam Traps & Operational Nuances​


  • The Trap: Confusing the structural composition of strata vs clusters in university exams.
  • The Golden Rule:
    • Stratified Sampling: Subgroups must be homogeneous within (members share identical traits) and heterogeneous between (strata differ markedly from each other). We sample from all strata.
    • Cluster Sampling: Subgroups must be heterogeneous within (each cluster is a mini-replica of the population) and homogeneous between (clusters resemble each other). We sample all units from some clusters.

9. Summary & Cheatsheet​

Core Data Taxonomies

  • Numerical: Discrete (integers, counts) vs Continuous (real line, measurements).
  • Categorical: Binary (2 states), Nominal (labels without rank), Ordinal (ordered ranks).
  • Data Formats: Cross-sectional (one time, many entities), Time series (one entity, repeated times), Panel (many entities across repeated times).

Sampling Rules

  • Probability: SRS, Systematic (k=N/nk = N/n), Stratified (homogeneity within), Cluster (heterogeneity within).
  • Non-Probability: Convenience, Purposive, Snowball, Quota. Cannot quantify confidence margins.
  • Error Law: Random error decreases with n\sqrt{n}; systematic bias is invariant to sample size.

info

Key Takeaways

  • Takeaway 1: Choosing the correct statistical tool depends strictly on knowing whether variables are nominal, ordinal, discrete, or continuous.
  • Takeaway 2: Probability sampling is the foundational prerequisite for all inferential tests, confidence intervals, and hypothesis tests.
  • Takeaway 3: Stratified sampling minimizes variance across diverse subgroups; systematic sampling simplifies physical or sequential data collection using periodic interval kk.

Next Section: Descriptive Statistics & Dispersion - Explore central tendency, variance, skewness, and kurtosis.


10. Active Recall & Practice​

Test your understanding by answering first, then clicking to reveal the underlying principles.

Checkpoint Quiz​

Interactive Checkpoint: Self-Test

A researcher wants to survey rural healthcare conditions. They randomly choose 10 districts out of 50 in a state, and interview every household in those 10 districts. Which sampling method is this?

Review Flashcards​

1. [TAXONOMY] Why is customer satisfaction rated on a 1 to 5 scale ordinal rather than interval?

  • While the numbers 1, 2, 3, 4, 5 are ordered, the subjective perceptual distance between "Poor (1)" and "Fair (2)" is not guaranteed to be mathematically identical to the distance between "Good (4)" and "Excellent (5)".
  • Therefore, order exists, but equal intervals do not.
2. [EXAM QUESTION] What happens to the standard error of the mean if the sample size is increased from 100 to 400?

  • Since SE=σn\text{SE} = \frac{\sigma}{\sqrt{n}}, increasing nn by a factor of 4 reduces the standard error by 4=2\sqrt{4} = 2.
  • The standard error is halved (cut by 50%50\%).
3. [ENGINEERING] When is Snowball Sampling methodologically justifiable in real-world data collection?

  • When the target population is concealed, marginalized, or has no formal sampling directory/frame (e.g., investigating individuals with rare medical conditions, specialized domain hackers, or underground communities).