Sampling and Estimation
Topic - In real-world data science, measuring an entire population is rarely feasible due to operational constraints. Instead, we use sampling distributions to estimate unknown population parameters. Point estimation produces a single best-guess statistic, while interval estimation constructs a confidence interval bounded by a margin of error.
1. Population vs. Sample & The Need for Sampling
- Population: The complete collection of all elements, individuals, or items under study (e.g., all manufactured batteries from a factory). Characteristics of a population are called parameters (mean , variance , proportion ).
- Sample: A representative subset drawn from the population. Numerical summaries computed from a sample are called statistics (sample mean , sample variance , sample proportion ).
The Practical Need to Sample
- Destructive Testing: Evaluating battery lifespan or the tensile strength of an alloy exhausts the product. Testing the whole population leaves zero inventory to sell.
- Resource Constraints: Gathering data on large populations (such as all citizens in a country) is cost-prohibitive and computationally intractable.
- Efficiency: An appropriately sized, unbiased random sample provides precise, mathematically bounded estimates with minimal variance.
2. Sampling Distributions & The Central Limit Theorem (CLT)
When drawing independent random samples of size from a population, the calculated sample statistics will naturally fluctuate. This variation across repeated samples is called sampling variability.
The probability distribution of these sample means is the Sampling Distribution of the Mean.
The Central Limit Theorem Formulation
The Central Limit Theorem (CLT)
Regardless of the underlying population distribution (whether Normal, Poisson, Binomial, or Uniform), as the sample size becomes sufficiently large (), the sampling distribution of the sample mean approaches a Normal Distribution:
Where is called the Standard Error of the Mean (SE).
Converting the sample mean to a standard normal variable:
If the population standard deviation is unknown, we substitute the sample standard deviation , yielding for large samples ().
3. Point Estimation
A point estimate uses a single statistic computed from sample observations to estimate an unknown population parameter.
| Parameter to Estimate | Unbiased Point Estimator | Mathematical Formula |
|---|---|---|
| Population Mean () | Sample Mean () | |
| Population Variance () | Sample Variance () | |
| Population Proportion () | Sample Proportion () |
Sampling Error
The point estimate will differ from the true population parameter due to random chance. The absolute magnitude of this discrepancy is the Sampling Error:
Increasing the sample size consistently drives down the standard error (), decreasing sampling error.
4. Interval Estimation & Confidence Intervals
Because a point estimate rarely matches the true parameter exactly, Interval Estimation provides a range of plausible values associated with a specified confidence level ().
Derivation for Population Mean ()
From the Central Limit Theorem, the probability that the standardized statistic falls between two critical values and is :
Rearranging the inequality:
Thus, the Confidence Interval is:
Where is the Margin of Error.
Critical Values for Common Confidence Levels
- 90% Confidence (α = 0.10)
- 95% Confidence (α = 0.05)
- 99% Confidence (α = 0.01)
Width vs. Confidence Tradeoff
As the confidence level increases (e.g., from 90% to 99%), grows larger, making the interval wider. A higher degree of certainty requires a wider net.
5. Confidence Intervals for Proportions
For categorical variables (such as defective items or churn status), the sampling distribution of the sample proportion follows an approximate normal distribution for large samples:
The Confidence Interval for Proportion is:
6. Determining Sample Size from Margin of Error
When designing an experiment or survey, a primary goal is to determine how many observations must be sampled to guarantee a maximum desired margin of error :
7. Numerical Problem Walkthroughs
Scenario A: Portfolio Point Estimate
Problem: A financial firm manages portfolios. A sample of portfolios is audited, and are found underperforming. Estimate the total number of underperforming portfolios in the firm.
- Sample Proportion:
- Total Estimate: portfolios.
Scenario B: Average Income Confidence Interval
Problem: A sample of twenty-seven-year-olds in London has an average income of GBP with population standard deviation GBP. Calculate the 95% Confidence Interval.
- Given: , , ,
- Standard Error:
- Margin of Error:
- Interval:
- Interpretation: If we repeat this sampling process many times, roughly 95% of the calculated intervals will contain the true mean income. (It is incorrect to say there is a 95% chance that the true mean is inside this specific fixed interval, as the parameter is a fixed constant, not a random variable).
8. Implementation Lab
import numpy as np
from scipy import stats
# 1. Confidence Interval for Mean (London Income Scenario)
n = 250
mean_val = 45000
std_val = 4000
conf_level = 0.95
# Calculate scale = std / sqrt(n)
ci_mean = stats.norm.interval(conf_level, loc=mean_val, scale=std_val / np.sqrt(n))
print(f"95% CI for Mean: {ci_mean}")
# 2. Confidence Interval for Proportion (Portfolios Scenario)
n_port = 13
x_port = 8
p_hat = x_port / n_port
se_prop = np.sqrt((p_hat * (1 - p_hat)) / n_port)
ci_prop = stats.norm.interval(0.99, loc=p_hat, scale=se_prop)
print(f"99% CI for Proportion: {ci_prop}")
Key Takeaways
- CLT Guarantees Normality: The sampling distribution of the mean is normal for , even if the original population distribution is non-normal.
- Confidence vs. Probability: The parameter is fixed. A 95% confidence interval means that across repeated sampling experiments, 95% of constructed intervals capture .
- Controlling Margin of Error: To cut the margin of error in half, you must quadruple the sample size ().
Next Section: Hypothesis Testing Foundations - learning to evaluate statistical claims, understand errors, and compute test power.