Skip to main content

Large Sample Tests

Topic - When the sample size is large (n≥30n \ge 30), the Central Limit Theorem guarantees that the sampling distribution of the sample mean follows a normal distribution. In this regime, the Z-test provides the standard analytical tool to test hypotheses regarding single population means and differences between two independent population means.


1. Assumptions for Large Sample Z-Tests​

To ensure that inference using the standard normal distribution is valid, the following conditions must be satisfied:

  1. Continuous Data: The metric of interest must be continuous numerical data.
  2. Simple Random Sampling: Observations are independent and identically distributed (i.i.d.).
  3. Large Sample Size (n≥30n \ge 30): By the Central Limit Theorem, the distribution of Xˉ\bar{X} is normal. If the population standard deviation σ\sigma is unknown, we substitute the sample standard deviation ss without needing Student's t-distribution.
  4. Normality Verification: For moderate samples, normality of the data can be validated using the Shapiro-Wilk test (scipy.stats.shapiro).

2. One-Sample Z-Test for Population Mean​

Tests whether a single population mean μ\mu equals a hypothesized baseline value μ0\mu_0.

Test Statistic Formulation​

Zcalc=Xˉ−μ0σ/nZ_{\text{calc}} = \frac{\bar{X} - \mu_0}{\sigma / \sqrt{n}}

If population standard deviation σ\sigma is unknown and n≥30n \ge 30, we substitute sample standard deviation ss:

Zcalc=Xˉ−μ0s/n,where s=1n−1∑i=1n(xi−xˉ)2Z_{\text{calc}} = \frac{\bar{X} - \mu_0}{s / \sqrt{n}}, \quad \text{where } s = \sqrt{\frac{1}{n-1}\sum_{i=1}^n (x_i - \bar{x})^2}

Decision Rules across All Three Tails​

Test TypeHypothesesCritical Value RuleP-Value RuleConfidence Interval Rule
Two-TailedH0:μ=μ0H_0: \mu = \mu_0
H1:μ≠μ0H_1: \mu \neq \mu_0
Reject H0H_0 if ∣Zcalc∣>zα/2\lvert Z_{\text{calc}} \rvert > z_{\alpha/2}Reject H0H_0 if p≤αp \le \alpha
(p=2⋅P(Z>∣Zcalc∣)p = 2 \cdot P(Z > \lvert Z_{\text{calc}} \rvert))
Reject H0H_0 if μ0∉Xˉ±zα/2sn\mu_0 \notin \bar{X} \pm z_{\alpha/2}\frac{s}{\sqrt{n}}
Left-TailedH0:μ≥μ0H_0: \mu \ge \mu_0
H1:μ<μ0H_1: \mu < \mu_0
Reject H0H_0 if Zcalc<−zαZ_{\text{calc}} < -z_{\alpha}Reject H0H_0 if p≤αp \le \alpha
(p=P(Z<Zcalc)p = P(Z < Z_{\text{calc}}))
N/A (One-sided bound)
Right-TailedH0:μ≤μ0H_0: \mu \le \mu_0
H1:μ>μ0H_1: \mu > \mu_0
Reject H0H_0 if Zcalc>zαZ_{\text{calc}} > z_{\alpha}Reject H0H_0 if p≤αp \le \alpha
(p=P(Z>Zcalc)p = P(Z > Z_{\text{calc}}))
N/A (One-sided bound)

3. Two-Sample Z-Test for Independent Means​

Compares the means of two independent populations (μ1\mu_1 vs. μ2\mu_2) from two samples (n1≥30,n2≥30n_1 \ge 30, n_2 \ge 30).

Hypotheses​

  • Two-tailed: H0:μ1=μ2  ⟺  μ1−μ2=0H_0: \mu_1 = \mu_2 \iff \mu_1 - \mu_2 = 0 against H1:μ1≠μ2H_1: \mu_1 \neq \mu_2
  • One-tailed: H0:μ1≤μ2H_0: \mu_1 \le \mu_2 against H1:μ1>μ2H_1: \mu_1 > \mu_2

Test Statistic​

Zcalc=(Xˉ1−Xˉ2)−(μ1−μ2)σ12n1+σ22n2Z_{\text{calc}} = \frac{(\bar{X}_1 - \bar{X}_2) - (\mu_1 - \mu_2)}{\sqrt{\frac{\sigma_1^2}{n_1} + \frac{\sigma_2^2}{n_2}}}

When σ12,σ22\sigma_1^2, \sigma_2^2 are unknown:

Zcalc=Xˉ1−Xˉ2s12n1+s22n2Z_{\text{calc}} = \frac{\bar{X}_1 - \bar{X}_2}{\sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}}

Equal Variance Check (Levene's Test)​

Before conducting two-sample tests, evaluate whether population variances are equal using Levene's test (scipy.stats.levene):

  • H0:σ12=σ22H_0: \sigma_1^2 = \sigma_2^2 vs. H1:σ12≠σ22H_1: \sigma_1^2 \neq \sigma_2^2
  • If p≥0.05p \ge 0.05, fail to reject H0H_0 and assume equal variances.

4. End-to-End Problem Scenarios from Course Material​

Scenario 1: Two-Tailed Test (Car Mileage Claim)​

Problem: A car manufacturer claims the mileage of their new car is μ=25 km/L\mu = 25\text{ km/L} with σ=2.5 km/L\sigma = 2.5\text{ km/L}. A sample of n=45n = 45 cars gives a mean of Xˉ=24 km/L\bar{X} = 24\text{ km/L}. Is there evidence to claim that the mean mileage differs from 25 km/L25\text{ km/L}? Use α=0.01\alpha = 0.01.

  1. Formulate Hypotheses: H0:μ=25vs.H1:μ≠25(Two-Tailed)H_0: \mu = 25 \quad \text{vs.} \quad H_1: \mu \neq 25 \quad (\text{Two-Tailed})
  2. Calculate Test Statistic: Zcalc=Xˉ−μ0σ/n=24−252.5/45=−12.5/6.708=−10.3727≈−2.683Z_{\text{calc}} = \frac{\bar{X} - \mu_0}{\sigma / \sqrt{n}} = \frac{24 - 25}{2.5 / \sqrt{45}} = \frac{-1}{2.5 / 6.708} = \frac{-1}{0.3727} \approx -2.683
  3. Method A (Critical Value): For α=0.01\alpha = 0.01, zα/2=z0.005=2.575z_{\alpha/2} = z_{0.005} = 2.575. Since ∣Zcalc∣=∣−2.683∣=2.683>2.575|Z_{\text{calc}}| = |-2.683| = 2.683 > 2.575, we Reject H0H_0.
  4. Method B (P-Value): p=2⋅P(Z<−2.683)=2⋅0.00366=0.0073p = 2 \cdot P(Z < -2.683) = 2 \cdot 0.00366 = 0.0073 Since p-value=0.0073<0.01p\text{-value} = 0.0073 < 0.01, we Reject H0H_0.
  5. Method C (Confidence Interval): CI99%=24±2.575⋅0.3727=24±0.960=[23.04,  24.96]\text{CI}_{99\%} = 24 \pm 2.575 \cdot 0.3727 = 24 \pm 0.960 = [23.04, \; 24.96] Since hypothesized mean μ0=25\mu_0 = 25 does not lie within [23.04,24.96][23.04, 24.96], we Reject H0H_0.
  6. Conclusion: There is statistically significant evidence at α=0.01\alpha = 0.01 to conclude the mean mileage is different from 25 km/L25\text{ km/L}.

Scenario 2: Left-Tailed Test (PVC Pipe Thickness)​

Problem: A sample of n=900n = 900 PVC pipes has a mean thickness of Xˉ=12.5 mm\bar{X} = 12.5\text{ mm}. Test whether the sample comes from a population with mean μ≥13 mm\mu \ge 13\text{ mm} against the claim that it is less than 13 mm13\text{ mm}, given σ=1 mm\sigma = 1\text{ mm} at α=0.05\alpha = 0.05.

  1. Formulate Hypotheses: H0:μ≥13vs.H1:μ<13(Left-Tailed)H_0: \mu \ge 13 \quad \text{vs.} \quad H_1: \mu < 13 \quad (\text{Left-Tailed})
  2. Calculate Test Statistic: Zcalc=Xˉ−μ0σ/n=12.5−131/900=−0.51/30=−0.5⋅30=−15.0Z_{\text{calc}} = \frac{\bar{X} - \mu_0}{\sigma / \sqrt{n}} = \frac{12.5 - 13}{1 / \sqrt{900}} = \frac{-0.5}{1 / 30} = -0.5 \cdot 30 = -15.0
  3. Decision: Critical value for left-tailed test at α=0.05\alpha = 0.05 is −z0.05=−1.645-z_{0.05} = -1.645. Since Zcalc=−15.0<−1.645Z_{\text{calc}} = -15.0 < -1.645 (and p-value≈0.000p\text{-value} \approx 0.000), we Reject H0H_0.
  4. Conclusion: Significant evidence indicates average thickness is less than 13 mm13\text{ mm}.

Scenario 3: Right-Tailed Test (Food Delivery Time)​

Problem: An e-commerce food app claims delivery time is at most 60 minutes (μ≤60\mu \le 60) with σ=30\sigma = 30 minutes. A sample of n=45n = 45 orders yields a mean of Xˉ=75\bar{X} = 75 minutes. Test the claim at α=0.05\alpha = 0.05.

  1. Formulate Hypotheses: H0:μ≤60vs.H1:μ>60(Right-Tailed)H_0: \mu \le 60 \quad \text{vs.} \quad H_1: \mu > 60 \quad (\text{Right-Tailed})
  2. Calculate Test Statistic: Zcalc=Xˉ−μ0σ/n=75−6030/45=154.472≈+3.35Z_{\text{calc}} = \frac{\bar{X} - \mu_0}{\sigma / \sqrt{n}} = \frac{75 - 60}{30 / \sqrt{45}} = \frac{15}{4.472} \approx +3.35
  3. Decision: Critical value for right-tailed test at α=0.05\alpha = 0.05 is +z0.05=+1.645+z_{0.05} = +1.645. p-value=P(Z>3.35)=0.0004p\text{-value} = P(Z > 3.35) = 0.0004. Since p-value<0.05p\text{-value} < 0.05 and Zcalc>1.645Z_{\text{calc}} > 1.645, we Reject H0H_0.
  4. Conclusion: Strong evidence that the mean delivery time exceeds 60 minutes.

Scenario 4: Two-Sample Test (Hemoglobin Levels)​

Problem: A blood study evaluates hemoglobin levels for males and females:

  • Males: n1=160n_1 = 160, Xˉ1=13 g/dL\bar{X}_1 = 13\text{ g/dL}, s1=4.1 g/dLs_1 = 4.1\text{ g/dL}
  • Females: n2=180n_2 = 180, Xˉ2=15 g/dL\bar{X}_2 = 15\text{ g/dL}, s2=3.5 g/dLs_2 = 3.5\text{ g/dL} Test whether population mean hemoglobin levels differ between men and women at α=0.01\alpha = 0.01.
  1. Formulate Hypotheses: H0:μ1=μ2vs.H1:μ1≠μ2(Two-Tailed)H_0: \mu_1 = \mu_2 \quad \text{vs.} \quad H_1: \mu_1 \neq \mu_2 \quad (\text{Two-Tailed})
  2. Calculate Test Statistic: SEdiff=s12n1+s22n2=4.12160+3.52180=16.81160+12.25180=0.1051+0.0681=0.1732≈0.4161\text{SE}_{\text{diff}} = \sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}} = \sqrt{\frac{4.1^2}{160} + \frac{3.5^2}{180}} = \sqrt{\frac{16.81}{160} + \frac{12.25}{180}} = \sqrt{0.1051 + 0.0681} = \sqrt{0.1732} \approx 0.4161 Zcalc=Xˉ1−Xˉ2SEdiff=13−150.4161=−20.4161≈−4.807Z_{\text{calc}} = \frac{\bar{X}_1 - \bar{X}_2}{\text{SE}_{\text{diff}}} = \frac{13 - 15}{0.4161} = \frac{-2}{0.4161} \approx -4.807
  3. Decision: Critical value at α=0.01\alpha = 0.01 is z0.005=2.575z_{0.005} = 2.575. Since ∣Zcalc∣=∣−4.807∣=4.807>2.575|Z_{\text{calc}}| = |-4.807| = 4.807 > 2.575 (and p-value<0.0001p\text{-value} < 0.0001), we Reject H0H_0.
  4. Conclusion: There is strong empirical evidence that average hemoglobin levels differ significantly between men and women.

5. Implementation Lab​

Execute large sample Z-tests in Python utilizing statsmodels and scipy.stats.

Open In Colab
import numpy as np
from scipy import stats
from statsmodels.stats.weightstats import ztest

# 1. One-Sample Z-Test Function
def manual_z_test(samp_mean, pop_mean, std_dev, n, alternative='two-sided'):
z_stat = (samp_mean - pop_mean) / (std_dev / np.sqrt(n))
if alternative == 'two-sided':
p_val = 2 * (1 - stats.norm.cdf(np.abs(z_stat)))
elif alternative == 'smaller':
p_val = stats.norm.cdf(z_stat)
elif alternative == 'larger':
p_val = 1 - stats.norm.cdf(z_stat)
return z_stat, p_val

# Car Mileage Scenario (Two-Tailed)
z_calc, p_val = manual_z_test(samp_mean=24, pop_mean=25, std_dev=2.5, n=45, alternative='two-sided')
print(f"Car Mileage Z-Stat: {z_calc:.3f}, P-Value: {p_val:.5f}")

# 2. Two-Sample Z-Test Function (Hemoglobin Scenario)
def two_sample_z_test(m1, s1, n1, m2, s2, n2):
se_diff = np.sqrt((s1**2 / n1) + (s2**2 / n2))
z_stat = (m1 - m2) / se_diff
p_val = 2 * (1 - stats.norm.cdf(np.abs(z_stat)))
return z_stat, p_val

z_hemo, p_hemo = two_sample_z_test(13, 4.1, 160, 15, 3.5, 180)
print(f"Hemoglobin Z-Stat: {z_hemo:.3f}, P-Value: {p_hemo:.6f}")

info

Key Takeaways

  • Z-Test Criterion: Appropriate whenever n≥30n \ge 30 due to asymptotic normality from the Central Limit Theorem.
  • Three Equivalent Paths: For two-tailed tests, Critical Value (∣Z∣>zα/2|Z| > z_{\alpha/2}), P-value (p≤αp \le \alpha), and Confidence Interval (μ0∉CI\mu_0 \notin \text{CI}) always agree.
  • Two-Sample Variance Pooling: If variances are unknown but sample sizes are large (n1,n2≥30n_1, n_2 \ge 30), substituting individual sample variances s12s_1^2 and s22s_2^2 is asymptotically exact.