Skip to main content

Hypothesis Testing Foundations

Topic - A hypothesis is a formal, testable claim regarding an unknown population parameter. Hypothesis testing provides an objective statistical protocol to assess whether empirical sample data provides strong enough evidence to refute a baseline status-quo assumption.


1. Formulating Hypotheses​

Hypothesis testing sets up a contest between two contrasting statements:

The Null Hypothesis (H0H_0)​

The hypothesis of no effect, no difference, or status quo. It represents the baseline assumption that is presumed true unless the sample evidence proves overwhelmingly otherwise.

  • Contains an equality condition (==, ≥\ge, or ≤\le).
  • Example: A manufacturing plant claims battery lifespan is 25,500 hours (H0:μ=25500H_0: \mu = 25500).

The Alternative Hypothesis (H1H_1 or HaH_a)​

The research hypothesis or claim of interest. It represents the claim that there is a genuine effect, difference, or change.

  • Directly contradicts the null hypothesis.
  • Contains a strict inequality (≠\neq, <<, or >>).
  • Example: A consumer watchdog claims the battery lifespan is less than 25,500 hours (H1:μ<25500H_1: \mu < 25500).

2. The Decision Paradigm: "Fail to Reject" vs. "Accept"​

In scientific hypothesis testing, we never "accept" the null hypothesis. We either:

  1. Reject H0H_0: The sample evidence is sufficiently extreme that the null hypothesis is highly unlikely to be true.
  2. Fail to Reject H0H_0: There is insufficient sample evidence to conclude that the null hypothesis is false.
tip

The Legal Analogy: Presumption of Innocence

In a criminal trial, the defendant is presumed innocent until proven guilty (H0H_0: Innocent, H1H_1: Guilty). If evidence is lacking, the jury returns a verdict of "Not Guilty", not "Innocent". Failing to prove guilt does not prove innocence; it merely reflects an absence of conclusive evidence beyond a reasonable doubt.


3. Errors in Hypothesis Testing​

Because sample statistics fluctuate due to random sampling variability, decisions are subject to two fundamental types of error:

The Error Matrix​

RealityDecision: Reject H0H_0Decision: Fail to Reject H0H_0
H0H_0 is Actually TrueType I Error (α\alpha)
False Alarm (convicting an innocent person)
Correct Decision (1−α1 - \alpha)
Confidence Level
H0H_0 is Actually FalseCorrect Decision (1−β1 - \beta)
Statistical Power
Type II Error (β\beta)
Missed Detection (letting a guilty person walk free)

Tradeoffs and Error Control​

  • α\alpha (Significance Level): The maximum probability of committing a Type I error that the researcher is willing to tolerate (typically set to 0.050.05 or 0.010.01).
  • β\beta: The probability of committing a Type II error.
  • The Inversion: For a fixed sample size nn, lowering α\alpha (making it harder to reject H0H_0) automatically increases β\beta (increasing the chance of missing a real effect). Both errors cannot be minimized simultaneously without increasing the sample size nn.

Power of the Test (1−β1 - \beta)​

The probability of correctly rejecting a false null hypothesis. High power (>80%> 80\%) indicates that the test is sensitive enough to detect true effects.


4. Number of Tails​

The mathematical direction in H1H_1 determines the placement of the rejection region under the null sampling distribution:


  • Hypothesis: H0:μ=μ0H_0: \mu = \mu_0 vs. H1:μ≠μ0H_1: \mu \neq \mu_0
  • Rejection Region: Split equally into both tails (α/2\alpha / 2 in each tail).
  • Use Case: Testing whether average output differs in either direction (e.g., machine filling bottles: both underfilling and overfilling are defects).

5. Three Decision-Making Criteria​

There are three mathematically equivalent methods for deciding whether to reject H0H_0:

1. The Critical Region (Rejection Region) Approach​

  • Determine critical cutoff zcritz_{\text{crit}} from the normal distribution table for the chosen α\alpha.
  • Two-tailed: Reject H0H_0 if ∣Zcalc∣>zα/2|Z_{\text{calc}}| > z_{\alpha/2}.
  • Left-tailed: Reject H0H_0 if Zcalc<−zαZ_{\text{calc}} < -z_{\alpha}.
  • Right-tailed: Reject H0H_0 if Zcalc>zαZ_{\text{calc}} > z_{\alpha}.

2. The P-Value Approach​

The pp-value is the probability, assuming H0H_0 is true, of observing a test statistic as extreme as, or more extreme than, the value obtained from our sample.

  • Decision Rule: If p-value≤α  ⟹  Reject H0\text{If } p\text{-value} \le \alpha \implies \text{Reject } H_0 If p-value>α  ⟹  Fail to Reject H0\text{If } p\text{-value} > \alpha \implies \text{Fail to Reject } H_0
  • "If the p-value is low, the null must go. If the p-value is high, the null will fly."

3. The Confidence Interval Approach (Two-Tailed Tests)​

Construct the 100(1−α)%100(1 - \alpha)\% confidence interval around the sample mean xˉ\bar{x}.

  • If the hypothesized population parameter μ0\mu_0 does not lie inside the confidence interval, Reject H0H_0 at significance level α\alpha.
  • If μ0\mu_0 falls within the interval, Fail to Reject H0H_0.

6. Synthesis & Review​

Testing Principles

  • Directionality: Direction of H1H_1 determines the critical tail. ≠\neq yields two tails; &lt; yields left tail; &gt; yields right tail.
  • Equivalence: The Critical Region, P-value, and Confidence Interval approaches always yield identical statistical decisions for two-tailed tests.

Common Traps

  • Never say 'Accept H0': Lack of evidence against the null is not proof of the null.
  • Do not change tails post-hoc: Formulate H0H_0 and H1H_1 before observing the sample data to avoid p-hacking.

info

Key Takeaways

  • Alpha (α\alpha): The false-positive threshold, set by design prior to experimentation.
  • P-Value: A measure of evidence against H0H_0. Smaller pp-values indicate stronger evidence against the null hypothesis.
  • Statistical Power (1−β1 - \beta): The ability of an experiment to detect an actual effect when one exists. Higher sample sizes increase power.

Next Section: Large Sample Tests - implementing one-sample and two-sample Z-tests in Python.