Hypothesis Testing Foundations
Topic - A hypothesis is a formal, testable claim regarding an unknown population parameter. Hypothesis testing provides an objective statistical protocol to assess whether empirical sample data provides strong enough evidence to refute a baseline status-quo assumption.
1. Formulating Hypotheses
Hypothesis testing sets up a contest between two contrasting statements:
The Null Hypothesis ()
The hypothesis of no effect, no difference, or status quo. It represents the baseline assumption that is presumed true unless the sample evidence proves overwhelmingly otherwise.
- Contains an equality condition (, , or ).
- Example: A manufacturing plant claims battery lifespan is 25,500 hours ().
The Alternative Hypothesis ( or )
The research hypothesis or claim of interest. It represents the claim that there is a genuine effect, difference, or change.
- Directly contradicts the null hypothesis.
- Contains a strict inequality (, , or ).
- Example: A consumer watchdog claims the battery lifespan is less than 25,500 hours ().
2. The Decision Paradigm: "Fail to Reject" vs. "Accept"
In scientific hypothesis testing, we never "accept" the null hypothesis. We either:
- Reject : The sample evidence is sufficiently extreme that the null hypothesis is highly unlikely to be true.
- Fail to Reject : There is insufficient sample evidence to conclude that the null hypothesis is false.
The Legal Analogy: Presumption of Innocence
In a criminal trial, the defendant is presumed innocent until proven guilty (: Innocent, : Guilty). If evidence is lacking, the jury returns a verdict of "Not Guilty", not "Innocent". Failing to prove guilt does not prove innocence; it merely reflects an absence of conclusive evidence beyond a reasonable doubt.
3. Errors in Hypothesis Testing
Because sample statistics fluctuate due to random sampling variability, decisions are subject to two fundamental types of error:
The Error Matrix
| Reality | Decision: Reject | Decision: Fail to Reject |
|---|---|---|
| is Actually True | Type I Error () False Alarm (convicting an innocent person) | Correct Decision () Confidence Level |
| is Actually False | Correct Decision () Statistical Power | Type II Error () Missed Detection (letting a guilty person walk free) |
Tradeoffs and Error Control
- (Significance Level): The maximum probability of committing a Type I error that the researcher is willing to tolerate (typically set to or ).
- : The probability of committing a Type II error.
- The Inversion: For a fixed sample size , lowering (making it harder to reject ) automatically increases (increasing the chance of missing a real effect). Both errors cannot be minimized simultaneously without increasing the sample size .
Power of the Test ()
The probability of correctly rejecting a false null hypothesis. High power () indicates that the test is sensitive enough to detect true effects.
4. Number of Tails
The mathematical direction in determines the placement of the rejection region under the null sampling distribution:
- Two-Tailed Test (≠)
- Left-Tailed Test (<)
- Right-Tailed Test (>)
- Hypothesis: vs.
- Rejection Region: Split equally into both tails ( in each tail).
- Use Case: Testing whether average output differs in either direction (e.g., machine filling bottles: both underfilling and overfilling are defects).
- Hypothesis: vs.
- Rejection Region: Entire significance area is concentrated in the left (lower) tail.
- Use Case: Testing for reductions (e.g., drug reducing recovery time, algorithm reducing API response latency).
- Hypothesis: vs.
- Rejection Region: Entire significance area is concentrated in the right (upper) tail.
- Use Case: Testing for increases (e.g., campaign increasing sales, engine increasing fuel efficiency).
5. Three Decision-Making Criteria
There are three mathematically equivalent methods for deciding whether to reject :
1. The Critical Region (Rejection Region) Approach
- Determine critical cutoff from the normal distribution table for the chosen .
- Two-tailed: Reject if .
- Left-tailed: Reject if .
- Right-tailed: Reject if .
2. The P-Value Approach
The -value is the probability, assuming is true, of observing a test statistic as extreme as, or more extreme than, the value obtained from our sample.
- Decision Rule:
- "If the p-value is low, the null must go. If the p-value is high, the null will fly."
3. The Confidence Interval Approach (Two-Tailed Tests)
Construct the confidence interval around the sample mean .
- If the hypothesized population parameter does not lie inside the confidence interval, Reject at significance level .
- If falls within the interval, Fail to Reject .
6. Synthesis & Review
Testing Principles
- Directionality: Direction of determines the critical tail. yields two tails; < yields left tail; > yields right tail.
- Equivalence: The Critical Region, P-value, and Confidence Interval approaches always yield identical statistical decisions for two-tailed tests.
Common Traps
- Never say 'Accept H0': Lack of evidence against the null is not proof of the null.
- Do not change tails post-hoc: Formulate and before observing the sample data to avoid p-hacking.
Key Takeaways
- Alpha (): The false-positive threshold, set by design prior to experimentation.
- P-Value: A measure of evidence against . Smaller -values indicate stronger evidence against the null hypothesis.
- Statistical Power (): The ability of an experiment to detect an actual effect when one exists. Higher sample sizes increase power.
Next Section: Large Sample Tests - implementing one-sample and two-sample Z-tests in Python.