Module 3 Sample Questions: Estimation & Hypothesis Testing
This document provides rigorous, textbook-style solutions for the Module 3 examination questions and real-world scenario problems. All formulas are rendered in LaTeX, accompanied by step-by-step mathematical derivations, visual normal distribution graphs, decision flowcharts, and formal academic inferences.
Question 1: Concept of Point Estimation (Average Income)
Question Details
What would be the average income of a 27-year-old in London? How would you calculate it?
Step-by-Step Solution
1. Theoretical Formulation
Let the random variable represent the annual income of a 27-year-old individual residing in London. The true arithmetic mean income across all such individuals in the target population is denoted by the population parameter :
where is the total population size of 27-year-olds in London.
2. The Operational Reality (Need for Sampling)
- The Trivial Approach: Conduct a complete census by querying every single 27-year-old in London and computing .
- Practical Infeasibility: A full census is cost-prohibitive, operationally difficult, time-consuming, and subject to non-response bias.
- The Inferential Solution: Draw a representative Simple Random Sample (SRS) of size from the population.
3. Mathematical Point Estimator
The sample mean () serves as the best point estimator for the population mean :
- Unbiasedness Property: The sample mean is an unbiased estimator of because its expected value equals the parameter:
- Consistency Property: By the Weak Law of Large Numbers, as sample size , the sample mean converges in probability to :
Academic Inference: Because obtaining the exact population parameter via census is practically unviable, we draw an unbiased random sample of size , calculate the sample mean , and report as the point estimate for the true average income of a 27-year-old in London.
Question 2: Point Estimate of Count via Sample Proportion
Question Details
A financial firm has created 50 portfolios. From them a sample of 13 portfolios was selected, out of which 8 were found to be underperforming. Estimate the number of underperforming portfolios.
Step-by-Step Derivation
1. Given Parameters
- Total population size of portfolios:
- Sample size audited:
- Number of underperforming portfolios in the sample:
2. Calculation of Sample Proportion ()
The sample proportion of underperforming portfolios is computed as:
3. Estimation of Total Underperforming Portfolios ()
To estimate the total number of underperforming portfolios in the entire population of , we scale the point estimate by the population size :
Since the count of portfolios must be a discrete non-negative integer, we round to the nearest whole unit:
Academic Inference: Based on the observed sample failure rate of approximately , the point estimate for the total number of underperforming portfolios across the firm's 50 investment portfolios is 31 portfolios.
Question 3: Margin of Error Calculation for Population Mean
Question Details
100 bags of coal were tested and had an average of 35% of ash with a standard deviation of 15%. Calculate the margin of error for a 90% confidence level.
Step-by-Step Derivation
1. Given Parameters
- Sample size:
- Sample mean ash content:
- Standard deviation:
- Confidence level:
2. Determination of Critical Value ()
For a two-sided confidence interval at , the tail probability in each half is :
3. Standard Error of the Mean (SE)
4. Computation of Margin of Error ()
The margin of error is defined as the maximum expected distance between the sample estimate and the true population mean:
(Note: If using , or ).
Academic Inference: At a 90% level of confidence, the margin of error is 2.47% (or 0.0247). This indicates that the true population mean ash content is expected to lie within percentage points of the sample estimate across 90% of repeated random samples.
Question 4: Sample Size Verification from Fixed Margin of Error
Question Details
For the previous question, for a margin of error 2.46%. Verify whether you get a sample size of 100.
Step-by-Step Derivation
1. Mathematical Formula for Sample Size
Starting from the margin of error equation:
Multiplying both sides by and dividing by :
Squaring both sides yields the sample size determination formula:
2. Verification Calculation
- Given margin of error:
- Standard deviation:
- Critical value for 90% confidence (): (or )
Using :
Using standard three-decimal precision and exact rounding:
Academic Inference: The algebraic verification confirms that a desired margin of error of 2.46% under a 90% confidence level and requires an exact sample size of bags.
Question 5: Confidence Interval for Population Mean (London Incomes)
Question Details
From a sample of 250 observations it is found that average income of a 27-year-old Londoner is £45,000 with a population standard deviation of £4000. Obtain the 95% confidence interval to estimate the average income. Furthermore, compare the width of 90%, 95%, and 99% confidence intervals.
Step-by-Step Derivation
1. Given Data
- Sample size:
- Sample mean income:
- Population standard deviation:
- Confidence level:
2. Critical Value Determination
For a two-tailed 95% confidence interval:
3. Standard Error Calculation
4. Margin of Error ()
5. Confidence Interval Boundaries
- Lower Limit ():
- Upper Limit ():
6. Visualizing Confidence Level vs. Interval Width
Academic Inference & Interpretation Trap:
- Width vs. Certainty: As confidence increases from 90% to 99%, the critical value grows (), creating a progressively wider interval. Higher certainty requires casting a broader net.
- Interpretation Trap: It is mathematically incorrect to state that . The population mean is an invariant constant, not a random variable. The 95% confidence reflects the long-run coverage rate of the statistical procedure.
Question 6: Confidence Interval for Population Proportion
Question Details
A financial firm has created 50 portfolios. From them a sample of 13 portfolios was selected, out of which 8 were found to be underperforming. Construct a 99% confidence interval to estimate the population proportion.
Step-by-Step Derivation
1. Given Data
- Sample size:
- Number of underperforming units:
- Sample proportion:
- Confidence level:
2. Critical Value Determination
For a two-tailed 99% confidence interval:
3. Standard Error of Proportion ()
4. Margin of Error ()
5. Confidence Interval Boundaries
- Lower Limit:
- Upper Limit:
Academic Inference: We are 99% confident that the true population proportion of underperforming portfolios across the firm lies between 26.79% and 96.28%. The wide span of this interval reflects the high uncertainty resulting from a very small sample size ().
Question 7: Destructive Testing & The Rationale for Hypothesis Testing
Question Details
A manufacturer produces batteries. He claims that the life of his batteries is 25,500 hours. Can we verify his claim?
Step-by-Step Solution
1. The Trivial Verification Approach
The naive way to verify this claim is to measure the operating lifespan of every single battery produced by running each unit until its energy is completely depleted.
2. The Operational Dilemma: Destructive Testing
Lifespan testing is inherently destructive testing. Running a battery to exhaustion consumes the product entirely.
3. The Statistical Alternative: Hypothesis Testing
To overcome this dilemma, we execute Statistical Hypothesis Testing:
- Formulate Statistical Claims:
- Null Hypothesis (): (Manufacturer claim is true)
- Alternative Hypothesis (): (Watchdog claim that batteries underperform)
- Representative Sampling: Extract a small, random sample of batteries (e.g., ).
- Controlled Depletion: Test only the sample units to failure to compute the sample mean and sample standard deviation .
- Inferential Decision: Evaluate the -statistic or -statistic. If the empirical evidence deviates significantly from hours beyond random sampling chance, we reject ; otherwise, we fail to reject .
Academic Inference: We cannot verify the claim deterministically across the entire population because testing is destructive. Instead, we statistically test the manufacturer's claim using a representative sample, balancing statistical rigor with inventory preservation.
Question 8: Formulation of Null and Alternative Hypotheses
Question Details
Write the null and alternative hypotheses for the following scenarios and classify the test type (number of tails):
- A bolt manufacturing company claims that the average number of bolts manufactured in a day is 60.
- The new virus testing kit takes less time than the standard testing kit that takes 34 hours for the results.
- An analyst wants to test whether Apple Inc. stock has outperformed in 2019 by 13.5%.
Step-by-Step Formulations
Scenario 1: Bolt Manufacturing
- Target Metric: Let denote the number of bolts manufactured daily, and denote the true mean daily production.
- Manufacturer Claim: Daily production equals 60.
- Hypotheses:
- Classification: Two-Tailed Test. A defect or deviation occurs if the machinery under-produces () or over-produces (). The rejection region is split equally into both tails ().
Scenario 2: Virus Testing Kit Processing Time
- Target Metric: Let denote the time (in hours) required to obtain diagnostic results, and denote the true mean result latency.
- Research Claim: The new kit reduces processing time below the standard 34 hours ().
- Hypotheses:
- Classification: Left-Tailed Test (One-Tailed). The alternative claim specifically focuses on reduction/improvement. The entire rejection region is concentrated in the lower (left) tail.
Scenario 3: Apple Inc. Stock Outperformance
- Target Metric: Let denote the percentage return/performance of Apple Inc. stock in 2019, and denote the true average performance.
- Analyst Claim: The stock outperformed the baseline target of 13.5% ().
- Hypotheses:
- Classification: Right-Tailed Test (One-Tailed). The claim tests strictly for an increase or outperformance above the benchmark. The rejection region is placed entirely in the upper (right) tail.
Question 9: Practice Identification of Test Directionality
Question Details
For each scenario, define the parameter , state and , and identify whether the test is two-tailed, left-tailed, or right-tailed:
- Mean diastolic blood pressure for a group of 85 adults is less than 90 mm.
- The tensile strength of an alloy is more than 500 Pa.
- The average heart rate of a healthy human is 85 beats per minute.
- The yield of BT cotton is more than 450 kg lint per hectare.
Classification Matrix
| Objective / Statement | Parameter Definition () | Null Hypothesis () | Alternative Hypothesis () | Test Classification | Critical Region |
|---|---|---|---|---|---|
| 1. Diastolic Blood Pressure | Mean adult diastolic blood pressure (mm Hg) | Left-Tailed | Lower tail () | ||
| 2. Alloy Tensile Strength | Mean breaking tensile strength (Pa) | Right-Tailed | Upper tail () | ||
| 3. Healthy Human Heart Rate | Mean healthy resting heart rate (bpm) | Two-Tailed | Both tails () | ||
| 4. BT Cotton Crop Yield | Mean cotton yield (kg lint/hectare) | Right-Tailed | Upper tail () |
Question 10: Two-Tailed Large Sample Z-Test (Car Mileage)
Question Details
A car manufacturing company claims that the mileage of their new car is 25 kmph with a standard deviation of 2.5 kmph. A random sample of 45 cars was drawn and recorded their mileage as per the standard procedure. From the sample, the mean mileage was seen to be 24 kmph. Is this evidence to claim that the mean mileage is different from 25 kmph? (Assume normality of data). Test the claim using:
- Critical region technique
- p-value technique
- Confidence interval technique (Use ).
Step-by-Step Solution
1. Hypothesis Formulation
- (Mean mileage is equal to 25 kmph)
- (Mean mileage is significantly different from 25 kmph)
- Test type: Two-Tailed Test, .
2. Given Sample Data & Standard Error
- , , , .
- Since , the Central Limit Theorem validates the Z-test.
3. Calculation of Test Statistic ()
4. Distribution & Rejection Graph
Three Decision Methods Evaluation
- 1. Critical Region Method
- 2. P-Value Method
- 3. Confidence Interval Method
- Critical Cutoff Value: For (two-tailed), .
- Decision Rule: Reject if .
- Comparison:
- Decision: Reject . The test statistic falls in the lower rejection tail.
- P-Value Definition: Total tail probability beyond the observed test statistic:
- Probability Evaluation:
- Decision Rule: Reject if .
- Comparison:
- Decision: Reject .
- Construct 99% CI:
- Decision Rule: Reject if hypothesized mean does not fall within the confidence interval.
- Comparison:
- Decision: Reject .
Academic Inference: All three evaluation criteria converge to the identical conclusion: Reject at the 1% significance level. There is strong empirical evidence that the true average mileage of the new car differs from 25 kmph (specifically, it is statistically significantly lower at 24 kmph).
Question 11: One-Sample Left-Tailed Z-Test (PVC Pipe Thickness)
Question Details
A sample of 900 PVC pipes is found to have an average thickness of 12.5 mm. Can we assume that the sample is coming from a normal population with mean 13 mm against that it is less than 13 mm? The population standard deviation is 1 mm. Test the hypothesis using the p-value method at 5% level of significance.
Step-by-Step Derivation
1. Hypothesis Formulation
- (Mean thickness is at least 13 mm)
- (Mean thickness is strictly less than 13 mm)
- Test type: Left-Tailed Test, .
2. Given Data
- Hypothesized mean:
- Population standard deviation:
- Sample size:
- Sample mean:
3. Standard Error Calculation
4. Test Statistic ()
5. Distribution & Left-Tail Rejection Graph
6. Statistical Decision
- Critical cutoff value for left-tailed test at is .
- Since and :
Academic Inference: We reject the null hypothesis at . There is overwhelming statistical evidence to conclude that the sample does not originate from a population with a mean thickness of 13 mm, and that the true average thickness is significantly less than 13 mm.
Question 12: One-Sample Right-Tailed Z-Test (Food Delivery Time)
Question Details
An e-commerce company claims that the mean delivery time of food items on their website in NYC is 60 minutes with a standard deviation of 30 minutes. A random sample of 45 customers ordered from the website, and the mean time for delivery was found to be 75 minutes. Is this enough evidence to claim that the mean time to get items delivered is more than 60 minutes? (Assume normality of data). Test the claim using the p-value technique at .
Step-by-Step Derivation
1. Hypothesis Formulation
- (Mean delivery time is at most 60 minutes)
- (Mean delivery time exceeds 60 minutes)
- Test type: Right-Tailed Test, .
2. Given Data
- Hypothesized mean:
- Population standard deviation:
- Sample size:
- Sample mean:
3. Standard Error Calculation
4. Test Statistic ()
5. Distribution & Right-Tail Rejection Graph
6. Statistical Decision
- Critical threshold for right-tailed test at is .
- Comparing probabilities:
- Decision: Reject .
Academic Inference: We reject the null hypothesis at . There is sufficient empirical evidence to conclude that the true mean delivery time for food items in NYC significantly exceeds 60 minutes.
Question 13: One-Sample Mean Test with Unknown (Protein Powder)
Question Details
The manager of a packaging process at a protein powder manufacturing plant wants to determine if the protein powder packing process is in control. The correct amount of protein powder per box is 350 grams on an average. A sample of 80 boxes was drawn which gave a mean of 354.5 grams with a standard deviation of 15. At 5% level of significance, is there evidence to suggest that the weight is different from 350 grams?
Step-by-Step Derivation
1. Hypothesis Formulation
- (Packaging process is in statistical control)
- (Packaging process is out of control)
- Test type: Two-Tailed Test, .
2. Given Data & Large-Sample Property
- Specified parameter:
- Sample size:
- Sample mean:
- Sample standard deviation: (population unknown)
- Since , by Slutsky's theorem and the Central Limit Theorem, the sample standard deviation can replace with the standard normal approximation:
3. Test Statistic Calculation
4. Critical Value and Decision Rule
- For (two-tailed), critical cutoff is .
- Decision rule: Reject if .
- Comparing test statistic to critical cutoff:
- P-value verification: .
- Decision: Reject .
Academic Inference: We reject the null hypothesis at the 5% level of significance. There is statistically significant evidence that the average weight of protein powder in the boxes differs from 350 grams (the machinery is overfilling boxes by an average of 4.5 grams), indicating the packaging process is currently out of statistical control.
Question 14: Two-Sample Z-Test for Independent Means (Hemoglobin Study)
Question Details
A study was carried out to understand the amount of hemoglobin in blood for males and females. A random sample of 160 males and 180 females have means of 13 g/dL and 15 g/dL. The two samples have standard deviations of 4.1 g/dL for male donors and 3.5 g/dL for female donors. Can it be said that the population means of hemoglobin are the same for men and women? Use .
Step-by-Step Derivation
1. Hypothesis Formulation
Let denote the population mean hemoglobin for males, and denote the population mean hemoglobin for females:
- (Male and female mean hemoglobin levels are equal)
- (Male and female mean hemoglobin levels differ)
- Test type: Two-Tailed Independent Two-Sample Test, .
2. Given Sample Statistics
- Male Cohort (Sample 1): , ,
- Female Cohort (Sample 2): , ,
3. Standard Error of Difference Between Means ()
4. Test Statistic ()
5. Decision Rule and Conclusion
- For a two-tailed test at , the critical cutoff is .
- Critical Region Comparison:
- P-Value Comparison:
- Decision: Reject .
Academic Inference: We reject the null hypothesis at . The difference of between male and female donors is highly statistically significant (). We conclude that the population mean hemoglobin concentrations for men and women are not equal.
Quick-Facts & Formula Sheet (Estimation & Large Sample Tests)
| Statistical Concept | Mathematical Formulation | Conditions / Notes |
|---|---|---|
| Point Estimate for Mean | Unbiased estimator; . | |
| Point Estimate for Proportion | Unbiased estimator for binomial trials. | |
| Sampling Error | Decreases as increases. | |
| Margin of Error (Mean) | Half-width of confidence interval. | |
| Required Sample Size | Inversely proportional to . | |
| Confidence Interval (Mean, known) | Valid by CLT for . | |
| Confidence Interval (Mean, unknown) | substituted for when . | |
| Confidence Interval (Proportion) | Large sample normal approximation. | |
| One-Sample Z-Test Statistic | Tested against or . | |
| Two-Sample Z-Test Statistic | Compares equality of two population means. | |
| Type I Error () | Significance level; false alarm rate. | |
| Type II Error () | Missed detection rate. | |
| Statistical Power | Probability of correctly rejecting false . |