Module 5: Data Analysis - Problem Solving & Solutions
This document provides a comprehensive, exam-oriented set of scenario-based problem-solving questions and step-by-step solutions for Module 5: Data Analysis in Research Methodology.
1. Discrete Probability Distributions
Question 1: Quality Control in Manufacturing (Binomial Distribution)
Context: An electronics manufacturer produces microchips. Historically, the probability of any individual microchip being defective is 3% (). A quality assurance engineer randomly selects a batch of 50 microchips for inspection ().
- State the Probability Mass Function (PMF) that models this scenario.
- Formulate the equation to find the probability of finding exactly 3 defective microchips in this batch.
- Formulate the equation to find the probability that all microchips in the batch are non-defective (i.e., 0 defective).
Solution
1. Probability Mass Function (PMF): The scenario follows a Binomial Distribution because it consists of a fixed number of independent Bernoulli trials () with a constant probability of success/failure ().
Where:
- (number of trials / sample size)
- (probability of finding a defective component)
- is the number of defective components observed.
2. Probability of exactly 3 defective microchips ():
3. Probability of 0 defective microchips ():
Since and :
Engineer's Perspective: Quality Thresholds
In manufacturing pipelines, represents the likelihood of an uninterrupted, clean batch. If exceeds a predefined risk budget (e.g., 5%), automated sorting systems flag the entire lot for secondary manual inspection. The Binomial model provides a defensible statistical threshold for acceptance sampling.
Question 2: Modeling Server Breaches (Poisson Distribution)
Context: A network server experiences random security breach attempts. Historically, these breaches occur independently at a constant average rate of 4 incidents per month (). The IT security team wants to allocate response resources based on incident probabilities.
- Identify the appropriate probability distribution and write down its Probability Mass Function (PMF).
- Calculate the probability of observing exactly 2 breaches in a given month.
- Calculate the probability of observing no breaches in a month. What does this indicate for security readiness?
Solution
1. Distribution and PMF: The scenario models the count of rare, independent events occurring within a fixed interval of time at a known average rate. It follows the Poisson Distribution.
Where:
- (average rate of occurrences per interval)
- is the observed count of breaches.
2. Probability of exactly 2 breaches ():
Using :
3. Probability of no breaches ():
Interpretation: There is only approximately a 1.8% chance that a month will pass with zero breach attempts. The security operations center (SOC) cannot assume quiet periods and must staff continuous monitoring rotations.
Question 3: Geometric & Hypergeometric Sampling
Context A (Lead Conversion): A sales representative has a 10% chance () of closing a deal on any independent phone call.
- Identify the distribution used to model the number of calls required until the first successful sale, write its PMF, and formulate the probability that the first successful deal occurs on the 5th call.
Context B (Battery Audit Without Replacement): A warehouse receives a shipment of high-capacity batteries. An inspection reveals that of them are depleted. An auditor selects a random sample of batteries without replacement.
- Identify the probability distribution that models this sampling process and state its general formula.
Solution
Part A: Geometric Distribution
- Distribution: Geometric Distribution (models the number of Bernoulli trials needed to achieve the first success).
- PMF:
- Probability that the first success is on call 5 ():
Part B: Hypergeometric Distribution
- Distribution: Hypergeometric Distribution (models sampling from a finite population without replacement where probability shifts after each draw).
- PMF:
Where:
- (total population size)
- (total defective items in population)
- (sample size drawn)
- is the number of defective items observed in the sample.
2. Continuous Probability Distributions
Question 4: Conveyor Belt Lifetime (Exponential Distribution)
Context: A manufacturing assembly line uses industrial conveyor belts. The time between operational failures follows an exponential distribution with an average mean time between failures () of 200 hours. The failure rate is strictly constant and independent of the belt's operating age.
- Determine the rate parameter (). State the core property that justifies using the exponential distribution.
- Formulate the Cumulative Distribution Function (CDF) and compute the probability that the belt breaks down in less than or equal to 150 hours.
- Calculate the probability that the belt runs for at least 300 hours without breakdown.
Solution
1. Rate Parameter () and Justifying Property: The rate parameter is the reciprocal of the mean:
- Property: The Memoryless Property. Mathematically, . The future survival probability does not depend on how long the belt has already been running.
2. Probability of failure within 150 hours (): The CDF of the exponential distribution is:
Substitute and :
3. Probability of lasting at least 300 hours ():
Engineer's Perspective: When Memorylessness Fails
Assuming memorylessness is valid only when failures are caused by purely external, Poisson-distributed shocks (e.g., sudden voltage spikes). If mechanical wear, abrasion, or chemical degradation occurs, the failure rate increases over time, and the Exponential model must be replaced with a Weibull distribution.
Question 5: Predictive Maintenance of Brake Pads (Weibull Distribution)
Context: An automotive research lab evaluates a fleet of brake pads. Time-to-failure data indicates that degradation fits a Weibull distribution with a shape parameter (wear-out phase) and a scale parameter .
- Explain the mechanical and statistical meaning of the shape parameter relative to and .
- Using the Weibull CDF , calculate the probability that a brake pad fails within the first 15,000 miles.
- Formulate the algebraic steps to compute the exact mileage by which 90% of the brake pads will have failed ().
Solution
1. Interpretation of Shape Parameter :
- (Decreasing Failure Rate): "Infant mortality" period; items with manufacturing defects fail early.
- (Constant Failure Rate): Reduces strictly to the Exponential Distribution ().
- (Increasing Failure Rate): Characteristic wear-out behavior. Since , the hazard rate increases as operating mileage increases due to mechanical friction and material thinning.
2. Probability of failure within 15,000 miles ():
Using :
3. Expected mileage for 90% component failure ():
Take natural logarithm () on both sides:
Question 6: Identification of Continuous Distributions
Match each scientific modeling scenario to the appropriate continuous distribution and state its Probability Density Function (PDF):
- Scenario 1: Modeling user click-through rates or class probabilities restricted to the bounded domain .
- Scenario 2: Modeling the total waiting time until successive events occur in a Poisson process.
- Scenario 3: Modeling a hardware component simulation where every value inside interval has equal likelihood.
Solution
-
Beta Distribution (Scenario 1):
- Use Case: Proportions and probabilities bounded strictly on .
- PDF:
-
Gamma Distribution (Scenario 2):
- Use Case: Time until Poisson events occur.
- PDF:
-
Uniform Distribution (Scenario 3):
- Use Case: Constant, equal density across bounded interval .
- PDF:
3. Sampling Techniques & Sample Size
Question 7: Determining Sample Size for Survey Research
Context: A public health researcher wants to estimate the proportion of city residents who favor a municipal health policy. The researcher establishes:
- Confidence Level: ()
- Margin of Error (): ()
- Preliminary estimate of the population proportion ():
- State the formula for the sample size for proportion estimation.
- Calculate the minimum sample size required.
- Explain what adjustment must be made if the expected population proportion is completely unknown prior to data collection.
Solution
1. Sample Size Formula:
Where:
- is the critical value for the specified confidence level ( for ).
- is the estimated proportion.
- is the acceptable margin of error ().
2. Calculation:
Rounding up to ensure the target margin of error is satisfied:
3. Unknown Proportion Adjustment: If is completely unknown, researchers set . The product represents the point of maximum variance, which guarantees a conservatively large sample size sufficient for any true proportion.
4. Hypothesis Testing & ANOVA
Question 8: Comparing Agricultural Yields (One-Way ANOVA)
Context: An agricultural team tests three distinct fertilizer formulations (: Fertilizer A, B, and C) on plant growth across total experimental plots (5 plots per group). The analysis yields:
- Sum of Squares Between groups ():
- Sum of Squares Within groups ():
- State the Null () and Alternative () hypotheses.
- Determine the degrees of freedom between groups () and within groups ().
- Compute the Mean Square Between () and Mean Square Within ().
- Compute the -statistic and describe the decision rule.
Solution
1. Hypotheses:
- Null Hypothesis (): All group population means are equal ().
- Alternative Hypothesis (): At least one group mean is significantly different from the others.
2. Degrees of Freedom:
- Between groups:
- Within groups:
3. Mean Squares:
4. -Statistic Calculation & Decision Rule:
- Decision Rule: Compare the calculated -statistic () against the critical value from the -distribution table .
- Conclusion: Since , we reject the null hypothesis at the significance level. The fertilizer types result in statistically significant differences in plant growth rates.
Question 9: Checking Die Fairness (Chi-Square Goodness-of-Fit Test)
Context: A six-sided die is rolled times to determine if it is biased. The observed counts across faces 1 through 6 are:
| Face | 1 | 2 | 3 | 4 | 5 | 6 |
|---|---|---|---|---|---|---|
| Observed Frequency () | 8 | 10 | 9 | 12 | 11 | 10 |
- Formulate the Null and Alternative hypotheses.
- Compute the Expected Frequency () for each face under the fairness assumption.
- Calculate the Chi-Square test statistic () and identify the degrees of freedom ().
Solution
1. Hypotheses:
- : The die is fair (the observed distribution matches a discrete uniform distribution where for all ).
- : The die is biased (the distribution differs significantly from a uniform distribution).
2. Expected Frequency: For a fair die rolled 60 times across 6 equal outcomes:
3. Chi-Square Statistic ():
| Face () | |||||
|---|---|---|---|---|---|
| 1 | 8 | 10 | -2 | 4 | 0.4 |
| 2 | 10 | 10 | 0 | 0 | 0.0 |
| 3 | 9 | 10 | -1 | 1 | 0.1 |
| 4 | 12 | 10 | 2 | 4 | 0.4 |
| 5 | 11 | 10 | 1 | 1 | 0.1 |
| 6 | 10 | 10 | 0 | 0 | 0.0 |
| Total | 60 | 60 | 0 | - | 1.0 |
- Degrees of Freedom: .
- Decision: The critical value . Since , we fail to reject . There is insufficient evidence to suggest the die is biased.
5. Model Verification vs. Validation
Question 10: Manufacturing Simulation Quality Assurance
Context: An engineering consultant builds a computer simulation model of an automated automotive assembly line to estimate daily throughput and identify bottlenecks before purchasing physical robotics.
- Distinguish between the primary question answered by Model Verification versus Model Validation.
- Identify two technical methods to verify the simulation model.
- Identify two empirical methods to validate the simulation model.
Solution
1. Conceptual Distinction:
- Verification: "Are we building the model correctly?" Evaluates whether the computer program and mathematical equations conform precisely to the software specifications and design blueprints without coding defects.
- Validation: "Are we building the right model?" Evaluates whether the simulation accurately mimics the actual physics, operations, and behavior of the real-world manufacturing line.
2. Methods for Model Verification:
- Unit and Component Testing: Writing automated tests for individual subroutines (e.g., verifying that a single robotic arm cycle time logic computes accurately given arbitrary load inputs).
- Code Reviews and Algorithm Walkthroughs: Peer-reviewing source code to ensure equations, event queues, and state transitions are bug-free.
3. Methods for Model Validation:
- Comparison with Historical/Real Data: Running the simulation with parameter inputs from an existing plant and comparing simulated cycle times against actual logged cycle times.
- Sensitivity Analysis: Perturbing input variables (e.g., introducing a 10% delay in part arrival) to verify that the simulated throughput drops in a manner consistent with physical manufacturing operations.
6. Machine Learning & Multivariate Analysis
Question 11: Machine Learning Algorithm Selection
Context: A data science team is tasked with four distinct analytical projects. For each scenario:
- Classify whether it requires Supervised or Unsupervised learning.
- State the specific Technique (Regression, Classification, Clustering, or Factor Analysis).
- Recommend an appropriate Algorithm / Method.
- Task A: Predicting the sale price of residential properties based on square footage, location coordinates, and construction year.
- Task B: Automatically grouping customer transactional profiles into distinct behavioral segments without pre-existing labels.
- Task C: Classifying customer support emails as either "Urgent Escalation" or "General Inquiry" based on historical labeled transcripts.
- Task D: Condensing a 40-item employee satisfaction survey into 3 primary underlying constructs representing morale, compensation, and leadership.
Solution
| Task | Learning Paradigm | Technique | Recommended Algorithm / Method |
|---|---|---|---|
| A (House Prices) | Supervised Learning | Regression (continuous target) | Multiple Linear Regression / Ridge Regression |
| B (Customer Segments) | Unsupervised Learning | Cluster Analysis | -Means Clustering / DBSCAN |
| C (Email Triaging) | Supervised Learning | Classification (discrete categorical) | Naive Bayes / Support Vector Machines (SVM) |
| D (Survey Condensation) | Unsupervised / Multivariate | Factor Analysis | Exploratory Factor Analysis (Principal Axis Factoring) |