Set A - Question 1
Question Details
A fitness center wants to analyze the performance improvement of two training programs (Program X and Program Y). The performance score (out of 100) represents the improvement in clients' fitness levels after 8 weeks of training.
Data:
- Program X: 55, 60, 65, 70, 75, 60, 80
- Program Y: 58, 62, 64, 68, 70, 62, 66
Descriptive Statistics & Consistency Analysis (15 marks)
- i. Calculate the mean, median, and mode (if any) for both programs to understand the central tendency of performance improvement. (5 marks)
- ii. Compute the range, sample variance, and sample standard deviation for both programs to measure the dispersion in performance scores. (7 marks)
- iii. Based on your analysis, which training program shows higher consistency in client performance improvement? Provide a statistical justification for your conclusion using the measures computed above. (3 marks)
Additional Concept Question (5 marks)
- B. Provide difference between correlation and covariance and discuss its significance. (5 marks)
Model Answer
A. Descriptive Statistics & Consistency Analysis (15 marks)
Preliminary Organization of Data
Both datasets comprise client evaluations. We order the raw scores in ascending sequence:
- Program X (Sorted):
- Program Y (Sorted):
i. Measures of Central Tendency (5 Marks)
1. Sample Mean ()
The sample mean is the arithmetic average:
-
Program X:
-
Program Y:
2. Median
The median is the value occupying the central position of an ordered sample. With an odd sample size :
- Program X: The ordered value is .
- Program Y: The ordered value is .
3. Mode
The mode is the observation occurring with greatest frequency:
- Program X: The score appears twice (), while all remaining observations occur once. Hence, (Unimodal).
- Program Y: The score appears twice (), while all remaining observations occur once. Hence, (Unimodal).
Central Tendency Summary
| Metric | Program X | Program Y | Comparison & Analytical Insight |
|---|---|---|---|
| Mean () | Program X achieves a slightly higher average score (+2.14 points). | ||
| Median | Medians are nearly identical, differing by only 1 point. | ||
| Mode | Both distributions are unimodal. |
ii. Measures of Dispersion (7 Marks)
1. Range
- Program X:
- Program Y:
2. Sample Variance () and Sample Standard Deviation ()
Applying Bessel's correction for sample variance with degrees of freedom :
Deviation Table for Program X ()
| Client () | Score () | Deviation | Squared Deviation |
|---|---|---|---|
| 1 | 55 | ||
| 2 | 60 | ||
| 3 | 60 | ||
| 4 | 65 | ||
| 5 | 70 | ||
| 6 | 75 | ||
| 7 | 80 | ||
| Total () |
Using the exact sum of squares identity:
- Sample Variance ():
- Sample Standard Deviation ():
Deviation Table for Program Y ()
| Client () | Score () | Deviation | Squared Deviation |
|---|---|---|---|
| 1 | 58 | ||
| 2 | 62 | ||
| 3 | 62 | ||
| 4 | 64 | ||
| 5 | 66 | ||
| 6 | 68 | ||
| 7 | 70 | ||
| Total () |
Using the exact sum of squares identity:
- Sample Variance ():
- Sample Standard Deviation ():
Dispersion Summary
| Metric | Program X | Program Y | Dispersion Comparison |
|---|---|---|---|
| Range | Program Y has less than half the total spread of Program X. | ||
| Sample Variance () | Program X variance is nearly that of Program Y. | ||
| Sample Standard Deviation () | Program Y scores cluster much closer to the sample mean. |
iii. Consistency Analysis & Statistical Justification (3 Marks)
1. Relative Dispersion: Coefficient of Variation ()
The Coefficient of Variation measures relative dispersion independently of scale:
-
Program X:
-
Program Y:
2. Statistical Conclusion & Justification
- Conclusion: Program Y exhibits significantly higher consistency in client performance improvement than Program X.
- Justification:
- Substantially Lower Dispersion: Program Y has a standard deviation of , compared to for Program X.
- Lower Coefficient of Variation: , confirming that Program Y has less than half the relative variability of Program X.
- Analytical Takeaway: Although Program X produces a marginally higher mean (+2.14 points), its outcomes are widely volatile (scores span 55 to 80). Program Y offers predictable, reliable progress for all participants (scores tightly concentrated between 58 and 70).
B. Difference Between Correlation and Covariance & Discussion of Significance (5 Marks)
1. Mathematical Definitions
- Sample Covariance: Quantifies the directional joint variability between two variables:
- Pearson's Correlation Coefficient (): The standardized covariance normalized by the product of standard deviations:
2. Comparative Matrix: Covariance vs. Correlation
| Feature | Covariance | Correlation () |
|---|---|---|
| Primary Role | Determines the direction of joint linear movement (+ or -). | Determines both the direction and the exact strength of linear association. |
| Units of Measurement | Expressed in product units of variables (e.g., ). | Dimensionless / Unit-free scalar ratio. |
| Bounded Range | Extends from to . | Strictly bounded in . |
| Scale Sensitivity | Scale-dependent: Changing measurement units changes numerical value. | Scale-invariant: Unchanged by linear transformations or unit conversions. |
| Strength Assessment | Cannot determine strength because magnitude depends on arbitrary units. | Magnitude directly reflects strength: is perfect linearity, is none. |
3. Significance in Statistical Modeling and Data Science
- Scale-Free Feature Comparison: Covariance cannot compare relationships across heterogeneous dimensions (e.g., income in rupees vs age in years). Correlation normalizes variables, enabling direct comparison of feature relevance.
- Detection of Multicollinearity: In predictive regression models, high bivariate correlation () identifies redundant predictors, alerting modelers to variance inflation and potential inversion instability.
- Portfolio Risk Optimization: In Markowitz portfolio theory, the covariance matrix controls portfolio variance, while the correlation matrix isolates pure relational dependency between financial assets.