Skip to main content

Set A - Question 1

Question Details​

A fitness center wants to analyze the performance improvement of two training programs (Program X and Program Y). The performance score (out of 100) represents the improvement in clients' fitness levels after 8 weeks of training.

Data:

  • Program X: 55, 60, 65, 70, 75, 60, 80
  • Program Y: 58, 62, 64, 68, 70, 62, 66

Descriptive Statistics & Consistency Analysis (15 marks)​

  • i. Calculate the mean, median, and mode (if any) for both programs to understand the central tendency of performance improvement. (5 marks)
  • ii. Compute the range, sample variance, and sample standard deviation for both programs to measure the dispersion in performance scores. (7 marks)
  • iii. Based on your analysis, which training program shows higher consistency in client performance improvement? Provide a statistical justification for your conclusion using the measures computed above. (3 marks)

Additional Concept Question (5 marks)​

  • B. Provide difference between correlation and covariance and discuss its significance. (5 marks)

Model Answer​

A. Descriptive Statistics & Consistency Analysis (15 marks)​

Preliminary Organization of Data​

Both datasets comprise n=7n = 7 client evaluations. We order the raw scores in ascending sequence:

  • Program X (Sorted): 55,60,60,65,70,75,8055, 60, 60, 65, 70, 75, 80
  • Program Y (Sorted): 58,62,62,64,66,68,7058, 62, 62, 64, 66, 68, 70

i. Measures of Central Tendency (5 Marks)​

1. Sample Mean (xˉ\bar{x})​

The sample mean is the arithmetic average: xˉ=∑i=1nxin\bar{x} = \frac{\sum_{i=1}^n x_i}{n}

  • Program X: ∑xi=55+60+60+65+70+75+80=465\sum x_i = 55 + 60 + 60 + 65 + 70 + 75 + 80 = 465 xˉX=4657≈66.4286≈66.43\bar{x}_X = \frac{465}{7} \approx 66.4286 \approx \mathbf{66.43}

  • Program Y: ∑yi=58+62+62+64+66+68+70=450\sum y_i = 58 + 62 + 62 + 64 + 66 + 68 + 70 = 450 yˉY=4507≈64.2857≈64.29\bar{y}_Y = \frac{450}{7} \approx 64.2857 \approx \mathbf{64.29}

2. Median​

The median is the value occupying the central position of an ordered sample. With an odd sample size n=7n = 7: Rank=n+12=7+12=4th ordered observation\text{Rank} = \frac{n + 1}{2} = \frac{7 + 1}{2} = 4^{\text{th}} \text{ ordered observation}

  • Program X: The 4th4^{\text{th}} ordered value is 6565.
  • Program Y: The 4th4^{\text{th}} ordered value is 6464.
3. Mode​

The mode is the observation occurring with greatest frequency:

  • Program X: The score 6060 appears twice (frequency=2\text{frequency} = 2), while all remaining observations occur once. Hence, ModeX=60\text{Mode}_X = \mathbf{60} (Unimodal).
  • Program Y: The score 6262 appears twice (frequency=2\text{frequency} = 2), while all remaining observations occur once. Hence, ModeY=62\text{Mode}_Y = \mathbf{62} (Unimodal).
Central Tendency Summary​
MetricProgram XProgram YComparison & Analytical Insight
Mean (xˉ\bar{x})66.4366.4364.2964.29Program X achieves a slightly higher average score (+2.14 points).
Median65.0065.0064.0064.00Medians are nearly identical, differing by only 1 point.
Mode60.0060.0062.0062.00Both distributions are unimodal.

ii. Measures of Dispersion (7 Marks)​

1. Range​

Range=xmax⁡−xmin⁡\text{Range} = x_{\max} - x_{\min}

  • Program X: RangeX=80−55=25\text{Range}_X = 80 - 55 = \mathbf{25}
  • Program Y: RangeY=70−58=12\text{Range}_Y = 70 - 58 = \mathbf{12}
2. Sample Variance (s2s^2) and Sample Standard Deviation (ss)​

Applying Bessel's correction for sample variance with degrees of freedom (n−1)=7−1=6(n - 1) = 7 - 1 = 6: s2=∑i=1n(xi−xˉ)2n−1,s=s2s^2 = \frac{\sum_{i=1}^n (x_i - \bar{x})^2}{n - 1}, \qquad s = \sqrt{s^2}

Deviation Table for Program X (xˉ=66.4286\bar{x} = 66.4286)​
Client (ii)Score (xix_i)Deviation (xi−xˉ)(x_i - \bar{x})Squared Deviation (xi−xˉ)2(x_i - \bar{x})^2
155−11.4286-11.4286130.6122130.6122
260−6.4286-6.428641.326541.3265
360−6.4286-6.428641.326541.3265
465−1.4286-1.42862.04082.0408
570+3.5714+3.571412.755112.7551
675+8.5714+8.571473.469473.4694
780+13.5714+13.5714184.1837184.1837
Total (∑\sum)4654650.00000.0000485.7143485.7143

Using the exact sum of squares identity: ∑xi2=552+602+602+652+702+752+802=31375\sum x_i^2 = 55^2 + 60^2 + 60^2 + 65^2 + 70^2 + 75^2 + 80^2 = 31375 ∑(xi−xˉ)2=31375−46527=31375−2162257=34007≈485.7143\sum (x_i - \bar{x})^2 = 31375 - \frac{465^2}{7} = 31375 - \frac{216225}{7} = \frac{3400}{7} \approx 485.7143

  • Sample Variance (sX2s_X^2): sX2=3400/76=340042=170021≈80.95s_X^2 = \frac{3400 / 7}{6} = \frac{3400}{42} = \frac{1700}{21} \approx \mathbf{80.95}
  • Sample Standard Deviation (sXs_X): sX=170021≈9.00s_X = \sqrt{\frac{1700}{21}} \approx \mathbf{9.00}

Deviation Table for Program Y (yˉ=64.2857\bar{y} = 64.2857)​
Client (ii)Score (yiy_i)Deviation (yi−yˉ)(y_i - \bar{y})Squared Deviation (yi−yˉ)2(y_i - \bar{y})^2
158−6.2857-6.285739.509839.5098
262−2.2857-2.28575.22445.2244
362−2.2857-2.28575.22445.2244
464−0.2857-0.28570.08160.0816
566+1.7143+1.71432.93882.9388
668+3.7143+3.714313.795913.7959
770+5.7143+5.714332.653132.6531
Total (∑\sum)4504500.00000.000099.428699.4286

Using the exact sum of squares identity: ∑yi2=582+622+622+642+662+682+702=29028\sum y_i^2 = 58^2 + 62^2 + 62^2 + 64^2 + 66^2 + 68^2 + 70^2 = 29028 ∑(yi−yˉ)2=29028−45027=29028−2025007=6967≈99.4286\sum (y_i - \bar{y})^2 = 29028 - \frac{450^2}{7} = 29028 - \frac{202500}{7} = \frac{696}{7} \approx 99.4286

  • Sample Variance (sY2s_Y^2): sY2=696/76=69642=1167≈16.57s_Y^2 = \frac{696 / 7}{6} = \frac{696}{42} = \frac{116}{7} \approx \mathbf{16.57}
  • Sample Standard Deviation (sYs_Y): sY=1167≈4.07s_Y = \sqrt{\frac{116}{7}} \approx \mathbf{4.07}
Dispersion Summary​
MetricProgram XProgram YDispersion Comparison
Range25.0025.0012.0012.00Program Y has less than half the total spread of Program X.
Sample Variance (s2s^2)80.9580.9516.5716.57Program X variance is nearly 5×5\times that of Program Y.
Sample Standard Deviation (ss)9.009.004.074.07Program Y scores cluster much closer to the sample mean.

iii. Consistency Analysis & Statistical Justification (3 Marks)​

1. Relative Dispersion: Coefficient of Variation (CVCV)​

The Coefficient of Variation measures relative dispersion independently of scale: CV=(sxˉ)×100%CV = \left( \frac{s}{\bar{x}} \right) \times 100\%

  • Program X: CVX=(8.997466.4286)×100%≈13.54%CV_X = \left( \frac{8.9974}{66.4286} \right) \times 100\% \approx \mathbf{13.54\%}

  • Program Y: CVY=(4.070864.2857)×100%≈6.33%CV_Y = \left( \frac{4.0708}{64.2857} \right) \times 100\% \approx \mathbf{6.33\%}

2. Statistical Conclusion & Justification​
  • Conclusion: Program Y exhibits significantly higher consistency in client performance improvement than Program X.
  • Justification:
    1. Substantially Lower Dispersion: Program Y has a standard deviation of 4.074.07, compared to 9.009.00 for Program X.
    2. Lower Coefficient of Variation: CVY(6.33%)<CVX(13.54%)CV_Y (6.33\%) < CV_X (13.54\%), confirming that Program Y has less than half the relative variability of Program X.
    3. Analytical Takeaway: Although Program X produces a marginally higher mean (+2.14 points), its outcomes are widely volatile (scores span 55 to 80). Program Y offers predictable, reliable progress for all participants (scores tightly concentrated between 58 and 70).

B. Difference Between Correlation and Covariance & Discussion of Significance (5 Marks)​

1. Mathematical Definitions​

  • Sample Covariance: Quantifies the directional joint variability between two variables: Cov(X,Y)=∑i=1n(xi−xˉ)(yi−yˉ)n−1\text{Cov}(X, Y) = \frac{\sum_{i=1}^n (x_i - \bar{x})(y_i - \bar{y})}{n - 1}
  • Pearson's Correlation Coefficient (rr): The standardized covariance normalized by the product of standard deviations: r=Cov(X,Y)sX⋅sY=∑i=1n(xi−xˉ)(yi−yˉ)∑i=1n(xi−xˉ)2∑i=1n(yi−yˉ)2r = \frac{\text{Cov}(X, Y)}{s_X \cdot s_Y} = \frac{\sum_{i=1}^n (x_i - \bar{x})(y_i - \bar{y})}{\sqrt{\sum_{i=1}^n (x_i - \bar{x})^2 \sum_{i=1}^n (y_i - \bar{y})^2}}

2. Comparative Matrix: Covariance vs. Correlation​

FeatureCovarianceCorrelation (rr)
Primary RoleDetermines the direction of joint linear movement (+ or -).Determines both the direction and the exact strength of linear association.
Units of MeasurementExpressed in product units of variables (e.g., kg×cm\text{kg} \times \text{cm}).Dimensionless / Unit-free scalar ratio.
Bounded RangeExtends from −∞-\infty to +∞+\infty.Strictly bounded in [−1,+1][-1, +1].
Scale SensitivityScale-dependent: Changing measurement units changes numerical value.Scale-invariant: Unchanged by linear transformations or unit conversions.
Strength AssessmentCannot determine strength because magnitude depends on arbitrary units.Magnitude directly reflects strength: ∣r∣=1\lvert r \rvert = 1 is perfect linearity, 00 is none.

3. Significance in Statistical Modeling and Data Science​

  1. Scale-Free Feature Comparison: Covariance cannot compare relationships across heterogeneous dimensions (e.g., income in rupees vs age in years). Correlation normalizes variables, enabling direct comparison of feature relevance.
  2. Detection of Multicollinearity: In predictive regression models, high bivariate correlation (∣r∣>0.8\lvert r \rvert > 0.8) identifies redundant predictors, alerting modelers to variance inflation and potential inversion instability.
  3. Portfolio Risk Optimization: In Markowitz portfolio theory, the covariance matrix controls portfolio variance, while the correlation matrix isolates pure relational dependency between financial assets.