Standard Experimental Designs & ANOVA
Topic - Experimental designs structure how treatments are allocated to experimental units to isolate factor effects from background error. The four classical designs - Completely Randomized Design (CRD), Randomized Block Design (RBD), Latin Square Design (LSD), and 2 k 2^k 2 k Full Factorial Design - progressively control for zero, one, two, or multiple interactive sources of variation through linear hypothesis testing and Analysis of Variance (ANOVA).
1. Comparative Architecture of Standard Designs
Selecting an experimental layout depends strictly on the physical homogeneity of the experimental material and the presence of external nuisance factors:
Design Blocking Structure ANOVA Type Sources of Variation Key Precondition CRD No blocks (homogeneous units) One-Way ANOVA Treatments, Error All experimental units are uniform. RBD 1 blocking factor Two-Way ANOVA Treatments, Blocks, Error One-directional gradient across units. LSD 2 blocking factors (Rows & Columns) Three-Way ANOVA Treatments, Rows, Columns, Error Exactly p × p p \times p p × p matrix; no interactions between factors. 2 k 2^k 2 k Full FactorialAll combinations of k k k factors at 2 levels Multi-Factor ANOVA Main Effects, Interactions, Error Evaluates joint variable interactions.
2. Completely Randomized Design (CRD)
Theoretical Model & Assumptions
CRD is the basic single-factor design where treatments are assigned completely at random to all units. It is appropriate when units are completely homogeneous (e.g., standard laboratory test tubes, incubator shelves, well-mixed chemicals).
Statistical Model: y i j = μ + τ i + ϵ i j \text{Statistical Model: } y_{ij} = \mu + \tau_i + \epsilon_{ij} Statistical Model: y ij = μ + τ i + ϵ ij
Where μ \mu μ is the overall mean, τ i \tau_i τ i is the effect of treatment i i i , and ϵ i j ∼ iid N ( 0 , σ 2 ) \epsilon_{ij} \overset{\text{iid}}{\sim} \mathcal{N}(0, \sigma^2) ϵ ij ∼ iid N ( 0 , σ 2 ) is the random experimental error.
Given t t t treatments each replicated r r r times (N = t × r N = t \times r N = t × r observations):
Grand Total (G G G ): G = ∑ i = 1 t ∑ j = 1 r y i j G = \sum_{i=1}^t \sum_{j=1}^r y_{ij} G = ∑ i = 1 t ∑ j = 1 r y ij
Correction Factor (C F CF C F ):
C F = G 2 N CF = \frac{G^2}{N} C F = N G 2
Total Sum of Squares (SS Total \text{SS}_{\text{Total}} SS Total ):
SS Total = ∑ i = 1 t ∑ j = 1 r y i j 2 − C F \text{SS}_{\text{Total}} = \sum_{i=1}^t \sum_{j=1}^r y_{ij}^2 - CF SS Total = ∑ i = 1 t ∑ j = 1 r y ij 2 − C F
Treatment Sum of Squares (SS Treat \text{SS}_{\text{Treat}} SS Treat ):
SS Treat = ∑ i = 1 t T i 2 r − C F \text{SS}_{\text{Treat}} = \sum_{i=1}^t \frac{T_i^2}{r} - CF SS Treat = ∑ i = 1 t r T i 2 − C F
Error Sum of Squares (SS Error \text{SS}_{\text{Error}} SS Error ):
SS Error = SS Total − SS Treat \text{SS}_{\text{Error}} = \text{SS}_{\text{Total}} - \text{SS}_{\text{Treat}} SS Error = SS Total − SS Treat
Standard CRD ANOVA Table
Source of Variation Degrees of Freedom (df \text{df} df ) Sum of Squares (SS \text{SS} SS ) Mean Square (MS \text{MS} MS ) Calculated F F F (F cal F_{\text{cal}} F cal ) Treatments t − 1 t - 1 t − 1 SS Treat \text{SS}_{\text{Treat}} SS Treat MS Treat = SS Treat t − 1 \text{MS}_{\text{Treat}} = \frac{\text{SS}_{\text{Treat}}}{t - 1} MS Treat = t − 1 SS Treat MS Treat MS Error \frac{\text{MS}_{\text{Treat}}}{\text{MS}_{\text{Error}}} MS Error MS Treat Error N − t = t ( r − 1 ) N - t = t(r - 1) N − t = t ( r − 1 ) SS Error \text{SS}_{\text{Error}} SS Error MS Error = SS Error N − t \text{MS}_{\text{Error}} = \frac{\text{SS}_{\text{Error}}}{N - t} MS Error = N − t SS Error - Total N − 1 N - 1 N − 1 SS Total \text{SS}_{\text{Total}} SS Total - -
Decision Rule: Compare F cal F_{\text{cal}} F cal with tabulated F α , ( t − 1 ) , ( N − t ) F_{\alpha, (t-1), (N-t)} F α , ( t − 1 ) , ( N − t ) . If F cal > F tab F_{\text{cal}} \gt F_{\text{tab}} F cal > F tab , reject H 0 H_0 H 0 (at least one treatment mean differs significantly).
Solved Problem: Fertilizer Yield Evaluation
Problem Statement: Three fertilizers (A , B , C A, B, C A , B , C ) are tested on a homogeneous field. Each treatment is replicated 4 times (t = 3 , r = 4 , N = 12 t=3, r=4, N=12 t = 3 , r = 4 , N = 12 ). Observed crop yields (kg) are:
Fertilizer A: 18 , 20 , 19 , 22 18, 20, 19, 22 18 , 20 , 19 , 22
Fertilizer B: 24 , 22 , 21 , 23 24, 22, 21, 23 24 , 22 , 21 , 23
Fertilizer C: 26 , 25 , 27 , 28 26, 25, 27, 28 26 , 25 , 27 , 28
Test at α = 0.05 \alpha = 0.05 α = 0.05 whether there are significant treatment differences.
Step 1: Compute Totals and Correction Factor
T A = 18 + 20 + 19 + 22 = 79 T_A = 18 + 20 + 19 + 22 = 79 T A = 18 + 20 + 19 + 22 = 79
T B = 24 + 22 + 21 + 23 = 90 T_B = 24 + 22 + 21 + 23 = 90 T B = 24 + 22 + 21 + 23 = 90
T C = 26 + 25 + 27 + 28 = 106 T_C = 26 + 25 + 27 + 28 = 106 T C = 26 + 25 + 27 + 28 = 106
G = 79 + 90 + 106 = 275 G = 79 + 90 + 106 = 275 G = 79 + 90 + 106 = 275
C F = G 2 N = 275 2 12 = 75625 12 ≈ 6302.0833 CF = \frac{G^2}{N} = \frac{275^2}{12} = \frac{75625}{12} \approx 6302.0833 C F = N G 2 = 12 27 5 2 = 12 75625 ≈ 6302.0833
Step 2: Sum of Squares
∑ y i j 2 = 18 2 + 20 2 + 19 2 + 22 2 + 24 2 + 22 2 + 21 2 + 23 2 + 26 2 + 25 2 + 27 2 + 28 2 = 6413 \sum y_{ij}^2 = 18^2 + 20^2 + 19^2 + 22^2 + 24^2 + 22^2 + 21^2 + 23^2 + 26^2 + 25^2 + 27^2 + 28^2 = 6413 ∑ y ij 2 = 1 8 2 + 2 0 2 + 1 9 2 + 2 2 2 + 2 4 2 + 2 2 2 + 2 1 2 + 2 3 2 + 2 6 2 + 2 5 2 + 2 7 2 + 2 8 2 = 6413
SS Total = 6413 − 6302.0833 = 110.9167 \text{SS}_{\text{Total}} = 6413 - 6302.0833 = 110.9167 SS Total = 6413 − 6302.0833 = 110.9167
SS Treat = 79 2 + 90 2 + 106 2 4 − C F = 6241 + 8100 + 11236 4 − 6302.0833 = 25577 4 − 6302.0833 = 6394.25 − 6302.0833 = 92.1667 \text{SS}_{\text{Treat}} = \frac{79^2 + 90^2 + 106^2}{4} - CF = \frac{6241 + 8100 + 11236}{4} - 6302.0833 = \frac{25577}{4} - 6302.0833 = 6394.25 - 6302.0833 = 92.1667 SS Treat = 4 7 9 2 + 9 0 2 + 10 6 2 − C F = 4 6241 + 8100 + 11236 − 6302.0833 = 4 25577 − 6302.0833 = 6394.25 − 6302.0833 = 92.1667
SS Error = SS Total − SS Treat = 110.9167 − 92.1667 = 18.7500 \text{SS}_{\text{Error}} = \text{SS}_{\text{Total}} - \text{SS}_{\text{Treat}} = 110.9167 - 92.1667 = 18.7500 SS Error = SS Total − SS Treat = 110.9167 − 92.1667 = 18.7500
Step 3: Mean Squares & F-Ratio
df Treat = 3 − 1 = 2 \text{df}_{\text{Treat}} = 3 - 1 = 2 df Treat = 3 − 1 = 2
df Error = 12 − 3 = 9 \text{df}_{\text{Error}} = 12 - 3 = 9 df Error = 12 − 3 = 9
MS Treat = 92.1667 2 = 46.0833 \text{MS}_{\text{Treat}} = \frac{92.1667}{2} = 46.0833 MS Treat = 2 92.1667 = 46.0833
MS Error = 18.7500 9 = 2.0833 \text{MS}_{\text{Error}} = \frac{18.7500}{9} = 2.0833 MS Error = 9 18.7500 = 2.0833
F cal = 46.0833 2.0833 ≈ 22.12 F_{\text{cal}} = \frac{46.0833}{2.0833} \approx 22.12 F cal = 2.0833 46.0833 ≈ 22.12
Step 4: Decision & Conclusion
Critical F 0.05 , 2 , 9 ≈ 4.26 F_{0.05, 2, 9} \approx 4.26 F 0.05 , 2 , 9 ≈ 4.26 .
Since F cal = 22.12 > 4.26 F_{\text{cal}} = 22.12 \gt 4.26 F cal = 22.12 > 4.26 , we reject the null hypothesis H 0 H_0 H 0 at the 5 % 5\% 5% level of significance.
Conclusion: The three fertilizers differ significantly in their mean crop yield.
3. Randomized Block Design (RBD)
Theoretical Model & Assumptions
RBD is an improvement over CRD when experimental units exhibit a known one-directional gradient (e.g., slope fertility, temperature gradient across an oven, different laboratory technicians). Units are divided into homogeneous groups called blocks . Each treatment appears exactly once in every block in a randomized sequence.
Statistical Model: y i j = μ + τ i + β j + ϵ i j \text{Statistical Model: } y_{ij} = \mu + \tau_i + \beta_j + \epsilon_{ij} Statistical Model: y ij = μ + τ i + β j + ϵ ij
Where τ i \tau_i τ i is the i i i -th treatment effect, β j \beta_j β j is the j j j -th block effect, and ϵ i j ∼ iid N ( 0 , σ 2 ) \epsilon_{ij} \overset{\text{iid}}{\sim} \mathcal{N}(0, \sigma^2) ϵ ij ∼ iid N ( 0 , σ 2 ) .
Given t t t treatments and r r r blocks (N = t × r N = t \times r N = t × r ):
Block Sum of Squares (SS Block \text{SS}_{\text{Block}} SS Block ):
SS Block = ∑ j = 1 r B j 2 t − C F \text{SS}_{\text{Block}} = \sum_{j=1}^r \frac{B_j^2}{t} - CF SS Block = ∑ j = 1 r t B j 2 − C F
Error Sum of Squares (SS Error \text{SS}_{\text{Error}} SS Error ):
SS Error = SS Total − SS Treat − SS Block \text{SS}_{\text{Error}} = \text{SS}_{\text{Total}} - \text{SS}_{\text{Treat}} - \text{SS}_{\text{Block}} SS Error = SS Total − SS Treat − SS Block
Standard RBD ANOVA Table
Source Degrees of Freedom (df \text{df} df ) Sum of Squares (SS \text{SS} SS ) Mean Square (MS \text{MS} MS ) Calculated F F F Treatments t − 1 t - 1 t − 1 SS Treat \text{SS}_{\text{Treat}} SS Treat MS Treat = SS Treat t − 1 \text{MS}_{\text{Treat}} = \frac{\text{SS}_{\text{Treat}}}{t - 1} MS Treat = t − 1 SS Treat MS Treat MS Error \frac{\text{MS}_{\text{Treat}}}{\text{MS}_{\text{Error}}} MS Error MS Treat Blocks r − 1 r - 1 r − 1 SS Block \text{SS}_{\text{Block}} SS Block MS Block = SS Block r − 1 \text{MS}_{\text{Block}} = \frac{\text{SS}_{\text{Block}}}{r - 1} MS Block = r − 1 SS Block MS Block MS Error \frac{\text{MS}_{\text{Block}}}{\text{MS}_{\text{Error}}} MS Error MS Block (Optional) Error ( t − 1 ) ( r − 1 ) (t - 1)(r - 1) ( t − 1 ) ( r − 1 ) SS Error \text{SS}_{\text{Error}} SS Error MS Error = SS Error ( t − 1 ) ( r − 1 ) \text{MS}_{\text{Error}} = \frac{\text{SS}_{\text{Error}}}{(t - 1)(r - 1)} MS Error = ( t − 1 ) ( r − 1 ) SS Error - Total N − 1 N - 1 N − 1 SS Total \text{SS}_{\text{Total}} SS Total - -
Solved Problem: Cholesterol Analysis Across 4 Diets in 4 Labs
Problem Statement: Measurement of cholesterol content (%) is performed in 4 different laboratories (A , B , C , D A, B, C, D A , B , C , D ) across 4 diet foods (D 1 , D 2 , D 3 , D 4 D_1, D_2, D_3, D_4 D 1 , D 2 , D 3 , D 4 ). Analyze the data using RBD at α = 0.05 \alpha = 0.05 α = 0.05 (Given critical F 0.05 ( 3 , 9 ) = 3.86 F_{0.05}(3, 9) = 3.86 F 0.05 ( 3 , 9 ) = 3.86 ).
Laboratory (Block) D 1 D_1 D 1 D 2 D_2 D 2 D 3 D_3 D 3 D 4 D_4 D 4 Block Total (B j B_j B j ) A 2 5 4 3 14 B 3 7 4 2 16 C 5 5 6 4 20 D 5 5 8 5 23 Treatment Total (T i T_i T i ) 15 22 22 14 G = 73 G = 73 G = 73
Step 1: Totals and Correction Factor
t = 4 , r = 4 , N = 16 t = 4, r = 4, N = 16 t = 4 , r = 4 , N = 16
G = 73 G = 73 G = 73
C F = 73 2 16 = 5329 16 = 333.0625 CF = \frac{73^2}{16} = \frac{5329}{16} = 333.0625 C F = 16 7 3 2 = 16 5329 = 333.0625
Step 2: Sum of Squares
∑ y i j 2 = 2 2 + 5 2 + 4 2 + 3 2 + 3 2 + 7 2 + 4 2 + 2 2 + 5 2 + 5 2 + 6 2 + 4 2 + 5 2 + 5 2 + 8 2 + 5 2 = 373 \sum y_{ij}^2 = 2^2 + 5^2 + 4^2 + 3^2 + 3^2 + 7^2 + 4^2 + 2^2 + 5^2 + 5^2 + 6^2 + 4^2 + 5^2 + 5^2 + 8^2 + 5^2 = 373 ∑ y ij 2 = 2 2 + 5 2 + 4 2 + 3 2 + 3 2 + 7 2 + 4 2 + 2 2 + 5 2 + 5 2 + 6 2 + 4 2 + 5 2 + 5 2 + 8 2 + 5 2 = 373
SS Total = 373 − 333.0625 = 39.9375 \text{SS}_{\text{Total}} = 373 - 333.0625 = 39.9375 SS Total = 373 − 333.0625 = 39.9375
SS Treat = 15 2 + 22 2 + 22 2 + 14 2 4 − C F = 225 + 484 + 484 + 196 4 − 333.0625 = 1389 4 − 333.0625 = 347.25 − 333.0625 = 14.1875 \text{SS}_{\text{Treat}} = \frac{15^2 + 22^2 + 22^2 + 14^2}{4} - CF = \frac{225 + 484 + 484 + 196}{4} - 333.0625 = \frac{1389}{4} - 333.0625 = 347.25 - 333.0625 = 14.1875 SS Treat = 4 1 5 2 + 2 2 2 + 2 2 2 + 1 4 2 − C F = 4 225 + 484 + 484 + 196 − 333.0625 = 4 1389 − 333.0625 = 347.25 − 333.0625 = 14.1875
SS Block = 14 2 + 16 2 + 20 2 + 23 2 4 − C F = 196 + 256 + 400 + 529 4 − 333.0625 = 1381 4 − 333.0625 = 345.25 − 333.0625 = 12.1875 \text{SS}_{\text{Block}} = \frac{14^2 + 16^2 + 20^2 + 23^2}{4} - CF = \frac{196 + 256 + 400 + 529}{4} - 333.0625 = \frac{1381}{4} - 333.0625 = 345.25 - 333.0625 = 12.1875 SS Block = 4 1 4 2 + 1 6 2 + 2 0 2 + 2 3 2 − C F = 4 196 + 256 + 400 + 529 − 333.0625 = 4 1381 − 333.0625 = 345.25 − 333.0625 = 12.1875
SS Error = 39.9375 − 14.1875 − 12.1875 = 13.5625 \text{SS}_{\text{Error}} = 39.9375 - 14.1875 - 12.1875 = 13.5625 SS Error = 39.9375 − 14.1875 − 12.1875 = 13.5625
Step 3: Mean Squares & Decision
df Treat = 4 − 1 = 3 \text{df}_{\text{Treat}} = 4 - 1 = 3 df Treat = 4 − 1 = 3
df Block = 4 − 1 = 3 \text{df}_{\text{Block}} = 4 - 1 = 3 df Block = 4 − 1 = 3
df Error = ( 4 − 1 ) ( 4 − 1 ) = 9 \text{df}_{\text{Error}} = (4 - 1)(4 - 1) = 9 df Error = ( 4 − 1 ) ( 4 − 1 ) = 9
MS Treat = 14.1875 3 ≈ 4.7292 \text{MS}_{\text{Treat}} = \frac{14.1875}{3} \approx 4.7292 MS Treat = 3 14.1875 ≈ 4.7292
MS Error = 13.5625 9 ≈ 1.5069 \text{MS}_{\text{Error}} = \frac{13.5625}{9} \approx 1.5069 MS Error = 9 13.5625 ≈ 1.5069
F cal = 4.7292 1.5069 ≈ 3.14 F_{\text{cal}} = \frac{4.7292}{1.5069} \approx 3.14 F cal = 1.5069 4.7292 ≈ 3.14
Comparison: F cal = 3.14 < F tab = 3.86 F_{\text{cal}} = 3.14 \lt F_{\text{tab}} = 3.86 F cal = 3.14 < F tab = 3.86 .
Conclusion: Fail to reject H 0 H_0 H 0 . At the 5 % 5\% 5% significance level, there is no statistically significant difference in cholesterol content among the four diet foods after accounting for laboratory blocking variations.
4. Latin Square Design (LSD)
Theoretical Model & Assumptions
When two independent sources of nuisance variability exist simultaneously (e.g., Machine differences along rows and Operator differences along columns), the Latin Square Design (LSD) isolates both.
A p × p p \times p p × p square arrangement is used where each of p p p treatments appears exactly once in each row and once in each column .
The number of treatments must equal the number of rows and columns (t = rows = columns = p t = \text{rows} = \text{columns} = p t = rows = columns = p ).
Strict Assumption: There are no interactions between rows, columns, and treatments.
Statistical Model: y i j k = μ + τ i + ρ j + γ k + ϵ i j k \text{Statistical Model: } y_{ijk} = \mu + \tau_i + \rho_j + \gamma_k + \epsilon_{ijk} Statistical Model: y ij k = μ + τ i + ρ j + γ k + ϵ ij k
Where τ i \tau_i τ i is treatment effect, ρ j \rho_j ρ j is row effect, and γ k \gamma_k γ k is column effect.
C F = G 2 p 2 CF = \frac{G^2}{p^2} C F = p 2 G 2
SS Total = ∑ y 2 − C F \text{SS}_{\text{Total}} = \sum y^2 - CF SS Total = ∑ y 2 − C F
SS Treat = ∑ T i 2 p − C F , SS Row = ∑ R j 2 p − C F , SS Col = ∑ C k 2 p − C F \text{SS}_{\text{Treat}} = \frac{\sum T_i^2}{p} - CF, \quad \text{SS}_{\text{Row}} = \frac{\sum R_j^2}{p} - CF, \quad \text{SS}_{\text{Col}} = \frac{\sum C_k^2}{p} - CF SS Treat = p ∑ T i 2 − C F , SS Row = p ∑ R j 2 − C F , SS Col = p ∑ C k 2 − C F
SS Error = SS Total − SS Treat − SS Row − SS Col \text{SS}_{\text{Error}} = \text{SS}_{\text{Total}} - \text{SS}_{\text{Treat}} - \text{SS}_{\text{Row}} - \text{SS}_{\text{Col}} SS Error = SS Total − SS Treat − SS Row − SS Col
df Treat = p − 1 , df Row = p − 1 , df Col = p − 1 , df Error = ( p − 1 ) ( p − 2 ) , df Total = p 2 − 1 \text{df}_{\text{Treat}} = p - 1, \quad \text{df}_{\text{Row}} = p - 1, \quad \text{df}_{\text{Col}} = p - 1, \quad \text{df}_{\text{Error}} = (p - 1)(p - 2), \quad \text{df}_{\text{Total}} = p^2 - 1 df Treat = p − 1 , df Row = p − 1 , df Col = p − 1 , df Error = ( p − 1 ) ( p − 2 ) , df Total = p 2 − 1
Solved Problem: 4x4 Latin Square Evaluation
Problem Statement: Four treatments (A , B , C , D A, B, C, D A , B , C , D ) are allocated in a 4 × 4 4 \times 4 4 × 4 Latin square to control for row and column effects (p = 4 , N = 16 p=4, N=16 p = 4 , N = 16 ).
Row 1: A ( 18 ) , B ( 22 ) , C ( 25 ) , D ( 28 ) ⟹ R 1 = 93 A(18), B(22), C(25), D(28) \implies R_1 = 93 A ( 18 ) , B ( 22 ) , C ( 25 ) , D ( 28 ) ⟹ R 1 = 93
Row 2: B ( 21 ) , C ( 24 ) , D ( 26 ) , A ( 20 ) ⟹ R 2 = 91 B(21), C(24), D(26), A(20) \implies R_2 = 91 B ( 21 ) , C ( 24 ) , D ( 26 ) , A ( 20 ) ⟹ R 2 = 91
Row 3: C ( 19 ) , D ( 23 ) , A ( 22 ) , B ( 25 ) ⟹ R 3 = 89 C(19), D(23), A(22), B(25) \implies R_3 = 89 C ( 19 ) , D ( 23 ) , A ( 22 ) , B ( 25 ) ⟹ R 3 = 89
Row 4: D ( 20 ) , A ( 21 ) , B ( 23 ) , C ( 27 ) ⟹ R 4 = 91 D(20), A(21), B(23), C(27) \implies R_4 = 91 D ( 20 ) , A ( 21 ) , B ( 23 ) , C ( 27 ) ⟹ R 4 = 91
Column totals: C 1 = 78 , C 2 = 90 , C 3 = 96 , C 4 = 100 C_1 = 78, C_2 = 90, C_3 = 96, C_4 = 100 C 1 = 78 , C 2 = 90 , C 3 = 96 , C 4 = 100 . Grand Total G = 364 G = 364 G = 364 .
Treatment totals: T A = 81 , T B = 91 , T C = 95 , T D = 97 T_A = 81, T_B = 91, T_C = 95, T_D = 97 T A = 81 , T B = 91 , T C = 95 , T D = 97 . Total sum of squares ∑ y 2 = 8408 \sum y^2 = 8408 ∑ y 2 = 8408 .
Test at α = 0.05 \alpha = 0.05 α = 0.05 whether treatment differences are significant (F 0.05 ( 3 , 6 ) = 4.76 F_{0.05}(3, 6) = 4.76 F 0.05 ( 3 , 6 ) = 4.76 ).
Step 1: Sum of Squares
C F = 364 2 16 = 132496 16 = 8281.0 CF = \frac{364^2}{16} = \frac{132496}{16} = 8281.0 C F = 16 36 4 2 = 16 132496 = 8281.0
SS Total = 8408.0 − 8281.0 = 127.0 \text{SS}_{\text{Total}} = 8408.0 - 8281.0 = 127.0 SS Total = 8408.0 − 8281.0 = 127.0
SS Treat = 81 2 + 91 2 + 95 2 + 97 2 4 − 8281.0 = 33276 4 − 8281.0 = 8319.0 − 8281.0 = 38.0 \text{SS}_{\text{Treat}} = \frac{81^2 + 91^2 + 95^2 + 97^2}{4} - 8281.0 = \frac{33276}{4} - 8281.0 = 8319.0 - 8281.0 = 38.0 SS Treat = 4 8 1 2 + 9 1 2 + 9 5 2 + 9 7 2 − 8281.0 = 4 33276 − 8281.0 = 8319.0 − 8281.0 = 38.0
SS Row = 93 2 + 91 2 + 89 2 + 91 2 4 − 8281.0 = 33132 4 − 8281.0 = 8283.0 − 8281.0 = 2.0 \text{SS}_{\text{Row}} = \frac{93^2 + 91^2 + 89^2 + 91^2}{4} - 8281.0 = \frac{33132}{4} - 8281.0 = 8283.0 - 8281.0 = 2.0 SS Row = 4 9 3 2 + 9 1 2 + 8 9 2 + 9 1 2 − 8281.0 = 4 33132 − 8281.0 = 8283.0 − 8281.0 = 2.0
SS Col = 78 2 + 90 2 + 96 2 + 100 2 4 − 8281.0 = 33400 4 − 8281.0 = 8350.0 − 8281.0 = 69.0 \text{SS}_{\text{Col}} = \frac{78^2 + 90^2 + 96^2 + 100^2}{4} - 8281.0 = \frac{33400}{4} - 8281.0 = 8350.0 - 8281.0 = 69.0 SS Col = 4 7 8 2 + 9 0 2 + 9 6 2 + 10 0 2 − 8281.0 = 4 33400 − 8281.0 = 8350.0 − 8281.0 = 69.0
SS Error = 127.0 − 38.0 − 2.0 − 69.0 = 18.0 \text{SS}_{\text{Error}} = 127.0 - 38.0 - 2.0 - 69.0 = 18.0 SS Error = 127.0 − 38.0 − 2.0 − 69.0 = 18.0
Step 2: LSD ANOVA Table
Source df \text{df} df SS \text{SS} SS MS \text{MS} MS F cal F_{\text{cal}} F cal Critical F 0.05 F_{0.05} F 0.05 Rows 3 2.0 0.6667 - - Columns 3 69.0 23.0000 - - Treatments 3 38.0 12.6667 12.6667 3.0 ≈ 4.222 \frac{12.6667}{3.0} \approx 4.222 3.0 12.6667 ≈ 4.222 4.76 Error ( 4 − 1 ) ( 4 − 2 ) = 6 (4-1)(4-2)=6 ( 4 − 1 ) ( 4 − 2 ) = 6 18.0 3.0000 - - Total 15 127.0 - - -
Decision: F cal = 4.222 < 4.76 F_{\text{cal}} = 4.222 \lt 4.76 F cal = 4.222 < 4.76 . Fail to reject H 0 H_0 H 0 .
Conclusion: After adjusting for row and column variations, treatments do not produce statistically significant differences at the 5 % 5\% 5% level.
5. Full Factorial Design (2 2 , 2 3 , 2 k 2^2, 2^3, 2^k 2 2 , 2 3 , 2 k )
In a 2 k 2^k 2 k factorial design, k k k factors are evaluated at 2 coded levels: Low (− 1 -1 − 1 ) and High (+ 1 +1 + 1 ).
Total combinations: 2 k 2^k 2 k . With r r r replicates, total runs N = r ⋅ 2 k N = r \cdot 2^k N = r ⋅ 2 k .
Contrast (C C C ): Sum of signed responses for an effect:
C = ∑ i = 1 2 k Sign i ⋅ ( Sum of replicates for run i ) C = \sum_{i=1}^{2^k} \text{Sign}_i \cdot (\text{Sum of replicates for run } i) C = ∑ i = 1 2 k Sign i ⋅ ( Sum of replicates for run i )
Estimated Effect:
Effect = C 2 k − 1 ⋅ r \text{Effect} = \frac{C}{2^{k-1} \cdot r} Effect = 2 k − 1 ⋅ r C
Sum of Squares for any Effect:
SS = C 2 r ⋅ 2 k \text{SS} = \frac{C^2}{r \cdot 2^k} SS = r ⋅ 2 k C 2
Solved Problem: 2 2 2^2 2 2 Chemical Reaction Factorial Analysis
Problem Statement: A chemical engineer studies the effect of Temperature (A A A ) and Pressure (B B B ) on product yield (%). Both factors are tested at two coded levels: Low (− 1 -1 − 1 ) and High (+ 1 +1 + 1 ) with r = 2 r=2 r = 2 replicates (N = 2 2 × 2 = 8 N = 2^2 \times 2 = 8 N = 2 2 × 2 = 8 ).
Run A B Interaction (A B = A × B AB = A \times B A B = A × B ) Run 1 (y i 1 y_{i1} y i 1 ) Run 2 (y i 2 y_{i2} y i 2 ) Treatment Total (y i + y_{i+} y i + ) 1 − 1 -1 − 1 − 1 -1 − 1 + 1 +1 + 1 42 43 85 2 + 1 +1 + 1 − 1 -1 − 1 − 1 -1 − 1 50 49 99 3 − 1 -1 − 1 + 1 +1 + 1 − 1 -1 − 1 46 45 91 4 + 1 +1 + 1 + 1 +1 + 1 + 1 +1 + 1 62 61 123 Grand Total G = 398 G = 398 G = 398
Step 1: Contrasts & Main Effects
Contrast for A: C A = − 85 + 99 − 91 + 123 = 46 C_A = -85 + 99 - 91 + 123 = 46 C A = − 85 + 99 − 91 + 123 = 46
A ^ = C A 2 2 − 1 ⋅ 2 = 46 4 = 11.5 \hat{A} = \frac{C_A}{2^{2-1} \cdot 2} = \frac{46}{4} = 11.5 A ^ = 2 2 − 1 ⋅ 2 C A = 4 46 = 11.5
Contrast for B: C B = − 85 − 99 + 91 + 123 = 30 C_B = -85 - 99 + 91 + 123 = 30 C B = − 85 − 99 + 91 + 123 = 30
B ^ = C B 4 = 30 4 = 7.5 \hat{B} = \frac{C_B}{4} = \frac{30}{4} = 7.5 B ^ = 4 C B = 4 30 = 7.5
Contrast for AB: C A B = + 85 − 99 − 91 + 123 = 18 C_{AB} = +85 - 99 - 91 + 123 = 18 C A B = + 85 − 99 − 91 + 123 = 18
A B ^ = C A B 4 = 18 4 = 4.5 \widehat{AB} = \frac{C_{AB}}{4} = \frac{18}{4} = 4.5 A B = 4 C A B = 4 18 = 4.5
Step 2: Sum of Squares (N = 8 N = 8 N = 8 )
SS A = C A 2 8 = 46 2 8 = 2116 8 = 264.5 \text{SS}_A = \frac{C_A^2}{8} = \frac{46^2}{8} = \frac{2116}{8} = 264.5 SS A = 8 C A 2 = 8 4 6 2 = 8 2116 = 264.5
SS B = C B 2 8 = 30 2 8 = 900 8 = 112.5 \text{SS}_B = \frac{C_B^2}{8} = \frac{30^2}{8} = \frac{900}{8} = 112.5 SS B = 8 C B 2 = 8 3 0 2 = 8 900 = 112.5
SS A B = C A B 2 8 = 18 2 8 = 324 8 = 40.5 \text{SS}_{AB} = \frac{C_{AB}^2}{8} = \frac{18^2}{8} = \frac{324}{8} = 40.5 SS A B = 8 C A B 2 = 8 1 8 2 = 8 324 = 40.5
SS Treat = 264.5 + 112.5 + 40.5 = 417.5 \text{SS}_{\text{Treat}} = 264.5 + 112.5 + 40.5 = 417.5 SS Treat = 264.5 + 112.5 + 40.5 = 417.5
∑ y 2 = 42 2 + 43 2 + 50 2 + 49 2 + 46 2 + 45 2 + 62 2 + 61 2 = 20220 \sum y^2 = 42^2 + 43^2 + 50^2 + 49^2 + 46^2 + 45^2 + 62^2 + 61^2 = 20220 ∑ y 2 = 4 2 2 + 4 3 2 + 5 0 2 + 4 9 2 + 4 6 2 + 4 5 2 + 6 2 2 + 6 1 2 = 20220
C F = 398 2 8 = 158404 8 = 19800.5 CF = \frac{398^2}{8} = \frac{158404}{8} = 19800.5 C F = 8 39 8 2 = 8 158404 = 19800.5
SS Total = 20220 − 19800.5 = 419.5 \text{SS}_{\text{Total}} = 20220 - 19800.5 = 419.5 SS Total = 20220 − 19800.5 = 419.5
SS Error = SS Total − SS Treat = 419.5 − 417.5 = 2.0 \text{SS}_{\text{Error}} = \text{SS}_{\text{Total}} - \text{SS}_{\text{Treat}} = 419.5 - 417.5 = 2.0 SS Error = SS Total − SS Treat = 419.5 − 417.5 = 2.0
Step 3: 2 2 2^2 2 2 Factorial ANOVA Table
Source df \text{df} df SS \text{SS} SS MS \text{MS} MS Calculated F F F Critical F 0.05 ( 1 , 4 ) F_{0.05}(1, 4) F 0.05 ( 1 , 4 ) Conclusion A (Temperature) 1 264.5 264.5 264.5 0.5 = 529.0 \frac{264.5}{0.5} = 529.0 0.5 264.5 = 529.0 7.708 Reject H 0 H_0 H 0 (p < 0.001 p \lt 0.001 p < 0.001 ) B (Pressure) 1 112.5 112.5 112.5 0.5 = 225.0 \frac{112.5}{0.5} = 225.0 0.5 112.5 = 225.0 7.708 Reject H 0 H_0 H 0 (p < 0.001 p \lt 0.001 p < 0.001 ) AB (Interaction) 1 40.5 40.5 40.5 0.5 = 81.0 \frac{40.5}{0.5} = 81.0 0.5 40.5 = 81.0 7.708 Reject H 0 H_0 H 0 (p < 0.001 p \lt 0.001 p < 0.001 ) Error (Residual) 2 2 ( 2 − 1 ) = 4 2^2(2 - 1) = 4 2 2 ( 2 − 1 ) = 4 2.0 MSE = 0.5 \text{MSE} = 0.5 MSE = 0.5 - - - Total 7 419.5 - - - -
Engineering Interpretation:
Increasing Temperature raises yield by an average of 11.5 % 11.5\% 11.5% . Increasing Pressure raises yield by 7.5 % 7.5\% 7.5% . The strong positive interaction (A B = + 4.5 % AB = +4.5\% A B = + 4.5% , F = 81.0 F = 81.0 F = 81.0 ) demonstrates synergistic behavior: the positive impact of temperature is amplified at high pressure.
6. Exam Traps & Operational Nuances
1. Degrees of Freedom Trap in LSD 2. Factorial Divisor Confusion
The Trap: Confusing error degrees of freedom in Latin Square Design.
The Reality: In a p × p p \times p p × p square, df Error = ( p − 1 ) ( p − 2 ) \text{df}_{\text{Error}} = (p - 1)(p - 2) df Error = ( p − 1 ) ( p − 2 ) .
Exam Consequence: For a 2 × 2 2 \times 2 2 × 2 Latin Square, df Error = ( 2 − 1 ) ( 2 − 2 ) = 0 \text{df}_{\text{Error}} = (2 - 1)(2 - 2) = 0 df Error = ( 2 − 1 ) ( 2 − 2 ) = 0 . For a 3 × 3 3 \times 3 3 × 3 square, df Error = ( 2 ) ( 1 ) = 2 \text{df}_{\text{Error}} = (2)(1) = 2 df Error = ( 2 ) ( 1 ) = 2 (too small to achieve statistical power). Hence, Latin Squares should practically be at least 4 × 4 4 \times 4 4 × 4 (p ≥ 4 p \ge 4 p ≥ 4 ).
The Trap: Using N N N instead of 2 k − 1 ⋅ r 2^{k-1} \cdot r 2 k − 1 ⋅ r when computing the estimated main effect from a contrast.
The Reality: The effect is the difference between two averages (High average minus Low average). Each average contains half the observations (2 k − 1 ⋅ r 2^{k-1} \cdot r 2 k − 1 ⋅ r ). In contrast, when computing Sum of Squares, the denominator is the total observations N = 2 k ⋅ r N = 2^k \cdot r N = 2 k ⋅ r .
7. Summary & Cheatsheet
CRD: SST = SSTreat + SSE \text{SST} = \text{SSTreat} + \text{SSE} SST = SSTreat + SSE (df : t − 1 , N − t \text{df}: t-1, N-t df : t − 1 , N − t ).RBD: SST = SSTreat + SSBlock + SSE \text{SST} = \text{SSTreat} + \text{SSBlock} + \text{SSE} SST = SSTreat + SSBlock + SSE (df Error = ( t − 1 ) ( r − 1 ) \text{df}_{\text{Error}} = (t-1)(r-1) df Error = ( t − 1 ) ( r − 1 ) ).LSD: SST = SSTreat + SSRow + SSCol + SSE \text{SST} = \text{SSTreat} + \text{SSRow} + \text{SSCol} + \text{SSE} SST = SSTreat + SSRow + SSCol + SSE (df Error = ( p − 1 ) ( p − 2 ) \text{df}_{\text{Error}} = (p-1)(p-2) df Error = ( p − 1 ) ( p − 2 ) ).Contrast: C = ∑ ( ± 1 ) ⋅ y i + C = \sum (\pm 1) \cdot y_{i+} C = ∑ ( ± 1 ) ⋅ y i + .Effect: Effect = C / ( 2 k − 1 r ) \text{Effect} = C / (2^{k-1} r) Effect = C / ( 2 k − 1 r ) .Sum of Squares: SS = C 2 / ( 2 k r ) \text{SS} = C^2 / (2^k r) SS = C 2 / ( 2 k r ) with df = 1 \text{df} = 1 df = 1 per effect.
Key Takeaways
Blocking Reduces MSE: Introducing blocks (RBD) or rows and columns (LSD) partitions nuisance variation out of the error sum of squares, lowering MSE \text{MSE} MSE and boosting the F F F -test sensitivity.
Synergy in Factorials: Full factorial designs are uniquely capable of discovering non-linear interactions where the effect of one parameter depends on another.
Exponential Scaling: As k k k increases, full factorial runs escalate exponentially (2 k 2^k 2 k ), motivating fractional factorial designs and Taguchi orthogonal arrays.
Next Section: Fractional Factorials & Taguchi Robust Design - Learn how fractional designs and Taguchi orthogonal arrays (L 9 L_9 L 9 ) dramatically reduce required runs while optimizing robust performance against noise.