Steyn's (2009) analysis of the extended Solomon design
Source:R/solomon_steyn.R
fit_solomon_steyn.Rd
Carries out the sequence of tests Steyn (2009) proposed for the Solomon
four-group design and for its extension to several interventions, the
extended Solomon design: k interventions and a control, each with and
without a pretest, giving 2(k + 1) groups. The sequence checks the
threats to internal validity the design can detect (nonequivalent groups,
history and maturation, testing, the pretest-intervention interaction,
instrumentation, regression to the mean, and attrition) before it
compares the interventions. It is a published
proposal, kept for replication and teaching. The package's recommended
analysis is
fit_solomon_glm() with control =, which estimates the
Solomon contrasts of every intervention in one model.
Usage
fit_solomon_steyn(
y_post,
treat,
pretested,
y_pre = NULL,
control = NULL,
alpha = 0.05,
include_unpretested_control = FALSE,
posthoc = c("scheffe", "holm"),
data = NULL
)Arguments
- y_post
Numeric posttest scores, with
NAfor participants who dropped out.- treat
The condition: a 0/1 (or logical) indicator for one intervention (1 = intervention), or a factor or character vector of conditions, with the control named by
control.- pretested
Pretest indicator coded 0/1 (or logical).
- y_pre
Optional numeric pretest scores,
NAfor unpretested participants. Without them the steps that use the pretest are skipped.- control
The control condition, when
treatis a factor or character vector.- alpha
Significance level of every test in the sequence. Default 0.05.
- include_unpretested_control
Logical. If
TRUE, the equivalence step adds a one-way ANOVA with the unpretested control's posttests (Of) as a further group. DefaultFALSE.- posthoc
The post hoc tests of pairs of groups:
"scheffe"(the default), Scheffé's (1953) tests, as in Steyn (2005); or"holm", pairwise t tests with Holm's (1979) adjustment. Both use a pooled SD. See the Operationalization section.- data
Optional data frame. When supplied, the other data arguments are looked up in it first, as bare column names (
y_post = post) or as strings (y_post = "post").
Value
An object of class solomon_steyn, a list with:
equivalence,testing,reliability,regression: tables of tests for steps 1, 3, 5, and 6.testingaddscomparison,term("Pretest","Intervention","Interaction"), andsum_sq;regressionaddsvar_pre,var_post, anddirection.history:tests, andpattern, the classification of step 2.classic:summary(one row per intervention: Test A, the path, and the conclusion) andfits, thefit_solomon_classic()results.attrition:counts(randomized, observed, and missing posttests in each group),dropouts(the two-by-two table of dropouts), andtests.effects:tests(E1 to E5),groups(for several interventions, whether each intervention group differs fromOdand fromOf), andhighest(the intervention with the highest combined posttest mean in E5, orNA; see the Operationalization section).path: the steps of step 8 that Steyn's sequence acts on.conclusions: one plain-language summary for each step.notes: skipped steps and data notes.n: counts of participants, posttests, and pretests in each group.conditions: the control and the interventions.settings:alpha,include_unpretested_control,posthoc,k, andpretest(whether pretest scores were supplied).
The tables of tests share the columns step, test, groups (Steyn's
labels with the condition names), n, estimate (the first group's
mean minus the second's for a t test; the mean change for a paired t
test; r for a correlation; the variance ratio; or the difference in
dropout proportions), statistic, reference ("t", "F",
"chisq", or "z"), df1, df2, p.value, p.adjusted (the post
hoc p-value of a pair of groups: Scheffé's, or Holm-adjusted, as
posthoc sets; NA for the other tests), and interpretation
(Steyn's reading, where he gives one). For a post hoc test, statistic
and p.value are the pairwise t with a pooled SD and its unadjusted
p-value; the decisions of the sequence read p.adjusted.
Steyn's sequence
Steyn labels the scores by group: Oa1, Ob1 are the pretest and
posttest of the pretested group of intervention 1 (EG1); Oc, Od the
pretest and posttest of the pretested control (CG1); Oe1 the posttest
of the unpretested group of intervention 1 (CG2.1); and Of the posttest
of the unpretested control (CG3). With one intervention the number is
dropped (Oa, Ob, Oe; EG, CG2). The steps below follow a
pre-publication draft of the article, dated March 2, 2009 (see Source).
In Steyn's order:
Equivalence after randomization. The pretests of the pretested groups are compared: a t test for one intervention, a one-way ANOVA otherwise. If they differ, Steyn advises post hoc tests to locate the difference, an explanation, and reconsidering whether to continue. With
include_unpretested_control = TRUE, a further one-way ANOVA addsOf, which Steyn suggests when the study is short or history and maturation are believed negligible.History and maturation. A paired t test of
OcandOd; a t test of all the pretests (the Time 1 value for the unpretested control) againstOf; and a one-way ANOVA ofOc,Od, andOf. IfOcdiffers fromOdandOf, Steyn reads history or maturation; ifOddiffers fromOcandOf, the pretest (a testing effect). Steyn (2009) presents this reading of the three sets of scores (Oc,Od, andOf) as new.Testing effect. The pretest main effect in the two-way between-groups ANOVA of the four posttest groups of each intervention and the control.
Pretest-intervention interaction. The decision sequence of Walton Braver and Braver (1988),
fit_solomon_classic(), for each intervention against the control.Test-retest reliability and instrumentation. The correlation of
OcandOd; the paired t test ofOcandOd; and a t test ofOcandOf.Regression to the mean. A chi-square test for the variance of
Odagainst that ofOc. Steyn reads an increase in variance as a sign of regression to the mean, a threat when groups were selected for extreme scores.Attrition. A z test of the dropout proportions of the intervention groups against the non-intervention groups (CG1 and CG3), and Steyn's chi-square on the two-by-two table of dropouts, intervention or not by pretested or not (df = 1). That chi-square asks whether intervention and pretesting are associated among the dropouts: whether the dropouts of the intervention groups were pretested more, or less, often than the dropouts of the non-intervention groups. It counts dropouts only, so it does not compare dropout rates, and it depends on how many participants each group started with. To compare dropout rates by group, see
attrition$countsandcheck_solomon_missing(), which counts the missing posttests in each group.Effects of the interventions. For one intervention: E1, a one-way ANOVA of the four posttest groups, and E2, a t test of
ObplusOeagainstOdplusOf; Steyn (2009) credits these two comparisons to his doctoral thesis (Steyn, 2001, as cited in Steyn, 2009). For several:E1, a one-way ANOVA of all 2(k + 1) posttest groups;
E2, post hoc tests of every pair of groups, summarized by whether each intervention group differs from
Odand fromOf;E3, a one-way ANOVA of the 2k intervention groups;
E4, for each intervention, a t test of
ObagainstOe;E5, when no E4 test is significant, a one-way ANOVA of the k interventions with their two groups combined, with post hoc tests, which Steyn uses to find which intervention had the greatest effect. When an E4 test is significant, Steyn cautions that internal validity is in question and the groups are not combined.
Every test is computed when the data allow it; path records the steps
Steyn's sequence acts on at alpha. Steps 1, 2, 4, 5, and 6 use the
pretest and are skipped, with a note, when y_pre is NULL.
Operationalization
Steyn's description leaves these choices open. solomonR made them as follows.
t tests. Independent-samples t tests are Student's pooled-variance tests, matching the equal-variance one-way ANOVAs.
Post hoc tests. Steyn (2009) names no post hoc test. In the study the article describes, Steyn (2005, Table 5.61, p. 153) used Scheffé tests after the one-way ANOVA of the eight posttest groups, so they are the default (
posthoc = "scheffe"). They are the method of Scheffé (1953). For a pair of groups among G groups with N scores in all, the Scheffé statistic is the squared pairwise t divided by G - 1, on G - 1 and N - G degrees of freedom, with the pooled error variance of the groups in the ANOVA. The pair differs atalphaexactly when Scheffé's criterion holds: |t| exceeds S, where S^2 is G - 1 times the upperalphapoint of F(G - 1, N - G) (Scheffé, 1953, pp. 87–89, 97). Scheffé's procedure protects every contrast among the groups, not only the pairs, so it is conservative when only pairs are compared; for pairs of equally precise means (equal group sizes), Scheffé recommended Tukey's method instead (pp. 89, 92–93, 96–97). Because the F test rejects exactly when some contrast is significant (pp. 87, 95–96), no pair can differ unless the ANOVA of the same groups is significant, and a significant ANOVA need not yield a pair that differs; step 2 then reports a pattern Steyn's rule does not cover.posthoc = "holm"gives pairwise t tests with the same pooled SD and Holm's (1979) adjustment, the default ofstats::pairwise.t.test(), which covers the pairwise comparisons only. In step 1 the post hoc tests follow a significant ANOVA of more than two groups; in step 2, and in step 8 with several interventions (E2 and E5), they are always computed.Two-way ANOVA. Type III sums of squares, with effect coding, one analysis for each intervention against the control.
History classification. Two sets "differ" when their post hoc p-value (Scheffé's, or Holm-adjusted) is below
alpha. A nonsignificant ANOVA gives "no evidence of history, maturation, or a testing effect";Ocdiffering fromOdand fromOf, which do not differ, gives "history or maturation";Oddiffering fromOcand fromOf, which do not differ, gives "the pretest (a testing effect)"; any other pattern is reported as one Steyn's rule does not cover.Scores used. The one-way ANOVA of
Oc,Od, andOfand the t test ofOcandOfuse every available score in each set. The paired t test, the correlation, and the variance test use the pretested controls with both scores.Variance test. The statistic is (n - 1) times the variance of
Oddivided by the variance ofOc, on n - 1 degrees of freedom, with a two-sided p-value (twice the smaller tail).Attrition. A participant with a missing posttest is a dropout. "Intervention" means any intervention. The z test pools the two proportions; neither test has a continuity correction. The chi-square is computed on the dropout counts, as Steyn describes; a note is recorded when an expected count is below 5.
"Differs from both control groups" (E2). An intervention group differs from both when its post hoc p-values against
Odand againstOfare both belowalpha. The sequence continues to E3 when at least one intervention group does."Greatest effect" (E5). Steyn does not say how the intervention with the greatest effect is identified. solomonR reports the intervention with the highest combined posttest mean (
highest) and the post hoc tests of every pair. The highest mean is the greatest effect only when the interventions raise the scores; when they lower them, as in Steyn (2005), the greatest effect is the lowest mean. Readhighestwith the direction of the outcome and the post hoc tests. In the English report of that study, Steyn and Mynhardt (2008, pp. 569–570) ranked the treatments by the difference between treated and untreated means and by d, from separate 2 x 2 analyses, without a test of the differences between treatments.Missing scores. Participants with a missing posttest stay in the data for the attrition step and are left out of the posttest analyses. The pretest analyses use the pretested participants with a pretest.
Tests that cannot be computed. A test of groups that are empty, too small, or without variation is left out, with a note in
notes.Step 4 uses
fit_solomon_classic(flow = "1988")with its defaults (ANCOVA as the pretested-groups test, and Test I).
Cautions
Steyn notes that multiple ANOVAs capitalize on chance. The sequence runs many tests, each at
alpha, with no control of the error rate across them.A nonsignificant test is not evidence that groups are equivalent, or that a threat is absent.
The one-way ANOVA of
Oc,Od, andOf, and its post hoc tests, treat the paired scores ofOcandOd, which come from the same participants, as independent. Scheffé's (1953, p. 87) tests take the covariances of the group means as known; here that ofOcandOdis taken to be zero.The chi-square test for the variance treats the variance of
Ocas a known value and ignores the pairing ofOcandOd.A difference between
ObandOe(E4) mixes a pretest main effect with pretest sensitization. Step 4 tests the interaction for each intervention separately; the joint model (fit_solomon_glm()) tests it for all the interventions in one model, with adjusted p-values for the comparisons.The package's recommended analysis is
fit_solomon_glm(control = ).
Source
This function follows a pre-publication draft of Steyn's (2009) article, dated March 2, 2009, which the author provided (R. Steyn, personal communication, September 30, 2026). The published article was not available for comparison. The draft has no page numbers, so none are cited.
Acknowledgment
We thank Renier Steyn for providing the draft of his article, and for noting that his model is intended for situations with a large amount of data and ample time (R. Steyn, personal communication, September 30, 2026).
References
Holm, S. (1979). A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics, 6(2), 65–70. https://www.jstor.org/stable/4615733
Scheffé, H. (1953). A method for judging all contrasts in the analysis of variance. Biometrika, 40(1–2), 87–104. https://doi.org/10.1093/biomet/40.1-2.87
Steyn, R. (2005). Self-evaluasie en die vorming van selfdoeltreffendheidspersepsies [Self-evaluation and the forming of self-efficacy perceptions] [Doctoral thesis, University of South Africa]. Unisa Institutional Repository. https://hdl.handle.net/10500/1745
Steyn, R. (2009). Re-designing the Solomon four-group: Can we improve on this exemplary model? Design Principles and Practices: An International Journal—Annual Review, 3(1), 383–394. https://doi.org/10.18848/1833-1874/CGP/v03i01/37588
Steyn, R., & Mynhardt, J. (2008). Factors that influence the forming of self-evaluation and self-efficacy perceptions. South African Journal of Psychology, 38(3), 563–573. https://doi.org/10.1177/008124630803800310
Walton Braver, M. C., & Braver, S. L. (1988). Statistical treatment of the Solomon four-group design: A meta-analytic approach. Psychological Bulletin, 104(1), 150–154. https://doi.org/10.1037/0033-2909.104.1.150
See also
fit_solomon_glm() for the recommended analysis,
fit_solomon_classic(), check_solomon_missing(),
report_solomon() for APA 7 text of the results, and steyn2005 for
the summary statistics of Steyn's (2005) eight-group study.
Examples
# Six groups: two interventions and a control.
fit <- fit_solomon_steyn(post_behavior, condition, pretested, pre_behavior,
control = "Control", data = mai2020)
fit
#> Steyn's (2009) analysis of the extended Solomon design (a published proposal; the package's recommended analysis is fit_solomon_glm())
#> Follows a pre-publication draft of the article.
#> Interventions: RP, GS; control: Control. With and without a pretest: 6 groups. alpha = 0.05.
#>
#> Groups
#> Group Condition Pretested Pretest Posttest n Posttests
#> EG1 RP yes Oa1 Ob1 35 24
#> EG2 GS yes Oa2 Ob2 33 23
#> CG1 Control yes Oc Od 50 27
#> CG2.1 RP no - Oe1 31 22
#> CG2.2 GS no - Oe2 27 15
#> CG3 Control no - Of 35 22
#>
#> 1. Equivalence after randomization
#> Groups Test Estimate Statistic p
#> Oa1, Oa2, Oc One-way ANOVA F(2, 115) = 0.78 0.459
#> No significant difference between the pretested groups at pretest.
#>
#> 2. History and maturation
#> Groups Test Estimate Statistic p p adj.
#> Od vs Oc, paired Paired t test -0.010 t(26) = -0.13 0.901
#> Of vs Oa1 + Oa2 + Oc t test (pooled variance) -0.142 t(138) = -1.75 0.083
#> Oc, Od, Of One-way ANOVA (scores as independent) F(2, 96) = 0.78 0.462
#> Oc vs Od Pairwise (pooled SD, Scheffe) 0.016 t(96) = 0.19 0.851 0.982
#> Oc vs Of Pairwise (pooled SD, Scheffe) 0.109 t(96) = 1.23 0.223 0.474
#> Od vs Of Pairwise (pooled SD, Scheffe) 0.094 t(96) = 0.94 0.351 0.646
#> Oc and Od do not differ (paired t test). The pretests and Of do not differ.
#> The one-way ANOVA of Oc, Od, and Of finds no evidence of history,
#> maturation, or a testing effect.
#>
#> 3. Testing effect: two-way between-groups ANOVA (Type III)
#> Comparison Term SS Statistic p
#> RP vs Control Pretest 0.066 F(1, 91) = 0.45 0.504
#> RP vs Control Intervention 0.031 F(1, 91) = 0.21 0.645
#> RP vs Control Interaction 0.508 F(1, 91) = 3.46 0.066
#> GS vs Control Pretest 0.063 F(1, 83) = 0.48 0.492
#> GS vs Control Intervention 0.186 F(1, 83) = 1.42 0.237
#> GS vs Control Interaction 0.031 F(1, 83) = 0.24 0.626
#> RP vs Control: no pretest main effect. GS vs Control: no pretest main
#> effect.
#>
#> 4. Pretest-intervention interaction: Tests A-I (Walton Braver & Braver, 1988)
#> Comparison Test A p Path
#> RP vs Control F(1, 91) = 3.46 0.066 A -> D -> E -> H -> I
#> GS vs Control F(1, 83) = 0.24 0.626 A -> D -> E -> H -> I
#> RP vs Control (path A -> D -> E -> H -> I): Historical pathway: no
#> treatment test in the selected A-I sequence reaches the specified alpha
#> level. GS vs Control (path A -> D -> E -> H -> I): Historical pathway: no
#> treatment test in the selected A-I sequence reaches the specified alpha
#> level.
#>
#> 5. Test-retest reliability and instrumentation
#> Step Groups Test Estimate Statistic p
#> Reliability Oc with Od Pearson r (test-retest) 0.247 t(25) = 1.28 0.213
#> Instrumentation Od vs Oc, paired Paired t test -0.010 t(26) = -0.13 0.901
#> Instrumentation Of vs Oc t test (pooled variance) -0.109 t(70) = -1.24 0.221
#> Test-retest r = 0.25 (n = 27). Oc and Od do not differ. Oc and Of do not
#> differ.
#>
#> 6. Regression to the mean: chi-square test for the variance
#> Groups Var(Oc) Var(Od) Ratio Statistic p
#> Od vs Oc, paired controls 0.095 0.126 1.33 chi2(26) = 34.59 0.242
#> The variance changed from 0.09496 (Oc) to 0.1263 (Od), a ratio of 1.33; the
#> change is not significant.
#>
#> 7. Attrition
#> Group Condition Randomized Observed Missing Rate
#> EG1 RP 35 24 11 31.4%
#> EG2 GS 33 23 10 30.3%
#> CG1 Control 50 27 23 46.0%
#> CG2.1 RP 31 22 9 29.0%
#> CG2.2 GS 27 15 12 44.4%
#> CG3 Control 35 22 13 37.1%
#>
#> Groups Test Estimate Statistic p
#> EG1 + EG2 + CG2.1 + CG2.2 vs CG1 + CG3 Two-proportion z test -0.090 z = -1.33 0.183
#> Dropouts: intervention by pretested Chi-square (2 x 2 dropouts) chi2(1) = 1.52 0.218
#> 78 of 211 participants (37.0%) have no posttest. Dropout was 33.3% with an
#> intervention and 42.4% without; the difference is not significant. Steyn's
#> chi-square: intervention and pretesting are not significantly associated
#> among the dropouts.
#>
#> 8. Effects of the interventions
#> Step Groups Test Estimate Statistic p p adj.
#> E1 Ob1, Ob2, Od, Oe1, Oe2, Of One-way ANOVA F(5, 127) = 1.34 0.251
#> E3 Ob1, Ob2, Oe1, Oe2 One-way ANOVA F(3, 80) = 1.83 0.148
#> E4 Ob1 vs Oe1 t test (pooled variance) -0.200 t(44) = -1.65 0.107
#> E4 Ob2 vs Oe2 t test (pooled variance) 0.016 t(36) = 0.13 0.898
#> E5 Ob1 + Oe1, Ob2 + Oe2 One-way ANOVA, groups combined F(1, 82) = 2.50 0.118
#> E5 Ob1 + Oe1 vs Ob2 + Oe2 Pairwise (pooled SD, Scheffe) -0.137 t(82) = -1.58 0.118 0.118
#> Steps not on the decision path are shown for completeness.
#>
#> E2: intervention groups against Od and Of (Scheffé tests)
#> Group Condition Mean p vs Od p vs Of Differs from both
#> Ob1 RP 2.929 0.707 0.986 no
#> Ob2 GS 3.168 0.998 0.882 no
#> Oe1 RP 3.129 1.000 0.968 no
#> Oe2 GS 3.152 1.000 0.953 no
#>
#> Decision path: E1
#> E1: the posttest groups do not differ, so Steyn's sequence finds no
#> evidence that the interventions had an effect.
# Pairwise t tests with Holm's adjustment as the post hoc tests.
fit_solomon_steyn(post_behavior, condition, pretested, pre_behavior,
control = "Control", posthoc = "holm", data = mai2020)
#> Steyn's (2009) analysis of the extended Solomon design (a published proposal; the package's recommended analysis is fit_solomon_glm())
#> Follows a pre-publication draft of the article.
#> Interventions: RP, GS; control: Control. With and without a pretest: 6 groups. alpha = 0.05.
#>
#> Groups
#> Group Condition Pretested Pretest Posttest n Posttests
#> EG1 RP yes Oa1 Ob1 35 24
#> EG2 GS yes Oa2 Ob2 33 23
#> CG1 Control yes Oc Od 50 27
#> CG2.1 RP no - Oe1 31 22
#> CG2.2 GS no - Oe2 27 15
#> CG3 Control no - Of 35 22
#>
#> 1. Equivalence after randomization
#> Groups Test Estimate Statistic p
#> Oa1, Oa2, Oc One-way ANOVA F(2, 115) = 0.78 0.459
#> No significant difference between the pretested groups at pretest.
#>
#> 2. History and maturation
#> Groups Test Estimate Statistic p p adj.
#> Od vs Oc, paired Paired t test -0.010 t(26) = -0.13 0.901
#> Of vs Oa1 + Oa2 + Oc t test (pooled variance) -0.142 t(138) = -1.75 0.083
#> Oc, Od, Of One-way ANOVA (scores as independent) F(2, 96) = 0.78 0.462
#> Oc vs Od Pairwise t (pooled SD, Holm) 0.016 t(96) = 0.19 0.851 0.851
#> Oc vs Of Pairwise t (pooled SD, Holm) 0.109 t(96) = 1.23 0.223 0.668
#> Od vs Of Pairwise t (pooled SD, Holm) 0.094 t(96) = 0.94 0.351 0.703
#> Oc and Od do not differ (paired t test). The pretests and Of do not differ.
#> The one-way ANOVA of Oc, Od, and Of finds no evidence of history,
#> maturation, or a testing effect.
#>
#> 3. Testing effect: two-way between-groups ANOVA (Type III)
#> Comparison Term SS Statistic p
#> RP vs Control Pretest 0.066 F(1, 91) = 0.45 0.504
#> RP vs Control Intervention 0.031 F(1, 91) = 0.21 0.645
#> RP vs Control Interaction 0.508 F(1, 91) = 3.46 0.066
#> GS vs Control Pretest 0.063 F(1, 83) = 0.48 0.492
#> GS vs Control Intervention 0.186 F(1, 83) = 1.42 0.237
#> GS vs Control Interaction 0.031 F(1, 83) = 0.24 0.626
#> RP vs Control: no pretest main effect. GS vs Control: no pretest main
#> effect.
#>
#> 4. Pretest-intervention interaction: Tests A-I (Walton Braver & Braver, 1988)
#> Comparison Test A p Path
#> RP vs Control F(1, 91) = 3.46 0.066 A -> D -> E -> H -> I
#> GS vs Control F(1, 83) = 0.24 0.626 A -> D -> E -> H -> I
#> RP vs Control (path A -> D -> E -> H -> I): Historical pathway: no
#> treatment test in the selected A-I sequence reaches the specified alpha
#> level. GS vs Control (path A -> D -> E -> H -> I): Historical pathway: no
#> treatment test in the selected A-I sequence reaches the specified alpha
#> level.
#>
#> 5. Test-retest reliability and instrumentation
#> Step Groups Test Estimate Statistic p
#> Reliability Oc with Od Pearson r (test-retest) 0.247 t(25) = 1.28 0.213
#> Instrumentation Od vs Oc, paired Paired t test -0.010 t(26) = -0.13 0.901
#> Instrumentation Of vs Oc t test (pooled variance) -0.109 t(70) = -1.24 0.221
#> Test-retest r = 0.25 (n = 27). Oc and Od do not differ. Oc and Of do not
#> differ.
#>
#> 6. Regression to the mean: chi-square test for the variance
#> Groups Var(Oc) Var(Od) Ratio Statistic p
#> Od vs Oc, paired controls 0.095 0.126 1.33 chi2(26) = 34.59 0.242
#> The variance changed from 0.09496 (Oc) to 0.1263 (Od), a ratio of 1.33; the
#> change is not significant.
#>
#> 7. Attrition
#> Group Condition Randomized Observed Missing Rate
#> EG1 RP 35 24 11 31.4%
#> EG2 GS 33 23 10 30.3%
#> CG1 Control 50 27 23 46.0%
#> CG2.1 RP 31 22 9 29.0%
#> CG2.2 GS 27 15 12 44.4%
#> CG3 Control 35 22 13 37.1%
#>
#> Groups Test Estimate Statistic p
#> EG1 + EG2 + CG2.1 + CG2.2 vs CG1 + CG3 Two-proportion z test -0.090 z = -1.33 0.183
#> Dropouts: intervention by pretested Chi-square (2 x 2 dropouts) chi2(1) = 1.52 0.218
#> 78 of 211 participants (37.0%) have no posttest. Dropout was 33.3% with an
#> intervention and 42.4% without; the difference is not significant. Steyn's
#> chi-square: intervention and pretesting are not significantly associated
#> among the dropouts.
#>
#> 8. Effects of the interventions
#> Step Groups Test Estimate Statistic p p adj.
#> E1 Ob1, Ob2, Od, Oe1, Oe2, Of One-way ANOVA F(5, 127) = 1.34 0.251
#> E3 Ob1, Ob2, Oe1, Oe2 One-way ANOVA F(3, 80) = 1.83 0.148
#> E4 Ob1 vs Oe1 t test (pooled variance) -0.200 t(44) = -1.65 0.107
#> E4 Ob2 vs Oe2 t test (pooled variance) 0.016 t(36) = 0.13 0.898
#> E5 Ob1 + Oe1, Ob2 + Oe2 One-way ANOVA, groups combined F(1, 82) = 2.50 0.118
#> E5 Ob1 + Oe1 vs Ob2 + Oe2 Pairwise t (pooled SD, Holm) -0.137 t(82) = -1.58 0.118 0.118
#> Steps not on the decision path are shown for completeness.
#>
#> E2: intervention groups against Od and Of (Holm-adjusted pairwise t tests)
#> Group Condition Mean p vs Od p vs Of Differs from both
#> Ob1 RP 2.929 1.000 1.000 no
#> Ob2 GS 3.168 1.000 1.000 no
#> Oe1 RP 3.129 1.000 1.000 no
#> Oe2 GS 3.152 1.000 1.000 no
#>
#> Decision path: E1
#> E1: the posttest groups do not differ, so Steyn's sequence finds no
#> evidence that the interventions had an effect.
# The four-group design: one intervention.
fit_solomon_steyn(y_post, treat, pretested, y_pre, data = solomon_example)
#> Steyn's (2009) analysis of the extended Solomon design (a published proposal; the package's recommended analysis is fit_solomon_glm())
#> Follows a pre-publication draft of the article.
#> Intervention: Treatment; control: Control. With and without a pretest: 4 groups. alpha = 0.05.
#>
#> Groups
#> Group Condition Pretested Pretest Posttest n Posttests
#> EG Treatment yes Oa Ob 30 30
#> CG1 Control yes Oc Od 30 30
#> CG2 Treatment no - Oe 30 30
#> CG3 Control no - Of 30 30
#>
#> 1. Equivalence after randomization
#> Groups Test Estimate Statistic p
#> Oa vs Oc t test (pooled variance) 0.067 t(58) = 0.02 0.981
#> No significant difference between the pretested groups at pretest.
#>
#> 2. History and maturation
#> Groups Test Estimate Statistic p p adj.
#> Od vs Oc, paired Paired t test 4.933 t(29) = 2.72 0.011
#> Of vs Oa + Oc t test (pooled variance) 1.500 t(88) = 0.65 0.516
#> Oc, Od, Of One-way ANOVA (scores as independent) F(2, 87) = 2.14 0.124
#> Oc vs Od Pairwise (pooled SD, Scheffe) -4.933 t(87) = -2.02 0.046 0.136
#> Oc vs Of Pairwise (pooled SD, Scheffe) -1.533 t(87) = -0.63 0.531 0.821
#> Od vs Of Pairwise (pooled SD, Scheffe) 3.400 t(87) = 1.39 0.167 0.383
#> Oc and Od differ (paired t test). The pretests and Of do not differ. The
#> one-way ANOVA of Oc, Od, and Of finds no evidence of history, maturation,
#> or a testing effect.
#>
#> 3. Testing effect: two-way between-groups ANOVA (Type III)
#> Comparison Term SS Statistic p
#> Treatment vs Control Pretest 180.075 F(1, 116) = 1.95 0.165
#> Treatment vs Control Intervention 216.008 F(1, 116) = 2.34 0.129
#> Treatment vs Control Interaction 27.075 F(1, 116) = 0.29 0.589
#> Treatment vs Control: no pretest main effect.
#>
#> 4. Pretest-intervention interaction: Tests A-I (Walton Braver & Braver, 1988)
#> Comparison Test A p Path
#> Treatment vs Control F(1, 116) = 0.29 0.589 A -> D -> E -> H -> I
#> Treatment vs Control (path A -> D -> E -> H -> I): Historical pathway: no
#> treatment test in the selected A-I sequence reaches the specified alpha
#> level.
#>
#> 5. Test-retest reliability and instrumentation
#> Step Groups Test Estimate Statistic p
#> Reliability Oc with Od Pearson r (test-retest) 0.456 t(28) = 2.71 0.011
#> Instrumentation Od vs Oc, paired Paired t test 4.933 t(29) = 2.72 0.011
#> Instrumentation Of vs Oc t test (pooled variance) 1.533 t(58) = 0.61 0.543
#> Test-retest r = 0.46 (n = 30). Oc and Od differ, a possible instrumentation
#> effect. Oc and Of do not differ.
#>
#> 6. Regression to the mean: chi-square test for the variance
#> Groups Var(Oc) Var(Od) Ratio Statistic p
#> Od vs Oc, paired controls 100.530 79.569 0.79 chi2(29) = 22.95 0.443
#> The variance changed from 100.5 (Oc) to 79.57 (Od), a ratio of 0.79; the
#> change is not significant.
#>
#> 7. Attrition
#> Group Condition Randomized Observed Missing Rate
#> EG Treatment 30 30 0 0.0%
#> CG1 Control 30 30 0 0.0%
#> CG2 Treatment 30 30 0 0.0%
#> CG3 Control 30 30 0 0.0%
#> No dropouts: every posttest was observed.
#>
#> 8. Effect of the intervention
#> Step Groups Test Estimate Statistic p
#> E1 Ob, Od, Oe, Of One-way ANOVA F(3, 116) = 1.53 0.211
#> E2 Ob + Oe vs Od + Of t test (pooled variance) 2.683 t(118) = 1.53 0.129
#>
#> Decision path: E1 -> E2
#> E1: the four posttest groups do not differ. E2: the intervention groups
#> (Ob, Oe) do not differ from the non-intervention groups (Od, Of).
#>
#> Notes
#> - Every posttest was observed, so the attrition tests were skipped.