solomonR aims to capture the published methodology for the Solomon four-group design, from Solomon (1949) to the present. This page makes that aim checkable, source by source. For each work in the Solomon-design section of the References page, it gives:
- what the work contributes to the analysis of the design;
- where solomonR implements or uses it;
- its status.
A check in the package’s repository
(tools/check-references.R) fails if a work on the Solomon
design is added to the bibliography without an entry here.
The status column uses six labels:
- Historical. The procedure is reproduced for teaching and replication. It is not the recommended analysis.
- Implemented. The procedure is available as a current method.
- Data. The published data or numbers are bundled and reproduced.
- Evidence. The empirical findings inform planning or guidance.
- Guidance. The recommendations shape the package’s advice and reports.
- Planned or Not yet reviewed. An open issue tracks the work.
Foundations (1949–1973)
| Source | Contribution to the analysis of the design | In solomonR | Status |
|---|---|---|---|
| Solomon (1949) | The three- and four-group designs. An inferred pretest for the unpretested groups, improvement scores, and the interaction term I (pp. 141–147). The fourth group shows the effect of outside events and time. |
fit_solomon_1949(), solomon1949; the
history check of fit_solomon_classic(); every Solomon
estimand |
Historical; Data |
| Campbell (1957) | The interaction of testing and treatment as sensitization (p. 302). The first recommendation of the 2 x 2 analysis of variance of the four posttests and of the t test of the unpretested control against the pretests, in place of the inferred-pretest gain analysis (p. 303). | Tests A–D and the history check of
fit_solomon_classic(); the history article |
Historical |
| Lana (1959) | An attitude experiment with the Solomon design that found no pretest-by-treatment interaction, analyzed by a 2 x 2 analysis of variance of the posttest means (pp. 295–298). |
lana1959, whose published ANOVA
solomon_from_summary() reproduces; the history article |
Data (reproduced); Evidence |
| Entwisle (1961) | Training experiments with pretested and unpretested pupils paired on sex and IQ: no overall pretest effect, but a Pretesting x IQ x Sex interaction (pp. 612–613). | The history article | Evidence |
| Campbell & Stanley (1963/1966) | The design as Design 5 (p. 24). The 2 x 2 analysis of variance of the posttests, then an analysis of covariance of the pretested groups (p. 25). The unpretested control against the pretests as a check of history and maturation (p. 25). The static-group comparison (p. 12). | Tests A–D and the history check of
fit_solomon_classic(); the caveat on nonrandomized designs
in baseline_solomon() and
report_solomon()
|
Historical |
| Bracht & Glass (1968) | Pretest sensitization, posttest sensitization, and the timing of measurement as threats to external validity (p. 439); a review finding the pretest effect most likely with self-reported attitudes and personality (p. 463). | The history article | Evidence |
| Solomon & Lessac (1968) | The four-group design in developmental studies: the inferred pretest judges absolute improvement or deterioration, and the interaction is a difference of differences in a 2 x 2 factorial (pp. 146–147). | The inferred pretest of fit_solomon_1949(); the history
article |
Historical |
| Lana (1969/2009) | A review of pretest sensitization: sensitization with pretests that involve learning or recall (pp. 101–103), an “overwhelming lack” of it with attitude pretests (pp. 103–104), and the 2 x 2 posttest analysis (pp. 99–100). | The history article | Evidence |
| Lessac & Solomon (1969) | The full report of the beagle experiment that Solomon and Lessac (1968, p. 149) summarized, with isolation for a year from 12 weeks of age (pp. 15–16). Analyses of variance of the posttests for isolation, pretesting, and their interaction, and Mann–Whitney U tests of posttests against the pretests of the pretested groups, which separate a loss of existing abilities from a failure to develop (pp. 18–23). The pretest protected the isolates on some tests and impaired them on another (pp. 19, 21–24). | The history article | Evidence |
| Huck & Sandler (1973) | The analysis of covariance of the pretested groups remains valid when pretesting has a main effect. After a significant interaction, the unpretested groups’ t test answers whether the treatment works (pp. 54–55). | Tests E and H of fit_solomon_classic()
|
Historical |
The test sequence and its critics (1982–2000)
| Source | Contribution to the analysis of the design | In solomonR | Status |
|---|---|---|---|
| Willson & Putnam (1982) | A meta-analysis of pretest effects and pretest-by-treatment interactions in 32 studies. | Planning values in “Planning a Solomon Study”; defaults discussed in
simulate_solomon()
|
Evidence |
| Walton Braver & Braver (1988) | The Test A–I decision sequence, ending in a Stouffer combination of the pretested and unpretested comparisons (Test I). |
fit_solomon_classic(flow = "1988"),
stouffer_solomon(), plot_classic_flow()
|
Historical |
| Sawilowsky & Markman (1990a) | A counterexample in which Test I misses an effect that one of its component tests detects. | The history article | Evidence |
| Braver & Walton Braver (1990) | The 1990 amendment: once Tests A and D are nonsignificant, every test through Test I is run. | fit_solomon_classic(flow = "1990") |
Historical |
| Sawilowsky & Markman (1990b) | The rejoinder, calling for systematic study of the procedure. | The history article | Evidence |
| Sawilowsky et al. (1994) | Monte Carlo evidence of an inflated experiment-wise Type I error for the sequence. | Reproduced in “Historical Tests: Replicating the Published Error Rates”; the caution printed with Test I | Data (reproduced) |
| Sawilowsky (1996) | Error rates with nonnormal data, the 1995 revision without Test D, and two methods of alpha allocation. |
fit_solomon_classic(flow = "1995", alpha_allocation = ...);
reproduced in the replication article |
Historical; Data (reproduced) |
| Kvalem et al. (1996) | A cluster-randomized school trial with a binary outcome, analyzed at the individual level. |
kvalem1996; tests of fisher_solomon() and
marginal_solomon(); the allocation in the cluster-level
perm_solomon() simulation |
Data |
| van Engelenburg (1999) | Full-information maximum likelihood for the design, treating the absent pretests as structurally missing. |
fit_solomon_ml(), validated by simulation |
Implemented |
| Sawilowsky (2000) | Rank-transform tests fail for interactions. | The reason the package has no rank-transform test of sensitization;
perm_solomon() is its distribution-free option |
Guidance |
Recent work (2002–2025)
| Source | Contribution to the analysis of the design | In solomonR | Status |
|---|---|---|---|
| McCarthy & Tucker (2002) | A nonrandomized eight-group design crossing two interventions with pretesting, analyzed as three overlapping four-group designs (pp. 637–640). | The planned comparisons of
fit_solomon_glm(control = , contrasts = ); the factorial
example in “Designs With Several Treatments” |
Guidance |
| Steyn (2005) | An eight-group study with three treatments and 1,723 participants, in which existing classes were allocated to the groups (pp. 103–106). Analyzed as overlapping four-group designs, then by a one-way analysis of variance of the eight posttests with Scheffé tests (pp. 151–153), the method of Scheffé (1953). |
steyn2005, whose published analyses
solomon_from_summary() and the package’s tests reproduce;
the default post hoc tests of fit_solomon_steyn()
|
Data (reproduced) |
| Morris (2008) | Effect sizes for pretest–posttest–control designs; the pooled-pretest-SD estimator and its variance. |
solomon_effect_sizes() for the pretested pair |
Implemented |
| Steyn & Mynhardt (2008) | The English report of the eight-group study in Steyn (2005), with participants said to be allocated at random (pp. 566–567). Each treatment analyzed with the two control groups by a 2 x 2 analysis of variance of the posttests, through Step 2 of the sequence of Walton Braver and Braver (1988), at the .01 level; no pretest-by-intervention interaction (pp. 568–569). |
steyn2005, whose group statistics reproduce its 2 x 2
analyses within rounding; the English source cited in
?steyn2005
|
Data (reproduced) |
| Steyn (2009) | The design extended to k treatments, with 2(k + 1) groups, and to repeated posttests. A sequence of tests of internal validity (equivalence, history and maturation, testing, the pretest-intervention interaction, reliability, regression to the mean, attrition) and of the treatments’ effects. |
fit_solomon_steyn(); the design of
fit_solomon_glm(control = ). solomonR follows a
pre-publication draft, which the author provided. |
Historical; Implemented |
| McCambridge et al. (2011) | A systematic review of Solomon studies of behavior change: too little evidence to settle whether assessment biases trials. | The getting-started guide and the planning, decision, and reporting articles | Evidence |
| Jordaan (2014) | A randomized four-group study with three posttest occasions, each analyzed separately by the sequence of Walton Braver and Braver (1988), with full cell statistics on every occasion (pp. 98, 112–127). |
jordaan2014, whose published analyses
solomon_from_summary() and fit_solomon_mmrm()
reproduce; “Worked Example: Repeated Posttests” |
Data (reproduced) |
| Edmonds & Kennedy (2017) | Solomon four-, six-, and eight-group designs, and the threats of nonrandomized designs (pp. 7–8, 93–101). |
baseline_solomon(); the nonrandomized wording of
report_solomon(); designs with several treatments in
fit_solomon_glm(control = )
|
Guidance |
| Mai et al. (2020) | A randomized six-group design analyzed as overlapping four-group designs, with published individual data. |
mai2020 and the worked example, which reproduce its
Tables 4, 5, and 7; the history check of
fit_solomon_classic(); the six-group example of
fit_solomon_glm(control = )
|
Data; Historical |
| French et al. (2021a) | The MERIT recommendations on measurement reactivity in trials, including when a Solomon design is warranted. | “Should I Use a Solomon Design?”; the history article | Guidance |
| French et al. (2021b) | The full MERIT report, including its recommendations on reporting measurement in trials. | The measurement items of report_solomon(); the
reporting and decision articles |
Guidance |
| El Karkri et al. (2025a) | A classroom Solomon study with full cell statistics and one intact class per condition. |
elkarkri2025a; tests of
solomon_from_summary() and baseline_solomon();
the class–condition check of validate_solomon()
|
Data |
| El Karkri et al. (2025b) | An analysis path for categorical outcomes, judging sensitization by the significance of the two simple effects. |
fisher_solomon(), with the caution of Gelman and Stern
(2006) |
Historical |
What the package adds
Some methods in solomonR are not specific to the Solomon design. The package applies general methods to the design’s estimands, and each function cites its sources:
- one model for all the groups with robust standard errors
(
fit_solomon_glm()), and, for designs with several treatments, omnibus tests and comparisons adjusted by Holm’s (1979) procedure; - randomization inference (
perm_solomon()); - equivalence tests of sensitization
(
equivalence_solomon()); - latent-variable models and measurement invariance
(
fit_solomon_sem_latent(),invariance_solomon()); - marginal contrasts for binary and count outcomes
(
marginal_solomon()); - clustered designs;
- multiple imputation of missing posttests, with a tipping-point
sensitivity analysis (
fit_solomon_mi(),tipping_point_solomon()); - designs with several posttest occasions
(
fit_solomon_mmrm()); - planning (
plan_solomon(),power_solomon()), analysis plans for preregistration (analysis_plan_solomon()), and reporting (report_solomon()).
These sources are in the “Methods references” section of the References page, and “How to Cite solomonR
and the Methods It Implements” lists them by function. Where the package
combines methods in a way no source describes, the documentation labels
the combination a solomonR extension. An example is
simulate_solomon(), which adds a pretesting effect to the
validated simulation model of power_solomon().
Sources read in a version other than the published one
Steyn (2009) has been read in a pre-publication draft, dated March 2,
2009, which the author provided (R. Steyn, personal communication,
September 30, 2026). The draft has no page numbers, so none are cited.
The published article was not available for comparison, and
fit_solomon_steyn() follows the draft (#45).
Sources named but not read
Steyn (2005), Steyn and Mynhardt (2008), and the draft of Steyn (2009; R. Steyn, personal communication, September 30, 2026) cite a research-methods textbook, Foundations of Behavioral Research, by Kerlinger (1986) and, in its fourth edition, by Kerlinger and Lee (2000). The textbook has not been read. Each point these works take from it about the Solomon design is stated in a primary source in the bibliography:
- The analysis is complicated. The design’s six sets of observations are hard to combine in a single statistic (Kerlinger & Lee, 2000, as cited in Steyn & Mynhardt, 2008, p. 568). Steyn (2005, p. 103) cites both editions for the same point, and the draft of Steyn (2009) cites them as offering little guidance on the calculations before it sets out its own sequence of tests. Campbell and Stanley (1963/1966, p. 25) had made the point: no single statistical procedure uses all six sets of observations at once. Walton Braver and Braver (1988, p. 150) named uncertainty about the design’s statistical treatment as perhaps the main reason for its underuse, and Sawilowsky et al. (1994, p. 363) repeat Campbell and Stanley’s point.
- The design controls the threats to internal validity. The draft of Steyn (2009) cites both editions for this point. Campbell and Stanley (1963/1966, p. 8, Table 1) mark Design 5 as controlling every threat to internal validity they list, and Huck and Sandler (1973, p. 54) say the design controls them all.
- Random assignment to all four groups. The draft of Steyn (2009) cites Kerlinger (1986) and Solomon (1949) for this point. Solomon (1949, pp. 140–141) allowed groups that were either matched or randomly drawn. Random assignment is part of Design 5 in Campbell and Stanley (1963/1966, p. 24), and Huck and Sandler (1973, p. 54) make it the first step of the design.
Steyn (2005, p. 109) also lists both editions among the literature’s approaches to the design, with Solomon (1949), before adopting the approach of Walton Braver and Braver (1988). The remaining citations are general rather than specific to the Solomon design: Kerlinger and Lee (2000, as cited in Steyn, 2005, p. 111) for the .05 significance level and for reporting effect sizes; and, in the draft of Steyn (2009), a definition of research design and the claim that well-conceived research gives representative observations. The procedures solomonR takes from Steyn’s work come from Walton Braver and Braver (1988), Scheffé’s (1953) method, and the draft’s own sequence of tests. APA Style cites a source known only through another work as a secondary source, so the reference list gives the works that cite the textbook, not the textbook.
Steyn’s (2001) doctoral thesis, from Potchefstroom University for
Christian Higher Education, has not been read either: no copy was found.
Stadler and Kotze (2006) report that it evaluated an outdoor
experiential learning program for members of the Public Order Police
with “a four-group design with a pre-, post- and post-post test”
(p. 27). Its quantitative results showed a positive influence on the
participants’ view of humanity (p. 27), but no significant difference in
general self-efficacy scores between the experimental and control groups
over the three occasions (p. 30). Steyn advised future researchers to
avoid the four-group design, because the design is too complex and its
interdependency makes the hypotheses difficult to accept without
reservation (Steyn, 2001, as cited in Stadler & Kotze, 2006, p. 27).
He later extended it to eight groups (Steyn, 2005) and argued for the
extension (Steyn, 2009). Steyn (2005, p. 109) lists the thesis among the
approaches to the Solomon four-group design. The draft of Steyn (2009)
cites it for repeating the posttest 1 week and 3 months after the
intervention, to show whether an effect is sustained, and for comparing
the treated groups’ posttests with the untreated groups’: a one-way
analysis of variance of the four posttests, and a t test of the two
treated groups combined against the two untreated groups combined.
fit_solomon_steyn() carries out both comparisons (E1 and
E2). None of these works reports the thesis’s sample sizes, assignment,
or results in enough detail to reproduce. Designs with several posttest
occasions are analyzed by fit_solomon_mmrm() (#96).
References
Bracht, G. H., & Glass, G. V. (1968). The external validity of experiments. American Educational Research Journal, 5(4), 437–474. https://doi.org/10.3102/00028312005004437
Braver, S. L., & Walton Braver, M. C. (1990). Meta-analysis for Solomon four-group designs reconsidered: A reply to Sawilowsky and Markman. Perceptual and Motor Skills, 71(1), 321–322. https://doi.org/10.2466/pms.1990.71.1.321
Campbell, D. T. (1957). Factors relevant to the validity of experiments in social settings. Psychological Bulletin, 54(4), 297–312. https://doi.org/10.1037/h0040950
Campbell, D. T., & Stanley, J. C. (1966). Experimental and quasi-experimental designs for research. Rand McNally. (Original work published 1963)
Edmonds, W. A., & Kennedy, T. D. (2017). An applied guide to research designs: Quantitative, qualitative, and mixed methods (2nd ed.). SAGE Publications. https://doi.org/10.4135/9781071802779
El Karkri, M., Quesada, A., & Romero-Ariza, M. (2025a). The dual impact of pretest sensitisation and the cognitive acceleration through science education programme in the Solomon four-group design. Brain Sciences, 16(1), Article 64. https://doi.org/10.3390/brainsci16010064
El Karkri, M., Quesada, A., & Romero-Ariza, M. (2025b). Methodological aspects of the Solomon four-group design: Detecting pre-test sensitisation and analysing qualitative and quantitative variables in education research. Review of Education, 13(1), Article e70050. https://doi.org/10.1002/rev3.70050
Entwisle, D. R. (1961). Interactive effects of pretesting. Educational and Psychological Measurement, 21(3), 607–620. https://doi.org/10.1177/001316446102100307
French, D. P., Miles, L. M., Elbourne, D., Farmer, A., Gulliford, M., Locock, L., Sutton, S., McCambridge, J., & MERIT Collaborative Group. (2021a). Reducing bias in trials due to reactions to measurement: Experts produced recommendations informed by evidence. Journal of Clinical Epidemiology, 139, 130–139. https://doi.org/10.1016/j.jclinepi.2021.06.028
French, D. P., Miles, L. M., Elbourne, D., Farmer, A., Gulliford, M., Locock, L., Sutton, S., McCambridge, J., & MERIT Collaborative Group. (2021b). Reducing bias in trials from reactions to measurement: The MERIT study including developmental work and expert workshop. Health Technology Assessment, 25(55), 1–72. https://doi.org/10.3310/hta25550
Gelman, A., & Stern, H. (2006). The difference between “significant” and “not significant” is not itself statistically significant. The American Statistician, 60(4), 328–331. https://doi.org/10.1198/000313006X152649
Holm, S. (1979). A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics, 6(2), 65–70. https://www.jstor.org/stable/4615733
Huck, S. W., & Sandler, H. M. (1973). A note on the Solomon 4-group design: Appropriate statistical analyses. The Journal of Experimental Education, 42(2), 54–55. https://doi.org/10.1080/00220973.1973.11011460
Jordaan, J. (2014). The development and evaluation of a life skills programme for young adult prisoners [Doctoral thesis, University of the Free State]. KovsieScholar. https://hdl.handle.net/11660/832
Kvalem, I. L., Sundet, J. M., Rivø, K. I., Eilertsen, D. E., & Bakketeig, L. S. (1996). The effect of sex education on adolescents’ use of condoms: Applying the Solomon four-group design. Health Education Quarterly, 23(1), 34–47. https://doi.org/10.1177/109019819602300103
Lana, R. E. (1959). Pretest-treatment interaction effects in attitudinal studies. Psychological Bulletin, 56(4), 293–300. https://doi.org/10.1037/h0044646
Lana, R. E. (2009). Pretest sensitization. In R. Rosenthal & R. L. Rosnow, Artifacts in behavioral research: Robert Rosenthal and Ralph L. Rosnow’s classic books (pp. 93–109). Oxford University Press. https://doi.org/10.1093/acprof:oso/9780195385540.003.0004 (Original work published 1969)
Lessac, M. S., & Solomon, R. L. (1969). Effects of early isolation on the later adaptive behavior of beagles: A methodological demonstration. Developmental Psychology, 1(1), 14–25. https://doi.org/10.1037/h0026778
Mai, N. N., Takahashi, Y., & Oo, M. M. (2020). Testing the effectiveness of transfer interventions using Solomon four-group designs. Education Sciences, 10(4), Article 92. https://doi.org/10.3390/educsci10040092
McCambridge, J., Butor-Bhavsar, K., Witton, J., & Elbourne, D. (2011). Can research assessments themselves cause bias in behaviour change trials? A systematic review of evidence from Solomon 4-group studies. PLoS ONE, 6(10), Article e25223. https://doi.org/10.1371/journal.pone.0025223
McCarthy, A. M., & Tucker, M. L. (2002). Encouraging community service through service learning. Journal of Management Education, 26(6), 629–647. https://doi.org/10.1177/1052562902238322
Morris, S. B. (2008). Estimating effect sizes from pretest-posttest-control group designs. Organizational Research Methods, 11(2), 364–386. https://doi.org/10.1177/1094428106291059
Sawilowsky, S. S. (1996, June 23). Controlling experiment-wise Type I error of meta-analysis in the Solomon four-group design [Paper presentation]. First International Conference on Multiple Comparisons, Tel Aviv, Israel. https://digitalcommons.wayne.edu/coe_tbf/29/
Sawilowsky, S. S. (2000). Review of the rank transform in designed experiments. Perceptual and Motor Skills, 90(2), 489–497. https://doi.org/10.2466/pms.2000.90.2.489
Sawilowsky, S. S., Kelley, D. L., Blair, R. C., & Markman, B. S. (1994). Meta-analysis and the Solomon four-group design. The Journal of Experimental Education, 62(4), 361–376. https://doi.org/10.1080/00220973.1994.9944140
Sawilowsky, S. S., & Markman, B. S. (1990a). Another look at the power of meta-analysis in the Solomon four-group design. Perceptual and Motor Skills, 71(1), 177–178. https://doi.org/10.2466/pms.1990.71.1.177
Sawilowsky, S. S., & Markman, B. S. (1990b). Rejoinder to Braver and Walton Braver. Perceptual and Motor Skills, 71(2), 424–426. https://doi.org/10.2466/pms.1990.71.2.424
Scheffé, H. (1953). A method for judging all contrasts in the analysis of variance. Biometrika, 40(1–2), 87–104. https://doi.org/10.1093/biomet/40.1-2.87
Solomon, R. L. (1949). An extension of control group design. Psychological Bulletin, 46(2), 137–150. https://doi.org/10.1037/h0062958
Solomon, R. L., & Lessac, M. S. (1968). A control group design for experimental studies of developmental processes. Psychological Bulletin, 70(3, Pt. 1), 145–150. https://doi.org/10.1037/h0026147
Stadler, K., & Kotze, M. E. (2006). The influence of a ropes course development programme on the self-concept and self-efficacy of young career officers. SA Journal of Industrial Psychology, 32(1), 25–32. https://doi.org/10.4102/sajip.v32i1.225
Steyn, R. (2005). Self-evaluasie en die vorming van selfdoeltreffendheidspersepsies [Self-evaluation and the forming of self-efficacy perceptions] [Doctoral thesis, University of South Africa]. Unisa Institutional Repository. https://hdl.handle.net/10500/1745
Steyn, R. (2009). Re-designing the Solomon four-group: Can we improve on this exemplary model? Design Principles and Practices: An International Journal—Annual Review, 3(1), 383–394. https://doi.org/10.18848/1833-1874/CGP/v03i01/37588
Steyn, R., & Mynhardt, J. (2008). Factors that influence the forming of self-evaluation and self-efficacy perceptions. South African Journal of Psychology, 38(3), 563–573. https://doi.org/10.1177/008124630803800310
van Engelenburg, G. (1999). Statistical analysis for the Solomon four-group design (Research Report 99-06). University of Twente. ERIC. https://eric.ed.gov/?id=ED435692
Walton Braver, M. C., & Braver, S. L. (1988). Statistical treatment of the Solomon four-group design: A meta-analytic approach. Psychological Bulletin, 104(1), 150–154. https://doi.org/10.1037/0033-2909.104.1.150
Willson, V. L., & Putnam, R. R. (1982). A meta-analysis of pretest sensitization effects in experimental design. American Educational Research Journal, 19(2), 249–258. https://doi.org/10.3102/00028312019002249