solomonR attributes every procedure it implements to its published source. The article “Coverage of the Published Methodology” maps each work on the Solomon design to what the package does with it. This page is the package’s canonical reference list: each reference in the help pages and articles is written exactly as it appears here, in APA Style (7th ed.), with DOIs verified against Crossref.
The first list is the literature on the Solomon design itself, each
entry followed by a note on what the work contributes and where solomonR
uses it. The second list holds the general methods references cited in
the documentation. To cite solomonR, run
citation("solomonR"), and cite the methods you used from
the lists below.
The Solomon design literature
Bracht, G. H., & Glass, G. V. (1968). The external validity of experiments. American Educational Research Journal, 5(4), 437–474. https://doi.org/10.3102/00028312005004437
Lists pretest sensitization, posttest sensitization, and the interaction of time of measurement and treatment among the threats to external validity (p. 439), and reviews the evidence on pretest sensitization (pp. 460–463): the effect is most likely when the outcome is a self-report of personality, attitude, or opinion, and the evidence for achievement is inconclusive (p. 463). In solomonR: the history article and the coverage article (#80).
Braver, S. L., & Walton Braver, M. C. (1990). Meta-analysis for Solomon four-group designs reconsidered: A reply to Sawilowsky and Markman. Perceptual and Motor Skills, 71(1), 321–322. https://doi.org/10.2466/pms.1990.71.1.321
The authors’ reply to Sawilowsky and Markman (1990a). It amends the 1988 decision flow: once Tests A and D are nonsignificant, all tests through Test I are carried out. In solomonR: the 1990 flow version,
fit_solomon_classic(flow = "1990")andplot_classic_flow(flow = "1990")(#50).
Campbell, D. T. (1957). Factors relevant to the validity of experiments in social settings. Psychological Bulletin, 54(4), 297–312. https://doi.org/10.1037/h0040950
Describes the interaction of testing and treatment as a sensitization that limits generalization to unpretested populations (p. 302), and calls the Solomon design “the new ideal design for social scientists” (p. 303). It rejects the gain-score analysis with an inferred pretest, which restricts the degrees of freedom and violates independence, and recommends a two-by-two analysis of variance of the four posttests, with a t test of the unpretested control posttest against the pretests for history and maturation (p. 303). In solomonR: the first statement of Tests A–D and of the history check of
fit_solomon_classic(); the history article.
Campbell, D. T., & Stanley, J. C. (1966). Experimental and quasi-experimental designs for research. Rand McNally. (Original work published 1963)
Places the Solomon design among the true experimental designs (pp. 13, 24), names the interaction of testing and treatment as a threat to external validity (p. 18), and recommends a 2 x 2 analysis of posttest scores, followed by ANCOVA of the pretested groups when the effects of pretesting are negligible (p. 25). Comparing the unpretested control posttest with the pretests estimates the combined effect of maturation and history (p. 25). Without randomization, a comparison of groups measured only after the treatment is a static-group comparison, whose groups cannot be certified equivalent (p. 12). Page numbers refer to the 1966 book. In solomonR: Tests A–D and the history check of
fit_solomon_classic();plot_solomon_design(); the caveat on the unpretested arms of nonrandomized designs inbaseline_solomon()andreport_solomon()(#58).
Edmonds, W. A., & Kennedy, T. D. (2017). An applied guide to research designs: Quantitative, qualitative, and mixed methods (2nd ed.). SAGE Publications. https://doi.org/10.4135/9781071802779
Chapter 6 (pp. 93–101) describes Solomon four-, six-, and eight-group designs with published examples, and treats nonrandomized designs and their threats, selection bias above all (pp. 7–8, 94). In solomonR:
baseline_solomon(), the nonrandomized wording and threats paragraph ofreport_solomon()(#58), and designs with several treatments infit_solomon_glm(control = )(#45).
El Karkri, M., Quesada, A., & Romero-Ariza, M. (2025a). The dual impact of pretest sensitisation and the cognitive acceleration through science education programme in the Solomon four-group design. Brain Sciences, 16(1), Article 64. https://doi.org/10.3390/brainsci16010064
A recent classroom Solomon study that reports full cell statistics and a sensitization effect, with one intact class per condition. In solomonR: the data set
elkarkri2025a, the known-result test and example ofsolomon_from_summary(), the pretest imbalance example ofbaseline_solomon()(#58), and an example of the class–condition confounding thatvalidate_solomon()detects (#46, #53).
El Karkri, M., Quesada, A., & Romero-Ariza, M. (2025b). Methodological aspects of the Solomon four-group design: Detecting pre-test sensitisation and analysing qualitative and quantitative variables in education research. Review of Education, 13(1), Article e70050. https://doi.org/10.1002/rev3.70050
A methodological guide to detecting pretest sensitization, including categorical outcomes. In solomonR:
fisher_solomon(), the historical categorical path for binary outcomes.
Entwisle, D. R. (1961). Interactive effects of pretesting. Educational and Psychological Measurement, 21(3), 607–620. https://doi.org/10.1177/001316446102100307
Two fourth-grade training experiments in which pretested and unpretested pupils were paired on sex and IQ. They found no overall pretest effect but a Pretesting x IQ x Sex interaction (pp. 612–613), and suggest that pretest effects depend on participants’ characteristics (p. 614) and may fade with time (p. 610). In solomonR: the history article.
French, D. P., Miles, L. M., Elbourne, D., Farmer, A., Gulliford, M., Locock, L., Sutton, S., McCambridge, J., & MERIT Collaborative Group. (2021a). Reducing bias in trials due to reactions to measurement: Experts produced recommendations informed by evidence. Journal of Clinical Epidemiology, 139, 130–139. https://doi.org/10.1016/j.jclinepi.2021.06.028
The MERIT recommendations for recognizing and reducing bias from measurement reactivity in trials, including when a Solomon design is warranted. In solomonR: the decision article “Should I Use a Solomon Design?” and the history article (#48).
French, D. P., Miles, L. M., Elbourne, D., Farmer, A., Gulliford, M., Locock, L., Sutton, S., McCambridge, J., & MERIT Collaborative Group. (2021b). Reducing bias in trials from reactions to measurement: The MERIT study including developmental work and expert workshop. Health Technology Assessment, 25(55), 1–72. https://doi.org/10.3310/hta25550
The full MERIT report, with the evidence behind the recommendations and a worked Solomon example with a binary outcome. In solomonR: the measurement items of
report_solomon(), and the reporting, decision, and history articles (#48, #52).
Huck, S. W., & Sandler, H. M. (1973). A note on the Solomon 4-group design: Appropriate statistical analyses. The Journal of Experimental Education, 42(2), 54–55. https://doi.org/10.1080/00220973.1973.11011460
Argues that the ANCOVA on the pretested groups remains valid when pretesting has a main effect, and that a significant interaction calls for a posttest t test on the unpretested groups, whose result generalizes to an unpretested population. In solomonR: Tests E and H of
fit_solomon_classic().
Jordaan, J. (2014). The development and evaluation of a life skills programme for young adult prisoners [Doctoral thesis, University of the Free State]. KovsieScholar. https://hdl.handle.net/11660/832
A randomized Solomon four-group design with three posttest occasions: after a six-month program for young adult male offenders, and 3 and 6 months later (pp. 86–87, 98). Each occasion was analyzed separately by the sequence of Walton Braver and Braver (1988). Of the nine Pretest x Treatment interactions on the Coping Strategy Indicator, two were significant, one after the program and one 6 months later (Tables 7.4–7.19, pp. 113–127). In solomonR: the data set
jordaan2014, whose published analyses are reproduced, and the article “Worked Example: Repeated Posttests” (#57).
Kvalem, I. L., Sundet, J. M., Rivø, K. I., Eilertsen, D. E., & Bakketeig, L. S. (1996). The effect of sex education on adolescents’ use of condoms: Applying the Solomon four-group design. Health Education Quarterly, 23(1), 34–47. https://doi.org/10.1177/109019819602300103
A cluster-randomized school trial in which the intervention appeared to work only in combination with the pretest. In solomonR: the data set
kvalem1996, the known-result tests offisher_solomon()andmarginal_solomon(), and the unbalanced allocation (30 intervention and 94 control classes) in the simulation study of cluster-levelperm_solomon()(#19).
Lana, R. E. (1959). Pretest-treatment interaction effects in attitudinal studies. Psychological Bulletin, 56(4), 293–300. https://doi.org/10.1037/h0044646
An experiment with the Solomon design on attitudes toward vivisection that found no pretest-by-treatment interaction (pp. 297–298), and a review of attitude studies that had not controlled for one (pp. 293–294). In solomonR: the data set
lana1959, whose published analysis of variance (Table 3, p. 297)solomon_from_summary()reproduces; the history article.
Lana, R. E. (2009). Pretest sensitization. In R. Rosenthal & R. L. Rosnow, Artifacts in behavioral research: Robert Rosenthal and Ralph L. Rosnow’s classic books (pp. 93–109). Oxford University Press. https://doi.org/10.1093/acprof:oso/9780195385540.003.0004 (Original work published 1969)
A review of the pretest-sensitization literature to 1969. The four-group design is analyzed as a two-by-two analysis of variance of the posttests (pp. 99–100). Pretests that involve learning or recall did sensitize, in either direction (pp. 101–103). An “overwhelming lack” of sensitization was found when the pretest measures existing attitudes or opinions (pp. 103–104), a conclusion at odds with Bracht and Glass (1968, p. 463). In solomonR: the history article and the coverage article.
Lessac, M. S., & Solomon, R. L. (1969). Effects of early isolation on the later adaptive behavior of beagles: A methodological demonstration. Developmental Psychology, 1(1), 14–25. https://doi.org/10.1037/h0026778
The full report of Lessac’s (1965) dissertation experiment, which Solomon and Lessac (1968, p. 149) had summarized as an application of the four-group design. One beagle from each of six litters was placed in each group; the two isolated groups spent a year in solid-walled cages from 12 weeks of age, and one isolated and one control group were pretested (pp. 15–16). Analyses of variance of the posttests tested isolation, pretesting, and their interaction, and Mann–Whitney U tests compared posttests with the pretests of the pretested groups, usually combined (pp. 18–22). Because the adult controls matched the puppies’ pretests on all but one test, the isolates’ deficits were read as a loss of existing abilities, not a failure to develop (p. 23). The pretest protected the isolates’ avoidance learning (pp. 21–23) but deepened their impairment when the runway’s start box was rotated (pp. 19, 24). Steyn and Mynhardt (2008, p. 567) cite it for the design. In solomonR: the history article (#96).
Mai, N. N., Takahashi, Y., & Oo, M. M. (2020). Testing the effectiveness of transfer interventions using Solomon four-group designs. Education Sciences, 10(4), Article 92. https://doi.org/10.3390/educsci10040092
A randomized six-group Solomon design (two interventions and a control) with published individual-level data. In solomonR: the
mai2020data (CC BY 4.0) and the worked-example article, which reproduce its Tables 4, 5, and 7; the history/maturation check offit_solomon_classic(), which follows its comparisons of O6 with O1 and O3; the six-group example offit_solomon_glm(control = )and of the article “Designs With Several Treatments” (#45).
McCambridge, J., Butor-Bhavsar, K., Witton, J., & Elbourne, D. (2011). Can research assessments themselves cause bias in behaviour change trials? A systematic review of evidence from Solomon 4-group studies. PLoS ONE, 6(10), Article e25223. https://doi.org/10.1371/journal.pone.0025223
A systematic review of Solomon studies with behavioral outcomes, finding too little high-quality evidence to settle whether assessment effects bias trials. In solomonR: the getting-started guide, the planning article, and the decision article.
McCarthy, A. M., & Tucker, M. L. (2002). Encouraging community service through service learning. Journal of Management Education, 26(6), 629–647. https://doi.org/10.1177/1052562902238322
A nonrandomized Solomon eight-group design crossing two interventions with pretesting. In solomonR: the planned comparisons of
fit_solomon_glm(control = , contrasts = )and the factorial example in “Designs With Several Treatments” (#45).
Morris, S. B. (2008). Estimating effect sizes from pretest-posttest-control group designs. Organizational Research Methods, 11(2), 364–386. https://doi.org/10.1177/1094428106291059
Compares effect sizes for pretest–posttest–control designs and recommends the pooled-pretest-SD estimator. In solomonR: dppc2 and its variance for the pretested pair in
solomon_effect_sizes(), tested against the paper’s worked example and theoretical variances (#53).
Sawilowsky, S. S. (1996, June 23). Controlling experiment-wise Type I error of meta-analysis in the Solomon four-group design [Paper presentation]. First International Conference on Multiple Comparisons, Tel Aviv, Israel. https://digitalcommons.wayne.edu/coe_tbf/29/
Extends the Type I error evidence on the historical test sequence to nonnormal data, reports the 1995 revision that removed Test D (p. 2), and proposes alpha allocations to control the error rate (Table 4). In solomonR:
fit_solomon_classic(flow = "1995"), thealpha_allocationoption, and the replication study (#50, #51).
Sawilowsky, S. S. (2000). Review of the rank transform in designed experiments. Perceptual and Motor Skills, 90(2), 489–497. https://doi.org/10.2466/pms.2000.90.2.489
Documents the failure of rank-transform tests for interactions. In solomonR: why the package has no rank-transform test for sensitization; its distribution-free option is randomization inference (
perm_solomon()).
Sawilowsky, S. S., Kelley, D. L., Blair, R. C., & Markman, B. S. (1994). Meta-analysis and the Solomon four-group design. The Journal of Experimental Education, 62(4), 361–376. https://doi.org/10.1080/00220973.1994.9944140
A Monte Carlo study showing inflated experiment-wise Type I error for the meta-analytic test sequence. In solomonR: the caution printed with Test I, the methods guide, and the replication of its Table 2 in “Historical Tests: Replicating the Published Error Rates” (#51).
Sawilowsky, S. S., & Markman, B. S. (1990a). Another look at the power of meta-analysis in the Solomon four-group design. Perceptual and Motor Skills, 71(1), 177–178. https://doi.org/10.2466/pms.1990.71.1.177
A counterexample in which Test I misses an effect that a constituent test detects. In solomonR: the history article and the historical-analysis vignette (#50).
Sawilowsky, S. S., & Markman, B. S. (1990b). Rejoinder to Braver and Walton Braver. Perceptual and Motor Skills, 71(2), 424–426. https://doi.org/10.2466/pms.1990.71.2.424
Continues the exchange and calls for systematic study of the procedure. In solomonR: the history article (#50).
Solomon, R. L. (1949). An extension of control group design. Psychological Bulletin, 46(2), 137–150. https://doi.org/10.1037/h0062958
Introduces the four-group design to separate pretest effects from treatment effects, first as a three-group design (pp. 141–145) and then, for field studies exposed to outside events, with a fourth group (pp. 146–148). In solomonR: the design itself; every estimand; the rationale of the history/maturation check; his own analysis,
fit_solomon_1949(), and his spelling data,solomon1949.
Solomon, R. L., & Lessac, M. S. (1968). A control group design for experimental studies of developmental processes. Psychological Bulletin, 70(3, Pt. 1), 145–150. https://doi.org/10.1037/h0026147
Applies the four-group design to isolation and enrichment studies of development. The combined pretest mean of the pretested groups serves as the best estimate for the unpretested groups, to judge whether they improved or deteriorated, and the design is treated as a two-by-two factorial whose interaction is a difference of differences (pp. 146–147). In solomonR: the inferred pretest of
fit_solomon_1949(); the history article.
Steyn, R. (2005). Self-evaluasie en die vorming van selfdoeltreffendheidspersepsies [Self-evaluation and the forming of self-efficacy perceptions] [Doctoral thesis, University of South Africa]. Unisa Institutional Repository. https://hdl.handle.net/10500/1745
The eight-group study that Steyn (2009) describes: three treatments and a control, each with and without a pretest, with 1,723 police trainees (pp. 103–106). Existing classes were allocated to the eight groups, not at random (pp. 105–106). It was analyzed as overlapping four-group designs by the sequence of Walton Braver and Braver (1988), then by a one-way analysis of variance of the eight posttests with Scheffé tests (pp. 151–153), the method of Scheffé (1953); because the test is particularly strict, the significance level was set at .05 (p. 151). The thesis is in Afrikaans; Steyn and Mynhardt (2008) report the study in English. In solomonR: the data set
steyn2005, whose published analyses of variance and Scheffé tests are reproduced from its group statistics; the default post hoc tests offit_solomon_steyn().
Steyn, R. (2009). Re-designing the Solomon four-group: Can we improve on this exemplary model? Design Principles and Practices: An International Journal—Annual Review, 3(1), 383–394. https://doi.org/10.18848/1833-1874/CGP/v03i01/37588
Extends the four-group design to k interventions, each given with and without a pretest, alongside a pretested and an unpretested control group: 2(k + 1) groups, six for two interventions and eight for three. It sets out a sequence of checks of internal validity (equivalence after randomization, history and maturation, the testing effect, the pretest-intervention interaction, test-retest reliability and instrumentation, regression to the mean, and attrition), followed by tests of the interventions’ effects. Edmonds and Kennedy (2017) cite it for analyses of four-, six-, and eight-group designs. solomonR follows a pre-publication draft of the article, dated March 2, 2009, which the author provided (R. Steyn, personal communication, September 30, 2026); the published article was not available for comparison. In solomonR: the design of
fit_solomon_glm(control = )andfit_solomon_steyn()(#45).
Steyn, R., & Mynhardt, J. (2008). Factors that influence the forming of self-evaluation and self-efficacy perceptions. South African Journal of Psychology, 38(3), 563–573. https://doi.org/10.1177/008124630803800310
The English report of the eight-group study in Steyn (2005): three treatments, each given with and without a pretest, a pretested and an unpretested control group, and 1,723 police trainees (pp. 566–567). It says that the participants were allocated to the groups at random (pp. 566–567) and does not mention the existing classes that the thesis describes. Each treatment was analyzed with the two control groups by a 2 x 2 analysis of variance of the posttests, through Step 2 of the sequence of Walton Braver and Braver (1988), at the .01 level (p. 568). No pretest-by-intervention interaction was significant, and the pretest main effect of the test-only treatment (p = .033) was judged not significant (pp. 568–569). It does not report the eight groups’ means or the one-way analysis of the eight posttests. In solomonR: the English source cited in
?steyn2005, whose group statistics reproduce its 2 x 2 analyses within rounding; for the Norms treatment its means and standard deviation follow Table 5.45 of the thesis (#96).
van Engelenburg, G. (1999). Statistical analysis for the Solomon four-group design (Research Report 99-06). University of Twente. ERIC. https://eric.ed.gov/?id=ED435692
Proposes full-information maximum likelihood for the Solomon design, treating the absent pretests as structurally missing. In solomonR:
fit_solomon_ml().
Walton Braver, M. C., & Braver, S. L. (1988). Statistical treatment of the Solomon four-group design: A meta-analytic approach. Psychological Bulletin, 104(1), 150–154. https://doi.org/10.1037/0033-2909.104.1.150
Proposes the Test A–I decision flow ending in a meta-analytic combination of the pretested and unpretested comparisons. In solomonR:
fit_solomon_classic(),plot_classic_flow(), andstouffer_solomon(); step 4 offit_solomon_steyn().
Willson, V. L., & Putnam, R. R. (1982). A meta-analysis of pretest sensitization effects in experimental design. American Educational Research Journal, 19(2), 249–258. https://doi.org/10.3102/00028312019002249
A meta-analysis of pretest effects and pretest-by-treatment interactions. In solomonR: the planning values in the article “Planning a Solomon Study” (#49).
Methods references
Bell, R. M., & McCaffrey, D. F. (2002). Bias reduction in standard errors for linear regression with multi-stage samples. Survey Methodology, 28(2), 169–181.
Bennett, S., Parpia, T., Hayes, R., & Cousens, S. (2002). Methods for the analysis of incidence rates in cluster randomized trials. International Journal of Epidemiology, 31(4), 839–846. https://doi.org/10.1093/ije/31.4.839
Brown, M. B., & Forsythe, A. B. (1974). Robust tests for the equality of variances. Journal of the American Statistical Association, 69(346), 364–367. https://doi.org/10.1080/01621459.1974.10482955
Byrne, B. M., Shavelson, R. J., & Muthén, B. (1989). Testing for the equivalence of factor covariance and mean structures: The issue of partial measurement invariance. Psychological Bulletin, 105(3), 456–466. https://doi.org/10.1037/0033-2909.105.3.456
Cameron, A. C., & Trivedi, P. K. (2013). Regression analysis of count data (2nd ed.). Cambridge University Press. https://doi.org/10.1017/CBO9781139013567
Carpenter, J. R., Bartlett, J. W., Morris, T. P., Wood, A. M., Quartagno, M., & Kenward, M. G. (2023). Multiple imputation and its application (2nd ed.). Wiley. https://doi.org/10.1002/9781119756118
Chan, A.-W., Tetzlaff, J. M., Altman, D. G., Laupacis, A., Gøtzsche, P. C., Krleža-Jerić, K., Hróbjartsson, A., Mann, H., Dickersin, K., Berlin, J. A., Doré, C. J., Parulekar, W. R., Summerskill, W. S. M., Groves, T., Schulz, K. F., Sox, H. C., Rockhold, F. W., Rennie, D., & Moher, D. (2013). SPIRIT 2013 statement: Defining standard protocol items for clinical trials. Annals of Internal Medicine, 158(3), 200–207. https://doi.org/10.7326/0003-4819-158-3-201302050-00583
Chen, F. F. (2007). Sensitivity of goodness of fit indexes to lack of measurement invariance. Structural Equation Modeling: A Multidisciplinary Journal, 14(3), 464–504. https://doi.org/10.1080/10705510701301834
Cheung, G. W., & Rensvold, R. B. (2002). Evaluating goodness-of-fit indexes for testing measurement invariance. Structural Equation Modeling: A Multidisciplinary Journal, 9(2), 233–255. https://doi.org/10.1207/S15328007SEM0902_5
Cro, S., Carpenter, J. R., & Kenward, M. G. (2019). Information-anchored sensitivity analysis: Theory and application. Journal of the Royal Statistical Society Series A: Statistics in Society, 182(2), 623–645. https://doi.org/10.1111/rssa.12423
Cumming, G., & Finch, S. (2001). A primer on the understanding, use, and calculation of confidence intervals that are based on central and noncentral distributions. Educational and Psychological Measurement, 61(4), 532–574. https://doi.org/10.1177/00131640121971374
Daniel, R., Zhang, J., & Farewell, D. (2021). Making apples from oranges: Comparing noncollapsible effect estimators and their standard errors after adjustment for different covariate sets. Biometrical Journal, 63(3), 528–557. https://doi.org/10.1002/bimj.201900297
DiCiccio, C. J., & Romano, J. P. (2017). Robust permutation tests for correlation and regression coefficients. Journal of the American Statistical Association, 112(519), 1211–1220. https://doi.org/10.1080/01621459.2016.1202117
Fitzmaurice, G. M., Laird, N. M., & Ware, J. H. (2011). Applied longitudinal analysis (2nd ed.). Wiley. https://doi.org/10.1002/9781119513469
Gail, M. H., Mark, S. D., Carroll, R. J., Green, S. B., & Pee, D. (1996). On design considerations and randomization-based inference for community intervention trials. Statistics in Medicine, 15(11), 1069–1092. https://doi.org/10.1002/(SICI)1097-0258(19960615)15:11%3C1069::AID-SIM220%3E3.0.CO;2-Q
Gelman, A., & Stern, H. (2006). The difference between “significant” and “not significant” is not itself statistically significant. The American Statistician, 60(4), 328–331. https://doi.org/10.1198/000313006X152649
Graham, J. W., Taylor, B. J., Olchowski, A. E., & Cumsille, P. E. (2006). Planned missing data designs in psychological research. Psychological Methods, 11(4), 323–343. https://doi.org/10.1037/1082-989X.11.4.323
Groenwold, R. H. H., White, I. R., Donders, A. R. T., Carpenter, J. R., Altman, D. G., & Moons, K. G. M. (2012). Missing covariate data in clinical research: When and when not to use the missing-indicator method for analysis. Canadian Medical Association Journal, 184(11), 1265–1269. https://doi.org/10.1503/cmaj.110977
Hayes, A. F., & Cai, L. (2007). Using heteroskedasticity-consistent standard error estimators in OLS regression: An introduction and software implementation. Behavior Research Methods, 39(4), 709–722. https://doi.org/10.3758/BF03192961
Hayes, R. J., & Moulton, L. H. (2017). Cluster randomised trials (2nd ed.). Chapman and Hall/CRC. https://doi.org/10.4324/9781315370286
Hedges, L. V. (1981). Distribution theory for Glass’s estimator of effect size and related estimators. Journal of Educational Statistics, 6(2), 107–128. https://doi.org/10.3102/10769986006002107
Henry, L., & Wickham, H. (2026). lifecycle: Manage the life cycle of your package functions (Version 1.0.5) [R package]. https://doi.org/10.32614/CRAN.package.lifecycle
Holm, S. (1979). A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics, 6(2), 65–70. https://www.jstor.org/stable/4615733
A sequentially rejective form of the Bonferroni test that controls the familywise error rate “for any combination of true hypotheses” (p. 65); it also describes the classical Bonferroni test. In solomonR: the default adjustment for families of comparisons in designs with several treatments,
fit_solomon_glm(adjust = ), and the post hoc tests offit_solomon_steyn()(#45).
Imbens, G. W., & Kolesár, M. (2016). Robust standard errors in small samples: Some practical advice. The Review of Economics and Statistics, 98(4), 701–712. https://doi.org/10.1162/REST_a_00552
Kelley, K. (2007). Confidence intervals for standardized effect sizes: Theory, application, and implementation. Journal of Statistical Software, 20(8), 1–24. https://doi.org/10.18637/jss.v020.i08
Laird, N. M., & Ware, J. H. (1982). Random-effects models for longitudinal data. Biometrics, 38(4), 963–974. https://doi.org/10.2307/2529876
Lakens, D. (2017). Equivalence tests: A practical primer for t tests, correlations, and meta-analyses. Social Psychological and Personality Science, 8(4), 355–362. https://doi.org/10.1177/1948550617697177
Lakens, D., Scheel, A. M., & Isager, P. M. (2018). Equivalence testing for psychological research: A tutorial. Advances in Methods and Practices in Psychological Science, 1(2), 259–269. https://doi.org/10.1177/2515245918770963
Liang, K.-Y., & Zeger, S. L. (1986). Longitudinal data analysis using generalized linear models. Biometrika, 73(1), 13–22. https://doi.org/10.1093/biomet/73.1.13
Lin, W. (2013). Agnostic notes on regression adjustments to experimental data: Reexamining Freedman’s critique. The Annals of Applied Statistics, 7(1), 295–318. https://doi.org/10.1214/12-AOAS583
Little, R. J., D’Agostino, R., Cohen, M. L., Dickersin, K., Emerson, S. S., Farrar, J. T., Frangakis, C., Hogan, J. W., Molenberghs, G., Murphy, S. A., Neaton, J. D., Rotnitzky, A., Scharfstein, D., Shih, W. J., Siegel, J. P., & Stern, H. (2012). The prevention and treatment of missing data in clinical trials. The New England Journal of Medicine, 367(14), 1355–1360. https://doi.org/10.1056/NEJMsr1203730
Little, R. J. A., & Rubin, D. B. (2019). Statistical analysis with missing data (3rd ed.). Wiley. https://doi.org/10.1002/9781119482260
Localio, A. R., Margolis, D. J., & Berlin, J. A. (2007). Relative risks and confidence intervals were easily computed indirectly from multivariable logistic regression. Journal of Clinical Epidemiology, 60(9), 874–882. https://doi.org/10.1016/j.jclinepi.2006.12.001
Long, J. S., & Ervin, L. H. (2000). Using heteroscedasticity consistent standard errors in the linear regression model. The American Statistician, 54(3), 217–224. https://doi.org/10.1080/00031305.2000.10474549
Lundberg, I., Johnson, R., & Stewart, B. M. (2021). What is your estimand? Defining the target quantity connects statistical evidence to theory. American Sociological Review, 86(3), 532–565. https://doi.org/10.1177/00031224211004187
MacKinnon, J. G., & White, H. (1985). Some heteroskedasticity-consistent covariance matrix estimators with improved finite sample properties. Journal of Econometrics, 29(3), 305–325. https://doi.org/10.1016/0304-4076(85)90158-7
Mallinckrodt, C. H., Lane, P. W., Schnell, D., Peng, Y., & Mancuso, J. P. (2008). Recommendations for the primary analysis of continuous endpoints in longitudinal clinical trials. Drug Information Journal, 42(4), 303–319. https://doi.org/10.1177/009286150804200402
Meredith, W. (1993). Measurement invariance, factor analysis and factorial invariance. Psychometrika, 58(4), 525–543. https://doi.org/10.1007/BF02294825
Morris, T. P., White, I. R., & Crowther, M. J. (2019). Using simulation studies to evaluate statistical methods. Statistics in Medicine, 38(11), 2074–2102. https://doi.org/10.1002/sim.8086
Murphy, K. R., & Myors, B. (1999). Testing the hypothesis that treatments have negligible effects: Minimum-effect tests in the general linear model. Journal of Applied Psychology, 84(2), 234–248. https://doi.org/10.1037/0021-9010.84.2.234
Nosek, B. A., Ebersole, C. R., DeHaven, A. C., & Mellor, D. T. (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600–2606. https://doi.org/10.1073/pnas.1708274114
Phipson, B., & Smyth, G. K. (2010). Permutation p-values should never be zero: Calculating exact p-values when permutations are randomly drawn. Statistical Applications in Genetics and Molecular Biology, 9(1), Article 39. https://doi.org/10.2202/1544-6115.1585
Pustejovsky, J. E., & Tipton, E. (2018). Small-sample methods for cluster-robust variance estimation and hypothesis testing in fixed effects models. Journal of Business & Economic Statistics, 36(4), 672–683. https://doi.org/10.1080/07350015.2016.1247004
Rajh-Weber, H., Huber, S. E., & Arendasy, M. (2025). A practice-oriented guide to statistical inference in linear modeling for non-normal or heteroskedastic error distributions. Behavior Research Methods, 57(12), Article 338. https://doi.org/10.3758/s13428-025-02801-4
Rosseel, Y. (2012). lavaan: An R package for structural equation modeling. Journal of Statistical Software, 48(2), 1–36. https://doi.org/10.18637/jss.v048.i02
Rubin, D. B. (1976). Inference and missing data. Biometrika, 63(3), 581–592. https://doi.org/10.1093/biomet/63.3.581
Sabanes Bove, D., Li, L., Dedic, J., Kelkhoff, D., Kunzmann, K., Lang, B. M., Stock, C., Wang, Y., James, D., Sidi, J., Leibovitz, D., Sjoberg, D. D., Krieger, N. I., Panagos, A., & Jones, J. (2026). mmrm: Mixed models for repeated measures (Version 0.3.18) [R package]. https://doi.org/10.32614/CRAN.package.mmrm
Satorra, A., & Bentler, P. M. (2001). A scaled difference chi-square test statistic for moment structure analysis. Psychometrika, 66(4), 507–514. https://doi.org/10.1007/BF02296192
Satterthwaite, F. E. (1946). An approximate distribution of estimates of variance components. Biometrics Bulletin, 2(6), 110–114. https://doi.org/10.2307/3002019
Scheffé, H. (1953). A method for judging all contrasts in the analysis of variance. Biometrika, 40(1–2), 87–104. https://doi.org/10.1093/biomet/40.1-2.87
Simultaneous intervals for every contrast among k means: the estimate plus or minus S standard errors, where S² is k − 1 times the upper α point of F on k − 1 and ν degrees of freedom. All the intervals hold together with probability 1 − α, including contrasts suggested by the data (pp. 87–89). A contrast is significant when its estimate exceeds S standard errors in absolute value, and the F test rejects exactly when some contrast is significant (pp. 87, 95–96). The means may rest on unequal numbers of observations (pp. 87, 96–98). When only the differences between pairs of equally precise means are of interest, Scheffé recommends Tukey’s method, whose intervals are shorter (pp. 89, 92–93, 96–97). Crossref gives the pages as 87–110; the printed article ends on p. 104. In solomonR: the default post hoc tests of
fit_solomon_steyn(), the Scheffé tests that Steyn (2005) used (#96).
Schuirmann, D. J. (1987). A comparison of the two one-sided tests procedure and the power approach for assessing the equivalence of average bioavailability. Journal of Pharmacokinetics and Biopharmaceutics, 15(6), 657–680. https://doi.org/10.1007/BF01068419
Stadler, K., & Kotze, M. E. (2006). The influence of a ropes course development programme on the self-concept and self-efficacy of young career officers. SA Journal of Industrial Psychology, 32(1), 25–32. https://doi.org/10.4102/sajip.v32i1.225
Not a study of the Solomon design: two pretested groups from two institutions, each selected at random within its institution, were measured before a three-day ropes course, on its last day, and 8 weeks later, in a design the authors call quasi-experimental (pp. 25, 27). It is the only published account found of Steyn’s (2001) doctoral study of an outdoor experiential learning program with members of the Public Order Police, which used “a four-group design with a pre-, post- and post-post test” (p. 27). In solomonR: the coverage article, which records Steyn (2001) as not read (#96).
Steiger, J. H. (2004). Beyond the F test: Effect size confidence intervals and tests of close fit in the analysis of variance and contrast analysis. Psychological Methods, 9(2), 164–182. https://doi.org/10.1037/1082-989X.9.2.164
Stouffer, S. A., Suchman, E. A., DeVinney, L. C., Star, S. A., & Williams, R. M., Jr. (1949). The American soldier: Adjustment during army life (Vol. 1). Princeton University Press.
Tipton, E. (2015). Small sample adjustments for robust variance estimation with meta-regression. Psychological Methods, 20(3), 375–393. https://doi.org/10.1037/met0000011
Van Breukelen, G. J. P. (2006). ANCOVA versus change from baseline had more power in randomized studies and more bias in nonrandomized studies. Journal of Clinical Epidemiology, 59(9), 920–925. https://doi.org/10.1016/j.jclinepi.2006.02.007
van Buuren, S. (2018). Flexible imputation of missing data (2nd ed.). CRC Press. https://doi.org/10.1201/9780429492259
Vandenberg, R. J., & Lance, C. E. (2000). A review and synthesis of the measurement invariance literature: Suggestions, practices, and recommendations for organizational research. Organizational Research Methods, 3(1), 4–70. https://doi.org/10.1177/109442810031002
van ’t Veer, A. E., & Giner-Sorolla, R. (2016). Pre-registration in social psychology—A discussion and suggested template. Journal of Experimental Social Psychology, 67, 2–12. https://doi.org/10.1016/j.jesp.2016.03.004
Venables, W. N., & Ripley, B. D. (2002). Modern applied statistics with S (4th ed.). Springer. https://doi.org/10.1007/978-0-387-21706-2
Welch, B. L. (1947). The generalization of “Student’s” problem when several different population variances are involved. Biometrika, 34(1–2), 28–35. https://doi.org/10.1093/biomet/34.1-2.28
White, I. R., Horton, N. J., Carpenter, J., & Pocock, S. J. (2011). Strategy for intention to treat analysis in randomised trials with missing outcome data. BMJ, 342, Article d40. https://doi.org/10.1136/bmj.d40
White, I. R., & Thompson, S. G. (2005). Adjusting for partially missing baseline measurements in randomized trials. Statistics in Medicine, 24(7), 993–1007. https://doi.org/10.1002/sim.1981
Wu, J., & Ding, P. (2021). Randomization tests for weak null hypotheses in randomized experiments. Journal of the American Statistical Association, 116(536), 1898–1913. https://doi.org/10.1080/01621459.2020.1750415
Zimmerman, D. W. (2004). A note on preliminary tests of equality of variances. British Journal of Mathematical and Statistical Psychology, 57(1), 173–181. https://doi.org/10.1348/000711004849222