Simulate a Solomon four-group study with known effects
Source:R/solomon_simulate.R
simulate_solomon.Rd
Generates individual data for a randomized Solomon four-group study with
chosen treatment, pretest, and sensitization effects, and attaches the
true value of every Solomon estimand. It is a solomonR teaching tool:
instructors can show that each analysis recovers the effects that were
built in, and researchers can check an analysis plan on data whose answer
is known. With several treatment effects it simulates a Solomon N-group
design (see "Designs with several treatments").
Usage
simulate_solomon(
n = 30,
delta = 0,
sens = 0,
pretest_effect = 0,
rho = 0.5,
sigma = 1,
mean = 0,
digits = NULL,
limits = NULL,
seed = NULL
)Arguments
- n
Participants per group: one number for all groups, or one number per group. For the four-group design, four numbers in the order pretested treatment, pretested control, unpretested treatment, unpretested control. For k treatments, 2(k + 1) numbers in the order given in "Designs with several treatments". Default 30 per group.
- delta
Treatment effect among unpretested participants. For a design with several treatments, one effect per treatment, named by treatment. Default 0.
- sens
Pretest x Treatment interaction (sensitization): the extra treatment effect among pretested participants. For a design with several treatments, one number for every treatment or one per treatment. Default 0.
- pretest_effect
Effect of taking the pretest on the posttest, among controls. Default 0.
- rho
Pretest-posttest correlation within the pretested groups. Default 0.5.
- sigma
Standard deviation of the scores within each group. Default 1.
- mean
Mean pretest and control-group posttest score. Default 0.
- digits
Optional number of decimal places to round scores to.
- limits
Optional lower and upper limits for scores, applied after rounding.
- seed
Optional integer seed. The global random number state is left unchanged.
Value
A data frame with y_post, treat, pretested, and y_pre
(missing by design in the unpretested groups), in the form the analysis
functions take. treat is 0/1 for the four-group design and a factor of
conditions for a design with several treatments. Its "truth" attribute
is a data frame of the true estimands, and its "settings" attribute
holds the arguments, with the group sizes named by group.
Details
Data-generating model. It is the model of power_solomon(), whose
rebuild was validated under a protocol posted on issue #18, with a
pretesting main effect and a location added. Every participant has a
latent baseline A from a standard normal distribution. Participants in the
pretested groups have a pretest score of mean + sigma * A. Every
participant's posttest score is
mean + delta * treat + pretest_effect * pretested + sens * treat * pretested + sigma * (rho * A + sqrt(1 - rho^2) * e),
where e is independent standard normal error. So rho is the
pretest-posttest correlation within each pretested group, and sigma is
the standard deviation of both scores within each group. The four groups
are generated in the package's order (pretested treatment, pretested
control, unpretested treatment, unpretested control), and the baselines
are drawn before the errors.
Relation to the package's other simulations. With mean = 0,
sigma = 1, and pretest_effect = 0, the data are the first data set
that power_solomon() simulates with the same seed. With n = 30, delta = 5,
pretest_effect = 2, rho = 0.6, sigma = 10, mean = 50,
seed = 20260915, digits = 0, and limits = c(0, 100), they are the
bundled solomon_example data.
True values. The "truth" attribute gives each estimand:
the treatment effect among unpretested participants,
delta;among pretested participants,
delta + sens;the Pretest x Treatment interaction,
sens;the equal-weighted average treatment effect,
delta + sens / 2;the pretest effect among controls,
pretest_effect.
Rounding and limits (digits, limits) make the scores look like test
scores, but they shift the true values slightly, and more so when many
scores reach a limit.
Choosing values. Willson and Putnam (1982) found an average pretest effect of about a fifth of a standard deviation in randomized studies; the article "Planning a Solomon Study" discusses planning values.
Designs with several treatments
A Solomon N-group design crosses k treatments and a control with
pretesting, giving 2(k + 1) groups: six for two treatments and eight for
three (Edmonds & Kennedy, 2017; Steyn, 2009). Give delta one effect per
treatment, named by treatment, for example delta = c(A = 0.5, B = 0.2).
Without names the treatments are called "T1", "T2", and so on; the
control is "Control". sens is one number for every treatment or one
per treatment, in the order of delta; a named sens is matched to the
treatments by name.
The model is the one above, with the effects delta[j] and sens[j] of
treatment j: each treatment differs from the control only by its own
effects. n is one size for every group, or 2(k + 1) sizes in this order:
the pretested treatments (in the order of delta), the pretested control,
the unpretested treatments, and the unpretested control. Sizes named by
group, as in the "settings" attribute (such as "Pretested, A"), or
named n1, n2, and so on for the groups in that order, may be given in
any order. Other names are refused. The groups are generated in the order
above, and the baselines are drawn before the errors.
The treat column is then a factor whose levels are the treatments and
"Control", so the data go directly to
fit_solomon_glm(y_post, treat, pretested, y_pre, control = "Control").
The "truth" attribute has one row for each treatment-control comparison
and contrast, with the columns comparison (such as "A vs Control"),
contrast, and true_value, in the order of the effects table of that
fit; a last row gives the pretest effect among controls. With one
treatment effect the data and the "truth" attribute are those described
above.
References
Edmonds, W. A., & Kennedy, T. D. (2017). An applied guide to research designs: Quantitative, qualitative, and mixed methods (2nd ed.). SAGE Publications. https://doi.org/10.4135/9781071802779
Steyn, R. (2009). Re-designing the Solomon four-group: Can we improve on this exemplary model? Design Principles and Practices: An International Journal—Annual Review, 3(1), 383–394. https://doi.org/10.18848/1833-1874/CGP/v03i01/37588
Willson, V. L., & Putnam, R. R. (1982). A meta-analysis of pretest sensitization effects in experimental design. American Educational Research Journal, 19(2), 249–258. https://doi.org/10.3102/00028312019002249
See also
power_solomon() and plan_solomon() for planning,
fit_solomon_glm() for the recommended analysis, and the article
"Teaching with solomonR".
Examples
d <- simulate_solomon(n = 40, delta = 0.5, sens = 0.4, rho = 0.6, seed = 1)
attr(d, "truth")
#> estimand true_value
#> 1 ATE (avg over pretest) 0.7
#> 2 Pretest x Treatment 0.4
#> 3 Treatment | pretested 0.9
#> 4 Treatment | unpretested 0.5
#> 5 Pretest effect | control 0.0
with(d, fit_solomon_glm(y_post, treat, pretested, y_pre))
#> Solomon GLM (unified model)
#> Formula: y ~ treat * pretested + pre_obs
#> Covariance: HC3 heteroskedasticity-consistent; t tests (df = 155)
#>
#> Term Est (SE) t df p 95% CI
#> (Intercept) -0.045 (0.136) -0.33 155 0.739 [-0.314, 0.223]
#> treat 0.656 (0.204) 3.22 155 0.002 [0.254, 1.058]
#> pretested -0.011 (0.200) -0.05 155 0.958 [-0.406, 0.384]
#> pre_obs 0.756 (0.125) 6.05 155 <.001 [0.509, 1.002]
#> treat:pretested 0.413 (0.280) 1.48 155 0.142 [-0.140, 0.966]
#>
#> Key contrasts Est (SE) t df p 95% CI Wald R2
#> ATE (avg over pretest) 0.862 (0.140) 6.16 155 <.001 [0.586, 1.139] 0.197
#> Pretest x Treatment 0.413 (0.280) 1.48 155 0.142 [-0.140, 0.966] 0.014
#> Treatment | pretested 1.069 (0.192) 5.56 155 <.001 [0.689, 1.449] 0.166
#> Treatment | unpretested 0.656 (0.204) 3.22 155 0.002 [0.254, 1.058] 0.063
#>
#> Wald R2: partial R-squared for conventional Gaussian OLS;
#> a Wald-based descriptive approximation when robust covariance is used.
# The bundled teaching data.
ex <- simulate_solomon(n = 30, delta = 5, pretest_effect = 2, rho = 0.6,
sigma = 10, mean = 50, digits = 0, limits = c(0, 100),
seed = 20260915)
attr(ex, "truth") <- attr(ex, "settings") <- NULL
identical(ex, solomon_example)
#> [1] TRUE
# A six-group design: two treatments and a control.
d6 <- simulate_solomon(n = 50, delta = c(A = 0.5, B = 0.2), sens = c(0.3, 0),
rho = 0.6, seed = 1)
attr(d6, "truth")
#> comparison contrast true_value
#> 1 A vs Control ATE (avg over pretest) 0.65
#> 2 B vs Control ATE (avg over pretest) 0.20
#> 3 A vs Control Pretest x Treatment 0.30
#> 4 B vs Control Pretest x Treatment 0.00
#> 5 A vs Control Treatment | pretested 0.80
#> 6 B vs Control Treatment | pretested 0.20
#> 7 A vs Control Treatment | unpretested 0.50
#> 8 B vs Control Treatment | unpretested 0.20
#> 9 Control Pretest effect | control 0.00
fit_solomon_glm(y_post, treat, pretested, y_pre, control = "Control", data = d6)
#> Solomon GLM (unified model), N-group design
#> Conditions: A, B; control: Control. With and without a pretest: 6 groups.
#> Formula: y ~ (treat_A + treat_B) * pretested + pre_obs
#> Covariance: HC3 heteroskedasticity-consistent; t tests (df = 293)
#>
#> Omnibus tests
#> Test Statistic p
#> Condition (avg over pretest) F(2, 293) = 15.00 <.001
#> Pretest x Condition F(2, 293) = 1.15 0.318
#> Condition | pretested F(2, 293) = 17.50 <.001
#> Condition | unpretested F(2, 293) = 3.07 0.048
#>
#> Contrasts
#> Comparison Contrast Est (SE) t df p p adj. 95% CI
#> A vs Control ATE (avg over pretest) 0.767 (0.144) 5.32 293 <.001 <.001 [0.483, 1.052]
#> B vs Control ATE (avg over pretest) 0.197 (0.130) 1.52 293 0.130 0.130 [-0.059, 0.453]
#> A vs Control Pretest x Treatment 0.438 (0.289) 1.52 293 0.130 0.261 [-0.130, 1.006]
#> B vs Control Pretest x Treatment 0.194 (0.260) 0.74 293 0.457 0.457 [-0.318, 0.705]
#> A vs Control Treatment | pretested 0.986 (0.174) 5.66 293 <.001 <.001 [0.644, 1.329]
#> B vs Control Treatment | pretested 0.294 (0.170) 1.73 293 0.084 0.084 [-0.040, 0.628]
#> A vs Control Treatment | unpretested 0.549 (0.230) 2.38 293 0.018 0.036 [0.096, 1.002]
#> B vs Control Treatment | unpretested 0.100 (0.197) 0.51 293 0.611 0.611 [-0.287, 0.488]
#>
#> p adj.: adjusted by Holm's (1979) procedure within each contrast, across the 2 comparisons.
#> Confidence intervals are not adjusted.
#>
#> Experimental: in the package's simulation study (issue #45), the omnibus
#> tests of Condition | pretested and Condition | unpretested rejected in up
#> to 6.9% of replications at the .05 level with three treatments and 10
#> participants per group. No other test, and no family of adjusted
#> comparisons, failed the study's rule. See ?fit_solomon_glm.