Worked Example: Repeated Posttests
Source:vignettes/articles/repeated-posttests-example.Rmd
repeated-posttests-example.RmdWhen a Solomon study measures the outcome on several occasions, a new question arises: does pretest sensitization persist, grow, or fade over time? Jordaan (2014) evaluated a six-month life skills program for young adult male offenders with a Solomon four-group design, and tested every group three times: at the end of the program, and 3 and 6 months later.
The thesis reports the size, mean, and standard deviation of every
group on every occasion, and solomonR bundles them as
jordaan2014 (see ?jordaan2014 for the design
and the measures). This example:
- reproduces the published analyses from those statistics;
- fits one model for all occasions with
fit_solomon_mmrm(); - asks what the published statistics can say about change in sensitization, and what they cannot.
It reanalyzes published group statistics to teach the package. For the program, the measures, and the author’s conclusions, read the thesis (Jordaan, 2014).
The design and the data
- Assignment. 120 offenders aged 21 to 25 with long sentences were assigned at random to the program or to the control condition, and half of each condition was assigned at random to be pretested (Jordaan, 2014, pp. 86–87).
- Occasions. All four groups were tested after the program and again 3 and 6 months later (p. 98).
- Measures. The three subscales of the Coping Strategy Indicator, each 11 items scored from 1 to 3. Higher scores on problem solving and on seeking social support, and lower scores on avoidance, indicate better coping (p. 92).
- Attrition. Transfers removed 17 offenders from the program groups and 7 from the control groups (p. 109). The 96 who remained completed every posttest.
The problem-solving scores, as means (standard deviations):
ps <- subset(jordaan2014, subscale == "Problem solving")
levels(ps$occasion) <- c("Pretest", "Posttest", "3 months", "6 months")
ps$label <- factor(ps$group, labels = c("Program, pretested (n = 22)",
"Program, not pretested (n = 21)",
"Control, pretested (n = 22)",
"Control, not pretested (n = 31)"))
tab <- with(ps, tapply(sprintf("%.2f (%.2f)", mean, sd), list(label, occasion), c))
tab[is.na(tab)] <- ""
knitr::kable(tab)| Pretest | Posttest | 3 months | 6 months | |
|---|---|---|---|---|
| Program, pretested (n = 22) | 27.00 (3.67) | 28.68 (4.10) | 29.00 (3.69) | 30.23 (3.49) |
| Program, not pretested (n = 21) | 29.52 (3.44) | 27.95 (4.52) | 28.62 (3.47) | |
| Control, pretested (n = 22) | 24.68 (3.67) | 25.82 (6.39) | 28.00 (5.55) | 28.64 (4.17) |
| Control, not pretested (n = 31) | 28.55 (5.47) | 28.39 (4.57) | 30.35 (2.65) |
Randomization makes the groups comparable as assigned. More offenders were lost from the program groups than from the control groups, so the 96 who remained are comparable only if the transfers were unrelated to the outcome. No analysis below can check or correct this: the offenders who were transferred have no posttest at all.
Reproduce the published analyses
The thesis analyzed each occasion separately by the decision sequence
of Walton Braver and Braver (1988). On each occasion it began with a 2 x
2 analysis of variance of the posttests, Pretest x Treatment.
solomon_from_summary() computes that analysis from group
statistics. For problem solving 6 months after the program:
six <- subset(ps, occasion == "6 months")
with(six, solomon_from_summary(n, mean, sd, treat = treat, pretested = pretested))
#> Solomon analysis from summary statistics
#> ----------------------------------------
#> Pooled error variance: 11.657 on 92 df (equal variances assumed)
#>
#> Two-way ANOVA on the posttest (Type III sums of squares)
#> Treatment SS = 0.115 df = 1 F = 0.01 p = 0.921
#> Pretest SS = 0.059 df = 1 F = 0.01 p = 0.944
#> Treatment x Pretest SS = 64.539 df = 1 F = 5.54 p = 0.021
#> Error SS = 1072.442 df = 92
#>
#> Contrasts with 95% confidence intervals
#> Test A: Pretest x Treatment 3.320 [0.518, 6.122], t(92) = 2.35, p = 0.021
#> Test B: Treatment | pretested 1.590 [-0.455, 3.635], t(92) = 1.54, p = 0.126
#> Test C: Treatment | unpretested -1.730 [-3.646, 0.186], t(92) = -1.79, p = 0.076
#> Test D: ATE (avg over pretest) -0.070 [-1.471, 1.331], t(92) = -0.10, p = 0.921
#> Pretest main effect -0.050 [-1.451, 1.351], t(92) = -0.07, p = 0.944The interaction test, F(1, 92) = 5.54, p = .021, is the published result: F = 5.556, p = .021 (Table 7.13, p. 121). The small difference comes from the published means and standard deviations, which are rounded to two decimals.
The thesis then compared the program with the control among pretested and among unpretested offenders (Table 7.14, p. 122). Those are Tests B and C above, but the thesis used two-group t tests, each with only its own two groups’ variance. Among unpretested offenders it found t(50) = −2.04, p = .046; with the variance pooled over all four groups, Test C gives t(92) = −1.79, p = .076.
The interaction tests of all three subscales on all three occasions, against the published F values (Tables 7.4–7.19, pp. 113–127):
published <- data.frame(
subscale = rep(levels(jordaan2014$subscale), each = 3),
occasion = rep(c("Posttest", "Follow-up 1", "Follow-up 2"), 3),
published_F = c(9.678, 0.266, 2.306, 0.819, 0.563, 5.556, 0.373, 0.688, 0.327)
)
interaction <- function(s, o) {
x <- subset(jordaan2014, subscale == s & occasion == o)
a <- with(x, solomon_from_summary(n, mean, sd, treat = treat, pretested = pretested))$anova
a[a$source == "Treatment x Pretest", c("F", "p.value")]
}
nine <- cbind(published, do.call(rbind, Map(interaction, published$subscale, published$occasion)))
nine$p_holm <- p.adjust(nine$p.value, method = "holm")
rownames(nine) <- NULL
knitr::kable(nine, digits = 3)| subscale | occasion | published_F | F | p.value | p_holm |
|---|---|---|---|---|---|
| Social support | Posttest | 9.678 | 9.648 | 0.003 | 0.023 |
| Social support | Follow-up 1 | 0.266 | 0.270 | 0.605 | 1.000 |
| Social support | Follow-up 2 | 2.306 | 2.304 | 0.132 | 0.927 |
| Problem solving | Posttest | 0.819 | 0.821 | 0.367 | 1.000 |
| Problem solving | Follow-up 1 | 0.563 | 0.568 | 0.453 | 1.000 |
| Problem solving | Follow-up 2 | 5.556 | 5.537 | 0.021 | 0.166 |
| Avoidance | Posttest | 0.373 | 0.367 | 0.546 | 1.000 |
| Avoidance | Follow-up 1 | 0.688 | 0.692 | 0.408 | 1.000 |
| Avoidance | Follow-up 2 | 0.327 | 0.321 | 0.573 | 1.000 |
Every published value is reproduced within rounding. In the thesis,
“Follow-up 1” and “Follow-up 2” are the tests 3 and 6 months after the
program. Two of the nine interactions are significant at .05: social
support after the program, and problem solving 6 months later. The
thesis reads both as signs of pretest sensitization (pp. 113, 122). It
ran the nine tests without adjustment. Adjusted together by Holm’s
(1979) procedure (p_holm), only the social support
interaction after the program stays below .05.
One model for all occasions
fit_solomon_mmrm() analyzes all occasions in one model:
a mixed model for repeated measures, specified as Mallinckrodt et
al. (2008) recommend for longitudinal trials. It needs each
participant’s scores, which the thesis does not report. The function
rebuild() below creates data whose group means and standard
deviations on every occasion are exactly the published ones
(MASS::mvrnorm() with empirical = TRUE). The
correlations between occasions are not reported either, so one must be
assumed; here it is 0.5 for every pair of occasions.
rebuild <- function(stats, r) {
post <- subset(stats, occasion != "Pretest")
occasions <- levels(droplevels(post$occasion))
do.call(rbind, lapply(split(post, post$group), function(g) {
g <- g[order(g$occasion), ]
R <- matrix(r, nrow(g), nrow(g))
diag(R) <- 1
y <- MASS::mvrnorm(g$n[1], mu = g$mean, Sigma = diag(g$sd) %*% R %*% diag(g$sd),
empirical = TRUE)
data.frame(id = paste(g$group[1], row(y), sep = "-"), treat = g$treat[1],
pretested = g$pretested[1],
occasion = factor(as.character(g$occasion)[col(y)], levels = occasions),
y_post = as.vector(y))
}))
}
set.seed(2014)
long <- rebuild(ps, r = 0.5)The published analyses used the posttests alone, so the model is fitted without the pretest:
fit <- fit_solomon_mmrm(y_post, treat, pretested, id, occasion, data = long)
fit
#> Solomon MMRM: mixed model for repeated measures (Mallinckrodt et al., 2008)
#> Occasions: Posttest, 3 months, 6 months
#> Covariance: unstructured (REML)
#> Degrees of freedom: Kenward-Roger
#>
#> Observed posttests by group and occasion:
#> Posttest 3 months 6 months
#> pretested treatment 22 22 22
#> pretested control 22 22 22
#> unpretested treatment 21 21 21
#> unpretested control 31 31 31
#>
#> Occasion Contrast Estimate SE df t
#> Posttest ATE (avg over pretest) 1.915 1.043 92 1.84
#> Posttest Pretest x Treatment 1.890 2.086 92 0.91
#> Posttest Treatment | pretested 2.860 1.522 92 1.88
#> Posttest Treatment | unpretested 0.970 1.427 92 0.68
#> 3 months ATE (avg over pretest) 0.280 0.956 92 0.29
#> 3 months Pretest x Treatment 1.440 1.911 92 0.75
#> 3 months Treatment | pretested 1.000 1.394 92 0.72
#> 3 months Treatment | unpretested -0.440 1.307 92 -0.34
#> 6 months ATE (avg over pretest) -0.070 0.705 92 -0.10
#> 6 months Pretest x Treatment 3.320 1.411 92 2.35
#> 6 months Treatment | pretested 1.590 1.029 92 1.54
#> 6 months Treatment | unpretested -1.730 0.965 92 -1.79
#> 6 months vs Posttest Change in Pretest x Treatment 1.430 1.870 92 0.76
#> p 95% CI
#> 0.070 [-0.157, 3.987]
#> 0.367 [-2.254, 6.034]
#> 0.063 [-0.163, 5.883]
#> 0.498 [-1.864, 3.804]
#> 0.770 [-1.618, 2.178]
#> 0.453 [-2.356, 5.236]
#> 0.475 [-1.770, 3.770]
#> 0.737 [-3.036, 2.156]
#> 0.921 [-1.471, 1.331]
#> 0.021 [ 0.518, 6.122]
#> 0.126 [-0.455, 3.635]
#> 0.076 [-3.646, 0.186]
#> 0.446 [-2.284, 5.144]The squared t statistics of the Pretest x Treatment rows are the F values of the separate analyses:
px <- fit$effects[fit$effects$contrast == "Pretest x Treatment", ]
data.frame(occasion = px$occasion, F = round(px$statistic^2, 3), df = round(px$df, 1),
published_F = c(0.819, 0.563, 5.556))
#> occasion F df published_F
#> 1 Posttest 0.821 92 0.819
#> 2 3 months 0.568 92 0.563
#> 3 6 months 5.537 92 5.556They agree because no posttest is missing. The model’s estimate of each group’s mean on an occasion is then the observed mean. Its estimate of the variance on that occasion is the variance pooled within the four groups, which is what the analysis of variance uses. So on each occasion the assumed correlation has no effect: any value gives the same contrasts, standard errors, and tests.
Did sensitization change over time?
For problem solving, the interaction was 1.89 points after the program, 1.44 at 3 months, and 3.32 at 6 months, and only the last was significant. That pattern does not show that sensitization grew. A significant result on one occasion and a nonsignificant one on another is not itself evidence of a difference between them (Gelman & Stern, 2006). The model tests the change directly, in the row “Change in Pretest x Treatment”: the interaction at 6 months minus the interaction after the program.
The standard error of that change depends on how strongly each offender’s scores at the two occasions are correlated, which the thesis does not report. So the analysis is repeated over a range of correlations:
change <- function(stats, r) {
set.seed(2014)
e <- fit_solomon_mmrm(y_post, treat, pretested, id, occasion, data = rebuild(stats, r))$effects
e[e$contrast == "Change in Pretest x Treatment", c("estimate", "std.error", "p.value")]
}
ss <- subset(jordaan2014, subscale == "Social support")
r <- c(0.2, 0.5, 0.8)
knitr::kable(cbind(subscale = rep(c("Problem solving", "Social support"), each = 3),
correlation = rep(r, 2),
rbind(do.call(rbind, lapply(r, change, stats = ps)),
do.call(rbind, lapply(r, change, stats = ss)))),
digits = 3, row.names = FALSE)| subscale | correlation | estimate | std.error | p.value |
|---|---|---|---|---|
| Problem solving | 0.2 | 1.43 | 2.281 | 0.532 |
| Problem solving | 0.5 | 1.43 | 1.870 | 0.446 |
| Problem solving | 0.8 | 1.43 | 1.337 | 0.288 |
| Social support | 0.2 | -3.26 | 2.239 | 0.149 |
| Social support | 0.5 | -3.26 | 1.777 | 0.070 |
| Social support | 0.8 | -3.26 | 1.141 | 0.005 |
The estimate does not depend on the correlation, but its standard error does.
- Problem solving. The interaction rose by 1.43 points. The rise is not significant whatever the correlation, so the data do not show that sensitization grew.
- Social support. The interaction fell by 3.26 points, from 5.79 after the program to 2.53 at 6 months. Whether that fall is significant depends on the correlation: p = .149 if it is 0.2, and p = .005 if it is 0.8.
Whether sensitization to the social support items faded cannot be settled from the published statistics.
What the example shows
- Separate occasions need only group statistics. Every per-occasion Solomon analysis can be reproduced from sizes, means, and standard deviations. With no posttest missing, one model for all occasions gives the same per-occasion results.
- Change over time needs the correlations. Whether sensitization persists, grows, or fades depends on how each participant’s scores on two occasions are related. When a study does not report the correlations, a reanalysis can only show the conclusion over a range of them, as here. When you report a Solomon study with several posttests, give the correlations between occasions, or the standard deviations of the change scores, so that others can reanalyze it.
- The model’s main advantage does not show here. It stays valid when later posttests are missing at random (see “Missing and Repeated Posttests”), but nothing was missing among the 96 offenders analyzed. No model can recover the 24 offenders who were transferred before any posttest.
-
The joint model is a solomonR extension. The
literature search documented on issue #57 found no
published method for analyzing all the posttest occasions of a Solomon
design in one model. Published studies, Jordaan (2014) among them,
analyzed each occasion separately.
fit_solomon_mmrm()applies the mixed model for repeated measures to the Solomon contrasts; its validation is in the article Longitudinal Designs: Validating the Repeated-Measures Analysis.
References
Gelman, A., & Stern, H. (2006). The difference between “significant” and “not significant” is not itself statistically significant. The American Statistician, 60(4), 328–331. https://doi.org/10.1198/000313006X152649
Holm, S. (1979). A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics, 6(2), 65–70. https://www.jstor.org/stable/4615733
Jordaan, J. (2014). The development and evaluation of a life skills programme for young adult prisoners [Doctoral thesis, University of the Free State]. KovsieScholar. https://hdl.handle.net/11660/832
Mallinckrodt, C. H., Lane, P. W., Schnell, D., Peng, Y., & Mancuso, J. P. (2008). Recommendations for the primary analysis of continuous endpoints in longitudinal clinical trials. Drug Information Journal, 42(4), 303–319. https://doi.org/10.1177/009286150804200402
Walton Braver, M. C., & Braver, S. L. (1988). Statistical treatment of the Solomon four-group design: A meta-analytic approach. Psychological Bulletin, 104(1), 150–154. https://doi.org/10.1037/0033-2909.104.1.150