Skip to contents

Analyzes an expert panel rated over successive Delphi rounds. For each item it reports consensus in every round and the stability of experts' ratings between consecutive rounds, and it fits expert_validity() to each round so the usual relevance evidence is available round by round.

Consensus and stability are different questions. Consensus asks whether enough experts agree now. Stability asks whether experts are still changing their answers. A Delphi can reach one without the other, so both are shown.

Usage

delphi_validity(
  ratings,
  expert_col = "expert",
  item_col = "item",
  round_col = "round",
  rating_col = "rating",
  lo,
  hi,
  agree_cut = NULL,
  consensus_threshold = NULL,
  stability = c("kappa", "lambda", "chisq_individual", "chisq_group", "percent_change"),
  kappa_weights = c("quadratic", "linear"),
  alpha = 0.05,
  B = 1000,
  seed = NULL
)

Arguments

ratings

A data frame with one row per expert, item, and round. Rows with a missing rating are ignored.

expert_col, item_col, round_col, rating_col

Column names in ratings.

lo, hi

Lowest and highest points of the rating scale, as whole numbers.

agree_cut

Rating at or above which an expert counts as agreeing. Defaults to hi - 1, the usual relevance cut on a 4-point scale.

consensus_threshold

Share of experts that must agree for consensus, between 0 and 1, fixed before the study. NULL (default) reports agreement descriptively.

stability

Stability statistic; see Details.

kappa_weights

"quadratic" (default) or "linear", used when stability = "kappa".

alpha

Significance level for the chi-square methods, and 1 - alpha is the interval level.

B

Bootstrap resamples for the kappa interval. Use 0 to skip it.

seed

Optional seed for the bootstrap.

Value

An object of class contentvalid_delphi and contentvalid_workflow. results has one row per item: its last round, n_experts there, prop_agree, consensus, and, for the last pair of consecutive rounds, prop_unchanged, stability with stability_low and stability_high where an interval exists, stability_p for the chi-square methods, and stable for the methods that make a decision. details holds consensus (every item and round), stability (every item and pair of rounds, including n_paired, min_expected for the chi-square methods, n_boot_usable for the kappa interval, and a note where a statistic is undefined or unreliable), panel (experts per round), and round_fits, the expert_validity() fit for each round.

Details

Consensus. An expert agrees with an item when their rating is at least agree_cut. prop_agree is the share of responding experts who agree, which on a relevance scale is the I-CVI. An item reaches consensus when prop_agree meets consensus_threshold. There is deliberately no default threshold: Diamond et al. (2014) recommend fixing it before the study, and the 75% median they report describes common practice rather than a validated cut-off. Without a threshold, items are reported as Descriptive only.

Stability is computed for each item and each pair of consecutive rounds, on the experts who rated the item in both. prop_unchanged, the share who kept their rating, is always reported. The stability argument chooses the statistic reported beside it:

  • "kappa" (default): weighted kappa between each expert's ratings in the two rounds (Holey et al., 2007), read as a trend with no cut-off. Quadratic weights (the default) make kappa the intraclass correlation of the two rounds' ratings (Fleiss & Cohen, 1973); linear weights count a two-point change twice a one-point change (Cohen, 1968). The interval is a percentile bootstrap over the experts; see Reading the kappa interval below.

  • "lambda": Chaffin and Talley's (1980) index of predictive association.

  • "chisq_individual": Chaffin and Talley's (1980) chi-square test on each expert's pair of ratings; a significant result is read as stable.

  • "chisq_group": Dajani, Sincoff and Talley's (1979) chi-square test on the two rounds' distributions; a non-significant result is read as stable.

  • "percent_change": the net change of Scheibe et al. (1975/2002), stable below 15%.

The alternatives are published but contested, so the printed output explains each one's limits. Stability never changes an item's status: the status rests on consensus in the item's last round, and stability is read beside it.

Items may enter or leave between rounds. An item's last round is the last one in which anyone rated it, and stability is computed only between consecutive rounds in which it was rated.

When a stability statistic is undefined

A stability statistic can be NA for two different reasons, and prop_unchanged tells them apart. When prop_unchanged is also NA, the item has no pair of consecutive rounds: it was rated in one round only. When prop_unchanged has a value, a pair exists but the statistic is undefined for that data, and details$stability$note says why.

The common case is the one that reads worst if reported bare. Kappa is chance-corrected, so when every paired rating in both rounds falls in one category the disagreement expected by chance is zero and kappa is 0/0. There prop_unchanged is 1: the panel could not have been more stable, and saying only "not estimable" would describe a defect that does not exist. Goodman- Kruskal lambda is undefined when the later round is unanimous, and the chi-square methods when a table has fewer than two occupied rows or columns.

Why kappa has no verbal labels

Landis and Koch (1977) introduced the familiar labels (slight, fair, moderate, substantial, almost perfect) and called their divisions clearly arbitrary. Kappa also falls when ratings converge on one category, which is what a Delphi aims for: in Holey et al. (2007), the statement experts agreed on most had the lowest kappa. A label would therefore tend to worsen as a panel succeeds. Read kappa as a trend, next to prop_unchanged.

Reading the kappa interval

The interval beside kappa is a percentile bootstrap: the experts are resampled with replacement, each carrying both of their ratings, kappa is recomputed in every resample, and the interval runs from the alpha / 2 to the 1 - alpha / 2 quantile of those values. The experts are the units the two rounds cross-classify, so this is the same resampling scheme Klar et al. (2002) describe for kappa, where the subjects each contribute one pair of ratings.

Two things about it are worth knowing before the interval is reported.

First, it is not exact at the sizes Delphi panels run to. Klar et al. (2002) simulated this interval and found that a nominal 95% interval covered the true value about 83% of the time with 20 units, 89% with 25, 91% with 30, and only reached about 94% from 40 up. Panels smaller than 40 therefore get an interval narrower than its label claims, and the shortfall grows as the panel shrinks. Report it as a rough indication of precision rather than as a test of a hypothesis about kappa.

Second, those simulations used an unweighted kappa on two categories, while this function computes a weighted kappa on an ordinal scale. The resampling scheme is the same, and nothing in it depends on the number of categories or on the weights, but its coverage has not been simulated for that case. Treat the coverage figures above as the shape of the problem rather than as exact numbers for a Delphi panel.

Set B = 0 to omit the interval and report kappa beside prop_unchanged alone.

References

Chaffin, W. W., & Talley, W. K. (1980). Individual stability in Delphi studies. Technological Forecasting and Social Change, 16(1), 67–73. doi:10.1016/0040-1625(80)90074-8

Cohen, J. (1968). Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit. Psychological Bulletin, 70(4), 213–220. doi:10.1037/h0026256

Dajani, J. S., Sincoff, M. Z., & Talley, W. K. (1979). Stability and agreement criteria for the termination of Delphi studies. Technological Forecasting and Social Change, 13(1), 83–90. doi:10.1016/0040-1625(79)90007-6

Diamond, I. R., Grant, R. C., Feldman, B. M., Pencharz, P. B., Ling, S. C., Moore, A. M., & Wales, P. W. (2014). Defining consensus: A systematic review recommends methodologic criteria for reporting of Delphi studies. Journal of Clinical Epidemiology, 67(4), 401–409. doi:10.1016/j.jclinepi.2013.12.002

Fleiss, J. L., & Cohen, J. (1973). The equivalence of weighted kappa and the intraclass correlation coefficient as measures of reliability. Educational and Psychological Measurement, 33(3), 613–619. doi:10.1177/001316447303300309

Holey, E. A., Feeley, J. L., Dixon, J., & Whittaker, V. J. (2007). An exploration of the use of simple statistics to measure consensus and stability in Delphi studies. BMC Medical Research Methodology, 7, 52. doi:10.1186/1471-2288-7-52

Klar, N., Lipsitz, S. R., Parzen, M., & Leong, T. (2002). An exact bootstrap confidence interval for kappa in small samples. Journal of the Royal Statistical Society: Series D (The Statistician), 51(4), 467–478. doi:10.1111/1467-9884.00331

Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159–174. doi:10.2307/2529310

Scheibe, M., Skutsch, M., & Schofer, J. (2002). Experiments in Delphi methodology. In H. A. Linstone & M. Turoff (Eds.), The Delphi method: Techniques and applications (pp. 257–281). https://www.foresight.pl/assets/downloads/publications/Turoff_Linstone.pdf (Original work published 1975)

See also

expert_validity() for a single round, and compare_rounds(), which accepts the fits in details$round_fits.

Examples

# Eight experts rate four statements for relevance (1-4) over three rounds.
r1 <- cbind(S1 = c(4, 4, 3, 4, 2, 4, 3, 4), S2 = c(2, 3, 2, 1, 3, 2, 2, 3),
            S3 = c(3, 4, 2, 3, 4, 1, 3, 2), S4 = c(4, 3, 4, 2, 3, 3, 4, 2))
r2 <- cbind(S1 = c(4, 4, 4, 4, 3, 4, 3, 4), S2 = c(2, 2, 2, 1, 3, 2, 2, 2),
            S3 = c(3, 3, 3, 3, 4, 2, 3, 3), S4 = c(4, 3, 4, 3, 3, 3, 4, 3))
r3 <- cbind(S1 = c(4, 4, 4, 4, 3, 4, 4, 4), S2 = c(2, 2, 2, 1, 2, 2, 2, 2),
            S3 = c(3, 3, 3, 3, 4, 3, 3, 3), S4 = c(4, 3, 4, 3, 4, 3, 4, 3))
long <- function(m, round) {
  data.frame(expert = paste0("E", seq_len(nrow(m))),
             item = rep(colnames(m), each = nrow(m)),
             round = round, rating = as.vector(m))
}
ratings <- rbind(long(r1, 1), long(r2, 2), long(r3, 3))

fit <- delphi_validity(ratings, lo = 1, hi = 4, consensus_threshold = 0.75,
                       B = 200, seed = 1)
fit
#> contentvalidR Delphi analysis
#> -----------------------------
#> Items: 4 | Experts: 8 | Rounds: 3 (1, 2, 3)
#> Experts per round: 8, 8, 8
#> Agreement: a rating of 3 or higher on the 1-4 scale. Consensus threshold:
#> 75%, fixed before the study.
#> Stability: weighted kappa (quadratic weights) between consecutive rounds
#> 
#> 3 of 4 items reached consensus in their last round.
#> No consensus: S2
#> 
#> Item-level evidence (last round, and the last pair of rounds)
#>  item     decision last round n agree unchanged kappa      95% CI
#>    S1    Consensus          3 8  1.00       .88   .60 [.00, 1.00]
#>    S2 No consensus          3 8   .00       .88   .67 [.00, 1.00]
#>    S3    Consensus          3 8  1.00       .88   .67 [.00, 1.00]
#>    S4    Consensus          3 8  1.00       .88   .75 [.00, 1.00]
#> 
#> agree: share of experts agreeing in the item's last round. unchanged: share
#> who kept their rating between the last two rounds.
#> 
#> Stability trend (kappa) by pair of rounds
#>  item 1->2 2->3
#>    S1  .67  .60
#>    S2  .67  .67
#>    S3  .60  .67
#>    S4  .72  .75
#> 
#> Share of experts who kept their rating, by pair of rounds
#>  item 1->2 2->3
#>    S1  .75  .88
#>    S2  .75  .88
#>    S3  .50  .88
#>    S4  .75  .88
#> 
#> In some resamples kappa was undefined because every resampled rating fell in
#> one category. Those intervals use the remaining resamples (n_boot_usable in
#> details$stability), so treat them as rough.
#> 
#> Read kappa as a trend across rounds, beside the share of experts who kept
#> their rating (unchanged), not against a cut-off: kappa falls as a panel
#> converges on one category, so a stable panel can show a low kappa (Holey et
#> al., 2007).
#> 
#> The kappa intervals resample the experts (Klar et al., 2002). With fewer than
#> about 40 experts they cover less than their stated 95%, so read them as rough
#> indications of precision, not as tests.
#> 
#> How the stability statistic works
#>   Stability is weighted kappa between each expert's ratings in consecutive
#>   rounds (Holey et al., 2007), with quadratic weights. A change of two scale
#>   points counts four times a change of one. With these weights kappa equals
#>   the intraclass correlation of the two rounds' ratings, so a shift of the
#>   whole panel counts as instability (Fleiss & Cohen, 1973). No verbal labels
#>   such as 'substantial' are shown, because kappa falls when ratings converge,
#>   which is what a Delphi aims for: Holey et al. saw a low kappa for their
#>   most-agreed statement.
#> 
#>   The intervals are percentile bootstraps that resample the experts, the
#>   units the two rounds cross-classify: the procedure Klar et al. (2002)
#>   describe for kappa. They evaluated it for an unweighted kappa on two
#>   categories, and a nominal 95% interval covered about 83% of the time with
#>   20 units and 91% with 30, reaching 94% only from 40 up.
#> 
#> What these columns mean
#>   agree -- Share of experts agreeing. Share of experts at or above the
#>       agreement cut in a round; consensus means reaching the preset
#>       threshold.
#>   unchanged -- Share of experts keeping their rating. Share of experts
#>       giving the same rating in two consecutive rounds (1 means nobody
#>       changed).
#>   kappa -- Weighted kappa between rounds. Chance-corrected agreement of
#>       each expert's ratings across two rounds; read it as a trend, not
#>       against a cut-off.
#> 
#> What the decisions mean
#>   Consensus -- reached the consensus threshold in its last round.
#>   No consensus -- did not reach the consensus threshold.
#> 
#> Full definitions: contentvalid_glossary(). To hide this key:
#> options(contentvalidR.show_key = FALSE).
#> 
#> Consensus is not correctness, and 'No consensus' is not an instruction to
#> drop an item. Read these results with the experts' comments.
fit$details$stability
#>   item from_round to_round n_paired prop_unchanged method     value     lower
#> 1   S1          1        2        8          0.750  kappa 0.6666667 0.0000000
#> 2   S1          2        3        8          0.875  kappa 0.6000000 0.0000000
#> 3   S2          1        2        8          0.750  kappa 0.6666667 0.0000000
#> 4   S2          2        3        8          0.875  kappa 0.6666667 0.0000000
#> 5   S3          1        2        8          0.500  kappa 0.6000000 0.0000000
#> 6   S3          2        3        8          0.875  kappa 0.6666667 0.0000000
#> 7   S4          1        2        8          0.750  kappa 0.7241379 0.3425676
#> 8   S4          2        3        8          0.875  kappa 0.7500000 0.0000000
#>       upper df p_value stable min_expected n_low_expected n_boot_usable note
#> 1 1.0000000 NA      NA     NA           NA             NA           196     
#> 2 1.0000000 NA      NA     NA           NA             NA           179     
#> 3 1.0000000 NA      NA     NA           NA             NA           200     
#> 4 1.0000000 NA      NA     NA           NA             NA           188     
#> 5 0.8227778 NA      NA     NA           NA             NA           199     
#> 6 1.0000000 NA      NA     NA           NA             NA           182     
#> 7 1.0000000 NA      NA     NA           NA             NA           200     
#> 8 1.0000000 NA      NA     NA           NA             NA           200     

# A published alternative, with its critique printed.
delphi_validity(ratings, lo = 1, hi = 4, consensus_threshold = 0.75,
                stability = "percent_change")
#> contentvalidR Delphi analysis
#> -----------------------------
#> Items: 4 | Experts: 8 | Rounds: 3 (1, 2, 3)
#> Experts per round: 8, 8, 8
#> Agreement: a rating of 3 or higher on the 1-4 scale. Consensus threshold:
#> 75%, fixed before the study.
#> Stability: net percent change (Scheibe et al., 1975/2002) between consecutive
#> rounds
#> 
#> 3 of 4 items reached consensus in their last round.
#> No consensus: S2
#> 
#> Item-level evidence (last round, and the last pair of rounds)
#>  item     decision last round n agree unchanged change stable
#>    S1    Consensus          3 8  1.00       .88    .12    yes
#>    S2 No consensus          3 8   .00       .88    .12    yes
#>    S3    Consensus          3 8  1.00       .88    .12    yes
#>    S4    Consensus          3 8  1.00       .88    .12    yes
#> 
#> agree: share of experts agreeing in the item's last round. unchanged: share
#> who kept their rating between the last two rounds.
#> 
#> Stability trend (change) by pair of rounds
#>  item 1->2 2->3
#>    S1  .12  .12
#>    S2  .25  .12
#>    S3  .38  .12
#>    S4  .25  .12
#> 
#> Share of experts who kept their rating, by pair of rounds
#>  item 1->2 2->3
#>    S1  .75  .88
#>    S2  .75  .88
#>    S3  .50  .88
#>    S4  .75  .88
#> 
#> Stability is the net change of Scheibe et al. (1975/2002): half the summed
#> differences between the two rounds' rating distributions, as a share of the
#> experts compared, with change below 15% read as stable. The authors say the
#> measure has no statistical theory behind it; the 15% cut-off came from the
#> movement they observed in one classroom Delphi. Experts swapping answers
#> cancel out, and in a small panel one expert is a large share: with 10 experts
#> one net change is already 10%.
#> 
#> What these columns mean
#>   agree -- Share of experts agreeing. Share of experts at or above the
#>       agreement cut in a round; consensus means reaching the preset
#>       threshold.
#>   unchanged -- Share of experts keeping their rating. Share of experts
#>       giving the same rating in two consecutive rounds (1 means nobody
#>       changed).
#>   change -- Net change in the rating distribution. Net change in the rating
#>       distribution between rounds (stable below .15 by its authors' rule).
#> 
#> What the decisions mean
#>   Consensus -- reached the consensus threshold in its last round.
#>   No consensus -- did not reach the consensus threshold.
#> 
#> Full definitions: contentvalid_glossary(). To hide this key:
#> options(contentvalidR.show_key = FALSE).
#> 
#> Consensus is not correctness, and 'No consensus' is not an instruction to
#> drop an item. Read these results with the experts' comments.