Carry content-validity decisions into empirical validation
Source:R/content_handoff.R
content_handoff.RdPackages the item decisions from a finished content-validity workflow so they can be carried into an empirical scale-development workflow without retyping item names or losing the record of why each item was kept.
The result holds the item names that survived content review, the construct each belongs to where the design defines one, a per-item evidence table, the statistics behind each decision, and the provenance of the analysis. It is plain data, so a downstream package can read it without contentvalidR being installed.
Usage
content_handoff(
fit,
keep = "Supported",
round = 1,
reverse_keyed = NULL,
response_scale = NULL
)Arguments
- fit
A fitted
contentvalid_sort,contentvalid_rating,contentvalid_expert, orcontentvalid_delphiobject.- keep
Statuses that travel forward, defaulting to
"Supported". Any of"Supported","Review","Insufficient data", or"Descriptive only".- round
Pretest round this analysis represents. One fit is one round, so this defaults to
1and matters only when stacking rounds by hand. It cannot be set for a Delphi fit, which dates each item by the round it settled in.- reverse_keyed
Names of the reverse-worded items,
character(0)if none is, orNULL(the default) to leave keying unrecorded. See "Instrument metadata".- response_scale
The lowest and highest answer respondents can give, such as
c(1, 5), orNULL(the default) to leave it unrecorded. This is the scale of the instrument, not the scale the panel rated on.
Value
An object of class contentvalid_handoff, cv_handoff, and list,
as described under "Object shape".
Details
Item-level workflows are accepted: sort_validity(), rating_validity(),
expert_validity(), and delphi_validity(). judge_validity() and
domain_validity() are refused, because their rows are judges and blueprint
cells rather than items, so there is no item set to carry forward.
Items that do not meet keep are not dropped from the record. They stay in
item_evidence with carried = FALSE, so a reader can see what was held
back and why. Review is not deletion.
Object shape (schema version 1)
The object has class c("contentvalid_handoff", "cv_handoff", "list"). A
consumer matches on "cv_handoff" and reads these fields:
itemscharacter vector of the carried item names, that is
item_evidence$item[item_evidence$carried], unique and in results order.scalesnamed list mapping each construct to its carried items, or
NULLwhen the design has no construct mapping. Expert relevance and essentiality rate a single item set with no construct column, so they produceNULL, as does a Delphi study. An item belongs to at most one scale, the one itsitem_evidence$scalenames: every workflow that maps items to constructs requires exactly one target per item, and stops otherwise.item_evidencedata frame with one row per reviewed item:
item,scale(NAwithout a construct mapping),carried,status,recommendation,n_judges,rule, andround. For a Delphi handoffrounddiffers between items; see "A Delphi handoff". From contentvalidR 0.7.0 it also carrieskeying,response_min, andresponse_max, described under "Instrument metadata".item_statisticsdata frame, one row per item per statistic:
item,statistic,value,criterion(NAwhen the method sets no explicit criterion), andround. Which statistics carry a criterion depends on the workflow rather than on the statistic alone: modified kappa carries 0.74 fromexpert_validity(), and none from a Delphi handoff, which decides on the consensus threshold. From contentvalidR 0.7.0 it also carriesnote, described under "The note column". From contentvalidR 0.5.0 it also carrieslower,upper,interval_method, andinterval_level, described under "Intervals".provenancelist with
schema_version,package,package_version,workflow,mode,keep,method,citation,settings,design, andcreated.panel_statisticsadded in contentvalidR 0.5.0. Data frame of panel-level statistics, with the same columns as
item_statisticslessitem, includingnotefrom 0.7.0. It holds the panel agreement coefficient whenexpert_validity()computed one, and has zero rows otherwise.
This shape is agreed with the nomologR package, which consumes it in
nomo_screen() and nomo_run(). Neither package depends on the other.
Fields and columns added within schema version 1 are optional for a reader,
which should check that they are present rather than assume it.
Reading the decisions
Take each item's decision from the handoff, not from its statistics.
carried says whether the item travels forward, and status says why: it
is one of the four values keep accepts. recommendation words the same
decision for a person, and like rule it is prose.
Don't re-derive a decision by comparing value with criterion, and don't
branch on provenance$package_version. A corrected rule can change a
decision between releases while the statistics stay the same, and the
handoff records the decision its producer made. contentvalidR 0.8.0 is the
example. For an item that 7 of 9 experts rated relevant, the I-CVI is .778
in both releases. 0.7.0 compared it with a rounded .78 and held the item
back. 0.8.0 applies Lynn's (1986) 7 of 9, stores 7/9 as the criterion, and
carries the item. There the value equals the criterion, so a recomputed
value >= criterion would hang on floating-point rounding, where the
package itself compares counts of experts. Read carried, which holds what
each release decided.
Keying is read the same way, from the field and not from a default:
keying is 1 for a forward-worded item, -1 for a reverse-worded one,
and NA when nobody said, as described under "Instrument metadata". Treat
NA as unknown, never as forward-worded.
What version 1 freezes
Schema version 1 is frozen as of contentvalidR 0.7.0. Code that reads a
handoff can rely on all of the following, in every release that reports
schema_version = 1:
The six top-level fields above, under those names.
In
item_evidence:item,scale,carried,status,recommendation,n_judges,rule,round,keying,response_min,response_max.In
item_statistics:item,statistic,value,criterion,round,lower,upper,interval_method,interval_level,note.In
panel_statistics: the same columns lessitem.In
provenance:schema_version,package,package_version,workflow,mode,keep,method,citation,settings,design,created.
Each of those columns keeps its name, its position, and its type. Every
handoff carries every column, including when a workflow has nothing to put
in one: a statistic with no interval carries NA in the four interval
columns rather than dropping them, and a workflow with no panel coefficient
returns a zero-row panel_statistics with the full set of columns. A reader
can therefore bind handoffs from different workflows without reconciling
their columns.
These are deliberately not frozen, and a reader should not depend on them:
The set of rows. Which items, which statistics, and how many of each depend on the workflow and on the data.
The values in the
statisticcolumn. They are labels for display, and may be reworded in a minor release; match on the workflow inprovenanceinstead.The text in
note,rule,recommendation, andcitation, which is prose for a human reader.The contents of
settingsanddesign, which mirror the fitted object and grow with it.
Neither is the printed output part of the schema. print() on a handoff is
written for a person, and its layout and wording may change in any release.
Read the fields.
New optional fields and columns may still be added within version 1, at the end of a data frame or list. A reader written against this section keeps working when that happens, provided it addresses columns by name.
If the schema ever changes
Renaming a field, removing one, changing a type, or changing what a field
means is a version 2 change, not a minor release. It would raise
provenance$schema_version to 2L, and version 1 would keep being
produced for at least one full release cycle so that readers have a
version to fall back on. The release notes would say what moved.
A reader should gate on the version rather than on the contentvalidR version:
if (!inherits(h, "cv_handoff") || h$provenance$schema_version != 1L) {
stop("this reader understands handoff schema version 1 only")
}No version 2 is planned.
The note column
Added in contentvalidR 0.7.0 to item_statistics and panel_statistics.
It says why a value or interval is absent or degenerate, in the producing
function's own words, so a reader need not re-derive method-specific
semantics. For example, a Delphi stability row may carry "Kappa is
undefined: every rating fell in the same category in both rounds."
Its contract, agreed with the nomologR maintainers:
It is display text only. Never match on it, branch on it, or parse it. Its wording may change in any minor release without a schema change.
It is always a character vector with no
NA.""means there is nothing to say, not that something is missing, so a row can carry a value and an empty note.It never replaces the values. Whether a statistic is undefined, and which of the two cases applies, stays readable from
valueandproportion unchangedas described under "When a stability statistic is NA". That inference is the supported way to decide anything.Objects from contentvalidR 0.6.0 and earlier have no such column, and a reader should treat its absence as every note being empty.
Intervals
Each statistic's interval travels with it, so a reader can tell a unanimous
four-judge panel from a unanimous twenty-judge one. lower and upper are
the bounds, interval_method names the method, and interval_level is the
confidence level, for example 0.95.
Aiken's V: the Penfield-Giacobbi score interval.
I-CVI and Psa: the method chosen with
proportion_ci, the Wilson score interval by default.Panel agreement: the item-resampling percentile bootstrap of
panel_agreement().
The four columns are NA together when a statistic has no interval. That
happens when the method defines none (Csv, HTC, HTD, CVR, the essential
count, modified kappa, IOC, and p-values), when intervals were switched off
with proportion_ci = "none", or when the statistic itself could not be
computed. NA there never stands for missing data.
The handoff reports intervals only. It does not turn them into priors or weights for a later analysis; that is a question for the consuming package.
A Delphi handoff
A Delphi study settles one item at a time: an item that reached consensus
early was set aside, and its last round came before the study's final round.
A Delphi handoff therefore carries each item's evidence from its own last
round, and round holds that round's index rather than one constant. It is
the only workflow where round varies within a handoff.
Each item carries the relevance evidence of its last round, taken from that
round's expert_validity() fit, with intervals: I-CVI against the
consensus threshold, Aiken's V, and modified kappa. Modified kappa
carries no criterion here, because a Delphi decides on the consensus
threshold rather than on the 0.74 rule that expert_validity() applies.
Two more statistics record whether the panel had stopped moving:
proportion unchanged, and the stability statistic that ran, named for its
method, such as weighted kappa (quadratic) or Goodman-Kruskal lambda.
The chi-square methods add stability p_value against alpha. Stability
travels as evidence beside the decision; it never decides what is carried,
exactly as it never sets an item's status in delphi_validity().
When a stability statistic is NA
A stability row is always present for a carried item, so an NA there is a
statement about the data rather than a missing record. There are two cases,
and they can be told apart from the object alone:
proportion unchangedis alsoNAThe item has no pair of consecutive rounds: it was rated in one round only, so there was nothing to compare.
proportion unchangedhas a valueA pair exists, but the statistic is undefined for that data.
In the second case, read proportion unchanged, which is often the more
informative number. An undefined kappa beside proportion unchanged of 1 is
perfect stability that kappa cannot express: kappa is chance-corrected, and
when every paired rating in both rounds falls in one category, the
disagreement expected by chance is zero, so kappa is 0/0. Reporting only
"not estimable" there would describe a defect that does not exist. Goodman-
Kruskal lambda is undefined when the later round is unanimous, and the
chi-square methods are undefined for a table with fewer than two occupied
rows or columns.
delphi_validity() states the reason in details$stability$note, and from
contentvalidR 0.7.0 the handoff carries that sentence in the note column
of item_statistics, described under "The note column".
Instrument metadata
Added in contentvalidR 0.7.0, at the request of the nomologR maintainers,
because two empirical computations cannot be done correctly without them.
keyingis1for a forward-worded item,-1for a reverse-worded one, andNAwhen nobody said. An even-odd consistency index must recode reverse-worded items before splitting the scale, or a perfectly consistent respondent looks careless. And a negative corrected item-total correlation means opposite things in the two cases: on an item that was never recoded it is a coding error, and on a correctly coded item it is evidence against the item.response_minandresponse_maxare the lowest and highest answers a respondent can give. Screening for out-of-range answers needs the scale's limits rather than the observed ones, because a category nobody used is still a legal answer, and long-string and within-person variability indices mean different things on a two-point and a seven-point scale.
Both come from the analyst, never from the fit. A content-validity panel
rates relevance or correspondence on its own scale, usually 1 to 4, which is
not the scale respondents will answer the items on. Copying a fit's lo and
hi into response_min and response_max would give a downstream reader
the wrong limits, and it would then reject every legitimate top-category
answer. So both default to NA, which means unknown, and a reader should
treat NA that way rather than assume a forward-worded item or an observed
range.
Supplying reverse_keyed is a statement about every item: the ones named
are reverse-worded and the rest are not. Pass character(0) to record that
you checked and none is. So keying is either NA for every item or for
none of them, and the same holds for the response scale.
Naming a reverse-worded item without response_scale gives a warning,
because such an item cannot be recoded without the scale's limits. A reader
that recodes would otherwise have to refuse later, where the problem is
harder to fix.
What a handoff does and does not establish
Surviving content review is evidence about relevance, representation, and expert judgment. It does not establish that an item will behave well empirically. An item can be clearly relevant and still correlate poorly with its construct or load on an unintended factor. That is what the downstream empirical analysis tests, which is why the item set travels with its evidence rather than as a bare list of names.
References
Lynn, M. R. (1986). Determination and quantification of content validity. Nursing Research, 35(6), 382–385. doi:10.1097/00006199-198611000-00017
See also
as.data.frame.contentvalid_workflow() for the full results table,
and content_report() for manuscript tables.
Examples
relevance <- matrix(
c(4,4,4,3, 4,4,3,4, 3,4,4,4, 2,2,1,2),
nrow = 4,
dimnames = list(NULL, paste0("Item", 1:4))
)
fit <- expert_validity(relevance, mode = "relevance", lo = 1, hi = 4,
agreement = "none")
handoff <- content_handoff(fit)
handoff
#> contentvalidR handoff (schema version 1)
#> ----------------------------------------
#> Workflow: expert-panel (relevance) | contentvalidR 0.10.0.9000 | 2026-09-28
#> Items carried forward: 3 of 4
#> Carried when status is: Supported
#> Constructs: none in this design; the panel rated one item set.
#> Intervals carried: Aiken's V (Penfield-Giacobbi score, 95%); I-CVI (Wilson
#> score, 95%)
#>
#> Held back
#> item decision
#> Item4 Review
#>
#> Carry these items into the empirical workflow once response data are
#> collected. In nomologR that is
#> nomo_screen(data, items = handoff)
#> which screens the items carried here. Passing the whole handoff, rather than
#> handoff$items, keeps the keying and the reasons for anything held back.
#>
#> Surviving content review is evidence about relevance, representation, and
#> expert judgment. It does not establish that an item will behave well
#> empirically: an item can be clearly relevant and still correlate poorly with
#> its construct or load on an unintended factor. Items held back are listed
#> above rather than deleted, so the record stays complete.
#>
handoff$items
#> [1] "Item1" "Item2" "Item3"
handoff$item_evidence
#> item scale carried status recommendation n_judges
#> 1 Item1 <NA> TRUE Supported Strong support 4
#> 2 Item2 <NA> TRUE Supported Strong support 4
#> 3 Item3 <NA> TRUE Supported Strong support 4
#> 4 Item4 <NA> FALSE Review Review 4
#> rule
#> 1 at least 4 of 4 experts rate the item relevant (I-CVI >= 1.00; Lynn, 1986); modified kappa > .74 for strong support (Polit, Beck & Owen, 2007)
#> 2 at least 4 of 4 experts rate the item relevant (I-CVI >= 1.00; Lynn, 1986); modified kappa > .74 for strong support (Polit, Beck & Owen, 2007)
#> 3 at least 4 of 4 experts rate the item relevant (I-CVI >= 1.00; Lynn, 1986); modified kappa > .74 for strong support (Polit, Beck & Owen, 2007)
#> 4 at least 4 of 4 experts rate the item relevant (I-CVI >= 1.00; Lynn, 1986); modified kappa > .74 for strong support (Polit, Beck & Owen, 2007)
#> round keying response_min response_max
#> 1 1 NA NA NA
#> 2 1 NA NA NA
#> 3 1 NA NA NA
#> 4 1 NA NA NA
handoff$item_statistics
#> item statistic value criterion round lower upper
#> 1 Item1 Aiken's V 0.91666667 NA 1 0.64612009 0.9851349
#> 2 Item2 Aiken's V 0.91666667 NA 1 0.64612009 0.9851349
#> 3 Item3 Aiken's V 0.91666667 NA 1 0.64612009 0.9851349
#> 4 Item4 Aiken's V 0.25000000 NA 1 0.08894167 0.5323053
#> 5 Item1 I-CVI 1.00000000 1.00 1 0.51010916 1.0000000
#> 6 Item2 I-CVI 1.00000000 1.00 1 0.51010916 1.0000000
#> 7 Item3 I-CVI 1.00000000 1.00 1 0.51010916 1.0000000
#> 8 Item4 I-CVI 0.00000000 1.00 1 0.00000000 0.4898908
#> 9 Item1 modified kappa 1.00000000 0.74 1 NA NA
#> 10 Item2 modified kappa 1.00000000 0.74 1 NA NA
#> 11 Item3 modified kappa 1.00000000 0.74 1 NA NA
#> 12 Item4 modified kappa -0.06666667 0.74 1 NA NA
#> interval_method interval_level note
#> 1 Penfield-Giacobbi score 0.95
#> 2 Penfield-Giacobbi score 0.95
#> 3 Penfield-Giacobbi score 0.95
#> 4 Penfield-Giacobbi score 0.95
#> 5 Wilson score 0.95
#> 6 Wilson score 0.95
#> 7 Wilson score 0.95
#> 8 Wilson score 0.95
#> 9 <NA> NA
#> 10 <NA> NA
#> 11 <NA> NA
#> 12 <NA> NA
# Carry items flagged for review as well, when the study protocol says so.
content_handoff(fit, keep = c("Supported", "Review"))$items
#> [1] "Item1" "Item2" "Item3" "Item4"