Shared walkthrough data from content review to empirical screening
Source:R/data.R
nomo_demo_walkthrough.RdSimulated responses of 400 students to twelve candidate items for a two-facet
Study Persistence scale. These are the shared teaching data of the joint
walkthrough with contentvalidR: the same twelve items went through an
expert item sort there, and the ten it carried forward are screened here.
The data are simulated. No participant was involved, and the construct, its
facets, and every item stem were written for the example.
Format
nomo_demo_walkthrough is a data frame with 400 rows and 14 columns:
- respondent
Respondent number, 1 to 400.
- cohort
A factor with levels
"A"and"B", 200 respondents each.- EF1, EF2, EF3, EF4, EF5, EF6
Effort Regulation items, integer responses 1 to 5, as answered (
EF2is not recoded).- TF1, TF2, TF3, TF4, TF5, TF6
Task Focus items, integer responses 1 to 5, as answered (
TF2is not recoded).
nomo_demo_walkthrough_items is a data frame with 12 rows and 5 columns:
item; facet ("EF" or "TF"); stem, the item's wording;
reverse_worded, a logical; and role, what the item was built to do.
An object of class data.frame with 12 rows and 5 columns.
Source
Generated by contentvalidR's data-raw/build-walkthrough-data.R
(seed 20260921), copied unchanged from its v0.9.0 release into this
package's data-raw/walkthrough/, and converted by
data-raw/nomo_demo_walkthrough.R.
Details
Each item was built to behave in a particular way at one stage or both, so
the two stages can be seen to disagree. nomo_demo_walkthrough_items states
what each item was built to do, in its role column:
EF1,TF1,TF3: ordinary items, which both stages accept.EF2,TF2: reverse-worded. A high answer means less persistence, so until they are recoded (6 minus the response) each correlates negatively with its own facet. That is a coding matter, not evidence against the item.EF3: endorsed by almost everyone, so its answers pile up at the top of the scale and every correlation it has is attenuated. An empirical screen flags it; it is also the only item about finishing required work, so dropping it narrows the domain the panel defined.EF4: passes content review (18 of 20 judges) and carries almost no common variance. Content review cannot see this.EF5,TF5: fail content review, so a handoff holds them back. They stay in the data so a reader can see what keeping them would have done.EF6: meets the content criterion by one judge, then behaves well.TF4: loads on both facets.TF6: shifted upward in cohort B (a classroom rule about phones), a known source of non-invariance between cohorts.
Population model, per respondent:
Two facets, Effort Regulation and Task Focus, standard normal with correlation .45.
Each item's latent response is its loadings times the facets plus normal error, scaled to unit variance, and cut into five categories at the normal quantiles .10, .30, .60, and .85.
Standardized loadings on EF:
EF1.72,EF2.68,EF3.60,EF4.15,EF5.55,EF6.60, andTF4.45.Loadings on TF:
TF1.70,TF2.66,TF3.62,TF4.40,TF5.58,TF6.64.EF3's latent response is shifted up by 1.75, andTF6's by .35 in cohort B.EF2andTF2are answered in the opposite direction (6 minus the category).
The content-review half of the example, the expert sort of these items, and
the handoff it produced are in contentvalidR. The handoff for these items is
stored with this package as
system.file("extdata", "content-handoff-walkthrough.rds", package = "nomologR").
Examples
str(nomo_demo_walkthrough)
#> 'data.frame': 400 obs. of 14 variables:
#> $ respondent: int 1 2 3 4 5 6 7 8 9 10 ...
#> $ cohort : Factor w/ 2 levels "A","B": 1 1 1 1 1 1 1 1 1 1 ...
#> $ EF1 : int 3 4 4 5 3 4 1 3 3 3 ...
#> $ EF2 : int 2 2 3 3 4 4 3 5 5 3 ...
#> $ EF3 : int 5 5 5 5 5 5 4 4 5 5 ...
#> $ EF4 : int 3 3 3 3 5 4 4 1 5 3 ...
#> $ EF5 : int 4 5 2 5 3 3 3 1 3 5 ...
#> $ EF6 : int 2 2 2 4 3 1 1 1 3 5 ...
#> $ TF1 : int 3 5 2 2 4 4 5 3 4 5 ...
#> $ TF2 : int 2 3 3 4 2 3 3 3 3 1 ...
#> $ TF3 : int 4 5 3 2 2 5 5 2 3 4 ...
#> $ TF4 : int 5 3 3 4 4 2 3 1 2 5 ...
#> $ TF5 : int 5 4 4 2 4 3 5 4 5 4 ...
#> $ TF6 : int 4 3 3 2 3 4 4 3 5 4 ...
nomo_demo_walkthrough_items[, c("item", "role")]
#> item role
#> 1 EF1 ordinary: the panel places it and it behaves as intended
#> 2 EF2 reverse-worded: behaves as intended once recoded
#> 3 EF3 flagged by an empirical screen and worth keeping anyway
#> 4 EF4 passes content review, then carries almost no common variance
#> 5 EF5 fails content review: the wording pulls judges toward test anxiety
#> 6 EF6 meets the content criterion by one judge, then behaves well
#> 7 TF1 ordinary: the panel places it and it behaves as intended
#> 8 TF2 reverse-worded: behaves as intended once recoded
#> 9 TF3 ordinary: the panel places it and it behaves as intended
#> 10 TF4 passes content review, then loads on both facets
#> 11 TF5 fails content review: a competing facet takes more assignments
#> 12 TF6 behaves differently in the two cohorts
# Answered as written, a reverse-worded item runs against its facet.
cor(nomo_demo_walkthrough$EF1, nomo_demo_walkthrough$EF2)
#> [1] -0.4950639
cor(nomo_demo_walkthrough$EF1, 6 - nomo_demo_walkthrough$EF2)
#> [1] 0.4950639