Skip to contents

Content review and an empirical screen answer different questions. An expert panel judges whether an item represents the construct as it was defined. Response data show how the item behaves: whether it varies, what it shares with the other items, and whether it works the same way in different groups. Neither answer contains the other, and an item can pass one and fail the other.

This article takes one item set across that boundary. A content review sorted twelve items and carried ten of them forward. Here those ten are screened on 400 simulated responses. The content-review half, meaning how the panel’s sort was analyzed and how the handoff was made, is the joint walkthrough in contentvalidR, One item set, both stages. Both articles use the same data. The data are simulated from a model that is documented in ?nomo_demo_walkthrough, so every claim below can be checked against the truth.

What content review decided

contentvalidR records its decisions in a handoff. This package stores the handoff for these items, so the article runs without contentvalidR installed:

handoff <- readRDS(
  system.file("extdata", "content-handoff-walkthrough.rds", package = "nomologR")
)
handoff$item_evidence[, c("item", "scale", "status", "carried", "keying")]
#>    item scale    status carried keying
#> 1   EF1    EF Supported    TRUE      1
#> 2   EF2    EF Supported    TRUE     -1
#> 3   EF3    EF Supported    TRUE      1
#> 4   EF4    EF Supported    TRUE      1
#> 5   EF5    EF    Review   FALSE      1
#> 6   EF6    EF Supported    TRUE      1
#> 7   TF1    TF Supported    TRUE      1
#> 8   TF2    TF Supported    TRUE     -1
#> 9   TF3    TF Supported    TRUE      1
#> 10  TF4    TF Supported    TRUE      1
#> 11  TF5    TF    Review   FALSE      1
#> 12  TF6    TF Supported    TRUE      1

The panel carried ten items. It held back EF5 and TF5, whose sorts did not meet the criterion. It placed the carried items in two facets:

  • Effort Regulation (EF), keeping going when the work is dull or hard;
  • Task Focus (TF), staying on the task rather than switching away.

Keying -1 marks EF2 and TF2 as reverse-worded: a high answer means less persistence.

The responses, and what each item was built to do

nomo_demo_walkthrough_items[, c("item", "role")]
#>    item                                                               role
#> 1   EF1           ordinary: the panel places it and it behaves as intended
#> 2   EF2                   reverse-worded: behaves as intended once recoded
#> 3   EF3            flagged by an empirical screen and worth keeping anyway
#> 4   EF4      passes content review, then carries almost no common variance
#> 5   EF5 fails content review: the wording pulls judges toward test anxiety
#> 6   EF6        meets the content criterion by one judge, then behaves well
#> 7   TF1           ordinary: the panel places it and it behaves as intended
#> 8   TF2                   reverse-worded: behaves as intended once recoded
#> 9   TF3           ordinary: the panel places it and it behaves as intended
#> 10  TF4                   passes content review, then loads on both facets
#> 11  TF5     fails content review: a competing facet takes more assignments
#> 12  TF6                             behaves differently in the two cohorts

A real study does not come with a role column. Here it is known because the data were simulated to contain specific problems, and it is the answer key for what follows.

Screening the data as they arrive

nomo_run() takes the handoff whole: its carried items and facets become the run’s scales, and each facet is screened on its own.

run <- nomo_run(nomo_demo_walkthrough, scales = handoff)
run
#> <nomo_run> Guided workflow
#> Status: PAUSED | Mode: teaching | Sample design: same sample
#> Exploratory N = 400 | Confirmatory N = 400 | Scales: 2
#> Completed: screen -> factors | Next: efa
#> 
#> Key evidence
#>   - Item audit: 10 items; flags: 10 review, 0 concern
#>   - Parallel analysis suggests: EF 1, TF 1
#> 
#> Researcher decision required: efa (EF, TF)
#>   Reason: The EFA factor count changes the fitted model. Retention evidence
#>   can inform that choice, but it does not authorize the pipeline to choose for
#>   the researcher.
#>   Options: Inspect the full `nomo_factors` result, compare plausible
#>   neighboring solutions when appropriate, and supply a positive integer for
#>   each scale.
#>   Consequence: No EFA is fitted until an explicit researcher factor-count
#>   decision is supplied.
#>   - EF: Parallel analysis currently suggests 1 factor; the retained plausible
#>     set is 1. The pipeline has not adopted a factor count.
#>   - TF: Parallel analysis currently suggests 1 factor; the retained plausible
#>     set is 1. The pipeline has not adopted a factor count.
#>   Example: decisions = list(factor_count = c(EF = <integer>, TF = <integer>))
#> 
#> No later stage has been run automatically while this consequential decision is
#> unresolved.

Every item is flagged. The Effort Regulation audit shows why:

summary(run$results$screen$EF)
#> <nomo_screen summary> Item and data audit
#> Cases: 400 | Items: 5 | Flags: 5 review, 0 concern
#> Items with missing responses: 0 | Constant: 0 | All missing: 0
#> Relationship eligible: 5
#> 
#> Item review
#>   Item  Type      Missing  Top share  Item-rest r  Flag
#>   EF1   discrete     0.0%      34.8%        0.173  review
#>   EF2   discrete     0.0%      28.5%       -0.505  review
#>   EF3   discrete     0.0%      78.5%        0.210  review
#>   EF4   discrete     0.0%      33.8%        0.168  review
#>   EF6   discrete     0.0%      32.5%        0.175  review
#>   Top share is the proportion of responses in the most common category.
#> 
#> Flagged items
#>   - EF1 (review): `EF1` has a corrected item-rest correlation of r = 0.17 (n =
#>     400), below the teaching reference.
#>   - EF2 (review): `EF2` has a negative corrected item-rest correlation (r =
#>     -0.50, n = 400). It is declared reverse-keyed, and this is the sign such
#>     an item shows before it is recoded: recoded on the declared 1 to 5 scale,
#>     its item-rest correlation is r = 0.50. The data were not recoded.
#>   - EF3 (review): `EF3` has a corrected item-rest correlation of r = 0.21 (n =
#>     400), below the teaching reference.
#>   - EF4 (review): `EF4` has a corrected item-rest correlation of r = 0.17 (n =
#>     400), below the teaching reference.
#>   - EF6 (review): `EF6` has a corrected item-rest correlation of r = 0.18 (n =
#>     400), below the teaching reference.
#> 
#> Flags are review aids, not decisions to keep or delete an item.

EF2 correlates negatively with the rest of its facet. The audit says what that most likely means: the item is declared reverse-keyed, and recoded as declared its correlation would be positive. The negative sign is a coding matter, not evidence against the item.

Left as answered, the item also makes the rest of the facet look bad. Every other EF item’s correlation with the rest of the scale includes EF2 running the wrong way, so all of them fall below the teaching reference. A screen of data that are not yet recoded cannot say much about any item in the facet.

nomologR never recodes data. Recoding is the researcher’s step, and it belongs in the analysis code where it can be seen.

Recoding as content review declared

walk <- nomo_demo_walkthrough
reversed <- handoff$item_evidence$item[handoff$item_evidence$keying %in% -1]
walk[reversed] <- 6L - walk[reversed]

The responses are now recoded, so the run should not treat any item as still needing it. reverse = character(0) says so, in place of the handoff’s keying:

run <- nomo_run(
  walk,
  scales = handoff,
  settings = list(screen = list(reverse = character(0), scale_range = c(1, 5)))
)
ef <- run$results$screen$EF
summary(ef)
#> <nomo_screen summary> Item and data audit
#> Cases: 400 | Items: 5 | Flags: 1 review, 0 concern
#> Items with missing responses: 0 | Constant: 0 | All missing: 0
#> Relationship eligible: 5
#> 
#> Item review
#>   Item  Type      Missing  Top share  Item-rest r  Flag
#>   EF1   discrete     0.0%      34.8%        0.559
#>   EF2   discrete     0.0%      28.5%        0.505
#>   EF3   discrete     0.0%      78.5%        0.339
#>   EF4   discrete     0.0%      33.8%        0.228  review
#>   EF6   discrete     0.0%      32.5%        0.470
#>   Top share is the proportion of responses in the most common category.
#> 
#> Flagged items
#>   - EF4 (review): `EF4` has a corrected item-rest correlation of r = 0.23 (n =
#>     400), below the teaching reference.
#> 
#> Flags are review aids, not decisions to keep or delete an item.

Recoded, EF2 correlates 0.50 with the rest of its facet, and one EF item remains flagged: EF4, at 0.23. Eighteen of twenty judges placed EF4 in Effort Regulation, and they were right about its wording. Keeping to a self-set study schedule reads as effort. In the responses, though, it shares almost nothing with the other items. This is what content review cannot see.

EF3 is not flagged here (0.34), but its distribution stands out: the “Top share” column shows most answers in a single category. Almost everyone finishes the assignments that count toward their grade. The next stage shows what that costs.

The measurement model

An exploratory factor analysis of the ten recoded items, with the two facets the panel defined:

efa <- nomo_efa(walk[handoff$items], factors = 2)
summary(efa)
#> <nomo_efa summary> Exploratory factor analysis
#> Cases: 400 | Items: 10 | Factors: 2 (researcher specified)
#> Correlation: pearson | Extraction: minres | Rotation: oblimin
#> Supporting adequacy: KMO 0.839 | Bartlett chi-square(45) = 810.20, p < .001
#> 
#> Item structure
#>   Item  Factor  Loading  Next factor  Loading  Communality  Flag
#>   EF1   F1        0.778  F2            -0.067        0.564
#>   EF2   F1        0.654  F2             0.014        0.436
#>   EF3   F1        0.370  F2             0.081        0.170  concern
#>   EF4   F1        0.279  F2             0.021        0.084  concern
#>   EF6   F1        0.567  F2             0.045        0.346  review
#>   TF1   F2        0.695  F1             0.020        0.495
#>   TF2   F2        0.561  F1             0.095        0.370  review
#>   TF3   F2        0.600  F1            -0.013        0.354  review
#>   TF4   F1        0.444  F2             0.320        0.425  review
#>   TF6   F2        0.640  F1            -0.084        0.369  review
#> 
#> Flagged items
#>   - EF3 (concern): primary loading |0.37| is below the 0.40 teaching
#>     reference; communality 0.17 is below the 0.40 teaching reference
#>   - EF4 (concern): primary loading |0.28| is below the 0.40 teaching
#>     reference; communality 0.08 is below the 0.40 teaching reference
#>   - EF6 (review): communality 0.35 is below the 0.40 teaching reference
#>   - TF2 (review): communality 0.37 is below the 0.40 teaching reference
#>   - TF3 (review): communality 0.35 is below the 0.40 teaching reference
#>   - TF4 (review): secondary loading |0.32| meets/exceeds the 0.30
#>     cross-loading reference
#>   - TF6 (review): communality 0.37 is below the 0.40 teaching reference
#> 
#> Factor correlations
#>   Factor 1  Factor 2      r
#>   F1        F2        0.441
#> 
#> Largest residual correlations
#>   Off-diagonal RMSR: 0.021
#>   Item 1  Item 2  Residual
#>   EF4     TF6       -0.063
#>   EF2     TF2       -0.046
#>   EF4     TF4        0.040
#>   TF3     TF4        0.037
#>   EF2     TF6        0.035
#> 
#> Numerical references trigger inspection, not automatic deletion or hidden
#> refitting.

Three items stand apart:

  • EF4 has almost no common variance, as the screen suggested.
  • EF3 loads weakly as well. Its answers are piled at the top of the scale, and a variable that barely varies cannot correlate strongly with anything.
  • TF4 loads on both factors. Working through a task without taking breaks is effort as much as focus. The panel saw one facet, and the data see two.

A confirmatory model with the panel’s structure locates the same strain:

cfa <- nomo_cfa(nomo_model(handoff$scales), data = walk)
cfa
#> <nomo_cfa> Confirmatory factor analysis
#> Cases: 400 of 400 used | Estimator: ML | Converged: yes
#> Fit: CFI 0.941 | TLI 0.922 | RMSEA 0.058 | SRMR 0.056
#> Loadings: 10 | Flags: 2 review, 0 concern
#> No parameter was freed and no model was refit automatically. summary() shows
#> the evidence.
head(nomo_table(cfa, "modification_indices"), 3)
#> # A tibble: 3 × 8
#>   lhs   op    rhs      mi    epc sepc.lv sepc.all sepc.nox
#>   <chr> <chr> <chr> <dbl>  <dbl>   <dbl>    <dbl>    <dbl>
#> 1 EF    =~    TF4   53.6   0.786   0.640    0.529    0.529
#> 2 EF1   ~~    TF4   11.4   0.164   0.164    0.214    0.214
#> 3 EF    =~    TF6    8.35 -0.321  -0.262   -0.210   -0.210

The largest modification index asks for TF4 to load on Effort Regulation too. It is the same item the exploratory analysis flagged, found from the other direction. As always in this package, the index locates strain; it does not free the parameter.

Two cohorts

The respondents came from two cohorts. Before comparing their means, check that the items measure the same thing in both:

inv <- nomo_invariance(
  nomo_model(handoff$scales),
  data = walk,
  group = "cohort",
  levels = c("configural", "metric", "scalar")
)
summary(inv)
#> <nomo_invariance summary> Measurement invariance
#> Indicators: continuous | Groups: A, B
#> Levels completed: configural -> metric -> scalar
#> 
#> Identification and sequence
#>   Continuous indicators use the conventional configural, metric, scalar, and
#>   strict sequence.
#> 
#> Fit by level
#>   Level       Constraints           Chi-square  df       p    CFI  RMSEA   SRMR
#>   configural  none                      114.09  68  < .001  0.941  0.058  0.057
#>   metric      loadings                  121.68  76  < .001  0.941  0.055  0.064
#>   scalar      loadings, intercepts      153.60  84  < .001  0.911  0.064  0.071
#> 
#> Changes from the preceding level
#>   Level   CFI change  RMSEA change  SRMR change  LRT chi-square  df       p
#>   metric      +0.001        -0.003       +0.007            7.58   8    .475
#>   scalar      -0.031        +0.010       +0.008           31.92   8  < .001
#> 
#> Largest equality-constraint score diagnostics (diagnostic only)
#>   Level   Constraint                    Score  df       p
#>   scalar  Intercept: TF6 (A vs. B)      24.30   1  < .001
#>   scalar  Intercept: TF4 (A vs. B)       6.40   1    .011
#>   metric  Loading: TF -> TF6 (A vs. B)   2.71   1    .100
#>   scalar  Intercept: EF1 (A vs. B)       2.55   1    .110
#>   metric  Loading: EF -> EF3 (A vs. B)   2.32   1    .127
#>   scalar  Intercept: TF2 (A vs. B)       2.24   1    .135
#>   scalar  Loading: EF -> EF3 (A vs. B)   2.23   1    .136
#>   metric  Loading: TF -> TF1 (A vs. B)   2.07   1    .150
#>   scalar  Loading: TF -> TF1 (A vs. B)   1.69   1    .194
#>   scalar  Intercept: EF2 (A vs. B)       1.49   1    .222
#> 
#> No single delta-CFI, delta-RMSEA, delta-SRMR, chi-square difference, or score
#> diagnostic is treated as a universal invariance rule.

Equal loadings hold; equal intercepts cost fit. The largest score diagnostic is TF6’s intercept. Putting a phone away while studying is a specific behavior, and the two cohorts faced different classroom rules about phones. An item can be content-valid and still not comparable across groups. Whether to release that intercept (see nomo_partial()) is a researcher decision, and it needs a rationale such as that one.

Where the two stages disagree

Item Content review Empirical evidence What it teaches
EF4 Carried: 18 of 20 judges Almost no common variance Evidence against the item that no panel could see. Revise it, or drop it with a rationale.
EF3 Carried: 18 of 20 judges Weak loading, from answers piled at the ceiling The screen is not the final word. EF3 is the only item about finishing required work; dropping it narrows the domain the panel defined.
TF4 Carried: 16 of 20 judges Loads on both facets A question about content. Revise the wording, or model the cross-loading with a rationale.
TF6 Carried: 18 of 20 judges Intercept differs by cohort Explain the difference before comparing cohort means.
EF2, TF2 Carried; reverse-worded Negative until recoded Coding, not evidence.
EF5, TF5 Held back Not screened They remain in nomo_demo_walkthrough, so you can see what keeping them would have done.

Neither stage overrules the other. EF4 is the case for not trusting a panel alone, and EF3 is the case for not trusting a screen alone.

What each stage establishes

Content review establishes whether items represent the construct as it was defined: whether the domain is covered, and whether each item belongs where it was written to belong. It cannot see how an item varies, what it shares with the others, or how it behaves in different groups.

An empirical screen establishes how the responses behave. It cannot see whether the domain is covered. A screen would happily keep a scale of ten well-behaved items that all measured one corner of the construct.

nomologR keeps both in the record. A report from a run that started from a handoff opens with the content review, quoting the panel’s decisions in contentvalidR’s words, and each empirical flag carries its own explanation. Neither is presented as a verdict on the other:

nomo_report(run, file = "walkthrough-report.html")