Compares observed-minus-expected scores between groups, or describes linked subgroup facet estimates after refitting. Residual differences do not isolate differential functioning and are returned without tests or classifications.
analyze_dif() is retained for compatibility with earlier package versions.
In many-facet workflows, prefer analyze_dff() as the primary entry point.
Usage
analyze_dff(
fit,
diagnostics,
facet,
group,
data = NULL,
focal = NULL,
method = c("residual", "refit"),
min_obs = 10,
p_adjust = "holm"
)
analyze_dif(...)Arguments
- fit
Output from
fit_mfrm().- diagnostics
Output from
diagnose_mfrm().- facet
Character scalar naming the facet whose elements are tested for differential functioning (for example,
"Criterion"or"Rater").- group
Character scalar naming the column in the data that defines the grouping variable (e.g.,
"Gender","Site").- data
Optional data frame containing at least the group column and the same person/facet/score columns used to fit the model. If
NULL(default), mfrmr tries to recover the data fromfit$prep$data. That slot only holds the columns thatfit_mfrm()actually modelled; if the grouping column was not among them (common for DIF screening), pass the original data frame viadata = <df>explicitly. The same applies when the fit object has been serialized without the prep slot.- focal
Optional character vector of group levels to treat as focal. If
NULL(default), all pairwise group comparisons are performed.- method
Analysis method:
"residual"(default) uses the fitted model's residuals without re-estimation;"refit"re-estimates the model within each group subset. The residual method is faster and avoids additional subgroup fits; neither method establishes adequate sample size.- min_obs
Minimum number of observations per cell (facet-level x group). Cells below this threshold are flagged as sparse and their statistics set to
NA. Default10.- p_adjust
Adjustment for refit screening tail areas; default
"holm". Retained but unused for residual comparisons, which return no p-values.- ...
Passed directly to
analyze_dff().
Value
An object of class mfrm_dff (with compatibility class mfrm_dif) with:
dif_table: data.frame of differential-functioning contrasts.cell_table: (residual method) per-cell detail table.summary: counts by method-appropriate screening classification.group_fits: (refit method) per-group facet estimates.gpcm_boundary: forGPCMfits, a capability-boundary table.config: list with facet, group, method, min_obs, p_adjust settings.
Details
Differential facet functioning (DFF) occurs when the difficulty or severity of a facet element differs across subgroups of the population, after controlling for overall ability. In an MFRM context this generalises classical DIF (which applies to items) to any facet: raters, criteria, tasks, etc.
Differential functioning is a threat to measurement fairness: if Criterion 1 is harder for Group A than Group B at the same ability level, the measurement scale is no longer group-invariant.
Differences between group ability distributions are not themselves DFF. Residual screens inherit the fitted model's ability and population assumptions. With fixed-standard-normal RSM/PCM MML, subgroup refits retain that population assumption; linking anchors do not estimate subgroup ability distributions. A residual difference can therefore reflect an inadequately represented group difference as well as differential facet functioning. For MML, observation expectations are evaluated at each Person's EAP ability, rather than integrated over the conditional ability distribution. This nonlinear substitution can produce different mean residuals across groups even under a correctly specified no-DFF model. A zero mean-residual contrast and absence of differential functioning are therefore distinct null hypotheses; correcting the SE alone does not make them equivalent.
Two methods are available:
Residual method (method = "residual"): Uses the existing fitted
model's observation-level residuals. For each facet-level \(\times\)
group cell, the observed and expected score sums are aggregated and
a standardized residual is computed as:
$$z = \frac{\sum (X_{obs} - E_{exp})}{\sqrt{\sum \mathrm{Var}}}$$
Pairwise contrasts compare the mean observed-minus-expected score in Group1
minus that in Group2. Positive values mean higher residual scores in Group1,
not greater rater leniency for that group. Differences use score units.
StdResidual is a descriptive scaling by model response variance, not a
t-statistic. No residual-contrast SE, confidence interval, p-value or
positive/negative classification is provided. Compatibility columns SE,
t, df, p_value and p_adjusted contain NA.
Refit method (method = "refit"): Subsets the data by group, refits
the MFRM model within each subset, anchors all non-target facets back to
the baseline calibration when possible, and compares the resulting
facet-level estimates. When linking, convergence, model-based MML precision,
and sparsity screens all pass, a conditional plug-in Welch statistic is also
shown:
$$t = \frac{\hat{\delta}_1 - \hat{\delta}_2}
{\sqrt{SE_1^2 + SE_2^2}}$$
This provides group-specific point estimates on a common scale when linking
anchors are available. However, the plug-in standard error conditions on
the baseline anchors and omits baseline-anchor uncertainty and cross-refit
covariance. Refit statistics therefore remain screening evidence and set
FormalInferenceEligible, SupportsFormalInference,
PrimaryReportingEligible, and ETS_Eligible to FALSE.
The subgroup fits replay the baseline response family, resolved score range,
step/slope facets, weighting, optimizer, MML engine, and numerical controls;
the non-target anchors are the intentional linking change. Refit analysis
fails closed when the baseline uses a user-specified active latent-regression
population model, facet interactions, or group-anchor constraints, because
those structures do not yet have a complete subgroup-replay and linking
contract. The automatically generated intercept-only population model used
for default GPCM-MML scale identification is replayed and is not treated as
a substantive covariate model.
When facet refers to an item-like facet (for example Criterion), this
recovers the familiar DIF case. When facet refers to raters or
prompts/tasks, the same machinery supports DRF/DPF-style analyses.
Refit contrasts are not assigned ETS A/B/C labels. Even when subgroup point estimates share a linked logit scale, a validated joint, bootstrap, or replicate covariance contract is required before formal refit inference.
Refit screening tail areas are adjusted using Holm's step-down procedure by
default, jointly across the returned facet-level/group-pair rows in this
call. Holm's family-wise error control requires valid unadjusted p-values;
the approximate screening tail areas here have not been shown to meet that
requirement. Adjustment therefore does not establish an error-rate guarantee
or formal inference eligibility. Alternative adjustments can be specified
via p_adjust; see stats::p.adjust() for their assumptions.
Choosing a method
Use method = "residual" to describe group differences in model residuals
without fitting subgroup models. This does not test differential functioning.
Use method = "refit" when you specifically want group-specific parameter
estimates and can tolerate extra computation. Agreement between methods is
not guaranteed by a universal per-group sample-size threshold: stability and
detection depend jointly on effect size, category support, response-pattern
overlap, linking strength, slope heterogeneity, and estimator behavior.
min_obs is only a cell-computability/sparsity guard; it is not evidence of
power, parameter stability, or sample-size adequacy.
Interpreting output
$dif_table: one row per facet-level x group-pair with contrast, a residual mean difference or linked subgroup difference. Residual rows have no SE, test or binary classification. IncludesMethod,N_Group1,N_Group2,EffectMetric,ClassificationSystem,ContrastBasis,SEBasis,StatisticLabel,ProbabilityMetric,DFBasis,ReportingUse,PrimaryReportingEligible, andsparsecolumns.$cell_table: (residual method only) per-cell detail with N, ObsScore, ExpScore, ObsExpAvg, StdResidual and an interpretation note.$summary: available/unavailable residual comparisons or linked- screening and insufficient-linking rows (method = "refit").$group_fits: (refit method only) list of per-group facet estimates and subgroup linking diagnostics.$gpcm_boundary: forGPCMfits, a capability-boundary table marking the DFF/DIF output as caveated screening evidence.
GPCM boundary
For GPCM, DFF/DIF rows are available as slope-aware screening
evidence over the fitted expected-score and residual scale. Keep
residual-method contrasts and interaction cells in screening language.
Refit contrasts require explicit subgroup linking and precision support for
conditional screening, but remain in screening language.
Saved results
Residual results, summaries and reports saved by earlier versions must be recomputed with this function using the fitted model and original data. This does not require refitting the model. Previously exported tables and figures should also be regenerated.
Typical workflow
Fit a model with
fit_mfrm(). ForRSM/PCMfairness review, prefermethod = "MML".Run
diagnose_mfrm()and, forRSM/PCM, preferdiagnostic_mode = "both"so legacy and strict marginal screens remain visible together.Run
analyze_dff(fit, diagnostics, facet = "Criterion", group = "Gender", data = my_data).Inspect
$dif_tablefor flagged levels and$summaryfor counts.Use
dif_interaction_table()when you need cell-level diagnostics.Use
plot_dif_heatmap()ordif_report()for communication.
Examples
# \donttest{
toy <- load_mfrmr_data("example_bias")
fit <- fit_mfrm(toy, "Person", c("Rater", "Criterion"), "Score",
method = "MML", model = "RSM", quad_points = 7, maxit = 30)
diag <- diagnose_mfrm(fit, residual_pca = "none", diagnostic_mode = "both")
dff <- analyze_dff(fit, diag, facet = "Rater", group = "Group", data = toy)
dff$summary
#> # A tibble: 2 × 2
#> Classification Count
#> <chr> <int>
#> 1 Residual contrast 4
#> 2 Unavailable 0
# Read the residual mean difference and the number of observations per group.
head(dff$dif_table[, c("Level", "Group1", "Group2", "Contrast",
"N_Group1", "N_Group2")])
#> # A tibble: 4 × 6
#> Level Group1 Group2 Contrast N_Group1 N_Group2
#> <chr> <chr> <chr> <dbl> <int> <int>
#> 1 R01 A B 0.146 48 48
#> 2 R02 A B -0.152 48 48
#> 3 R03 A B -0.164 48 48
#> 4 R04 A B 0.00347 48 48
# The residual contrast is an observed-minus-expected average contrast
# between groups. It is useful for screening, but it is not an ETS
# A/B/C logit-delta classification.
dff_refit <- analyze_dff(fit, diag, facet = "Rater", group = "Group",
data = toy, method = "refit")
unique(dff_refit$dif_table$ClassificationSystem)
#> [1] "descriptive"
# Look for: "descriptive". Linked refit contrasts remain screening-only
# because their plug-in uncertainty omits anchor uncertainty and
# cross-refit covariance.
sc <- subset_connectivity_report(fit, diagnostics = diag)
plot(sc, type = "design_matrix", draw = FALSE)
if ("ScaleLinkStatus" %in% names(dff_refit$dif_table)) {
unique(dff_refit$dif_table$ScaleLinkStatus)
}
#> [1] "weak_link"
# Look for: "linked" in `ScaleLinkStatus` confirms the focal and
# reference groups share enough common elements for a comparable
# contrast; "demoted_*" rows lose linking under the refit branch
# and should be read as exploratory.
# }
