
Compare two or more fitted MFRM models
Source:R/api-estimation.R, R/api-extended-comparison.R
compare_mfrm.RdProduce a side-by-side comparison of multiple fit_mfrm() results using
AIC, Person-based BIC, Sclove SABIC, log-likelihood, and free-parameter
counts. An ordinary RSM paired with a testlet or shared-rater fit instead
follows a descriptive facet-effect route, with no automatic ranking.
For ordinary models, when exactly
two models are supplied and the current conservative nesting review passes,
a likelihood-ratio test is included.
Arguments
- ...
Two or more
mfrm_fitobjects, or exactly one ordinarymfrm_fitand onefit_mfrm_testlet()orfit_mfrm_random_rater()result. The latter route is descriptive; see the extended-model section.- labels
Optional character vector of labels for each model. If
NULL, labels are generated from model/method combinations.- warn_constraints
Logical. If
TRUE(the default), emit a warning when models use different score coding or constraint settings. Ranking is suppressed whenever this basis differs, regardless of whether the warning is displayed.- nested
Logical. Set to
TRUEonly when the supplied models are known to be nested and fitted with the same likelihood basis on the same observations. The default isFALSE, in which case no likelihood-ratio test is reported. WhenTRUE, the function still runs a conservative structural nesting review and computes the LRT only for supported nesting patterns.- response_diagnostics
Optional list of two saved
mfrm_response_diagnostics()outputs in the same order as the fits. Supported for an ordinary RSM paired with a testlet or shared-rater fit. This adds descriptive predictive comparisons without computing integrals.- person_scores
Optional list of two saved
score_mfrm_persons()results in fit order for an ordinary versus extended RSM comparison. Both must score the same complete source roster, return the same Person IDs and use the same conditional interval level. No scoring is run here.- object, x
An extended-model comparison returned by
compare_mfrm().
Value
With one extended model, an mfrm_extended_comparison containing models,
checks, centered effects, notes, and source/omission metadata.
With ordinary models only, an object of class mfrm_comparison (named list) with:
table: data.frame of model-level statistics includingLogLik,Deviance,Npar(and equal compatibility aliasnpar),AIC,BIC,SABIC, criterion deltas and candidate-set weights,ResponseRows,WeightedResponseTotal,Persons,ICSampleSize, formula and integration identities, integration tier/selectability, stored-value consistency fields, convergence fields,ICComparable,SABICComparable, andConstraintComparable.lrt: data.frame with likelihood-ratio test result (only when two models are supplied andnested = TRUE). ContainsChiSq,df,p_value.evidence_ratios: data.frame of pairwise Akaike-weight ratios (Model1, Model2, EvidenceRatio).NULLwhen weights cannot be computed.preferred: named list with the preferred model label by each criterion.comparison_basis: list describing whether IC and LRT comparisons were considered comparable. Includes a conservativenesting_reviewpluslrt_status/lrt_reasonso withheld LRTs are explicit rather than silently absent.
Details
GPCM MML and RSM/PCM MML with an estimated normal population use a separate
local-solution check for information criteria, without refitting. It checks
the retained likelihood, terminal gradient and positive unregularized
observed information. ICFitEligible, ICFitBasis and ICFitReview record
that decision; ICComparable also requires the common data/likelihood and
integration checks. This can allow IC comparison while InferenceReady
remains false for intervals or tests. It shares the joint-information
calculation and workspace budget with confint.mfrm_fit(). Unknown, unstable or unavailable results
remain ineligible with a reason. A local check does not prove a global
maximum or negligible integration error; examine starting-value and
mml_quadrature_sensitivity() results when the decision is close.
Positive but ill-conditioned information may pass with a caution after
numerical refinement, unregularized inversion and curvature-scaled gradient
checks. A warning is emitted and retained in ICFitCaution, ICFitReview
and the requested LRT's interpretation; it does not establish
finite-sample accuracy for model ranking or tests.
Models should be fit to the same data (same rows, same person/facet columns) for the comparison to be meaningful. The function checks that observation counts match and warns otherwise.
Information-criterion ranking is reported only when all candidates are
eligible solutions under the package's current MML contract, use the same
prepared observations, score coding, constraints, formula contract, and
integration-evaluation identity, are eligible under the weighting policy,
and have a selectable integration tier. For fixed-facet MML,
ICSampleSize is the number of independent
Persons and ICSampleSizeBasis is "person_count"; response rows and their
weighted total are retained separately. Explicit all-unit weights are
eligible, whereas every non-unit observation-weight fit fails closed under
the current package contract.
Canonical AIC, BIC, and SABIC are recomputed from the retained
objective, optimizer-vector dimension, and Person count. A stale stored
value, a legacy object without the current contract identity, a JML fit, or
an ineligible weighted fit cannot enter ranking. Earlier raw values may be
shown only as explicitly labelled LegacyAIC and LegacyBIC fields.
Integration selection guard: raw canonical criteria are retained at
every valid MML quadrature count, but automatic comparison is screening-only
below 15 points and review-only at 15–30 points. ICSelectable and
ICIntegrationSelectable become TRUE at q>=31. q=31 is a starting grid,
not a universal guarantee: close or consequential comparisons should be
reevaluated at a denser shared grid (normally q>=61). The guard also applies
to nested = TRUE, so a coarse-grid LRT cannot bypass it.
Nesting: Two models are nested when one is a special case of the other obtained by imposing equality constraints. The most common nesting in MFRM is RSM (shared thresholds) inside PCM (item-specific thresholds). Models that differ only in estimation method (MML vs JML) on the same specification are not nested in the usual sense, and their information criteria do not share the common MML contract. Do not use either the LRT or IC ranking as a direct cross-method comparison.
In the current mfrmr model space, the automatic nesting review is
intentionally conservative. It currently supports the following
restrictions under shared data and shared constraints:
RSMnested insidePCMwhen thePCMfit has an explicitstep_facet;PCMnested insideGPCMwith the same step facet, population design, other facet/step constraints and interactions. The only additional parameters must be G-1 relative log-slope contrasts for G slope levels. The GPCM slope owner may differ from the shared step owner;same-family additive-vs-interaction comparisons when the smaller fit's
facet_interactionsset is a subset of the larger fit's set.
Cross-method comparisons, comparisons that change anchors/dummying/centering, and same-family comparisons that do not add fixed interaction terms are not automatically promoted to LRT claims.
Automatic nesting also requires a shared population specification: either both fits use the fixed standard-normal population or both use the same estimated-normal design, with the same declared columns and design values aligned by Person. Coefficient and variance estimates may differ. Row and column permutations are aligned; other recodings, changed designs, and unavailable design metadata require separate review. Passing this check does not grant inference readiness to an estimated-population fit. For a matched PCM/GPCM pair, the local-solution checks described above supply the numerical requirement independently of slope-interval availability.
The likelihood-ratio test (LRT) is reported only when exactly two
models are supplied, nested = TRUE, the structural nesting review passes, and the
difference in the number of parameters is positive:
$$\Lambda = -2 (\ell_{\mathrm{restricted}} - \ell_{\mathrm{full}}) \sim \chi^2_{\Delta p}$$
The LRT is asymptotically valid only under its regularity assumptions. With small samples or boundary/singular conditions, the reference chi-square p-value can be incorrect. At large Person counts, a very small practical improvement can also become statistically significant. Read the p-value as formal nested-fit evidence, not as an automatic practical model preference or evidence that a subscore is useful.
Ordinary versus extended RSMs
Exactly one ordinary RSM MML fit and one testlet/shared-rater fit return an
mfrm_extended_comparison, with models, checks, effects and notes.
The information-criterion and LRT sections describe ordinary mfrm_fit
comparisons only. This route does not rank models, compute difference
intervals, or provide an AIC/BIC preference. nested = TRUE is refused.
Both fits must use identical observed rating events, including repeated
events, category coding, unit weights and fixed facets except for the
deliberately changed rater treatment. Use unanchored additive severity
facets, noncenter_facet = "Person", no facet shrinkage, and matching
person/score column roles. Category recoding must preserve the same
consecutive adjacent-category scale. These checks cannot be disabled by
warn_constraints = FALSE.
Match the population assumption: ordinary population_formula = ~1 with
both extension defaults estimates a common normal ability SD; the ordinary
default and extension person_sd = 1 impose known N(0,1). Covariate-dependent
populations and fixed nonunit SDs have no matching ordinary route here.
Population intercepts and free step locations may use different origins;
they are recorded, not compared as substantive mean differences.
The source roster and omitted counts must agree. New ordinary fits retain
prep$omitted_data and prep$omitted_input_rows. When scores were omitted,
their identities must match, including repeated events. Older fits without
this provenance require refitting from the complete assigned-score roster.
Missing IDs, excluded weights and population-data omissions are unsupported.
Effects are centered at the unweighted mean of the same complete set of
levels within each facet. This preserves pairwise contrasts while removing
arbitrary facet origins. SourceReference/SourceComparison retain raw
values and CenterReference/CenterComparison retain the subtracted means.
Difference is comparison minus reference. Conditional rater modes and
fixed coefficients are labeled separately; shrinkage is not evidence of
greater accuracy or rater quality. Failed numerical checks or a missing
level estimate withhold affected differences without dropping source rows.
Native inference readiness remains explicit even when descriptive numerical
checks pass. No SE or confidence interval for model differences is implied.
With response_diagnostics, responses retains matched category
probabilities, means, full mixture variances and grouped descriptive
Infit/Outfit. Both saved diagnostics must use the same probability target,
selected event multiset and shared group_by identifier columns. Matching
uses event contents, including repeated-event multiplicities and missing
scores, rather than assuming row numbers agree. Identical repeated events
match in source occurrence order. responses$events and responses$rows
retain the matching and original row numbers. Missing or unavailable rows
remain explicit and differences involving them are withheld.
All observed source events condition each model; selecting a Person's rows
does not remove other Persons from shared-rater inference. These are
same-data descriptive comparisons with calibration held fixed, not
held-out predictive performance. A smaller Infit/Outfit is not evidence
of improvement. Do not substitute ordinary plug-in diagnostics or their
reference cutoffs. See plot.mfrm_extended_comparison() for predictive
paired and difference displays.
With person_scores, persons$table compares conditional EAPs after
subtracting each fitted population mean, retaining original estimates,
origins, posterior SDs and aligned conditional endpoints. Unit Rasch slopes
retain the logit unit; centering removes the arbitrary origin, not shrinkage
or changes in the assumed population variance. Prior-only and unavailable
differences are withheld. Separate conditional intervals are not intervals
for model differences or tests of Person differences. Use metric = "person"
in plot.mfrm_extended_comparison() for a paired/difference display.
The default facet-only output does not compare Person scores, steps, predictive Infit/Outfit or replacement-rater predictions. For testlets, zero local variance retains the fixed facets; zero shared-rater variance removes rater differences and is not the fixed-rater model. Numerical zero-reduction agreement does not establish model adequacy or variance-boundary inference.
Use plot(comparison, style = "paired") for matched effects, or
style = "difference" for differences versus means. See
plot.mfrm_extended_comparison() for display and accessibility controls.
Attach saved results with mfrm_results(extended_fit, comparison = comparison)
for static reports, tables, figures and replay. No comparison/report step
fits, scores or resamples either model.
Information-criterion diagnostics
In addition to the canonical criteria, the function computes:
Delta_AIC / Delta_BIC / Delta_SABIC: difference from the minimum value in the supplied candidate set. A Delta < 2 is typically considered negligible; 4–7 suggests moderate evidence; > 10 indicates strong evidence against the higher-scoring model (Burnham & Anderson, 2002).
AkaikeWeight / BICWeight / SABICWeight: relative candidate-set weights derived from
exp(-0.5 * Delta)and normalized only across the supplied models. They are not posterior probabilities, probabilities of model truth, or guarantees that the candidate set is adequate.Evidence ratios: pairwise ratios of Akaike weights, quantifying relative support within this candidate set. A ratio of 5 means only that one normalized Akaike weight is five times the other; it does not mean that one model is five times more likely to be true.
AIC, BIC, and SABIC answer different approximation/penalty questions. SABIC
is a sensitivity criterion, not a universal tie-breaker. At 22 or fewer
Persons its Sclove penalty is non-positive, so SABICSelectable and
SABICComparable are FALSE and no automatic SABIC preference is returned.
What this comparison means
compare_mfrm() is a same-basis model-comparison helper. Its strongest
claims apply only when the models were fit to the same response data,
under a compatible likelihood basis, and with compatible constraint
structure.
What this comparison does not justify
Do not treat AIC/BIC/SABIC differences as primary evidence when
table$ICComparableisFALSE.Do not infer numerical adequacy merely because all models share the same coarse quadrature identity. Below q=31, raw criteria are diagnostic only and automatic deltas, weights, preferences, and LRT are suppressed.
Do not interpret the LRT unless
nested = TRUEand the structural nesting review incomparison_basis$nesting_reviewpasses.Same-family additive-vs-interaction fits are considered nested only when all other structural settings match and the smaller model's
facet_interactionsset is a subset of the larger model's set.Do not assume that
nested = TRUEoverrides the package's conservative nesting boundary; unsupported relations remain unsupported.PCM is the unit-slope reduction of an aligned GPCM.
nested = TRUEtests equal relative slopes only after population, constraint, free-dimension and numerical checks pass. The relation is recorded asPCM_in_GPCM. Default PCM uses a fixed standard-normal population, whereas default GPCM estimates its mean and variance: changing onlymodeldoes not isolate slope differences. Supplypopulation_formula = ~1and the sameperson_datato both fits for an estimated-normal comparison. Unit slopes are interior positive values; the reference is ordinary asymptotic chi-square with G-1 degrees of freedom, not a boundary mixture. IC eligibility alone does not establish nesting or qualify slope intervals.Do not compare models fit to different datasets, different score codings, or materially different constraint systems as if they were commensurate.
At large Person counts, a small systematic likelihood gain can dominate an IC difference; boundary or singular fits can also invalidate routine asymptotics. Evaluate effect size, stability, interpretability, and score consequences separately.
Interpreting output
Lower AIC/BIC values indicate relative support only when
table$ICComparableisTRUE; use SABIC ranking only whentable$SABICComparableis alsoTRUE.A significant LRT p-value is formal evidence against the restricted model under the stated assumptions; practical gain, score utility, and dimensional interpretation require separate checks.
preferredidentifies the minimum criterion within the supplied candidate set; it is not an unconditional model verdict.evidence_ratiosgives pairwise Akaike-weight ratios (returned only when Akaike weights can be computed for at least two models).When comparing more than two models, interpret evidence ratios cautiously—they do not adjust for multiple comparisons.
How to read the main outputs
table: first-pass comparison table; start withICComparable,ICSelectable,ICIntegrationTier,Model,Method,ICStatus,AIC,BIC, andSABIC.comparison_basis: records whether IC and LRT claims are defensible for the supplied models. Inspectcomparison_basis$nesting_review$relationandreasonbefore reading any LRT output.lrt: nested-model test summary, present only when the requested and reviewed conditions are met.preferred: candidate preferred by each criterion when those summaries are available.
Recommended next step
Inspect comparison_basis before writing conclusions. If comparability is
weak, treat the result as descriptive and revise the model setup (for
example, explicit step_facet, common data, or common constraints) before
using IC or LRT results in reporting.
Typical workflow
Fit two models with
fit_mfrm()(e.g., PCM and GPCM) on the same prepared rows, explicit step owner, constraints, and MML quadrature setting. Use at least 31 common quadrature points for selectable ICs.Compare with
compare_mfrm(fit_pcm, fit_gpcm)and start by checkingcomparison$table$ICComparable.Inspect
summary(comparison)for AIC/BIC/SABIC diagnostics and the reasons for withholding ranking or tests. GPCM IC ranking has separate solution checks; request the matched PCM/GPCM test withnested = TRUE.Use
build_weighting_review(fit_pcm, fit_gpcm)to inspect which selected facet levels and information shares were reweighted by the slopes.
References
Burnham, K. P., & Anderson, D. R. (2002). Model selection and multimodel inference: A practical information-theoretic approach (2nd ed.). Springer.
Akaike, H. (1974). A new look at the statistical model identification. IEEE Transactions on Automatic Control, 19(6), 716-723.
Schwarz, G. (1978). Estimating the dimension of a model. Annals of Statistics, 6(2), 461-464.
Sclove, S. L. (1987). Application of model-selection criteria to some problems in multivariate analysis. Psychometrika, 52(3), 333-343.
Examples
# \donttest{
# Load one dataset for both models
library(mfrmr)
toy <- load_mfrmr_data("example_operational")
# RSM: shared category thresholds
fit_rsm <- fit_mfrm(
data = toy,
person = "Person",
facets = c("Rater", "Criterion"),
score = "Score",
method = "MML",
model = "RSM"
)
# PCM: separate category thresholds for each criterion
fit_pcm <- fit_mfrm(
data = toy,
person = "Person",
facets = c("Rater", "Criterion"),
score = "Score",
method = "MML",
model = "PCM",
step_facet = "Criterion"
)
# Check that the fitted models can be compared
comparison <- compare_mfrm(fit_rsm, fit_pcm, labels = c("RSM", "PCM"))
comparison$table[, c("Label", "Converged", "ICComparable", "ICSelectable")]
#> # A tibble: 2 × 4
#> Label Converged ICComparable ICSelectable
#> <chr> <lgl> <lgl> <lgl>
#> 1 RSM TRUE TRUE TRUE
#> 2 PCM TRUE TRUE TRUE
# Smaller AIC/BIC indicates better relative support among these models
comparison$table[, c("Label", "AIC", "Delta_AIC", "BIC", "Delta_BIC")]
#> # A tibble: 2 × 5
#> Label AIC Delta_AIC BIC Delta_BIC
#> <chr> <dbl> <dbl> <dbl> <dbl>
#> 1 RSM 712. 0 729. 0
#> 2 PCM 713. 0.375 737. 7.86
# Delta is the difference from the lowest value of that criterion
# For close or consequential comparisons, check a denser shared grid with
# mml_quadrature_sensitivity() before choosing a model
# For a PCM/GPCM equal-slope test, see the matched-population example in
# vignette("mfrmr-gpcm-scope"); request compare_mfrm(..., nested = TRUE).
# }