Skip to contents

Produce a side-by-side comparison of multiple fit_mfrm() results using AIC, Person-based BIC, Sclove SABIC, log-likelihood, and free-parameter counts. An ordinary RSM paired with a testlet or shared-rater fit instead follows a descriptive facet-effect route, with no automatic ranking. For ordinary models, when exactly two models are supplied and the current conservative nesting review passes, a likelihood-ratio test is included.

Usage

compare_mfrm(
  ...,
  labels = NULL,
  warn_constraints = TRUE,
  nested = FALSE,
  response_diagnostics = NULL,
  person_scores = NULL
)

# S3 method for class 'mfrm_extended_comparison'
summary(object, ...)

# S3 method for class 'mfrm_extended_comparison'
print(x, ...)

Arguments

...

Two or more mfrm_fit objects, or exactly one ordinary mfrm_fit and one fit_mfrm_testlet() or fit_mfrm_random_rater() result. The latter route is descriptive; see the extended-model section.

labels

Optional character vector of labels for each model. If NULL, labels are generated from model/method combinations.

warn_constraints

Logical. If TRUE (the default), emit a warning when models use different score coding or constraint settings. Ranking is suppressed whenever this basis differs, regardless of whether the warning is displayed.

nested

Logical. Set to TRUE only when the supplied models are known to be nested and fitted with the same likelihood basis on the same observations. The default is FALSE, in which case no likelihood-ratio test is reported. When TRUE, the function still runs a conservative structural nesting review and computes the LRT only for supported nesting patterns.

response_diagnostics

Optional list of two saved mfrm_response_diagnostics() outputs in the same order as the fits. Supported for an ordinary RSM paired with a testlet or shared-rater fit. This adds descriptive predictive comparisons without computing integrals.

person_scores

Optional list of two saved score_mfrm_persons() results in fit order for an ordinary versus extended RSM comparison. Both must score the same complete source roster, return the same Person IDs and use the same conditional interval level. No scoring is run here.

object, x

An extended-model comparison returned by compare_mfrm().

Value

With one extended model, an mfrm_extended_comparison containing models, checks, centered effects, notes, and source/omission metadata. With ordinary models only, an object of class mfrm_comparison (named list) with:

  • table: data.frame of model-level statistics including LogLik, Deviance, Npar (and equal compatibility alias npar), AIC, BIC, SABIC, criterion deltas and candidate-set weights, ResponseRows, WeightedResponseTotal, Persons, ICSampleSize, formula and integration identities, integration tier/selectability, stored-value consistency fields, convergence fields, ICComparable, SABICComparable, and ConstraintComparable.

  • lrt: data.frame with likelihood-ratio test result (only when two models are supplied and nested = TRUE). Contains ChiSq, df, p_value.

  • evidence_ratios: data.frame of pairwise Akaike-weight ratios (Model1, Model2, EvidenceRatio). NULL when weights cannot be computed.

  • preferred: named list with the preferred model label by each criterion.

  • comparison_basis: list describing whether IC and LRT comparisons were considered comparable. Includes a conservative nesting_review plus lrt_status / lrt_reason so withheld LRTs are explicit rather than silently absent.

Details

GPCM MML and RSM/PCM MML with an estimated normal population use a separate local-solution check for information criteria, without refitting. It checks the retained likelihood, terminal gradient and positive unregularized observed information. ICFitEligible, ICFitBasis and ICFitReview record that decision; ICComparable also requires the common data/likelihood and integration checks. This can allow IC comparison while InferenceReady remains false for intervals or tests. It shares the joint-information calculation and workspace budget with confint.mfrm_fit(). Unknown, unstable or unavailable results remain ineligible with a reason. A local check does not prove a global maximum or negligible integration error; examine starting-value and mml_quadrature_sensitivity() results when the decision is close. Positive but ill-conditioned information may pass with a caution after numerical refinement, unregularized inversion and curvature-scaled gradient checks. A warning is emitted and retained in ICFitCaution, ICFitReview and the requested LRT's interpretation; it does not establish finite-sample accuracy for model ranking or tests.

Models should be fit to the same data (same rows, same person/facet columns) for the comparison to be meaningful. The function checks that observation counts match and warns otherwise.

Information-criterion ranking is reported only when all candidates are eligible solutions under the package's current MML contract, use the same prepared observations, score coding, constraints, formula contract, and integration-evaluation identity, are eligible under the weighting policy, and have a selectable integration tier. For fixed-facet MML, ICSampleSize is the number of independent Persons and ICSampleSizeBasis is "person_count"; response rows and their weighted total are retained separately. Explicit all-unit weights are eligible, whereas every non-unit observation-weight fit fails closed under the current package contract.

Canonical AIC, BIC, and SABIC are recomputed from the retained objective, optimizer-vector dimension, and Person count. A stale stored value, a legacy object without the current contract identity, a JML fit, or an ineligible weighted fit cannot enter ranking. Earlier raw values may be shown only as explicitly labelled LegacyAIC and LegacyBIC fields.

Integration selection guard: raw canonical criteria are retained at every valid MML quadrature count, but automatic comparison is screening-only below 15 points and review-only at 15–30 points. ICSelectable and ICIntegrationSelectable become TRUE at q>=31. q=31 is a starting grid, not a universal guarantee: close or consequential comparisons should be reevaluated at a denser shared grid (normally q>=61). The guard also applies to nested = TRUE, so a coarse-grid LRT cannot bypass it.

Nesting: Two models are nested when one is a special case of the other obtained by imposing equality constraints. The most common nesting in MFRM is RSM (shared thresholds) inside PCM (item-specific thresholds). Models that differ only in estimation method (MML vs JML) on the same specification are not nested in the usual sense, and their information criteria do not share the common MML contract. Do not use either the LRT or IC ranking as a direct cross-method comparison.

In the current mfrmr model space, the automatic nesting review is intentionally conservative. It currently supports the following restrictions under shared data and shared constraints:

  • RSM nested inside PCM when the PCM fit has an explicit step_facet;

  • PCM nested inside GPCM with the same step facet, population design, other facet/step constraints and interactions. The only additional parameters must be G-1 relative log-slope contrasts for G slope levels. The GPCM slope owner may differ from the shared step owner;

  • same-family additive-vs-interaction comparisons when the smaller fit's facet_interactions set is a subset of the larger fit's set.

Cross-method comparisons, comparisons that change anchors/dummying/centering, and same-family comparisons that do not add fixed interaction terms are not automatically promoted to LRT claims.

Automatic nesting also requires a shared population specification: either both fits use the fixed standard-normal population or both use the same estimated-normal design, with the same declared columns and design values aligned by Person. Coefficient and variance estimates may differ. Row and column permutations are aligned; other recodings, changed designs, and unavailable design metadata require separate review. Passing this check does not grant inference readiness to an estimated-population fit. For a matched PCM/GPCM pair, the local-solution checks described above supply the numerical requirement independently of slope-interval availability.

The likelihood-ratio test (LRT) is reported only when exactly two models are supplied, nested = TRUE, the structural nesting review passes, and the difference in the number of parameters is positive:

$$\Lambda = -2 (\ell_{\mathrm{restricted}} - \ell_{\mathrm{full}}) \sim \chi^2_{\Delta p}$$

The LRT is asymptotically valid only under its regularity assumptions. With small samples or boundary/singular conditions, the reference chi-square p-value can be incorrect. At large Person counts, a very small practical improvement can also become statistically significant. Read the p-value as formal nested-fit evidence, not as an automatic practical model preference or evidence that a subscore is useful.

Ordinary versus extended RSMs

Exactly one ordinary RSM MML fit and one testlet/shared-rater fit return an mfrm_extended_comparison, with models, checks, effects and notes. The information-criterion and LRT sections describe ordinary mfrm_fit comparisons only. This route does not rank models, compute difference intervals, or provide an AIC/BIC preference. nested = TRUE is refused.

Both fits must use identical observed rating events, including repeated events, category coding, unit weights and fixed facets except for the deliberately changed rater treatment. Use unanchored additive severity facets, noncenter_facet = "Person", no facet shrinkage, and matching person/score column roles. Category recoding must preserve the same consecutive adjacent-category scale. These checks cannot be disabled by warn_constraints = FALSE.

Match the population assumption: ordinary population_formula = ~1 with both extension defaults estimates a common normal ability SD; the ordinary default and extension person_sd = 1 impose known N(0,1). Covariate-dependent populations and fixed nonunit SDs have no matching ordinary route here. Population intercepts and free step locations may use different origins; they are recorded, not compared as substantive mean differences.

The source roster and omitted counts must agree. New ordinary fits retain prep$omitted_data and prep$omitted_input_rows. When scores were omitted, their identities must match, including repeated events. Older fits without this provenance require refitting from the complete assigned-score roster. Missing IDs, excluded weights and population-data omissions are unsupported.

Effects are centered at the unweighted mean of the same complete set of levels within each facet. This preserves pairwise contrasts while removing arbitrary facet origins. SourceReference/SourceComparison retain raw values and CenterReference/CenterComparison retain the subtracted means. Difference is comparison minus reference. Conditional rater modes and fixed coefficients are labeled separately; shrinkage is not evidence of greater accuracy or rater quality. Failed numerical checks or a missing level estimate withhold affected differences without dropping source rows. Native inference readiness remains explicit even when descriptive numerical checks pass. No SE or confidence interval for model differences is implied.

With response_diagnostics, responses retains matched category probabilities, means, full mixture variances and grouped descriptive Infit/Outfit. Both saved diagnostics must use the same probability target, selected event multiset and shared group_by identifier columns. Matching uses event contents, including repeated-event multiplicities and missing scores, rather than assuming row numbers agree. Identical repeated events match in source occurrence order. responses$events and responses$rows retain the matching and original row numbers. Missing or unavailable rows remain explicit and differences involving them are withheld.

All observed source events condition each model; selecting a Person's rows does not remove other Persons from shared-rater inference. These are same-data descriptive comparisons with calibration held fixed, not held-out predictive performance. A smaller Infit/Outfit is not evidence of improvement. Do not substitute ordinary plug-in diagnostics or their reference cutoffs. See plot.mfrm_extended_comparison() for predictive paired and difference displays.

With person_scores, persons$table compares conditional EAPs after subtracting each fitted population mean, retaining original estimates, origins, posterior SDs and aligned conditional endpoints. Unit Rasch slopes retain the logit unit; centering removes the arbitrary origin, not shrinkage or changes in the assumed population variance. Prior-only and unavailable differences are withheld. Separate conditional intervals are not intervals for model differences or tests of Person differences. Use metric = "person" in plot.mfrm_extended_comparison() for a paired/difference display.

The default facet-only output does not compare Person scores, steps, predictive Infit/Outfit or replacement-rater predictions. For testlets, zero local variance retains the fixed facets; zero shared-rater variance removes rater differences and is not the fixed-rater model. Numerical zero-reduction agreement does not establish model adequacy or variance-boundary inference.

Use plot(comparison, style = "paired") for matched effects, or style = "difference" for differences versus means. See plot.mfrm_extended_comparison() for display and accessibility controls. Attach saved results with mfrm_results(extended_fit, comparison = comparison) for static reports, tables, figures and replay. No comparison/report step fits, scores or resamples either model.

Information-criterion diagnostics

In addition to the canonical criteria, the function computes:

  • Delta_AIC / Delta_BIC / Delta_SABIC: difference from the minimum value in the supplied candidate set. A Delta < 2 is typically considered negligible; 4–7 suggests moderate evidence; > 10 indicates strong evidence against the higher-scoring model (Burnham & Anderson, 2002).

  • AkaikeWeight / BICWeight / SABICWeight: relative candidate-set weights derived from exp(-0.5 * Delta) and normalized only across the supplied models. They are not posterior probabilities, probabilities of model truth, or guarantees that the candidate set is adequate.

  • Evidence ratios: pairwise ratios of Akaike weights, quantifying relative support within this candidate set. A ratio of 5 means only that one normalized Akaike weight is five times the other; it does not mean that one model is five times more likely to be true.

AIC, BIC, and SABIC answer different approximation/penalty questions. SABIC is a sensitivity criterion, not a universal tie-breaker. At 22 or fewer Persons its Sclove penalty is non-positive, so SABICSelectable and SABICComparable are FALSE and no automatic SABIC preference is returned.

What this comparison means

compare_mfrm() is a same-basis model-comparison helper. Its strongest claims apply only when the models were fit to the same response data, under a compatible likelihood basis, and with compatible constraint structure.

What this comparison does not justify

  • Do not treat AIC/BIC/SABIC differences as primary evidence when table$ICComparable is FALSE.

  • Do not infer numerical adequacy merely because all models share the same coarse quadrature identity. Below q=31, raw criteria are diagnostic only and automatic deltas, weights, preferences, and LRT are suppressed.

  • Do not interpret the LRT unless nested = TRUE and the structural nesting review in comparison_basis$nesting_review passes.

  • Same-family additive-vs-interaction fits are considered nested only when all other structural settings match and the smaller model's facet_interactions set is a subset of the larger model's set.

  • Do not assume that nested = TRUE overrides the package's conservative nesting boundary; unsupported relations remain unsupported.

  • PCM is the unit-slope reduction of an aligned GPCM. nested = TRUE tests equal relative slopes only after population, constraint, free-dimension and numerical checks pass. The relation is recorded as PCM_in_GPCM. Default PCM uses a fixed standard-normal population, whereas default GPCM estimates its mean and variance: changing only model does not isolate slope differences. Supply population_formula = ~1 and the same person_data to both fits for an estimated-normal comparison. Unit slopes are interior positive values; the reference is ordinary asymptotic chi-square with G-1 degrees of freedom, not a boundary mixture. IC eligibility alone does not establish nesting or qualify slope intervals.

  • Do not compare models fit to different datasets, different score codings, or materially different constraint systems as if they were commensurate.

  • At large Person counts, a small systematic likelihood gain can dominate an IC difference; boundary or singular fits can also invalidate routine asymptotics. Evaluate effect size, stability, interpretability, and score consequences separately.

Interpreting output

  • Lower AIC/BIC values indicate relative support only when table$ICComparable is TRUE; use SABIC ranking only when table$SABICComparable is also TRUE.

  • A significant LRT p-value is formal evidence against the restricted model under the stated assumptions; practical gain, score utility, and dimensional interpretation require separate checks.

  • preferred identifies the minimum criterion within the supplied candidate set; it is not an unconditional model verdict.

  • evidence_ratios gives pairwise Akaike-weight ratios (returned only when Akaike weights can be computed for at least two models).

  • When comparing more than two models, interpret evidence ratios cautiously—they do not adjust for multiple comparisons.

How to read the main outputs

  • table: first-pass comparison table; start with ICComparable, ICSelectable, ICIntegrationTier, Model, Method, ICStatus, AIC, BIC, and SABIC.

  • comparison_basis: records whether IC and LRT claims are defensible for the supplied models. Inspect comparison_basis$nesting_review$relation and reason before reading any LRT output.

  • lrt: nested-model test summary, present only when the requested and reviewed conditions are met.

  • preferred: candidate preferred by each criterion when those summaries are available.

Inspect comparison_basis before writing conclusions. If comparability is weak, treat the result as descriptive and revise the model setup (for example, explicit step_facet, common data, or common constraints) before using IC or LRT results in reporting.

Typical workflow

  1. Fit two models with fit_mfrm() (e.g., PCM and GPCM) on the same prepared rows, explicit step owner, constraints, and MML quadrature setting. Use at least 31 common quadrature points for selectable ICs.

  2. Compare with compare_mfrm(fit_pcm, fit_gpcm) and start by checking comparison$table$ICComparable.

  3. Inspect summary(comparison) for AIC/BIC/SABIC diagnostics and the reasons for withholding ranking or tests. GPCM IC ranking has separate solution checks; request the matched PCM/GPCM test with nested = TRUE.

  4. Use build_weighting_review(fit_pcm, fit_gpcm) to inspect which selected facet levels and information shares were reweighted by the slopes.

References

  • Burnham, K. P., & Anderson, D. R. (2002). Model selection and multimodel inference: A practical information-theoretic approach (2nd ed.). Springer.

  • Akaike, H. (1974). A new look at the statistical model identification. IEEE Transactions on Automatic Control, 19(6), 716-723.

  • Schwarz, G. (1978). Estimating the dimension of a model. Annals of Statistics, 6(2), 461-464.

  • Sclove, S. L. (1987). Application of model-selection criteria to some problems in multivariate analysis. Psychometrika, 52(3), 333-343.

Examples

# \donttest{
# Load one dataset for both models
library(mfrmr)
toy <- load_mfrmr_data("example_operational")

# RSM: shared category thresholds
fit_rsm <- fit_mfrm(
  data = toy,
  person = "Person",
  facets = c("Rater", "Criterion"),
  score = "Score",
  method = "MML",
  model = "RSM"
)

# PCM: separate category thresholds for each criterion
fit_pcm <- fit_mfrm(
  data = toy,
  person = "Person",
  facets = c("Rater", "Criterion"),
  score = "Score",
  method = "MML",
  model = "PCM",
  step_facet = "Criterion"
)

# Check that the fitted models can be compared
comparison <- compare_mfrm(fit_rsm, fit_pcm, labels = c("RSM", "PCM"))
comparison$table[, c("Label", "Converged", "ICComparable", "ICSelectable")]
#> # A tibble: 2 × 4
#>   Label Converged ICComparable ICSelectable
#>   <chr> <lgl>     <lgl>        <lgl>       
#> 1 RSM   TRUE      TRUE         TRUE        
#> 2 PCM   TRUE      TRUE         TRUE        

# Smaller AIC/BIC indicates better relative support among these models
comparison$table[, c("Label", "AIC", "Delta_AIC", "BIC", "Delta_BIC")]
#> # A tibble: 2 × 5
#>   Label   AIC Delta_AIC   BIC Delta_BIC
#>   <chr> <dbl>     <dbl> <dbl>     <dbl>
#> 1 RSM    712.     0      729.      0   
#> 2 PCM    713.     0.375  737.      7.86
# Delta is the difference from the lowest value of that criterion
# For close or consequential comparisons, check a denser shared grid with
# mml_quadrature_sensitivity() before choosing a model
# For a PCM/GPCM equal-slope test, see the matched-population example in
# vignette("mfrmr-gpcm-scope"); request compare_mfrm(..., nested = TRUE).
# }