Inspect how many rating rows can be used, how often each score category
occurs, and whether facet levels connect through shared persons. This
function prepares descriptive checks; it does not fit a model or change
the data object supplied by the caller. Each row should represent one
rating, and person, facets, and score name its columns.
Usage
describe_mfrm_data(
data,
person,
facets,
score,
weight = NULL,
rating_min = NULL,
rating_max = NULL,
keep_original = FALSE,
missing_codes = NULL,
include_person_facet = FALSE,
include_agreement = TRUE,
rater_facet = NULL,
context_facets = NULL,
agreement_top_n = NULL,
expected_design = NULL,
min_linking_persons = 2L,
category_policy = NULL
)Arguments
- data
A data.frame in long format (one row per rating event).
- person
Column name for person IDs.
- facets
Character vector of facet column names.
- score
Column name for observed score.
- weight
Optional weight/frequency column name.
- rating_min
Optional minimum category value. Supply with
rating_maxto retain unused boundary categories in the intended score support.- rating_max
Optional maximum category value. Supply with
rating_minto retain unused boundary categories in the intended score support.- keep_original
Keep original category values. Use this with
rating_min/rating_maxwhen the intended scale has unused intermediate categories such as1, 2, 4, 5on a 1-5 scale. New code can instead usecategory_policy = "preserve".- missing_codes
Optional.
NULL(default) is a no-op;TRUEor"default"activates the FACETS / SPSS / SAS convention (c("99", "999", "-1", "N", "NA", "n/a", ".", "")) for the score column while preserving person/facet identifiers; supply a character vector to apply a custom code set across all model columns. Replacement counts are returned in themissing_recodingcomponent when supported by the calling helper. Seerecode_missing_codes()for the standalone version.- include_person_facet
If
TRUE, include person-level rows infacet_level_summary.- include_agreement
If
TRUE, include an observed-score agreement bundle (summary/pairs/settings) for a selected non-person facet.- rater_facet
Optional facet name used to identify repeated scorers for agreement summaries. If
NULL, a rater-like name such asRater,Judge, orScoreris inferred. No agreement analysis is run when such a name is absent; set this argument explicitly only when another facet genuinely represents repeated scorers.- context_facets
Optional facets used to define matched contexts for agreement. If
NULL, all remaining facets (includingPerson) are used.- agreement_top_n
Optional maximum number of agreement pair rows.
- expected_design
Optional data frame declaring the planned assignment roster. It must contain the columns named by
personandfacets, with one row per planned Person x facet cell. Extra columns are ignored. When supplied, observed cells are compared with the roster so planned omissions can be distinguished from cells that were never assigned.- min_linking_persons
Positive integer used as a descriptive sparse-link flag. A facet level observed for fewer than this many distinct persons is counted in
linkage_summary$SparseLevels. This is a review threshold, not a model-acceptance rule.- category_policy
Optional explicit category choice:
"collapse"maps gaps in the observed categories to consecutive scores;"preserve"keeps the intended ladder, declared withrating_minandrating_max. This changes the fitted category steps, not just labels.NULL(default) useskeep_original, whose default isFALSE("collapse"). Supplying both choices is allowed only when they agree. Preservation does not estimate unsupported steps: fitting stops if a retained internal category has no observations. Use the same policy indescribe_mfrm_data()andreview_mfrm_anchors().
Value
A list of class mfrm_data_description with:
overview: one-row run-level summary includingCategoryPolicyandScoreRecoded. The former records the selected category handling; the latter indicates whether original score values actually changed. A"collapse"policy can leave a contiguous scale unchanged. Inspectscore_support$score_mapfor the mapping.missing_by_column: missing counts in selected input columnsmissing_rate_summary: per-column missingness rate summary (one row per input column, with raw and proportion-of-N columns)score_descriptives: output frompsych::describe()for scoreweight_descriptives: output frompsych::describe()for weightscore_distribution: weighted and raw score frequencies over the prepared score support. Unused boundary categories are retained when the rating range was supplied explicitly; unused intermediate categories requirekeep_original = TRUE.facet_level_summary: per-level usage and score summariesfacet_crosstabs: pairwise observation-count crosstabs between non-person facets (named list keyed"facetA__facetB") for optional downstream coverage displayslinkage_summary: person-facet connectivity diagnosticsstructural_missingness: declared-design comparison bundle containing a one-row summary, missing expected cells, unexpected observed cells, per-facet level coverage, and settingsdesign_connectivity: component counts for each observed Person-facet graph and, when declared, each expected Person-facet graphdesign_components: component-level counts and facet-level labels; person labels are included only wheninclude_person_facet = TRUEduplicate_cell_summary: counts of duplicate Person x facet cellsduplicate_cell_detail: duplicate-cell keys and row countsagreement: observed-score agreement bundle for the selected scorer facetrow_retention: row counts before and after preparation filterspreparation_notes: structured notes for row drops, ID trimming, and design conditions detected during preparationmissing_recoding: per-column counts of declared missing-code values replaced withNAbefore row filteringscore_support: minimal prepared score-support metadata used bysummary(ds)$caveats
Details
Set rating_min and rating_max from the rubric, including categories
nobody received. Use keep_original = TRUE to preserve its category
structure in the review. Numeric descriptives of score and weight use
psych::describe().
Key data-quality checks to perform before fitting:
Sparse categories: review categories with little weighted support because their threshold estimates may be imprecise. Do not collapse categories solely from a package warning; also consider the rubric, intended score interpretation, and category diagnostics after fitting.
Unlinked elements: inspect
design_connectivityfor the observed Person-facet graph. More than one component means that the levels of that facet are not connected through shared persons. This facet-specific check is conservative and does not by itself prove full model identification.Extreme scores: MML uses a person distribution to obtain posterior person scores, including for persons with all-minimum or all-maximum scores. Non-person facets remain fixed effects: an extreme rater or criterion is not given a prior or automatically shrunk by choosing MML. Review the fitted boundary and precision evidence before interpretation.
Interpreting output
Recommended order:
overview: confirms retained ratings (Observations), persons, facets, and category span. Userow_retentionto compare input and retainedRows;DroppedRowscounts exclusions during preparation.missing_by_column: countsNAvalues in the input columns. A missing score or required identifier excludes that rating row, not automatically the person's other ratings. The package does not fill missing ratings. Whenmissing_codesis supplied, these counts still describe the original input; inspectmissing_recodingandpreparation_notesas well.structural_missingness: compares observed rating cells withexpected_design, when supplied. Without a declared roster, structural missingness is reported as not assessed rather than assumed to be zero.score_distribution: checks sparse/unused score categories. Skew can be substantively expected, but weakly supported or unused categories need explicit interpretation.facet_level_summaryandlinkage_summary: checks per-level support, shared-person counts, and sparse levels. Usedesign_connectivityfor the separate graph-component result.agreement: optional observed agreement summary for the selected scorer facet (exact agreement, correlation, and mean differences per pair).
data_review <- describe_mfrm_data(...) saves all these checks.
review <- summary(data_review) provides a compact view; its missingness
table is named review$missing, while the original full table is
data_review$missing_by_column. Summary previews use top_n = 10 by
default. Use the original tables to inspect all rows or categories.
If the input needs attention
Column name not found: run
names(ratings)and match spelling, spaces, and case inperson,facets, andscore.Unexpected row loss: inspect
row_retention,missing_by_column, andpreparation_notes. Resolve unintended missing IDs and invalid score text in the input data. For documented score markers such as99or., userecode_missing_codes()with explicitcolumnsandcodes.Repeated person-by-facet cells: inspect
duplicate_cell_detail. Correct accidental duplicates; include a task or occasion facet when ratings represent distinct events in the design.Unused category or disconnected design: inspect
score_distributionanddesign_connectivity. Review the rubric and assignments before changing the model. A retained internal zero-count category stops fitting; extra optimizer iterations cannot supply the missing category information.
After editing or recoding ratings, rerun this review on the corrected data
and pass that same data to fit_mfrm(), using the same columns, score
bounds, and keep_original setting. If you use missing_codes within the
review instead of recoding first, supply the same option to the fit:
reviewing does not modify the original data. For CSV import, column mapping,
and a complete worked example, see
vignette("mfrmr-workflow", package = "mfrmr").
Typical workflow
Run
data_review <- describe_mfrm_data(...)on the rating data.Inspect row retention, category counts, and design connectivity.
Correct input issues, repeat the review, and fit the reviewed data with
fit_mfrm().
Examples
library(mfrmr)
toy <- load_mfrmr_data("example_operational")
head(toy)
#> Study Person Rater Criterion Score Group
#> 1 OperationalExample P001 R01 Language 4 A
#> 2 OperationalExample P001 R01 Organization 2 A
#> 3 OperationalExample P001 R02 Content 4 A
#> 4 OperationalExample P001 R02 Language 3 A
#> 5 OperationalExample P001 R02 Organization 2 A
#> 6 OperationalExample P002 R01 Content 3 A
# Check the data before fitting; the intended score categories are 1 to 4
data_review <- describe_mfrm_data(
data = toy,
person = "Person",
facets = c("Rater", "Criterion"),
score = "Score",
rating_min = 1,
rating_max = 4,
category_policy = "preserve"
)
data_review$row_retention # Input and retained rows; check DroppedRows
#> Stage Rows DroppedRows
#> 1 input_selected_columns 282 0
#> 2 after_missing_and_weight_filter 282 0
#> DroppedReason
#> 1
#> 2 missing values or non-positive weights
data_review$missing_by_column # Missing input values in each model column
#> # A tibble: 4 × 2
#> Column Missing
#> <chr> <int>
#> 1 Person 0
#> 2 Rater 0
#> 3 Criterion 0
#> 4 Score 0
data_review$score_distribution # RawN is the number of ratings per category
#> # A tibble: 4 × 4
#> Score RawN WeightedN Percent
#> <int> <int> <dbl> <dbl>
#> 1 1 62 62 22.0
#> 2 2 96 96 34.0
#> 3 3 78 78 27.7
#> 4 4 46 46 16.3
data_review$design_connectivity # Components = 1 means connected for that facet
#> Basis Facet PersonNodes FacetLevelNodes Edges Components
#> 1 observed Rater 48 6 96 1
#> 2 observed Criterion 48 3 144 1
#> LargestComponentPersons LargestComponentLevels LargestComponentPercent
#> 1 48 6 100
#> 2 48 3 100
#> Connected
#> 1 TRUE
#> 2 TRUE
# Here all 282 rows are retained, and all four categories have observations
# Save a compact summary when you want the overview and review notes
review <- summary(data_review)
review$overview
#> Observations TotalWeight Persons Facets Categories RatingMin RatingMax
#> 1 282 282 48 2 4 1 4
#> RatingRangeSource RatingMinSource RatingMaxSource CategoryPolicy ScoreRecoded
#> 1 declared declared declared preserve FALSE
review$notes
#> [1] "No missing values were detected in selected input columns."
#> [2] "Structural missingness was not assessed because `expected_design` was not supplied. Absent rows cannot be distinguished from cells that were never assigned."
# For the next fit, use data = toy; data_review is a set of checks, not ratings
