Classical Cultural Consensus Analysis

Werner Hertzog

Scope And Data

Romney implements classical consensus procedures for respondent-by-item data. Rows are informants and columns are judgments about a cultural domain. Remove identifiers and summary rows before analysis. Missing responses must be NA, not extra response categories or numeric sentinel codes.

Consensus concerns the answers shared within the sampled domain and group; it does not establish that those answers are objectively true. A good numerical solution does not by itself establish independent informants, comparable knowledge across items, or a single shared answer pattern. Interpret the diagnostics alongside the study’s ethnographic context.

There are two broad classical approaches: a formal knowledge-and-guessing model for categorical responses and an informal analysis of quantitative respondent profiles. Binary covariance agreement is an alternative within the categorical approach. This package retains the names formal, informal, and covariance for its three analysis settings.

Formal Model

Let \(X_{ij}\) denote respondent \(i\)’s answer to item \(j\), \(T_j\) its unknown cultural answer, \(c_i\) respondent competence, and \(m\) the number of options. A respondent knows the answer with probability \(c_i\) and otherwise guesses uniformly among all options. Consequently,

\[ P(X_{ij}=k\mid T_j=t,c_i)= \begin{cases} c_i+(1-c_i)/m,&k=t,\\ (1-c_i)/m,&k\ne t. \end{cases} \]

Competence is therefore not the raw proportion of correct responses. Even someone with \(c_i=0\) gives the correct answer with probability \(1/m\).

For distinct respondents, if \(p_{ih}\) is their proportion of matching responses, the chance-corrected agreement estimate is

\[a_{ih}=\frac{m p_{ih}-1}{m-1}.\]

Under the model, its expectation is \(c_i c_h\). Each item must use the same set of response options. Declare unobserved options explicitly rather than underestimating the number of choices.

sim <- simulate_consensus_data(24, 60, n_answers = 4,
                               competence = 0.6, seed = 7)
formal <- consensus(sim$responses, method = "formal", answer_levels = 1:4)
formal
#> Consensus analysis (formal)
#>   first factor SS: 8.535
#>   second factor SS: 0.682 
#>   ratio 1/2:       12.51
#>   mean competence: 0.588
#>   negatives first: 0
#>   extraction:     minres
#>   answer key:     estimated
head(formal$answer_key$key)
#> question_1 question_2 question_3 question_4 question_5 question_6 
#>          2          3          3          4          3          2

With answer prior \(q_{kj}\), the conditional posterior is

\[ P(T_j=k\mid X,\widehat{c})= \frac{q_{kj}\prod_{i\in O_j}P(X_{ij}\mid T_j=k,\widehat{c}_i)} {\sum_{\ell=1}^{m}q_{\ell j}\prod_{i\in O_j} P(X_{ij}\mid T_j=\ell,\widehat{c}_i)}, \]

where \(O_j\) contains respondents with an observed answer. Calculation uses log likelihoods to avoid underflow. answerkey_formal() accepts custom priors, whereas consensus() uses uniform answer priors for the formal key. Posterior probabilities condition on estimated competence and do not include uncertainty from estimating it.

Contradictory answers from respondents assigned competence one can make every candidate impossible. The function then warns and returns NA probabilities and keys, rather than falling back to the prior. Ties also have NA keys. Entirely missing items retain their prior probabilities but have no_data status and no inferred key.

Binary Covariance Model

For binary responses, define pairwise counts \(n_{11},n_{10},n_{01},n_{00}\) over \(n\) jointly observed items. With assumed true-item proportion \(\pi\), the agreement coefficient is

\[ a_{ih}=\frac{n_{11}n_{00}-n_{10}n_{01}} {n(n-1)\pi(1-\pi)}. \]

The numerator divided by \(n(n-1)\) is the sample covariance of the two binary profiles. The second allowable category is treated as true; inferred categories are sorted. The default \(\pi=0.5\) concerns the unknown answer key, not an individual respondent’s frequency of saying yes.

path <- system.file("extdata", "synthetic_yesno_36x103.csv", package = "Romney")
binary_data <- as.matrix(read.csv(path, row.names = 1))
binary <- consensus(binary_data, method = "covariance", prior = 0.5)
binary
#> Consensus analysis (covariance)
#>   first factor SS: 8.968
#>   second factor SS: 1.128 
#>   ratio 1/2:       7.948
#>   mean competence: 0.461
#>   negatives first: 0
#>   extraction:     minres
#>   answer key:     estimated
head(binary$answer_key$key)
#> item_1 item_2 item_3 item_4 item_5 item_6 
#>      1      1      0      0      0      1

answerkey_covariance() chooses the category with the greater sum of respondent competences among observed responses. Its weighted_proportions are normalized voting weights, not posterior probabilities. To reproduce the scale of UCINET detail tables, weighted_frequencies multiplies these proportions by the total respondent count, even when an item has missing answers. Neither output is an estimate of respondent response-bias parameters.

Informal Model

Rankings and numerical ratings can be compared with Pearson correlation,

\[a_{ih}=\operatorname{cor}(X_i,X_h).\]

Informal competence measures correspondence with a shared response pattern, not a probability of knowing an answer. Numeric ranks and ratings are treated as scores: arbitrary recoding of an ordinal scale can change Pearson correlations. Constant respondent profiles or zero variation on overlapping items make correlations undefined.

path <- system.file("extdata", "synthetic_ordinal_36x103_1to5.csv", package = "Romney")
ordinal_data <- as.matrix(read.csv(path, row.names = 1))
informal <- consensus(ordinal_data, method = "informal")
informal
#> Consensus analysis (informal)
#>   first factor SS: 27.394
#>   second factor SS: 0.451 
#>   ratio 1/2:       60.728
#>   mean competence: 0.871
#>   negatives first: 0
#>   extraction:     minres
#>   answer key:     estimated
head(informal$answer_key$key)
#>   item_1   item_2   item_3   item_4   item_5   item_6 
#> 3.087166 3.940723 4.064580 4.759537 4.746451 2.899934

This implementation reports original-scale weighted means,

\[\widehat{T}_j= \frac{\sum_{i\in O_j}\widehat{c}_i X_{ij}} {\sum_{i\in O_j}\widehat{c}_i}.\]

It retains signed weights and requires a positive total weight for each item. These means are not standardized regression factor scores; negative weights can produce estimates outside the observed response range.

Extraction And Interpretation

Romney uses psych::fa() with unrotated minimum-residual extraction. In a factor representation, off-diagonal agreements are approximated by \(a_{ih}\approx\sum_f L_{if}L_{hf}\). The diagonal of the input matrix is set to one; the fitted communalities are estimated during extraction. At least two factors are extracted to form the first-to-second ratio, even when cultures = 1 returns only first-factor loadings.

The factor_ss_loadings field contains sums of squared loadings, with eigenvalues retained as a historical alias. raw_eigenvalues contains eigenvalues of the original agreement matrix. These quantities need not match. The legacy argument cultures controls returned dimensions; it does not identify distinct cultural groups or fit a mixture of cultures.

A ratio above three and a non-negative first factor are commonly used diagnostics, not significance tests. Inspect residuals, communalities, negative loadings, missingness, and extraction warnings as well. Probability-scale competence thresholds are not applied to the informal model. No bootstrap intervals or formal multi-culture selection test are implemented in this version.

formal$criteria
#>         eigen_ratio_gt_3   mean_competence_gt_0_5 no_negative_first_factor 
#>                     TRUE                     TRUE                     TRUE
formal$diagnostics$extracted_factors
#> [1] 2
formal$diagnostics$positive_semidefinite
#> [1] TRUE
range(formal$diagnostics$pairwise_items)
#> [1] 60 60

Signed loadings are never silently clipped in reported competence. If a formal or covariance first-factor loading is outside \([0,1]\), the default competence_policy = "strict" skips answer-key estimation with a warning. Explicit "truncate" clips only the weights used for a sensitivity calculation and does not remedy a violation of the model.

Missing responses are omitted pairwise, not imputed. Pairwise deletion can produce a non-positive-semidefinite matrix. The factor library may issue smoothing or other warnings, which remain visible and are retained in diagnostics$factor_warnings. Extraction errors stop the analysis; they do not trigger another estimator. At least three respondents are required, but this is an identification minimum, not a recommended study size.

Validation And Citation

The three bundled synthetic datasets are fixed software-comparison fixtures, not empirical anthropological findings. Complete available UCINET 6.832 tables are compared for all respondents and items. Agreement matches at the printed precision, categorical keys match exactly, and ordinal weighted means differ by less than 0.001. Factor statistics and competence estimates are close but not identical. The relatively larger ordinal ratio discrepancy is reported explicitly.

The installed validation directory contains full CSV comparisons and environment metadata:

system.file("validation", "ucinet-summary.csv", package = "Romney")
#> [1] "/private/var/folders/r5/ns58l7x54qzd37rp8cly8k3m0000gp/T/RtmpO8mQ9l/Rinst34d23b65f70/Romney/validation/ucinet-summary.csv"

There are no bundled UCINET posterior-probability tables, second-factor loading tables, or independent ANTHROPAC runs. Those are limits on the cross-software validation claims. Mathematical and simulation tests provide separate checks of the implementation. The original reference inputs are preserved because their simulation seeds are not recorded.

Use citation("Romney") to cite the software. Cite the methodological references appropriate to the analysis as well.

References

Romney, A. K., Weller, S. C., and Batchelder, W. H. (1986). Culture as consensus: A theory of culture and informant accuracy. American Anthropologist, 88(2), 313–338. doi:10.1525/aa.1986.88.2.02a00020.

Romney, A. K., Batchelder, W. H., and Weller, S. C. (1987). Recent Applications of Cultural Consensus Theory. American Behavioral Scientist, 31(2), 163–177. doi:10.1177/000276487031002003.

Weller, S. C. (2007). Cultural consensus theory: Applications and frequently asked questions. Field Methods, 19(4), 339–368. doi:10.1177/1525822X07303502.