From marginal study summaries to synthetic patients

summary2joint estimates a joint latent Gaussian distribution from marginal summaries in repeated independent studies of a common population. It does not identify arbitrary dependence from one set of margins and does not recover the original patient records.

Input

Use a named list specifying continuous, binary, and ordinal variables. For each study, supply its size and a named list of means/sample SDs, binary event counts, and ordered category counts. All summaries in a study refer to the same people. Unreported variables can be omitted, but every pair needs repeated joint reporting.

library(summary2joint)
set.seed(31)
variables <- list(age = list(type = "continuous"),
                  response = list(type = "binary"),
                  severity = list(type = "ordinal", levels = 3L))
studies <- lapply(seq_len(80), function(i) {
  z <- matrix(rnorm(300), 100, 3)
  z[, 2] <- 0.4 * z[, 1] + sqrt(0.84) * z[, 2]
  age <- 55 + 8 * z[, 1]
  response <- as.integer(z[, 2] > 0)
  severity <- findInterval(z[, 3], c(-Inf, -0.5, 0.5, Inf))
  list(n = 100L, summaries = list(
    age = list(mean = mean(age), sd = sd(age)),
    response = list(events = sum(response)),
    severity = list(counts = tabulate(severity, 3))))
})
fit <- fit_summary_copula(studies, variables)
fit
## Summary copula fit: 80 studies; 3 variables
## Converged: TRUE   Boundary: FALSE   Inference available: TRUE 
##             age response severity
## age       1.000    0.200   -0.174
## response  0.200    1.000    0.109
## severity -0.174    0.109    1.000
fit$diagnostics[c("converged", "boundary", "inference_ok")]
## $converged
## [1] TRUE
## 
## $boundary
## [1] FALSE
## 
## $inference_ok
## [1] TRUE

Probability and generation

joint_probability(fit, lower = c(age = 60, response = 1, severity = 2))
##    estimate         se      lower     upper
## 1 0.1032815 0.02197137 0.06746762 0.1549475
head(simulate_summary_copula(fit, n = 100, seed = 42))
##        age response severity
## 1 65.95292        1        1
## 2 50.54048        1        3
## 3 57.92819        0        3
## 4 60.07592        1        3
## 5 58.25576        0        1
## 6 54.19182        1        1
confint(fit)
##   variable1 variable2   estimate        se      lower     upper studies
## 1       age  response  0.1997753 0.1911061 -0.1854615 0.5317844      80
## 2       age  severity -0.1743854 0.1374007 -0.4251265 0.1012130      80
## 3  response  severity  0.1090878 0.1726777 -0.2288678 0.4235758      80

The correlation matrix is on the latent normal scale, not generally the Pearson correlation of the observed discrete variables. Probability intervals propagate uncertainty in both the margins and dependence. Fits are ordinary serializable R objects; save them using saveRDS().

Independent groups and limitations

If scientific knowledge supports independent groups, supply a partition to fit_summary_copula_groups(). Each group must contain at least two variables. Use joint_probability_groups() for the fitted grouped distribution; simulate_summary_copula() works for both classes. Grouping is a modeling assumption, not an automatic selection procedure.

Always inspect convergence, boundaries, and inference diagnostics. A full-rank sandwich requires more studies than parameters, but that condition alone does not ensure accurate intervals. Rare events and few studies can cause undercoverage. Between-study heterogeneity can confound within-patient dependence. This package assumes a common population, aligned definitions, independent participants across studies, and reporting independent of measurements. It does not model missing patient data or rounded counts. Details and numerical controls are in ?fit_summary_copula.