gson

CRAN checks Dependencies

Provides a lightweight container and exchange format for gene set collections. A ‘GSON’ object stores which genes belong to which gene set, together with gene set and gene names, the identifier types in use, species, versions and source metadata. A collection can be built from data frames, read from and written to the ‘gson’ JavaScript Object Notation (JSON) format and the ‘GMT’ format, subset by gene set, merged across sources, validated, and resolved to the web addresses of the databases it comes from, so that a collection gathered by one package can be analysed by another.

A gene set collection answers one question: which genes belong to which gene set. GSON holds that answer together with the metadata needed to read it: what the identifiers are, which organism, which release, where the sets came from. The file it writes is plain JSON, so a collection can be read outside R too.

:arrow_double_down: Installation

# the released version, from CRAN
install.packages("gson")

# the development version, from GitHub
# install.packages("remotes")
remotes::install_github("YuLab-SMU/gson")

Build a collection

Three tables make up a GSON. gsid2gene is required and holds one row per gene set and gene. gsid2name and gene2name are optional and map gene set IDs and gene IDs to readable names. The rest is metadata: keytype names the identifier type of the genes, gsidtype the identifier domain of the gene set IDs, and species, gsname, version, accessed_date, urlpattern and info say where the collection came from.

library(gson)

gsid2gene <- data.frame(
    gsid = c("GO:0007049", "GO:0007049", "GO:0006915"),
    gene = c("101", "102", "103")
)

gsid2name <- data.frame(
    gsid = c("GO:0007049", "GO:0006915"),
    name = c("cell cycle", "apoptotic process")
)

x <- gson(
    gsid2gene,
    gsid2name = gsid2name,
    species = "Homo sapiens",
    gsname = "example",
    version = "2026-06-30",
    keytype = "ENTREZID",
    gsidtype = "GO"
)

x
#> >> Gene Set: example
#> >> 3 genes annotated by 2 gene sets.
#> >> Species: Homo sapiens
#> >> Version: 2026-06-30
#> >> Gene set ID type: GO

gson() keeps the columns of the three tables character. A factor would retain the levels it no longer uses, and a collection is read by split()-ing it by gene set ID, so one unused level becomes one empty gene set that length(), names() and both writers then report as real.

Read and write

f <- tempfile(fileext = ".gson")
write.gson(x, f)
cat(readLines(f), sep = "\n")
#> {
#>   "gsid2gene": {
#>     "GO:0007049": ["101", "102"],
#>     "GO:0006915": ["103"]
#>   },
#>   "gsid2name": {
#>     "gsid": ["GO:0007049", "GO:0006915"],
#>     "name": ["cell cycle", "apoptotic process"]
#>   },
#>   "gene2name": null,
#>   "schema_version": ["1.0"],
#>   "species": ["Homo sapiens"],
#>   "gsname": ["example"],
#>   "version": ["2026-06-30"],
#>   "accessed_date": null,
#>   "keytype": ["ENTREZID"],
#>   "gsidtype": ["GO"],
#>   "urlpattern": null,
#>   "info": null
#> }
y <- read.gson(f)
identical(x, y)
#> [1] TRUE

read.gson() opens a .gson, a .gson.gz, and a file written by any other program that follows the format. A file it cannot make sense of is answered by name:

read.gson("no-such-file.gson")
#> Error:
#> ! the gson file no-such-file.gson does not exist

A file that exists but is not a collection is answered by the member that is wrong and the shape it holds. tests/testthat/test-io.R covers the shapes a hand-written .gson turns up with.

GMT is the other side of the same object. write.gmt() gives one line per gene set, holding the ID, a description and the genes; read.gmt() reads a plain GMT back as a two-column table.

gmt <- tempfile(fileext = ".gmt")
write.gmt(x, gmt)
cat(readLines(gmt), sep = "\n")
#> GO:0006915   apoptotic process   103
#> GO:0007049   cell cycle  101 102

A membership row with no gene set ID or no gene cannot go into either format, so both writers drop those rows and say how many. validate_gson() reports them instead, on an object you can still look at.

Read a GMT file

wpfile <- system.file(
    "extdata",
    "wikipathways-20220310-gmt-Homo_sapiens.gmt",
    package = "gson"
)

wp <- read.gmt.wp(wpfile, output = "GSON")
wp
#> >> Gene Set: WikiPathways
#> >> 7781 genes annotated by 724 gene sets.
#> >> Species: Homo sapiens
#> >> Version: WikiPathways_20220310
#> >> Gene set ID type: WP

read.gmt.wp() is the WikiPathways reader. It takes the species and the release out of the file name, records gsidtype = "WP", and sets urlpattern to the address WikiPathways serves its pathways from. Your own GMT follows none of those conventions, so read it with read.gmt() and build the object:

term2gene <- read.gmt(gmt)
head(term2gene)
#>         term gene
#> 1 GO:0006915  103
#> 2 GO:0007049  101
#> 3 GO:0007049  102

from_gmt <- gson(
    term2gene,
    species = "Homo sapiens",
    gsname = "my sets",
    version = "1.0",
    keytype = "ENTREZID"
)

from_gmt
#> >> Gene Set: my sets
#> >> 3 genes annotated by 2 gene sets.
#> >> Species: Homo sapiens
#> >> Version: 1.0

gson() renames the first two columns of the table it is given to gsid and gene, so the term column of read.gmt() becomes the gene set ID here.

Look a gene set up online

gson_url(wp, c("WP100", "WP106"))
#>                                              WP100 
#> "https://www.wikipathways.org/pathways/WP100.html" 
#>                                              WP106 
#> "https://www.wikipathways.org/pathways/WP106.html"

gson_url() substitutes a gene set ID for the {gsid} token of the pattern the object declares. With no gsid it covers every gene set, in the order names(x) holds them, and browseGS(wp, "WP100") opens the same address in the browser. gson ships no table of database websites, since it has no way to check one, so an address comes only from the object:

mine <- gson(gsid2gene, gsname = "my sets", gsidtype = "GO")
gson_url(mine)
#> Warning: this GSON carries no 'urlpattern', so its gene sets have no URL
#> GO:0007049 GO:0006915 
#>         NA         NA
browseGS(mine, "GO:0007049")
#> Error:
#> ! this GSON carries no 'urlpattern', so its gene sets have no URL to browse

Subset

x[i] keeps the gene sets i selects and returns a GSON. A filtered collection stays a collection, rather than three tables you filtered by hand and kept in step yourself.

wp2 <- wp[c("WP100", "WP106")]
wp2
#> >> Gene Set: WikiPathways
#> >> 35 genes annotated by 2 gene sets.
#> >> Species: Homo sapiens
#> >> Version: WikiPathways_20220310
#> >> Gene set ID type: WP

Gene set IDs, positions and logical vectors all select, x[] returns x, and x[[i]] still returns the genes of one gene set. The three tables are pruned together, so gene2name loses any gene that is no longer a member of something kept, and every metadata slot carries over to the result. Most subsets come from the gene set names:

keep <- wp@gsid2name[grepl("apoptosis", wp@gsid2name$name, ignore.case = TRUE), ]
wp[keep$gsid]
#> >> Gene Set: WikiPathways
#> >> 211 genes annotated by 9 gene sets.
#> >> Species: Homo sapiens
#> >> Version: WikiPathways_20220310
#> >> Gene set ID type: WP

A GSON has no second dimension, so wp[, 2] is an error, and so is an index out of bounds or a gene set ID the object does not hold.

Combine collections

gson_union() merges several GSON objects, or one GSONList, into a single collection. Two objects that name the same source are two parts of one collection, so a gene set both carry joins into one set holding every gene either gives it:

monday <- gson(
    data.frame(gsid = c("GO:0007049", "GO:0006915"), gene = c("101", "103")),
    gsid2name = data.frame(gsid = c("GO:0007049", "GO:0006915"),
                           name = c("cell cycle", "apoptotic process")),
    gsname = "example", gsidtype = "GO", species = "Homo sapiens",
    version = "2026-06-30", keytype = "ENTREZID"
)

tuesday <- gson(
    data.frame(gsid = c("GO:0006915", "GO:0008150"), gene = c("103", "104")),
    gsname = "example", gsidtype = "GO", species = "Homo sapiens",
    version = "2026-06-30", keytype = "ENTREZID"
)

merged <- gson_union(monday, tuesday)
merged
#> >> Gene Set: example
#> >> 3 genes annotated by 3 gene sets.
#> >> Species: Homo sapiens
#> >> Version: 2026-06-30
#> >> Gene set ID type: GO
merged@gsid2gene
#>         gsid gene
#> 1 GO:0007049  101
#> 2 GO:0006915  103
#> 3 GO:0008150  104

Two objects from different sources that both use the same gene set ID are refused:

custom <- gson(
    data.frame(gsid = "GO:0007049", gene = "999"),
    gsname = "my sets", gsidtype = "GO", species = "Homo sapiens",
    version = "1.0", keytype = "ENTREZID"
)

gson_union(monday, custom)
#> Error:
#> ! gene set ID(s) used by more than one source: GO:0007049 -- merge with .conflict = "prefix" to qualify every ID with its source, or .conflict = "rename" to qualify only these

The merged table would say that one ID names two sets of genes, and a result read from it would be wrong without looking wrong. Ask for what you want instead: .conflict = "rename" qualifies only the colliding IDs, .conflict = "prefix" qualifies all of them, and the gsid2name rows follow the IDs they describe.

merged <- gson_union(monday, custom, .conflict = "rename")
merged
#> >> Gene Set: example + my sets
#> >> 3 genes annotated by 3 gene sets.
#> >> Species: Homo sapiens
#> >> Version: 2026-06-30 + 1.0
#> >> Gene set ID type: GO
names(merged)
#> [1] "example:GO:0007049" "GO:0006915"         "my sets:GO:0007049"

species, keytype and schema_version have to agree between the objects, because one membership table cannot mean two organisms, two gene identifier types or two file schemas. gsname, gsidtype, version, accessed_date and info are joined with + when the sources name themselves differently. urlpattern is kept only when every object declares the same one and no ID moved, because a pattern that resolves part of the collection to the wrong page is worse than no pattern.

A GSONList is a list of collections. gson_union() takes one as its input, and x[1] of a GSONList is still a GSONList.

Check a collection

gson() refuses what the class cannot hold: a table that is not a data frame, one that lacks the documented columns, identifiers that are not character. validate_gson() reports what a valid object can still get wrong. With error = FALSE it returns the report as a character vector instead of throwing it.

validate_gson(x)
#> [1] TRUE

broken <- gson(
    data.frame(gsid = c("GO:0007049", "GO:0007049"), gene = c("101", "101")),
    gsid2name = data.frame(gsid = "GO:0007049", name = "cell cycle"),
    gsidtype = "KEGG", keytype = "UNKNOWN", species = "Homo sapiens"
)

validate_gson(broken, error = FALSE)
#> [1] "gsid2gene contains duplicate gsid-gene memberships"                                         
#> [2] "keytype 'UNKNOWN' names no identifier type; leave it NULL when the type is not known"       
#> [3] "gsidtype 'KEGG' is not the prefix of the gene set IDs, which begin GO -- as in 'GO:0007049'"

There is deliberately no dictionary of allowed keytype or gsidtype values. Producers in this family write kegg_orthology, eggNOG_OG and hmdb_id, none of them a value any annotation package offers, and an enum held by the format would reject collections that work perfectly well. What gson can verify is an object against itself, so a gsidtype is compared with the prefixes the gene set IDs of that object actually carry, and nothing is ever rewritten.

Use it for enrichment

enrichit takes the object as it is; ora_gson(), gsea_gson() and nsea_gson() all have a gson argument that wants a GSON.

geneList <- sort(setNames(rnorm(30), as.character(101:130)), decreasing = TRUE)

gsea <- enrichit::gsea_gson(geneList, gson = x)
ora  <- enrichit::ora_gson(gene = c("101", "102"), gson = x,
                           pvalueCutoff = 1, minGSSize = 1)

A tool that wants the bare tables finds them in the slots. gsid2name is the map from gene set ID to label that puts a readable name on a result, and keytype is what tells an annotation package how to read the genes.

head(x@gsid2gene)
#>         gsid gene
#> 1 GO:0007049  101
#> 2 GO:0007049  102
#> 3 GO:0006915  103
head(x@gsid2name)
#>         gsid              name
#> 1 GO:0007049        cell cycle
#> 2 GO:0006915 apoptotic process

The gson file format

{
  "gsid2gene": {"GO:0007049": ["101", "102"], "GO:0006915": ["103"]},
  "gsid2name": {"gsid": ["GO:0007049", "GO:0006915"],
                "name": ["cell cycle", "apoptotic process"]},
  "gene2name": null,
  "schema_version": ["1.0"],
  "species": ["Homo sapiens"],
  "gsname": ["example"],
  "version": ["2026-06-30"],
  "accessed_date": null,
  "keytype": ["ENTREZID"],
  "gsidtype": ["GO"],
  "urlpattern": null,
  "info": null
}

Because gsid2gene is keyed by gene set ID, reading a file back groups the memberships by gsid. The gene sets keep the order they were written in, and so do the genes inside each set; the row order of a table that was not already grouped by gene set ID does not survive. Row names do not survive either, and JSON has no factor type, which is why gson() holds the columns as character.

:writing_hand: Authors

Guangchuang YU https://yulab-smu.top

School of Basic Medical Sciences, Southern Medical University

:sparkling_heart: Contributing

We welcome any contributions! By participating in this project you agree to follow the conventions the rest of the clusterProfiler family uses. Please open an issue at https://github.com/YuLab-SMU/gson/issues before a large change.