TransfereGov publishes four open data APIs on
api-publica.transferegov.gestao.gov.br, covering
seventy-four tables and about 7.5 million rows between them. This
vignette shows how to find your way around them and retrieve data
without downloading more than you meant to.
The code here is not run when the vignette is built, because it would call the government’s servers.
tg_modules()
#> # A tibble: 4 × 5
#> module label tables max_page_size url
#> <chr> <chr> <int> <int> <chr>
#> 1 especiais Special transfers 23 200 https://api-publ…
#> 2 fundoafundo Fund-to-fund transfers 20 1000 https://api-publ…
#> 3 parcerias Partnerships 17 200 https://api-publ…
#> 4 ted Decentralized credit 14 1000 https://api-publ…especiais covers special transfers, the
mechanism created by Constitutional Amendment 105/2019 that lets an
individual parliamentary amendment send money straight to a
municipality’s account, with no agreement to sign.
fundoafundo covers fund-to-fund transfers,
from a federal fund to a state or municipal one — health, social
assistance, education. parcerias covers
partnerships with civil society organizations, from the
program that announces them through to the bank statement of the account
they are paid from. ted covers decentralized
credit, the termos de execução descentralizada through
which one federal body passes budget to another to carry out a program
on its behalf.
What is not here: SICONV agreement data, which is published as CSV downloads rather than as an API.
tg_tables("parcerias")
#> # A tibble: 17 × 6
#> module table path columns params
#> <chr> <chr> <chr> <int> <int>
#> 1 parcerias analise_proposta analise-proposta 7 5
#> 2 parcerias beneficiario_emenda_… beneficiario_emenda_… 16 14
#> 3 parcerias cronograma_desembolso cronograma-desembolso 7 7
#> …table is the name to pass to tg_get();
path is the endpoint it maps to. The two differ wherever
the endpoint uses a hyphen, and either spelling is accepted.
tg_fields() describes the columns, and
tg_params() the filters:
tg_fields("parcerias", "proposta")
#> # A tibble: 43 × 5
#> field r_type api_type nested description
#> <chr> <chr> <chr> <chr> <chr>
#> 1 id_proposta double integer NA Identificador único da prop…
#> 2 id_programa double integer NA Identificador do programa a…
#> …
tg_params("parcerias", "proposta")These two are not the same set. A column is what comes back; a parameter is what you can filter on. Most columns are both, but not all.
Both work offline: the schema is frozen into the package from the
APIs’ own OpenAPI documents. tg_schema_date() reports
when.
Each filter is named after a parameter, and parameters combine with AND:
propostas <- tg_get(
"parcerias", "proposta",
sg_uf_recebedor = "PE",
situacao_proposta = "Aprovada",
.limit = 100
)That is almost the entire filtering vocabulary. These services compare for equality — there is no greater-than and no pattern match — and they publish no ordering or column-selection parameter.
The exception is on identifiers. Some parameters take several values
in one request and match any of them. tg_params() marks
them as multiple, with the most each accepts — 100 in
especiais, 200 elsewhere:
params <- tg_params("parcerias", "parceria")
params[params$multiple, c("param", "max_values")]
#> # A tibble: 2 × 2
#> param max_values
#> <chr> <int>
#> 1 id_parceria 200
#> 2 id_proposta 200
tg_get("parcerias", "parceria", id_proposta = c(1, 2))The OpenAPI documents do not say which parameters these are, and the names do not either, so the package asked the service when it froze the schema.
For any other parameter, query each value and bind:
library(purrr)
nordeste <- c("PE", "PB", "AL", "RN", "CE", "SE", "BA", "PI", "MA")
propostas <- list_rbind(map(
nordeste,
\(uf) tg_get("parcerias", "proposta", sg_uf_recebedor = uf, .limit = Inf)
))Many parameters accept only a fixed set of values, and the package knows which:
params <- tg_params("parcerias", "proposta")
params[lengths(params$values) > 0, c("param", "values")]
#> # A tibble: 5 × 2
#> param values
#> <chr> <list>
#> 1 sg_uf_recebedor <chr [27]>
#> 2 situacao_proposta <chr [5]>
#> 3 in_situacao_analise <chr [4]>
#> …
params$values[[match("situacao_proposta", params$param)]]
#> [1] "Em Análise" "Rejeitada" "Aprovada" "Em Elaboração"
#> [5] "Inativada"A value outside the set fails before the request is made:
This is the one thing worth internalizing about these APIs.
They ignore a query parameter they do not recognize.
No warning, no 400 — the request succeeds and returns the unfiltered
table. The parameter on /proposta is
situacao_proposta; write in_situacao_proposta,
which is what the sibling /parceria endpoint calls its own
version, and you get every one of the 89,415 proposals instead of the
85,041 that are approved. Nothing in the response says so.
So the package refuses to send a name the frozen schema does not know:
tg_count("parcerias", "proposta", in_situacao_proposta = "Aprovada")
#> Error in `tg_count()`:
#> ! Unknown filter: "in_situacao_proposta".
#> ✖ The API ignores a parameter it does not recognize and returns every row, so
#> this would look like a query that matched nothing in particular.
#> ℹ Did you mean "situacao_proposta"?If the API gains a parameter after the packaged schema was built,
turn the check off with
options(transferegovr.validate = FALSE) — and know what you
are trading away.
For the same reason a repeated parameter is refused rather than sent:
these services keep the last occurrence and discard the rest silently,
so sg_uf_recebedor = "PE", sg_uf_recebedor = "PB" would
quietly mean "PB".
Columns are typed from the frozen schema rather than inferred from the values, so a column that happens to be entirely null on one page does not come back logical while the next page returns it as character.
propostas <- tg_get("parcerias", "proposta", .limit = 5)
class(propostas$dt_proposta)
#> [1] "Date"
class(propostas$vl_total_planejamento_gastos)
#> [1] "numeric"Integers are returned as double. These documents declare no
format, so int32 and int64 cannot be told apart, and
identifiers here genuinely exceed .Machine$integer.max —
cd_parceria reaches 202500037062, which as an integer would
be NA.
Some tables have no endpoint of their own: the API folds them into their parent as an array. Those arrive as list columns.
programas <- tg_get("parcerias", "programa", .limit = 20)
fields <- tg_fields("parcerias", "programa")
fields$field[!is.na(fields$nested)]
#> [1] "ufs_habilitadas" "programa_atende_a" "categorias_despesa"
#> [4] "resultados_esperados" "indicadores_programa"
tg_fields("parcerias", "programa", nested = "ufs_habilitadas")
#> # A tibble: 3 × 5
#> field r_type api_type nested description
#> <chr> <chr> <chr> <chr> <chr>
#> 1 nm_uf character string NA NA
#> 2 sg_uf character string NA NA
#> 3 cd_ibge double integer NA NATo flatten one:
library(dplyr)
library(tidyr)
programas |>
select(id_programa, ufs_habilitadas) |>
unnest_longer(ufs_habilitadas) |>
unnest_wider(ufs_habilitadas)There are 5 such columns in fundoafundo and 13 in
parcerias; especiais has none.
Each module reports when it was last loaded. It is the only freshness
signal these APIs give — they send no ETag,
Cache-Control or Last-Modified header, which
is also why the package caches responses itself rather than relying on
HTTP caching.
vignette("pagination") — collecting a large table
without losing rows.vignette("joining-tables") — how the tables fit
together.