rbiogeme lets an R user write a Biogeme model
specification in R while the native Biogeme engine performs compilation,
likelihood evaluation, differentiation, integration, optimization,
simulation, and reporting. The R package does not create a second
numerical engine and does not require the user to manipulate Python
objects.
This guide is intentionally self-contained. All expressions in the examples below are R expressions, and all names passed to the model are preserved when the specification is compiled by the bridge.
The code chunks are shown as runnable examples but are not evaluated while the package documentation is built. Estimation and simulation require the user’s configured Python environment and can create native output files. Any operation that creates persistent native files requires an explicit output directory; the package never silently writes those files to the current working directory.
Install a source tarball with ordinary R tools. When working from a
checkout, build the tarball first with R CMD build . and
install the resulting file.
The package requires R 4.3 or later and Python 3.12 or later. The
recommended path is to let biogeme_setup() provision the
native requirement. If you manage Python yourself, pass the interpreter
to biogeme_setup() before the first operation that
initializes Python:
library(rbiogeme)
biogeme_setup(
python = "/absolute/path/to/python",
biogeme_requirement = "biogeme==3.3.5"
)
# This reports the versions visible to reticulate.
biogeme_diagnostics()If you do not already manage Python, the shorter recommended path is:
The default native requirement is biogeme==3.3.5. With
no explicit interpreter, biogeme_setup() asks reticulate to
provision an isolated managed environment and installs the requirement
there on first use. With an explicit python, it selects
that existing interpreter and verifies it; it does not install packages
into a user-managed environment. Configuration is session-wide. Calling
biogeme_config() or biogeme_setup() with a
different interpreter after Python has already been initialized raises
an error so that a model cannot accidentally use a different interpreter
from the one selected.
biogeme_check() is the recommended first troubleshooting
step. It reports whether the configured runtime is ready and gives a
corrective action for each failure. Use
biogeme_diagnostics() when the detailed version list is
needed.
For a shorter first-run check after setup, use
biogeme_check(). It catches runtime initialization
failures, verifies the minimum R and Python versions, confirms that
Biogeme can be imported, and gives a corrective action for each
failure:
check <- biogeme_check()
if (!check$ready) {
print(check)
stop("The rbiogeme environment is not ready.")
}There are two supported setup paths. If you do not already manage
Python, omit python and let biogeme_setup()
provision biogeme==3.3.5. If you already have a Python
environment, install Biogeme into it before starting R. For example,
outside R:
python3.12 -m venv /path/to/rbiogeme-venv
/path/to/rbiogeme-venv/bin/python -m pip install "biogeme==3.3.5"
Then select that interpreter before any operation that initializes Python:
Selecting an existing interpreter does not install Biogeme into it.
If configuration fails because Python has already been initialized,
restart R and call biogeme_config() before constructing a
database or model.
Biogeme databases contain numeric columns. A database copies its input data at construction time, so later changes to the original data frame do not change the model specification.
data <- data.frame(
choice = c(1, 2, 1, 2, 1, 2, 1, 2),
time = c(10, 8, 12, 7, 11, 9, 13, 8),
cost = c(5, 7, 6, 8, 5, 7, 6, 9),
income = c(1, 2, 1, 3, 2, 1, 3, 2)
)
database <- biogeme_database("demo", data)
# variable() is symbolic: it refers to a database column, not to an R vector.
time <- variable("time")
cost <- variable("cost")
income <- variable("income")
# biogeme_beta() creates a named native parameter. The name is part of the
# equivalence contract and will appear unchanged in the result.
b_time <- biogeme_beta("b_time", start = 0)
b_cost <- biogeme_beta("b_cost", start = 0)
asc_2 <- biogeme_beta("asc_2", start = 0)
utility_1 <- b_time * time + b_cost * cost
utility_2 <- asc_2 + b_time * time + b_cost * costArithmetic operators build a neutral expression tree. They do not
calculate a likelihood in R. Constants such as 0 are
accepted wherever a scalar expression is expected.
Comparisons and logical operators are symbolic too. They are useful for availability, filtering, piecewise definitions, and subsets:
available_2 <- (cost < 10) & (income >= 1)
not_available_2 <- !(available_2)
either_condition <- (time < 9) | (cost > 8)
# Native-safe mathematical primitives are available as expression functions.
safe_probability <- logzero(logit_probability(
utilities = list(`1` = utility_1, `2` = utility_2),
availability = list(`1` = 1, `2` = available_2),
alternative = variable("choice")
))Use logzero() for a native numerically safe logarithm.
Other commonly used functions include normal_cdf(),
normal_pdf(), safe_exp(), sqrt(),
abs(), biogeme_min(),
biogeme_max(), Elem(),
piecewise(), boxcox(), and
derive().
logit_model() is the shortest route for a
cross-sectional multinomial logit model. The names of
utilities identify alternatives. If those names are integer
strings, they are also used as the observed choice codes.
model <- logit_model(
database = database,
choice = "choice",
utilities = list(
`1` = utility_1,
`2` = utility_2
),
availability = list(
`1` = 1,
`2` = available_2
)
)
# Check the native specification before running the optimizer.
validation <- validate_model(model)
stopifnot(validation$valid)
# A fresh temporary output directory prevents an old YAML or iteration file
# from being reused during an equivalence test.
output_directory <- tempfile("rbiogeme-demo-")
dir.create(output_directory)
fit <- estimate(
model,
model_name = "rbiogeme_demo",
control = biogeme_control(
output_directory = output_directory,
generate_html = FALSE,
generate_yaml = FALSE,
save_iterations = FALSE
)
)estimate() always performs a fresh native estimation.
The result is an R object containing serialized native results, not a
live Python result object. The usual R methods expose the central
post-estimation information:
The names in coef(fit) are the names supplied to
biogeme_beta(). Do not rename or reorder them when
comparing an R fit with a native fit.
Native report and checkpoint files are opt-in and must have an explicit destination. For example, a user-facing run can choose a project directory:
output_directory <- "/absolute/path/to/my-biogeme-results"
fit_with_reports <- estimate(
model,
model_name = "my_model_with_reports",
control = biogeme_control(
output_directory = output_directory,
generate_html = TRUE,
generate_yaml = TRUE,
save_iterations = TRUE
)
)Use tempdir() instead when the files are only
intermediate artifacts in a test or a short demonstration. The
command-line examples in inst/examples/ follow the same
rule through their required --output=/path/to/output
argument.
Before estimation, validate_model(model) performs native
specification validation without running the optimizer or writing
estimation results. After fitting a logit model,
predict(fit) evaluates one native probability column per
alternative. Scenario predictions use
predict(fit, newdata = data.frame(...)); the supplied data
frame must contain the variables referenced by the model.
Specialized constructors are convenient, but every model can be
expressed with biogeme_model(). This is the general
interface for a complete likelihood, an optional probability, named
simulation expressions, weights, panels, draws, subsets, and parameter
overrides.
probability <- logit_probability(
utilities = list(`1` = utility_1, `2` = utility_2),
availability = list(`1` = 1, `2` = available_2),
alternative = variable("choice")
)
generic_model <- biogeme_model(
database = database,
formula = logzero(probability),
probability = probability,
simulations = list(
probability = probability,
time_cost_ratio = time / cost
)
)
generic_fit <- estimate(
generic_model,
model_name = "rbiogeme_generic",
control = biogeme_control(
output_directory = output_directory,
generate_html = FALSE,
generate_yaml = FALSE,
save_iterations = FALSE
)
)
simulated <- simulate(
generic_model,
beta = generic_fit,
control = biogeme_control(output_directory = output_directory)
)
as.data.frame(simulated)The formula is the expression used for estimation.
probability documents the corresponding probability when a
generic model has one, while simulations is a named list of
expressions evaluated by native Biogeme at the supplied estimates.
Database transformations can remain symbolic until the bridge compiles the complete model. This keeps derived-variable and filtering semantics in native Biogeme and avoids an R callback during numerical evaluation.
database_with_ratio <- biogeme_database_define_variable(
database,
name = "time_cost_ratio",
expression = variable("time") / variable("cost")
)
database_without_high_cost <- biogeme_database_remove(
database_with_ratio,
condition = variable("cost") > 8
)
biogeme_database_columns(database_without_high_cost)
biogeme_database_nrow(database_without_high_cost)
biogeme_database_filtered_row_count(database_without_high_cost)
biogeme_database_row_ids(database_without_high_cost)
# Native derived columns and filters are materialized explicitly when their
# resulting data frame or row count is needed in R.
materialized <- biogeme_database_materialize(database_without_high_cost)
as.data.frame(materialized)For panels, the panel identifier must be a numeric column and observations for each individual must already be contiguous. The constructor validates that ordering before the model is compiled:
Native estimation can create YAML, HTML, iteration, NetCDF, checkpoint, or diagnostic files depending on the operation. Use an operation-specific temporary directory in tests and examples. Disable each file type that is not part of the behavior under test:
clean_control <- biogeme_control(
output_directory = tempfile("rbiogeme-output-"),
seed = 1234,
generate_html = FALSE,
generate_yaml = FALSE,
save_iterations = FALSE
)For operations where loading an old result is the feature being
demonstrated, use estimate_or_load() explicitly and set
force deliberately. Ordinary estimate() calls
do not silently recycle a previous result.
| Goal | Start with |
|---|---|
| Estimate a standard choice model | logit_model() and estimate() |
| Specify a custom likelihood | biogeme_model() |
| Work with panel observations | biogeme_panel_database() and
panel_likelihood_trajectory() |
| Evaluate probabilities or scenarios | predict() and simulate() |
| Use Bayesian, MDCEV, Monte Carlo, catalog, hybrid-choice, or sampling features | The corresponding specialized constructor and
advanced-models |
The reference manual is available through R’s normal help system. The most useful entry points are:
If a first model does not run, use this order:
biogeme_check() and resolve every
ERROR row.biogeme_database_columns(database) and compare the
result with the names used by variable().validate_model(model) to check the compiled native
specification without running the optimizer.biogeme_config(debug = TRUE) in a fresh R session
if the bridge error still needs its native traceback.For more specialized model families, continue with
modeling-workflows and advanced-models.