Getting started with rbiogeme

rbiogeme contributors

rbiogeme lets an R user write a Biogeme model specification in R while the native Biogeme engine performs compilation, likelihood evaluation, differentiation, integration, optimization, simulation, and reporting. The R package does not create a second numerical engine and does not require the user to manipulate Python objects.

This guide is intentionally self-contained. All expressions in the examples below are R expressions, and all names passed to the model are preserved when the specification is compiled by the bridge.

The code chunks are shown as runnable examples but are not evaluated while the package documentation is built. Estimation and simulation require the user’s configured Python environment and can create native output files. Any operation that creates persistent native files requires an explicit output directory; the package never silently writes those files to the current working directory.

Installation and Python configuration

Install a source tarball with ordinary R tools. When working from a checkout, build the tarball first with R CMD build . and install the resulting file.

install.packages(
  "/path/to/rbiogeme_0.1.2.tar.gz",
  repos = NULL,
  type = "source"
)

The package requires R 4.3 or later and Python 3.12 or later. The recommended path is to let biogeme_setup() provision the native requirement. If you manage Python yourself, pass the interpreter to biogeme_setup() before the first operation that initializes Python:

library(rbiogeme)

biogeme_setup(
  python = "/absolute/path/to/python",
  biogeme_requirement = "biogeme==3.3.5"
)

# This reports the versions visible to reticulate.
biogeme_diagnostics()

If you do not already manage Python, the shorter recommended path is:

library(rbiogeme)
check <- biogeme_setup()
stopifnot(check$ready)

The default native requirement is biogeme==3.3.5. With no explicit interpreter, biogeme_setup() asks reticulate to provision an isolated managed environment and installs the requirement there on first use. With an explicit python, it selects that existing interpreter and verifies it; it does not install packages into a user-managed environment. Configuration is session-wide. Calling biogeme_config() or biogeme_setup() with a different interpreter after Python has already been initialized raises an error so that a model cannot accidentally use a different interpreter from the one selected.

biogeme_check() is the recommended first troubleshooting step. It reports whether the configured runtime is ready and gives a corrective action for each failure. Use biogeme_diagnostics() when the detailed version list is needed.

For a shorter first-run check after setup, use biogeme_check(). It catches runtime initialization failures, verifies the minimum R and Python versions, confirms that Biogeme can be imported, and gives a corrective action for each failure:

check <- biogeme_check()
if (!check$ready) {
  print(check)
  stop("The rbiogeme environment is not ready.")
}

There are two supported setup paths. If you do not already manage Python, omit python and let biogeme_setup() provision biogeme==3.3.5. If you already have a Python environment, install Biogeme into it before starting R. For example, outside R:

python3.12 -m venv /path/to/rbiogeme-venv
/path/to/rbiogeme-venv/bin/python -m pip install "biogeme==3.3.5"

Then select that interpreter before any operation that initializes Python:

biogeme_setup(python = "/path/to/rbiogeme-venv/bin/python")

Selecting an existing interpreter does not install Biogeme into it. If configuration fails because Python has already been initialized, restart R and call biogeme_config() before constructing a database or model.

Data and expressions

Biogeme databases contain numeric columns. A database copies its input data at construction time, so later changes to the original data frame do not change the model specification.

data <- data.frame(
  choice = c(1, 2, 1, 2, 1, 2, 1, 2),
  time = c(10, 8, 12, 7, 11, 9, 13, 8),
  cost = c(5, 7, 6, 8, 5, 7, 6, 9),
  income = c(1, 2, 1, 3, 2, 1, 3, 2)
)

database <- biogeme_database("demo", data)

# variable() is symbolic: it refers to a database column, not to an R vector.
time <- variable("time")
cost <- variable("cost")
income <- variable("income")

# biogeme_beta() creates a named native parameter. The name is part of the
# equivalence contract and will appear unchanged in the result.
b_time <- biogeme_beta("b_time", start = 0)
b_cost <- biogeme_beta("b_cost", start = 0)
asc_2 <- biogeme_beta("asc_2", start = 0)

utility_1 <- b_time * time + b_cost * cost
utility_2 <- asc_2 + b_time * time + b_cost * cost

Arithmetic operators build a neutral expression tree. They do not calculate a likelihood in R. Constants such as 0 are accepted wherever a scalar expression is expected.

Comparisons and logical operators are symbolic too. They are useful for availability, filtering, piecewise definitions, and subsets:

available_2 <- (cost < 10) & (income >= 1)
not_available_2 <- !(available_2)
either_condition <- (time < 9) | (cost > 8)

# Native-safe mathematical primitives are available as expression functions.
safe_probability <- logzero(logit_probability(
  utilities = list(`1` = utility_1, `2` = utility_2),
  availability = list(`1` = 1, `2` = available_2),
  alternative = variable("choice")
))

Use logzero() for a native numerically safe logarithm. Other commonly used functions include normal_cdf(), normal_pdf(), safe_exp(), sqrt(), abs(), biogeme_min(), biogeme_max(), Elem(), piecewise(), boxcox(), and derive().

A complete multinomial logit model

logit_model() is the shortest route for a cross-sectional multinomial logit model. The names of utilities identify alternatives. If those names are integer strings, they are also used as the observed choice codes.

model <- logit_model(
  database = database,
  choice = "choice",
  utilities = list(
    `1` = utility_1,
    `2` = utility_2
  ),
  availability = list(
    `1` = 1,
    `2` = available_2
  )
)

# Check the native specification before running the optimizer.
validation <- validate_model(model)
stopifnot(validation$valid)

# A fresh temporary output directory prevents an old YAML or iteration file
# from being reused during an equivalence test.
output_directory <- tempfile("rbiogeme-demo-")
dir.create(output_directory)

fit <- estimate(
  model,
  model_name = "rbiogeme_demo",
  control = biogeme_control(
    output_directory = output_directory,
    generate_html = FALSE,
    generate_yaml = FALSE,
    save_iterations = FALSE
  )
)

estimate() always performs a fresh native estimation. The result is an R object containing serialized native results, not a live Python result object. The usual R methods expose the central post-estimation information:

summary(fit)
coef(fit)
vcov(fit)
logLik(fit)
nobs(fit)

The names in coef(fit) are the names supplied to biogeme_beta(). Do not rename or reorder them when comparing an R fit with a native fit.

Choosing an output directory

Native report and checkpoint files are opt-in and must have an explicit destination. For example, a user-facing run can choose a project directory:

output_directory <- "/absolute/path/to/my-biogeme-results"
fit_with_reports <- estimate(
  model,
  model_name = "my_model_with_reports",
  control = biogeme_control(
    output_directory = output_directory,
    generate_html = TRUE,
    generate_yaml = TRUE,
    save_iterations = TRUE
  )
)

Use tempdir() instead when the files are only intermediate artifacts in a test or a short demonstration. The command-line examples in inst/examples/ follow the same rule through their required --output=/path/to/output argument.

Before estimation, validate_model(model) performs native specification validation without running the optimizer or writing estimation results. After fitting a logit model, predict(fit) evaluates one native probability column per alternative. Scenario predictions use predict(fit, newdata = data.frame(...)); the supplied data frame must contain the variables referenced by the model.

Generic likelihoods and simulations

Specialized constructors are convenient, but every model can be expressed with biogeme_model(). This is the general interface for a complete likelihood, an optional probability, named simulation expressions, weights, panels, draws, subsets, and parameter overrides.

probability <- logit_probability(
  utilities = list(`1` = utility_1, `2` = utility_2),
  availability = list(`1` = 1, `2` = available_2),
  alternative = variable("choice")
)

generic_model <- biogeme_model(
  database = database,
  formula = logzero(probability),
  probability = probability,
  simulations = list(
    probability = probability,
    time_cost_ratio = time / cost
  )
)

generic_fit <- estimate(
  generic_model,
  model_name = "rbiogeme_generic",
  control = biogeme_control(
    output_directory = output_directory,
    generate_html = FALSE,
    generate_yaml = FALSE,
    save_iterations = FALSE
  )
)

simulated <- simulate(
  generic_model,
  beta = generic_fit,
  control = biogeme_control(output_directory = output_directory)
)
as.data.frame(simulated)

The formula is the expression used for estimation. probability documents the corresponding probability when a generic model has one, while simulations is a named list of expressions evaluated by native Biogeme at the supplied estimates.

Database operations

Database transformations can remain symbolic until the bridge compiles the complete model. This keeps derived-variable and filtering semantics in native Biogeme and avoids an R callback during numerical evaluation.

database_with_ratio <- biogeme_database_define_variable(
  database,
  name = "time_cost_ratio",
  expression = variable("time") / variable("cost")
)

database_without_high_cost <- biogeme_database_remove(
  database_with_ratio,
  condition = variable("cost") > 8
)

biogeme_database_columns(database_without_high_cost)
biogeme_database_nrow(database_without_high_cost)
biogeme_database_filtered_row_count(database_without_high_cost)
biogeme_database_row_ids(database_without_high_cost)

# Native derived columns and filters are materialized explicitly when their
# resulting data frame or row count is needed in R.
materialized <- biogeme_database_materialize(database_without_high_cost)
as.data.frame(materialized)

For panels, the panel identifier must be a numeric column and observations for each individual must already be contiguous. The constructor validates that ordering before the model is compiled:

panel_data <- data.frame(
  person = c(1, 1, 2, 2, 3, 3),
  choice = c(1, 2, 2, 1, 1, 2),
  time = c(10, 8, 9, 11, 12, 7)
)

panel_database <- biogeme_panel_database(
  name = "demo_panel",
  data = panel_data,
  panel_id = "person"
)

biogeme_database_is_panel(panel_database)

Reproducible output

Native estimation can create YAML, HTML, iteration, NetCDF, checkpoint, or diagnostic files depending on the operation. Use an operation-specific temporary directory in tests and examples. Disable each file type that is not part of the behavior under test:

clean_control <- biogeme_control(
  output_directory = tempfile("rbiogeme-output-"),
  seed = 1234,
  generate_html = FALSE,
  generate_yaml = FALSE,
  save_iterations = FALSE
)

For operations where loading an old result is the feature being demonstrated, use estimate_or_load() explicitly and set force deliberately. Ordinary estimate() calls do not silently recycle a previous result.

Finding the right starting point

Goal Start with
Estimate a standard choice model logit_model() and estimate()
Specify a custom likelihood biogeme_model()
Work with panel observations biogeme_panel_database() and panel_likelihood_trajectory()
Evaluate probabilities or scenarios predict() and simulate()
Use Bayesian, MDCEV, Monte Carlo, catalog, hybrid-choice, or sampling features The corresponding specialized constructor and advanced-models

The reference manual is available through R’s normal help system. The most useful entry points are:

?rbiogeme
?biogeme_check
?biogeme_model
?estimate
?simulate

If a first model does not run, use this order:

  1. Run biogeme_check() and resolve every ERROR row.
  2. Run biogeme_database_columns(database) and compare the result with the names used by variable().
  3. Run validate_model(model) to check the compiled native specification without running the optimizer.
  4. Use biogeme_config(debug = TRUE) in a fresh R session if the bridge error still needs its native traceback.

For more specialized model families, continue with modeling-workflows and advanced-models.