This vignette is a fail-closed workflow for coding agents that use GERDA in a research project. It separates dataset selection, acquisition, schema checks, joins, and provenance so that an agent cannot silently continue after choosing the wrong dataset or losing observations.
Do not guess a dataset name from a prose description. Query the catalog and require one result before downloading:
catalog <- gerda_data_list(print_table = FALSE)
candidates <- subset(
catalog,
election_type == "federal" &
geographic_level == "municipality" &
boundary == "2025"
)
candidates[, c(
"data_name", "year_start", "year_end", "boundary", "formats"
)]
#> # A tibble: 1 × 5
#> data_name year_start year_end boundary formats
#> <chr> <int> <int> <chr> <chr>
#> 1 federal_muni_harm_25 1990 2025 2025 csv,rds
stopifnot(nrow(candidates) == 1L)
data_name <- candidates$data_name[[1]]The boundary choice is part of the research design.
*_unharm files use the geography in force at each election;
*_harm files map observations to a common geography. A
suffix such as _21, _23, or _25
identifies the target boundary vintage. An agent should record why that
vintage matches the analysis.
Use RDS for smaller downloads, make load failures fatal, and opt into caching when repeated runs are expected:
elections <- load_gerda_web(
data_name,
file_format = "rds",
on_error = "stop",
cache = TRUE
)
stopifnot(is.data.frame(elections), nrow(elections) > 0L)The cache improves speed but does not pin an immutable upstream version. For a reproducible project, save the data actually analyzed and its checksum inside the project, subject to the project’s data-storage policy:
dir.create("data/derived", recursive = TRUE, showWarnings = FALSE)
snapshot_path <- file.path("data/derived", paste0(data_name, ".rds"))
saveRDS(elections, snapshot_path)
provenance <- list(
package = "gerda",
package_version = as.character(utils::packageVersion("gerda")),
dataset = data_name,
format = "rds",
retrieved_at_utc = format(Sys.time(), tz = "UTC", usetz = TRUE),
snapshot_path = snapshot_path,
snapshot_md5 = unname(tools::md5sum(snapshot_path))
)
saveRDS(provenance, "data/derived/gerda_provenance.rds")load_gerda_web() reads final data artifacts from the GERDA data
repository. Use that repository when an audit requires source files,
cleaning scripts, or lineage below the package’s final artifacts. Files
named *_raw are the most source-near GERDA representation;
they are not a substitute for inspecting the original source and build
pipeline.
Geographic identifiers must remain character vectors. Re-read the original source with explicit character column types if an identifier became numeric; padding a damaged numeric value later can conceal leading-zero loss.
For the selected municipal panel, a minimal structural audit is:
required <- c("ags", "election_year")
stopifnot(all(required %in% names(elections)))
if (!is.character(elections$ags) ||
any(!is.na(elections$ags) & !grepl("^[0-9]{8}$", elections$ags))) {
stop("ags must contain eight-digit character identifiers")
}
if (any(!stats::complete.cases(elections[required]))) {
stop("Missing geographic or time keys require investigation")
}
if (anyDuplicated(elections[required])) {
stop("ags + election_year is not unique; identify additional row keys")
}
expected_years <- sort(unique(elections$election_year))
years_by_unit <- split(elections$election_year, elections$ags)
units_with_gaps <- names(Filter(
function(x) !identical(sort(unique(x)), expected_years),
years_by_unit
))The correct row key varies by dataset. Candidate, election-round,
party-long, and stimme files need additional key columns.
Never resolve duplicate keys by calling distinct() unless
the duplication has been explained. Inspect missingness in every
analysis variable before filtering or estimation; do not silently drop
rows with missing outcomes, treatments, identifiers, or years.
The two enrichment helpers enforce a many-to-one left-join contract.
They reject numeric or malformed geographic identifiers, duplicate
reference keys, columns that would be overwritten, and any join that
expands the input row count. Set unmatched = "error" in
agent-run pipelines:
enriched <- elections |>
add_gerda_covariates(unmatched = "error") |>
add_gerda_census(unmatched = "error")
join_report <- gerda_join_diagnostics(enriched)
print(join_report)
stopifnot(
all(join_report$input_rows == join_report$output_rows),
all(join_report$unexpected_unmatched_rows == 0L)
)
saveRDS(join_report, "data/derived/gerda_join_diagnostics.rds")For INKAR covariates, election years outside 1995–2022 are retained
with missing joined values and reported as
outside_coverage_rows; they are not treated as unexpected
matches. unmatched = "error" still stops on valid-year keys
that fail to match and on missing join keys in any year. The other modes
are "warn" (the interactive default) and
"ignore". Capture diagnostics immediately after the join
pipeline because unrelated data-frame transformations may remove custom
attributes.
A successful key match does not imply complete covariate values. Individual INKAR and Census variables can still be missing, so inspect payload missingness separately. Census values are a single 2022 snapshot and are time-invariant when attached to elections from other years. Exact identifier matches also do not by themselves establish that the election and covariate boundary vintages are substantively comparable.
INKAR covariates and Census 2022 are bundled with the package and require no download. Their codebooks are machine-readable data frames:
covariates <- gerda_covariates()
covariate_codebook <- gerda_covariates_codebook()
census <- gerda_census()
census_codebook <- gerda_census_codebook()
stopifnot(
!anyDuplicated(covariates[c("county_code", "year")]),
!anyDuplicated(census["ags"]),
all(names(covariates) %in% covariate_codebook$variable),
all(names(census) %in% census_codebook$variable)
)Before handing the project back, an agent should leave the exact
dataset name and boundary rationale, package version, saved snapshot and
checksum, complete row-key definition, missing-key and panel-gap checks,
join diagnostics, variable missingness summary, and the GERDA citation
returned by citation("gerda").