Release notes
Upgrading? See Upgrading for the steps and the breaking changes.
Changelog
0.6.1 (unreleased)
- Fix: backups to HTTPS S3 endpoints failed in 0.6.0 (AWS, R2, Hetzner, Backblaze…): every backup job panicked with “Could not automatically determine the process-level CryptoProvider”. The 0.6 dependencies enable two rustls crypto providers (aws-lc-rs and ring) and none was chosen; pgbx now installs aws-lc-rs explicitly in the extension and the CLI. Plain-http endpoints were not affected, which is why the test suites (local S3 over http) missed it.
- New test
tests/https_s3_e2e.sh: real https calls to AWS S3 with dummy keys, from the extension’s backup job and the CLI; it reproduces the 0.6.0 panic and passes on 0.6.1. Run bytests/run_all.sh. - No schema change (
pgbx--0.6.0--0.6.1.sqlis empty); the worker updates every database’s extension version itself.
0.6.0
-
Windows: one x64 build (
pgbx-windows-amd64.zip). Windows on ARM runs it through its built-in x64 emulation, andinstall.ps1installs it there; the separate arm64 Windows build is dropped (its cross-build of the TLS/crypto C libraries fails). -
Breaking: pgbx’s built-in SSH is gone (ADR 0003).
--ssh,--ssh-port,--ssh-jump,--tunnel-idleandpgbx tunnelare removed (they now fail with a pointer to the ssh adapter); doctor / logs / diagnose / setup no longer run over SSH. Use the ssh example adapter:pgbx profile add prod --adapter ssh target=ops@db1 user=app 'password=$PGPASSWORD'. -
Breaking: profiles live in
config.yaml. Aprofiles.jsonis migrated automatically on first use (SSH profiles becomeadapter: sshprofiles, direct onesurl:profiles); pgbx says so once and keeps the old file asprofiles.json.migrated. -
Connection adapters (ADR 0003): a profile is a connection string or an adapter, any program that hands pgbx a connection string over a two-message stdin protocol (own process group / Job Object, 30 s ready timeout, stop + 5 s grace, Ctrl-C stops it). One-off commands start and stop it;
pgbx servekeeps one per connection. Examples for ssh, aws (SSM), gcp (Cloud SQL Auth Proxy) and azure inadapters/, shipped in the release archives and copied to<config dir>/adaptersby the installers. -
config.yaml(0600) withdefault,secrets,adapters,profiles;pgbx profile edit;--adapter/--adapter-command/key=valuesettings onprofile add;--url/PGBX_URLfor one run. -
$VAR/${VAR}references anywhere in a profile or url, expanded at run time ($$=$) from the environment, then thesecrets:source (a .env file or a handler command run asCMD NAME). pgbx stores no secrets: literal passwords are refused, and every output masks passwords and expanded secrets.pg_restoregets the password through its environment. -
TLS to Postgres for every CLI connection, with libpq’s
sslmode:disable,allow/prefer(the default: TLS when the server offers it),require,verify-ca,verify-full, andsslrootcert(a PEM file, else~/.postgresql/root.crt, else the built-in Mozilla roots plus the OS store). From the connection string, then the profile’ssslmode:/sslrootcert:keys, thenPGSSLMODE/PGSSLROOTCERT. Servers that force TLS (Azure flexible server, RDS withrds.force_ssl) and the aws / azure adapters’sslmode=requireURLs (iam_auth,entra_auth) now connect;pg_restoregets the same settings. rustls, so still no OpenSSL and one static binary. A connection that fails says why (invalid peer certificate: UnknownIssuer, …). Behaviour change: a server withssl=onis now reached over TLS by default, aspsqldoes. -
Resource caps (ADR 0001 §3): pg_dump / pg_restore run at
pgbx.job_nice(10) and, on Linux, IO prioritypgbx.job_ionice(best-effort-7), set before exec; their connections are namedpgbx_dump/pgbx_restore/pgbx_verify. -
Never queued behind DDL: pg_dump waits at most
pgbx.dump_lock_timeout(5s) for its table locks; on a timeout the backup stays queued and is retried withpgbx.defer_backoff, untilpgbx.max_defer(4h, first backup 15min, never past one schedule interval). Then it runsforcedwithpgbx.dump_lock_timeout_forced(60s) andpgbx.dump_compression_busy, and fails + alerts if it still cannot lock. -
pg_restore runs with
synchronous_commit=off(pgbx.restore_synchronous_commit). -
Bandwidth caps
pgbx.upload_kbps/pgbx.download_kbps(KiB/s, 0 = unlimited). -
doctor():
long_running_job(pgbx.doctor_long_job, 1h). -
backup_now()/verify_now()return the job already queued instead of adding another (pgbx.coalesce_manual, on); newpgbx.cancel(job_id)cancels a queued job, or a running one: its child is killed, the upload aborted (nothing left in S3), a half-restored database dropped, no alert; statecancelled. -
One server-wide job queue (ADR 0001 §0): the worker supervises jobs as child processes instead of blocking on them, picks restore > manual backup > scheduled backup > restore test > prune, then oldest, then round-robin over databases;
pgbx.max_concurrent_jobs(1) enforced by advisory-lock slots in the admin database;pgbx.restore_lane(on) gives restores their own slot.pgbx.overrun_policy(skip, orcatch_up) withpgbx.overrun_max_gap(1.5 intervals): a dump longer than its interval no longer runs back to back (params.skipped_slots). doctor():dump_longer_than_interval. Restart recovery and orphan-upload abort keep working (crash e2e). -
pgbx.server_queue(admin database) and status()running_job,next_job,queue_position,waiting_reasonsay what runs and why a job waits; CLIpgbx jobs [cancel ID --yes]; a queue card inpgbx ui;pgbx uiand--waitknow thecancelledstate. -
Time estimates (§4): a NOTICE on
backup_now()/verify_now()/restore()with start, duration, size, confidence and the bottleneck;pgbx.job_eta(job_id); live progress (history.bytesafter every 16 MiB) in status()job_progress/job_eta,pgbx jobsand the UI; a daily cpu probe and network / disk rates inpgbx.server_capacity; doctor():capacity,eta_accuracy. Settingspgbx.eta_samples,pgbx.eta_default_mbps,pgbx.eta_calibrate. -
Quiet-window suggestion (§2): activity per hour of the week (UTC) learned from
pg_stat_databasedeltas intopgbx.activity_hourly(pgbx.activity_sampling,pgbx.activity_decay);pgbx.suggest_window(), status()suggested_schedule, CLIpgbx schedule suggest [--apply], a suggestion card with copy-to-apply buttons in the UI, doctor():schedule_in_quiet_window(pgbx.doctor_busy_ratio,pgbx.suggest_min_days). Never applied by itself. -
Load gate (§1), default shadow: one load sample per poll (active sessions, tps, long writers, replica lag, load average;
pgbx.busy_*),params.would_deferon jobs it would have held back, never a delay. Per databaseconfigure(load_gate => 'on')(newload_gateargument) defers scheduled backups / restore tests withpgbx.defer_backoffup topgbx.max_defer, then runs themforced. Human jobs get a NOTICE that they compete with the app (pgbx.gate_manual_jobs). CLIpgbx load [--gate ...], a load card in the UI, doctor():load_gate,forced_backups_7d. -
bench/load_overhead.sh+bench/RESULTS.md: pgbench TPS and latency with no backup, an uncapped and a capped one. -
Fix: multipart uploads cut off by a crash are aborted at the next worker start even where the S3 listing of open uploads cannot be parsed (RustFS): the worker keeps its own list (
pgbx_open_uploadsin the data directory). -
Fix: a shutdown during a restore download no longer hangs (the writer returned
Interrupted, whichwrite_allretries). -
Client-side encryption (off by default):
pgbx.encryption_key_file(32 random bytes, base64 or hex, chmod 600) encrypts every dump and its roles file with AES-256-GCM in 4 MiB authenticated frames, streamed (no temp files, constant memory). Restores, restore tests andpgbx db-restore --from-s3 --key-filedecrypt in the stream; tampering, truncation, reordered frames and a wrong key fail loudly; older unencrypted dumps still restore.pgbx decryptfor download links (they serve the ciphertext).history.params.encrypted. -
Roles with every backup:
<ts>.globals.sql.zstnext to each dump (pg_dumpall --globals-only, no passwords unlesspgbx.backup_role_passwords = on) plus the roles the database references.restore(..., with_roles => true, roles => 'referenced'|'all'),pgbx db-restore [--from-s3] --with-roles [--roles R]: missing roles are created, existing ones never changed, owners kept; running it twice changes nothing. Pruning deletes each dump’s roles file. -
Notifications:
pgbx.notify = 'slack:NAME, telegram:NAME, webhook:NAME, email:NAME', URLs and tokens only inpgbx.notify_secrets_file(chmod 600); one message per incident plus one “OK again”, sent from the job’s thread. A URL written intopgbx.notifyis refused and the error never repeats it.pgbx.alert_commandunchanged. -
Prometheus:
GET /metricsonpgbx ui,pgbx metrics;server_overviewgainedlast_backup_bytes,failures_total,queued_jobs,last_verify_ok,last_backup_encrypted. -
GFS retention:
set_retention(gfs => '7d,4w,12m')('off'clears it; a span beyondpgbx.max_days_limitis refused),pgbx retention --gfs(changing or clearing an existing spec needs--yes);status()shows it. -
Point-in-time restore (optional, whole server, no pgBackRest):
pgbx setup pitr [--yes](aliassetup --pitr; ALTER SYSTEM archive_mode / archive_command /pgbx.pitr, refuses a foreign archive_command, one restart).pgbx wal-push(archive_command: zstd + sha256, never overwrites a different checksum, async spool with parallel look-ahead, drops WAL pastpgbx.wal_queue_maxinstead of filling the disk and records the gap) andpgbx wal-get(restore_command: verifies every file, parallel prefetch). Base backups (pg_basebackup→ zstd → parallel multipart) are jobs of the server-wide queue (kindbase_backupin the admin database:pgbx jobs, cancel, prompt shutdown), onpgbx.pitr_schedule, kept forpgbx.pitr_retentiontogether with their WAL; a gap is healed by a base backup queued automatically.pgbx pitr status|list|backup-now|restore --time TS|latest --target DIR(never starts Postgres; a copy gets archive_mode=off and its own server_name; in place only with--yes-replace-whole-serveron a stopped server, moved aside). WAL archiving incidents and gaps alert throughpgbx.alert_commandandpgbx.notify, from a thread. SQLpitr_status(),pitr_backup_now(),pgbx.pitr_state; history kindsbase_backup,wal_gap,wal_archive, triggerwal_gap; doctor() rowspitr archiving,pitr base backups,pitr gaps; a PITR card inpgbx ui. Settingspgbx.pitr,pgbx.pitr_schedule,pgbx.pitr_retention,pgbx.wal_queue_max,pgbx.wal_gap_margin,pgbx.wal_alert_after,pgbx.wal_alert_size,pgbx.work_dir,pgbx.cli_path. Designs ported from pgBackRest (MIT, seeNOTICE). -
S3 credentials without a keys file (EC2 instance role, for the AWS Marketplace AMI):
pgbx.credentials_file = 'aws-default'(or empty) uses the AWS default chain instead of static keys: envAWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY/AWS_SESSION_TOKEN(CLI only), web identity (EKS IRSA), container credentials (ECS, EKS Pod Identity), then the EC2 instance role through IMDSv2 (IMDSv1 is never used;AWS_EC2_METADATA_SERVICE_ENDPOINTis honoured). Temporary credentials are cached and renewed 5 minutes before they expire or on a 403ExpiredToken, between multipart parts and download resumes, so long transfers run across a refresh. Everywhere pgbx talks to S3: the worker and its jobs (which still read no settings: the job’s own thread resolves the credentials, so the poll loop never waits for a metadata service),wal-push/wal-get/pitr,backups --from-s3anddb-restore --from-s3(--credentials-filemay now be left out, meaningaws-default).pgbx setup server --credentials aws-defaultwrites the setting and no keys file. doctor(): thecredentials filerow is nows3 credentialsand names the source in use (file,env,web-identity,ecs,instance-role, with the role and expiry), or why none works and the IAM policy to attach;s3 settingsno longer requirespgbx.credentials_file. No key, secret or token is ever logged or stored. A non-empty file path behaves exactly as before. -
SQL signature changes:
set_retention(int, int, text),restore(text, timestamptz, bool, text); old calls keep working through the defaults. -
Fix: the startup cleanup of orphaned multipart uploads runs on a thread of its own (an unreachable S3 no longer holds up the worker’s first poll) and only aborts the worker’s own objects (
.dump,.globals.sql.zst) older than 10 minutes, so it can never cut off a base backup anotherpgbxprocess is uploading. -
Fix: a shutdown waits at most 5 s for running jobs (was 20 s), so a job stuck in an S3 request that does not answer no longer delays Postgres; the job is marked interrupted at the next start. A job child its thread has not reaped yet is killed and reaped before the worker exits.
-
Fix: no child process of pgbx is ever left to the postmaster. Where Postgres runs as PID 1 (containers), an orphan that died by a signal made the postmaster restart the whole server (seen when a base backup was cancelled: pg_basebackup died of SIGPIPE).
pgbx pitr backupreaps pg_basebackup on SIGTERM; the detached wal-push / wal-get background processes only ever exit 0 or 1. -
Schema update script
pgbx--0.5.0--0.6.0.sql(the worker applies it by itself) covers all of the above; an updated database equals a fresh 0.6.0 install.
0.5.0
First public release.
- Per-database backups: every database (and
template1) gets the extension;pg_dump/pg_restorestreamed to and from S3. Noarchive_modeor extra package needed. pgbx backups --from-s3andpgbx db-restore --from-s3 [--backup KEY | --time TS]— restore onto a new server without the extension (its objects are skipped), streamed with resumable download.- PostgreSQL 13–18.
pgbx.dump_compressiondefaults toauto(zstd:3 with pg_dump 16+, gzip level 6 before); the newest installed pg_dump/pg_restore is used. - The worker runs
ALTER EXTENSION pgbx UPDATEitself when a database ortemplate1is older than the library. - doctor(): extension loaded, s3 settings, credentials file, database backups, restore tests, workers, archive_mode (info), replication_slots.
- Docker images take a
PG_MAJORbuild arg.