<!--
CoderLegion · Article · Post 02.1
Serie: Software auditable (S1)
Autor: Ignacio Badenes (@yosoyignicion / IgnicionDev)
Idioma: bilingüe ES + EN · Read time: ~9 min
Tags: python, database, testing, opensource
Agradecimiento: comentario de @Mike Dabydeen (Líder de Comunidad)
-->
¿Tu backup sobrevivirá a la próxima versión? La prueba que me faltaba
Serie: Software auditable — Post 02.1
Read time: ~9 min · Etiquetas: python database testing opensource
Creí haber cerrado el problema
Hay una tranquilidad que no es ingeniería: es la sensación de haber terminado algo porque ya no duele. El Post 02 terminaba con una frase que me gustaba —recuperar es el trabajo— y con un test que restauraba una copia en una máquina limpia. Lo di por resuelto.
No lo estaba. Había confundido mi tranquilidad con una garantía.
Esto no va de arreglar un if. Va de un error de perspectiva: miré el sistema en el eje del espacio (¿funciona en otra máquina?) y me olvidé del eje del tiempo (¿funcionará cuando el software haya cambiado?). Un backup es, por definición, una promesa que se hace al futuro. Y yo la estaba validando solo en el presente.
El comentario que volvió a mirar donde yo no miraba
Mike Dabydeen, líder de la comunidad, leyó el post y escribió lo que yo no había querido ver. Su frase central: mi prueba demostraba que la copia de hoy se restaura con el código de hoy. Justo el caso que nadie necesita. Quien de verdad depende de un backup quiere restaurar una copia de hace un año con la versión que exista dentro de un año.
Cito sus dos ideas, porque las dos cambiaron el diseño:
"La prueba en máquina limpia demuestra que la copia de hoy se restaura con el código de hoy. Lo que todavía no cubre es que una copia hecha con la 1.3.0 se restaure con la versión que publiques dentro de un año, y ese es justo el caso de quien de verdad necesita su respaldo. Yo guardaría en el repo un paquete real generado por la 1.3.0 como accesorio y lo restauraría en CI en cada versión."
"Sobre el error del -wal, hay otra opción: restaurar con la misma API de copia de seguridad de SQLite, pero en sentido inverso... la copia mantiene abierta una transacción de escritura en el destino durante todo el proceso y la revierte si no termina, y el -wal que encuentra es el de esa misma base, así que no hay archivos huérfanos que borrar."
Tenía razón en las dos. Y la primera duele más, porque no es un bug: es una promesa sin testigo.
Acto I — El tiempo como testigo
La trampa del test que pasa
Mi test de recuperación era honesto y, aun así, insuficiente. Hacía esto:
# El test que probaba el presente
bundle = create_backup(source) # se genera con el código actual
restore_backup(bundle, clean) # y se restaura con el código actual
assert count_scans(clean) == 2
Generación y restauración usaban el mismo código. Un test cerrado sobre sí mismo: si mañana cambio el esquema y rompo la ruta de restauración, este test seguiría verde mientras las copias de los usuarios dejarían de restaurarse. La prueba de mi promesa era la promesa mirándose en un espejo.
Una promesa solo es real si sobrevive a quien la hizo.
La solución: una escalera de fixtures
Congelé un paquete real generado por la 1.3.0 y lo commité como golden file. En cada versión, CI debe restaurarlo en una máquina limpia:
tests/fixtures/backups/
├── aethernet-1.3.0-schema5.tar.gz # el paquete real, ~5 KB
├── aethernet-1.3.0-schema5.expected.json # su procedencia y las aserciones
└── README.md # cómo ampliar la escalera
El .expected.json guarda lo que hace verificable la prueba: sqlite_version, page_size, db_sha256 y los conteos que deben sobrevivir (escaneos, dispositivos, eventos). El test recorre todos los paquetes de la carpeta, no uno:
@pytest.mark.parametrize("bundle", sorted(FIXTURES.glob("*.tar.gz")), ids=lambda p: p.name)
def test_historical_backup_restores_on_clean_machine(bundle, tmp_path):
expected = json.loads(_expected_path(bundle).read_text())
result = restore_backup(bundle, _clean_paths(tmp_path)) # máquina limpia
repo = Repository(Database(result.db_path))
assert repo.counts()["wifi_scans"] == expected["counts"]["wifi_scans"]
# ...informe regenerado, config restaurada, esquema migrado a v5
La regla de mantenimiento la dejó escrita el propio Mike: cuando cambie el esquema o el formato, añade un paquete nuevo y conserva los anteriores. Lo automaticé en un generador, scripts/gen_backup_fixture.py, que usa siempre el código actual y congela reloj y hostname para que el artefacto no filtre la máquina de quien lo creó.
Aquí está lo importante: MIGRATIONS es append-only; una copia del esquema v5 se migra sola hacia adelante. Así, la escalera deja de ser una promesa en un CHANGELOG y pasa a ser un test que falla el día que alguien rompa una migración.
Acto II — La elegancia de hacer menos
Del borrado defensivo a la primitiva correcta
En el Post 02 arreglé el -wal que vaciaba la base restaurada a base de borrar los ficheros auxiliares huérfanos:
_remove_sidecars(paths.db_path) # fuera el -wal/-shm
os.replace(staged_db, paths.db_path)
_remove_sidecars(paths.db_path) # y otra vez, por si acaso
Funciona, pero es una solución por fuerza bruta: limpio el desorden que otra cosa dejó. Mike propuso bajar un nivel y usar la misma primitiva de SQLite, en sentido inverso. En vez de renombrar un fichero y rezar por los sidecars, dejo que SQLite escriba la base preparada dentro de la base destino:
def _replace_db(staged_db: Path, target: Path) -> None:
# Backup inverso: origen -> destino. SQLite abre una transacción de
# escritura sobre el destino y la revierte si no termina.
if _page_size(target) in (0, _page_size(staged_db)):
try:
_consistent_copy(staged_db, target)
return
except sqlite3.OperationalError as exc:
if "locked" in str(exc).lower() or "busy" in str(exc).lower():
raise BackupError("la base destino está en uso; detén el daemon antes de restaurar")
# Fallback: reemplazo atómico con limpieza de sidecars
_remove_sidecars(target)
os.replace(staged_db, target)
_remove_sidecars(target)
Las ventajas son las que él describió: no hay ventana en la que la base no exista, no quedan -wal/-shm huérfanos (el WAL que aparece es el de la propia base) y, si algo falla a mitad, el destino se revierte. Verifiqué el comportamiento con un experimento antes de escribir una línea: la copia inversa sobre una base WAL con conexión viva deja los datos correctos y no deja sidecar huérfano.
Hay elegancia en hacer menos y confiar en el motor que ya conoce sus invariantes.
La condición que no se puede ignorar
El backup inverso tiene un requisito real: si el destino está en modo WAL, ambos ficheros deben tener el mismo page_size. Si no coinciden, SQLite aborta con un OperationalError. Eso lo convierte en una restricción de diseño, no en una nota al pie:
- Si los
page_size coinciden → backup inverso (limpio, con rollback).
- Si no coinciden → fallback atómico conservando la limpieza de sidecars.
- Si el destino está bloqueado por un escritor vivo → se aborta con aviso, en lugar de competir por el fichero y corromperlo en silencio.
Ese último caso es el más importante para el usuario: restore --force sobre un daemon activo ya no es una ruleta. Si no puedo garantizar la transición, prefiero decir que no.
Lo que esto cambia (Takeaway)
Las dos observaciones de Mike son el mismo invariante visto desde dos ejes:
- Temporal: bytes del pasado frente al código del futuro.
- De concurrencia: el escritor del presente frente al restaurador.
Y hay una forma de leerlo que lo unifica: restaurar no es copiar ficheros, es una transición de estado. Cuando lo tratas como tal, el diseño se ordena solo — aparecen los guardarraíles, aparece el page_size, aparece el rechazo honesto — y el resto es escribir el test que lo demuestre.
Un test verde hoy es una promesa. Un test verde dentro de un año es una garantía.
Gracias
Otra vez, el mejor code review de Aethernet vino de la comunidad. Gracias, @Mike Dabydeen: no señalaste un bug, señalaste un punto ciego. Primero el tiempo, después la primitiva. Ambas correcciones están en la main por tu lectura.
Conclusión objetiva
La versión 1.3.0 de Aethernet incorpora ahora una garantía hacia adelante: un paquete real conservado como golden file que CI restaura en cada versión, un generador reproducible para ampliar la escalera y una restauración que usa el backup inverso de SQLite con fallback seguro. Todo con ruff + mypy --strict (0 errores) + vulture + 177 tests en verde (sin migración de esquema nueva: sigue en v5).
No es una promesa de marketing: es un test que puedes leer y un comando que puedes ejecutar.
# ampliar la escalera al cambiar el esquema o el formato
python scripts/gen_backup_fixture.py
# la garantía hacia adelante, corriendo en CI
pytest tests/test_backup_compat.py
# restaurar una copia vieja (aquí o en otra máquina)
aethernet restore aethernet-1.3.0-schema5.tar.gz
Y tú
¿Cómo verificas hoy que un backup de tu software seguirá restaurándose dentro de un año? ¿Confías en la compatibilidad hacia adelante de tus formatos o, como yo, creías tener un test que en realidad solo probaba el presente? ¿Y hasta dónde llevas la primitiva: ficheros, transacciones o capas más altas? Cuéntamelo en los comentarios; leo y respondo a todo.
English version below · El original está en español. If you're an English speaker, scroll down to the translation and tell me how you verify your backups in the discussion. I reply to everyone.
Will your backup survive the next version? The test I was missing
Series: Auditable software — Post 02.1
Read time: ~9 min · Tags: python database testing opensource
I thought I had closed the problem
There's a kind of calm that isn't engineering: the feeling of having finished something because it no longer hurts. Post 02 ended with a line I liked —recovering is the job— and with a test that restored a copy on a clean machine. I called it done.
It wasn't. I had confused my calm with a guarantee.
This isn't about fixing an if. It's a failure of perspective: I looked at the system along the axis of space (does it work on another machine?) and forgot the axis of time (will it still work once the software has changed?). A backup is, by definition, a promise made to the future. And I was validating it only in the present.
Mike Dabydeen, a community leader, read the post and wrote what I hadn't wanted to see. His central point: my test proved that today's copy restores with today's code. Exactly the case nobody needs. Whoever truly depends on a backup wants to restore a copy from a year ago with the version that exists a year from now.
I quote his two ideas, because both changed the design:
"The clean-machine test proves that today's copy restores with today's code. What it still doesn't cover is that a copy made with 1.3.0 restores with the version you publish a year from now — and that is exactly the case of whoever truly needs their backup. I'd keep a real package generated by 1.3.0 in the repo as an artifact and restore it in CI on every version."
"About the -wal error, there's another option: restore using the same SQLite backup API, but in reverse... the copy keeps a write transaction open on the destination for the whole process and rolls it back if it doesn't finish, and the -wal it finds is that same database's, so there are no orphan files to delete."
He was right on both. And the first one stings more, because it isn't a bug: it's a promise with no witness.
Act I — Time as a witness
The trap of the test that passes
My recovery test was honest and still insufficient. It did this:
# The test that proved the present
bundle = create_backup(source) # generated with today's code
restore_backup(bundle, clean) # and restored with today's code
assert count_scans(clean) == 2
Generation and restore used the same code. A test closed in on itself: if tomorrow I change the schema and break the restore path, this test would stay green while users' copies stopped restoring. The proof of my promise was the promise looking at itself in a mirror.
A promise is only real if it outlives the one who made it.
The fix: a ladder of fixtures
I froze a real package generated by 1.3.0 and committed it as a golden file. On every version, CI must restore it on a clean machine:
tests/fixtures/backups/
├── aethernet-1.3.0-schema5.tar.gz # the real package, ~5 KB
├── aethernet-1.3.0-schema5.expected.json # its provenance and assertions
└── README.md # how to extend the ladder
The .expected.json stores what makes the test verifiable: sqlite_version, page_size, db_sha256 and the counts that must survive (scans, devices, events). The test walks every package in the folder, not one:
@pytest.mark.parametrize("bundle", sorted(FIXTURES.glob("*.tar.gz")), ids=lambda p: p.name)
def test_historical_backup_restores_on_clean_machine(bundle, tmp_path):
expected = json.loads(_expected_path(bundle).read_text())
result = restore_backup(bundle, _clean_paths(tmp_path)) # clean machine
repo = Repository(Database(result.db_path))
assert repo.counts()["wifi_scans"] == expected["counts"]["wifi_scans"]
# ...report regenerated, config restored, schema migrated to v5
The maintenance rule was written by Mike himself: when the schema or the format changes, add a new package and keep the previous ones. I automated it in a generator, scripts/gen_backup_fixture.py, that always uses the current code and freezes clock and hostname so the artifact doesn't leak the machine that built it.
Here's the key part: MIGRATIONS is append-only, so a v5 copy migrates itself forward. The ladder stops being a promise in a CHANGELOG and becomes a test that fails the day someone breaks a migration.
Act II — The elegance of doing less
From defensive deletion to the right primitive
In Post 02 I fixed the -wal that wiped the restored database by deleting the orphan sidecar files:
_remove_sidecars(paths.db_path) # drop the -wal/-shm
os.replace(staged_db, paths.db_path)
_remove_sidecars(paths.db_path) # and again, just in case
It works, but it's brute force: I clean up the mess something else left. Mike suggested dropping one level and using SQLite's own primitive, in reverse. Instead of renaming a file and hoping about the sidecars, I let SQLite write the prepared database into the destination:
def _replace_db(staged_db: Path, target: Path) -> None:
# Reverse backup: source -> destination. SQLite opens a write
# transaction on the destination and rolls it back if it doesn't finish.
if _page_size(target) in (0, _page_size(staged_db)):
try:
_consistent_copy(staged_db, target)
return
except sqlite3.OperationalError as exc:
if "locked" in str(exc).lower() or "busy" in str(exc).lower():
raise BackupError("destination database is in use; stop the daemon before restoring")
# Fallback: atomic replace with sidecar cleanup
_remove_sidecars(target)
os.replace(staged_db, target)
_remove_sidecars(target)
The advantages are the ones he described: there's no window where the database doesn't exist, no orphan -wal/-shm remain (the WAL that shows up belongs to the database itself), and if something fails halfway, the destination rolls back. I verified the behavior with an experiment before writing a single line: the reverse copy onto a WAL database with a live connection leaves the data correct and the orphan sidecar gone.
There's elegance in doing less and trusting the engine that already knows its own invariants.
The condition you can't ignore
The reverse backup has a real requirement: if the destination is in WAL mode, both files must share the same page_size. If they don't, SQLite aborts with an OperationalError. That turns it into a design constraint, not a footnote:
- If the
page_size matches → reverse backup (clean, with rollback).
- If it doesn't match → atomic fallback keeping the sidecar cleanup.
- If the destination is locked by a live writer → abort with a warning, instead of racing for the file and corrupting it silently.
That last case matters most to the user: restore --force over an active daemon is no longer a coin flip. If I can't guarantee the transition, I'd rather say no.
What this changes (Takeaway)
Mike's two observations are the same invariant seen from two axes:
- Temporal: bytes from the past against the code of the future.
- Concurrency: the writer of the present against the restorer.
And there's a way to read it that unifies both: restoring isn't copying files, it's a state transition. When you treat it as such, the design sorts itself out — the guardrails appear, the page_size appears, the honest refusal appears — and all that's left is to write the test that proves it.
A green test today is a promise. A green test a year from now is a guarantee.
Thanks
Once again, the best code review of Aethernet came from the community. Thank you, @Mike Dabydeen: you didn't point at a bug, you pointed at a blind spot. First time, then the primitive. Both fixes are on main because of your read.
Objective conclusion
Version 1.3.0 of Aethernet now ships a forward guarantee: a real package kept as a golden file that CI restores on every version, a reproducible generator to extend the ladder, and a restore that uses SQLite's reverse backup with a safe fallback. All of it with ruff + mypy --strict (0 errors) + vulture + 177 tests green (no new schema migration: still v5).
It isn't a marketing promise: it's a test you can read and a command you can run.
# extend the ladder when the schema or format changes
python scripts/gen_backup_fixture.py
# the forward guarantee, running in CI
pytest tests/test_backup_compat.py
# restore an old copy (here or on another machine)
aethernet restore aethernet-1.3.0-schema5.tar.gz
Over to you
How do you verify today that a backup of your software will still restore a year from now? Do you trust your formats' forward compatibility, or — like me — did you believe you had a test that was really only proving the present? And how far down do you take the primitive: files, transactions, or higher layers? Tell me in the comments; I read and reply to everyone.
Project license: MIT · No secrets or keys in this post. · Projects: github.com/yosoyignicion · Aethernet code: github.com/yosoyignicion/Aethernet · Release v1.3.0