<!--
CoderLegion · Article (Discussion-style) · Post 03
Serie: — (standalone)
Autor: Ignacio Badenes (@yosoyignicion / IgnicionDev)
Idioma: bilingüe ES + EN · Read time: ~8 min
Tags: python, security, testing, opensource
Caso real: pass-encrypt-env (uso personal, Linux Mint) — https://github.com/yosoyignicion/pass-encrypt-env
Nota: mini-app en estado "uso personal"; la adaptación/portabilidad se planifica en días próximos.
-->
Mi Definition of Done: 5 gates que convierten "creo que funciona" en "puedo demostrarlo"
Nota honesta: esto no es un producto, es una mini-app de uso personal en mi Linux Mint. Funciona y la uso a diario; su adaptación para terceros la planifico para los próximos días.
Read time: ~8 min · Etiquetas: python security testing opensource
Un commit. Una hora. Y no es un juguete
El proyecto del que hablo hoy tiene un solo commit y tardó una hora en funcionar de punta a punta. Lo uso cada mañana. Y no es un juguete.
Lo digo porque la pregunta interesante no fue cómo lo programé rápido, sino cómo supe que estaba terminado. Ese es el punto: "terminado" no es una sensación, es una decisión que puedes defender. Cuando esa decisión tiene una definición clara, programar una hora deja de ser un atajo y pasa a ser consecuencia.
Este post es esa definición. La llamo Definition of Done y son cinco gates.
Qué hace, sin tecnicismos
Imagina un llavero. Guarda tus claves API y credenciales cifradas dentro de tu propio ordenador y se las sirve a tus aplicaciones durante toda la sesión después de un único desbloqueo al iniciar sesión. Nada sube a la nube. No hay telemetría. No hay cuenta que crear.
En la práctica, mis scripts y programas ya no llevan claves escritas dentro ni dependen de un .env en claro: consultan el llavero, que ya está abierto por mí, y arrancan sin volver a pedirme nada. Si me voy, el llavero se cierra; las claves desaparecen de la memoria.
Se llama pass-encrypt-env. Corre en mi Linux Mint y, hoy, es para mí.
Lo que es hoy (y lo que todavía no)
Ser claro con el estado real es parte del método. Un proyecto "terminado" también declara sus límites.
- Es: una herramienta personal, funcional, que uso a diario en mi equipo.
- No es: un producto instalable por cualquiera sin fricción, ni algo auditado por terceros.
- Viene: en los próximos días planifico la adaptación —empaquetado, arranque en otras distribuciones y documentación para quien no sea yo— como un esfuerzo aparte y planificado.
Y hay una razón de fondo para separar esa adaptación del desarrollo: no quiero declarar "lista" una app porque pase en mi máquina. La declaro lista cuando pasa los gates en cualquier máquina. Es exactamente la misma lógica que un buen test: no vale que funcione una vez, tiene que fallar cuando corresponda.
Por qué exijo evidencia y no sensaciones
Soy relativamente nuevo en el mundo de la IA agéntica, aunque llevo años gestionando sistemas bajo protocolo: control de accesos, mínima exposición, registros. Y una disciplina que aprendí ahí —y que el prompting me recordó— es esta:
No das algo por resuelto porque parezca resuelto; lo das por resuelto cuando puedes demostrarlo.
El hábito de protocolo es la mínima exposición: asumes que algo fallará y diseñas para que, cuando falle, no arrastre contigo todo lo demás. Aplicado al software, eso se traduce en una Definition of Done: una lista corta de condiciones que convierten "yo creo que funciona" en "aquí está la prueba".
En este proyecto, esa lista lleva cinco puertas.
Los 5 gates
Cada gate existe porque previene un fallo concreto, no por completitud.
1) Tests que pueden fallar
No me sirve "tengo tests"; me sirve que fallen cuando rompo algo. En pass-encrypt-env hay 19 tests y uno de los que más valor tiene es este, porque vigila una promesa sobre el disco:
def test_no_plaintext_on_disk(tmp_path):
path = tmp_path / "vault.enc"
v, unlocked, _ = _new()
vault.set_secret(unlocked.payload, "TOKEN", "super-secret-value")
vault.reseal(v, unlocked.dek, unlocked.payload)
vault.save_vault(v, path)
blob = path.read_text()
assert "super-secret-value" not in blob # el secreto no está en claro
assert (path.stat().st_mode & 0o777) == 0o600 # y el fichero no es legible por terceros
Ese test no comprueba que "va bien": comprueba que la promesa del producto se cumple. El fallo que previene es el peor de todos: creerte seguro y no serlo.
2) Lint limpio
ruff check .
No es estética. Un lint en verde elimina de antemano la clase de error más tonta —imports sin usar, variables fantasma, comparaciones sospechosas— antes de que la ejecutes. Es ruido que no llega a producción porque ni siquiera llega al commit.
3) Escaneo de seguridad
bandit -q -c bandit.yaml -r src
bandit busca patrones inseguros en el código: criptografía mal usada, subprocess peligrosos, aleatoriedad no criptográfica. En una app cuyo único trabajo es guardar secretos, este gate no es opcional. El fallo que previene es el uso accidental de una primitiva parecida a la correcta.
4) CI reproducible
Los tres gates anteriores corren solos en GitHub Actions, sobre Python 3.11 y 3.12:
strategy:
fail-fast: false
matrix:
python-version: ["3.11", "3.12"]
steps:
# ...
- name: Gate (ruff + bandit + pytest)
run: ./scripts/ci.sh python
Y aquí está la parte importante: ese ./scripts/ci.sh es el mismo comando que ejecuto en local. No hay una verdad para mi máquina y otra para el servidor. El fallo que previene este gate es "en mi máquina funciona", que no es un bug: es una mentira cómoda.
5) Smoke end-to-end
Los unit tests verdes no garantizan que el sistema entero arranque. Para eso hay un smoke test headless que levanta el agente, habla con él por su canal interno y comprueba el comportamiento observable:
python3 -m pass_encrypt_env.cli env --if-unlocked >/dev/null 2>&1 || rc=$?
echo "env --if-unlocked (bloqueada) -> exit $rc (esperado 3)"
[ "$rc" -eq 3 ]
Aun con piezas perfectas, el conjunto puede no montarse. El smoke test es el gate que mira el sistema, no las partes.
Todos los gates caben en un comando:
./scripts/ci.sh # esto corre en local y esto corre en CI. La misma verdad.
Qué protege y qué no
Un proyecto terminado con honestidad publica su techo, no lo esconde. Este es el mío:
| Amenaza | ¿Mitigado? |
Robo offline de vault.enc | Sí — sin master + TOTP no hay datos |
| Manipulación del fichero | Sí — AEAD detecta la alteración |
| Alguien sin tu TOTP | Sí — segundo factor obligatorio |
| Pérdida de master o TOTP | Sí — clave de recuperación |
| Proceso del mismo UID con la bóveda abierta | No — techo del diseño cero-prompt |
root / malware ya dentro | No — fuera de alcance |
Que la última fila diga "No" no es una debilidad: es una promesa cumplible. Prefiero un alcance pequeño y cierto que uno grande y falso.
Detalle técnico (para quien quiera el fondo)
El "cómo", para quien lea esto con la ingeniería delante:
- Cifrado en reposo:
vault.enc con Argon2id como función de derivación (64 MiB, 3 pasadas, 4 lanes) y AES-256-GCM como cifrado autenticado (confidencialidad e integridad).
- Doble envoltura de la clave de datos (DEK):
K1 = Argon2id(master) · K2 = HKDF(K1 ‖ totp_secret) abre la DEK · K3 = HKDF(recovery_key) guarda una copia para recuperación. El secreto TOTP viaja cifrado: sin un TOTP válido no se obtiene la DEK.
- Agente de sesión: escucha en un socket Unix en
$XDG_RUNTIME_DIR/passenv/, permisos 0600 y validación de UID del par (SO_PEERCRED). Se endurece con PR_SET_DUMPABLE=0 (anti-ptrace), sin volcados de memoria y mlockall best-effort.
- CLI
passenv: init, set/get/list/rm, export, exec, change-master, backup, restore.
- Modelo de amenazas completo en
docs/THREAT_MODEL.md, con el techo del diseño cero-prompt declarado a propósito.
Lo que esto cambia (Takeaway)
El gate no es un freno; es lo que te deja pisar el acelerador sin miedo. Un proyecto de una hora y un commit no llega "terminado de verdad" por suerte: llega porque las condiciones para llamarlo terminado ya estaban escritas antes de empezar a escribir código.
"Creo que funciona" es una sensación. "Aquí está la prueba" es un proyecto.
La checklist
- [ ] Tests que fallan cuando rompo algo (no que existan)
- [ ] Lint limpio
- [ ] Escaneo de seguridad sin hallazgos
- [ ] CI reproducible (mismo comando en local que en remoto)
- [ ] Smoke end-to-end en verde
Y tú
¿Qué requisitos le exiges tú a un proyecto para darlo por terminado? ¿Cuál añadirías a mis cinco gates? Y, ya que el mío está en estado de uso personal: ¿tú publicarías una herramienta así o la dejarías privada hasta "adaptarla"? Cuéntamelo en los comentarios; leo y respondo a todo.
English version below · El original está en español. If you're an English speaker, scroll down to the translation and tell me what your own Definition of Done looks like. I reply to everyone.
My Definition of Done: 5 gates that turn "I think it works" into "here's the proof"
Honest note: this isn't a product, it's a personal-use mini-app on my Linux Mint. It works and I use it daily; adapting it for other people is a planned effort for the coming days.
Read time: ~8 min · Tags: python security testing opensource
One commit. One hour. And it's not a toy
The project I'm talking about has a single commit and took one hour to work end to end. I use it every morning. And it isn't a toy.
The interesting question wasn't how I built it fast, but how I knew it was done. That's the point: "done" isn't a feeling, it's a decision you can defend. Once that decision has a clear definition, writing code for an hour stops being a shortcut and becomes a consequence.
This post is that definition. I call it a Definition of Done, and it's five gates.
What it does, without jargon
Picture a keyring. It keeps your API keys and credentials encrypted inside your own computer and serves them to your applications for the whole session after a single unlock at login. Nothing goes to the cloud. There's no telemetry. There's no account to create.
In practice, my scripts and apps no longer carry keys inside them or depend on a plaintext .env: they ask the keyring, already unlocked on my behalf, and start without asking me again. If I leave, the keyring closes; the keys disappear from memory.
It's called pass-encrypt-env. It runs on my Linux Mint and, today, it's for me.
What it is today (and what it isn't yet)
Being clear about the real state is part of the method. A "finished" project also declares its limits.
- It is: a personal, working tool that I use daily on my machine.
- It isn't: something anyone can install without friction, nor something audited by third parties.
- It's coming: in the coming days I'm planning the adaptation — packaging, booting on other distributions and documentation for someone who isn't me — as a separate, planned effort.
There's a deeper reason to separate that adaptation from development: I don't want to call an app "ready" because it passes on my machine. I call it ready when it passes the gates on any machine. It's the same logic as a good test: it doesn't count if it works once, it has to fail when it should.
Why I demand evidence, not feelings
I'm relatively new to the world of agentic AI, though I've spent years managing systems under protocol: access control, minimal exposure, logs. And one discipline I learned there — which prompting reminded me of — is this:
You don't call something solved because it looks solved; you call it solved when you can prove it.
The protocol habit is minimal exposure: you assume something will fail and you design so that, when it fails, it doesn't drag everything else down with it. Applied to software, that becomes a Definition of Done: a short list of conditions that turn "I think it works" into "here's the proof".
In this project, that list has five doors.
The 5 gates
Each gate exists because it prevents a concrete failure, not for completeness.
1) Tests that can fail
I don't care about "I have tests"; I care that they fail when I break something. pass-encrypt-env has 19 tests, and one of the most valuable guards a promise about the disk:
def test_no_plaintext_on_disk(tmp_path):
path = tmp_path / "vault.enc"
v, unlocked, _ = _new()
vault.set_secret(unlocked.payload, "TOKEN", "super-secret-value")
vault.reseal(v, unlocked.dek, unlocked.payload)
vault.save_vault(v, path)
blob = path.read_text()
assert "super-secret-value" not in blob # the secret isn't in cleartext
assert (path.stat().st_mode & 0o777) == 0o600 # and the file isn't readable by others
That test doesn't check that "it works": it checks that the product's promise holds. The failure it prevents is the worst of all: believing you're safe when you aren't.
2) Clean lint
ruff check .
It isn't aesthetics. A green lint removes the dumbest class of error —unused imports, ghost variables, suspicious comparisons— before you even run it. It's noise that never reaches production because it never reaches the commit.
3) Security scan
bandit -q -c bandit.yaml -r src
bandit looks for insecure patterns in the code: misused cryptography, dangerous subprocess, non-cryptographic randomness. In an app whose only job is to store secrets, this gate isn't optional. The failure it prevents is the accidental use of a primitive that looks like the right one.
4) Reproducible CI
The previous three gates run on their own in GitHub Actions, across Python 3.11 and 3.12:
strategy:
fail-fast: false
matrix:
python-version: ["3.11", "3.12"]
steps:
# ...
- name: Gate (ruff + bandit + pytest)
run: ./scripts/ci.sh python
And here's the important part: that ./scripts/ci.sh is the same command I run locally. There's no one truth for my machine and another for the server. The failure this gate prevents is "it works on my machine", which isn't a bug: it's a comfortable lie.
5) End-to-end smoke test
Green unit tests don't guarantee the whole system boots. For that there's a headless smoke test that starts the agent, talks to it over its internal channel and checks the observable behavior:
python3 -m pass_encrypt_env.cli env --if-unlocked >/dev/null 2>&1 || rc=$?
echo "env --if-unlocked (locked) -> exit $rc (expected 3)"
[ "$rc" -eq 3 ]
Even with perfect pieces, the whole may fail to assemble. The smoke test is the gate that looks at the system, not the parts.
Every gate fits in one command:
./scripts/ci.sh # this runs locally and this runs in CI. The same truth.
What it protects and what it doesn't
An honestly finished project publishes its ceiling, it doesn't hide it. This is mine:
| Threat | Mitigated? |
Offline theft of vault.enc | Yes — without master + TOTP there's no data |
| File tampering | Yes — AEAD detects the alteration |
| Someone without your TOTP | Yes — second factor required |
| Loss of master or TOTP | Yes — recovery key |
| A process of the same UID with the vault open | No — ceiling of the prompt-less design |
root / malware already inside | No — out of scope |
That the last row says "No" isn't a weakness: it's a promise you can keep. I'd rather have a small and certain scope than a large and false one.
Technical detail (for those who want the depth)
The "how", for whoever reads this with the engineering in front of them:
- Encryption at rest:
vault.enc with Argon2id as the key-derivation function (64 MiB, 3 passes, 4 lanes) and AES-256-GCM as authenticated encryption (confidentiality and integrity).
- Two-layer wrapping of the data key (DEK):
K1 = Argon2id(master) · K2 = HKDF(K1 ‖ totp_secret) unwraps the DEK · K3 = HKDF(recovery_key) keeps a recovery copy. The TOTP secret travels encrypted: without a valid TOTP you don't get the DEK.
- Session agent: it listens on a Unix socket at
$XDG_RUNTIME_DIR/passenv/, mode 0600 and UID validation of the peer (SO_PEERCRED). It's hardened with PR_SET_DUMPABLE=0 (anti-ptrace), no memory dumps and best-effort mlockall.
passenv CLI: init, set/get/list/rm, export, exec, change-master, backup, restore.
- Full threat model in
docs/THREAT_MODEL.md, with the prompt-less design's ceiling declared on purpose.
What this changes (Takeaway)
The gate isn't a brake; it's what lets you hit the accelerator without fear. A one-hour, one-commit project doesn't end up "truly finished" by luck: it does because the conditions to call it finished were written before the code.
"I think it works" is a feeling. "Here's the proof" is a project.
The checklist
- [ ] Tests that fail when I break something (not that they exist)
- [ ] Clean lint
- [ ] Security scan with no findings
- [ ] Reproducible CI (same command locally and remotely)
- [ ] End-to-end smoke test green
Over to you
What requirements do you demand from a project before calling it done? Which one would you add to my five gates? And since mine is in a personal-use state: would you publish a tool like this or keep it private until you "adapt" it? Tell me in the comments; I read and reply to everyone.
Project license: MIT · No secrets or keys in this post. · Project: github.com/yosoyignicion/pass-encrypt-env · Profile: github.com/yosoyignicion