Building a Safer Contributor Verification VPS Without Breaking Production
Hardening a live VPS is rarely about applying a single security checklist.
The harder problem is this:
How do you reduce attack surface on a server that is already running production services, contributor verification workloads, Docker environments, databases, reverse proxies, and experimental integrations — without breaking any of them?
That was the situation on a MyZubster VPS we are evolving into a more structured Contributor Verification Node.
The objective was deliberately conservative:
- identify what is actually running
- reduce unnecessary public binds
- preserve all live workloads
- distinguish repository architecture from runtime architecture
- document every conclusion with reproducible technical evidence
No mass restarts. No destructive Git operations. No speculative rewiring.
Just one dependency at a time.
The VPS was doing too many jobs at once
The server was already hosting a mix of services:
- nginx
- MyZubster production gateway
- PM2 applications
- Docker contributor environments
- MongoDB
- Monero services
- Bitcoin verification services
- IPFS
- Qdrant
- Ollama
- n8n
- contributor bridges and pilot services
This creates a common infrastructure trap.
A port may look unnecessary.
A process may look duplicated.
A service may appear safe to restart.
But if you do not first understand the dependency graph, a cleanup command can become an outage.
That immediately ruled out broad actions such as:
pm2 restart all
or:
git reset --hard
docker system prune
ufw reset
The working rule became:
observe
→ identify ownership
→ map dependencies
→ change one thing
→ verify
→ persist
Start from sockets, not assumptions
The first meaningful artifact was a listener inventory.
Using:
ss -lntp
we classified services into two groups.
Some were already correctly restricted:
127.0.0.1:5003 MyZubster Gateway
127.0.0.1:8787 BTC verifier
127.0.0.1:27017 MongoDB
127.0.0.1:6333 Qdrant
127.0.0.1:5678 n8n
127.0.0.1:8092 contributor bridge
Others were bound globally:
*:5002
0.0.0.0:5005
0.0.0.0:5173
A global bind does not automatically mean “publicly exposed.”
A firewall may still block traffic.
But it does mean the application is willing to accept traffic on every interface, which increases ambiguity and weakens least-privilege design.
So the next question was not:
Can we close this port?
It was:
Who actually needs to reach this service?
Hardening port 5173
The frontend process was listening on:
0.0.0.0:5173
But nginx already proxied to it locally:
location / {
proxy_pass http://localhost:5173;
}
That gave us the actual runtime graph:
Internet
↓
nginx :443
↓
127.0.0.1:5173
There was no need for the Node process itself to listen on every interface.
The original code contained:
server.listen(PORT, '0.0.0.0', () => {
We changed only the bind address:
server.listen(PORT, '127.0.0.1', () => {
Then checked syntax:
node --check server-static.js
and restarted only that PM2 process:
pm2 restart myzubster-web
Not the entire PM2 stack.
Verify before persisting
After the restart:
ss -lntp | grep ':5173'
showed:
127.0.0.1:5173
The local service still returned 200.
Then nginx HTTP returned the expected redirect.
Finally HTTPS was tested directly against the local nginx instance using the real virtual host:
curl -k -sS -o /dev/null \
--resolve myzubster.com:443:127.0.0.1 \
-w 'HTTPS nginx -> %{http_code}\n' \
https://myzubster.com/
Result:
HTTPS nginx -> 200
Only after that did we run:
pm2 save
This sequencing matters.
You do not persist a new runtime state before proving that the change is healthy.
Hardening port 5002
Port 5002 belonged to another Node application.
It was listening globally because the server was started without an explicit host:
httpServer.listen(PORT, () => {
Before changing it, we checked:
- nginx references
- frontend references
- active TCP connections
- WebSocket usage
- service routes
The frontend already targeted the service internally:
const API_TARGET = 'http://127.0.0.1:5002';
There was no nginx route sending external traffic directly to 5002.
No active external dependency was visible.
So we changed the listener to:
httpServer.listen(PORT, '127.0.0.1', () => {
Again:
backup
→ syntax check
→ restart one process
→ verify socket
→ verify health
The service remained reachable locally.
The frontend proxy returned a 404 for one test route rather than 502.
That difference was useful.
A 404 meant:
upstream reachable, route missing
A 502 would have meant:
upstream unreachable
That distinction prevented us from treating a valid hardening step as a regression.
The interesting part: a broken TAZ contract
During the dependency review, we found a frontend WebSocket client that constructed:
`${protocol}//${window.location.host}/api/taz/live`
The TAZ interface also expected:
GET /api/taz/dashboard
GET /api/taz/xmr/summary
WS /api/taz/live
At first this looked like something that might have been affected by our network changes.
So we tested the production gateway directly.
The results were:
404 Cannot GET /api/taz/dashboard
404 Cannot GET /api/taz/xmr/summary
404 Cannot GET /api/taz/live
This was important.
The feature was already incomplete.
The hardening work had not broken it.
That leads to one of the most useful lessons from the whole exercise:
A failure discovered during infrastructure work is not automatically caused by the infrastructure change.
Always reproduce the behavior against the actual upstream service.
Nginx was doing exactly what it should
The production nginx configuration routed:
location / {
proxy_pass http://localhost:5173;
}
location /api/ {
proxy_pass http://localhost:5003;
}
So /api/taz/live already reached the production gateway.
The gateway returned the 404.
That meant blindly adding WebSocket proxy headers such as:
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
would not solve the real problem.
A reverse proxy cannot invent an application endpoint.
There was already a realtime subsystem
Repository inspection revealed something more interesting.
MyZubster already contains a realtime implementation based on Socket.IO.
Its server path is:
/realtime
and supporting HTTP routes are mounted under:
/api/realtime
The realtime code exposes an explicit attachment function:
attachRealtimeServer(httpServer)
This changed the architecture question completely.
Instead of asking:
How do we build a new WebSocket service for TAZ?
the better question became:
Can TAZ reuse the existing realtime stack?
Maintaining two separate realtime systems would add unnecessary complexity unless there is a strong protocol reason.
Source code and production runtime were not the same thing
There was another subtle mismatch.
The repository's normal backend startup path includes:
attachRealtimeServer(server)
But the active systemd service starts:
/root/myzubster/scripts/start-gateway-systemd.js
and that startup path currently creates the server with:
const server = app.listen(port, host, ...)
without obviously attaching the realtime server.
So the repository may support realtime while production does not actually expose it.
This is a recurring infrastructure lesson:
Repository architecture is not runtime architecture.
A feature being present in source code does not prove that the active service manager actually executes it.
The authoritative chain is closer to:
systemd / PM2
→ executable
→ startup script
→ imported modules
→ bound socket
→ reverse proxy
→ externally observed response
That chain matters more than a repository grep.
Do not fabricate telemetry to make a dashboard look alive
The TAZ frontend expects values like:
drinksServed
xmrReceived
robot.status
transactions
During repository searches, we found unrelated simulation and robot data.
It would have been technically easy to wire some of that into the dashboard.
We deliberately did not.
A live interface must not display simulation values as if they were production evidence.
Our rule was:
no authoritative source → no authoritative metric
A degraded dashboard is more honest than a beautiful dashboard with fabricated telemetry.
What changed
The hardening work produced two concrete improvements:
MyZubsterWeb
0.0.0.0:5173
→
127.0.0.1:5173
and:
Urban Lab
*:5002
→
127.0.0.1:5002
The main gateway was already:
127.0.0.1:5003
nginx HTTPS continued to work.
Docker verification environments were not disturbed.
PM2 state was persisted only after verification.
No destructive Git operation touched the live checkout.
What we deliberately did not do
A large part of safe infrastructure work is knowing what not to touch.
We did not:
- restart every PM2 application
- reset firewall configuration
- expose additional ports
- modify nginx before proving nginx was the problem
- connect TAZ to an unrelated WebSocket server
- feed simulation values into production telemetry
- reset or pull the live Git checkout
- treat code presence as runtime proof
Those non-actions were part of the engineering result.
The Contributor Verification Node model
The broader goal is to evolve the VPS toward a structured Contributor Verification Node.
That should not mean giving every contributor unrestricted shell access.
A safer pattern looks like:
SSH identity
↓
dedicated Unix account
↓
dedicated workspace
↓
container / sandbox
↓
explicit mounts
↓
explicit network permissions
↓
reproducible verifier
↓
immutable evidence
The node should prove that a technical checkpoint can be reproduced.
It should not silently become a shared production administration box.
Current state
At the current checkpoint:
5173 localhost-only ✅
5002 localhost-only ✅
5003 localhost-only ✅
nginx HTTPS ✅
PM2 persistence ✅
TAZ HTTP endpoints missing
TAZ live WebSocket missing
realtime subsystem present in source
production realtime wiring under review
This is not a flashy infrastructure milestone.
But it is a useful one.
We now know more precisely:
- what is actually listening
- which services need external exposure
- which assumptions in the frontend are currently false
- which realtime infrastructure already exists
- where production startup differs from repository design
That is enough to make the next change smaller and safer.
Final takeaway
The most valuable part of production hardening is often not the final configuration.
It is the discipline of reducing uncertainty.
For every service, ask:
What process owns it?
Who actually needs to reach it?
Does it need to bind globally?
Which reverse proxy reaches it?
What does production really execute?
How will I prove the change did not break anything?
The result is not just a smaller attack surface.
It is a better mental model of the system.
And once the runtime model is accurate, contributor verification becomes much easier to isolate, reproduce, and trust.