From a GitHub OAuth 404 to a More Reliable MyZubster Runtime
Sometimes a production problem starts with a very small symptom.
In our case, it was a GitHub OAuth 404.
What looked like an authentication issue eventually led us through a much broader review of MyZubster’s development and runtime environment: GitHub account migration, repository ownership, systemd, PM2, MongoDB initialization, multiple Mongoose runtimes, realtime authorization, Bitcoin payment logic, and an aging Jest test suite.
This is the story of how one error turned into a useful infrastructure cleanup.
What is MyZubster?
MyZubster is an evolving open-source ecosystem combining community participation, digital experiences, Marketplace and Seller workflows, AI assistance, experimental projects, a Metaverse interface, and open-source contribution.
Useful links:
The project is actively evolving, and part of the current work is making sure the codebase, infrastructure, runtime behavior, and tests all describe the same system.
The original problem: GitHub OAuth
The first visible issue was an unreliable GitHub authentication flow that could end in a 404.
At the same time, the previous development account had accumulated other operational limitations affecting API access, OAuth integrations, Actions, and repository workflows.
Instead of trying to keep building on top of an unstable setup, we separated two concerns:
- the historical/public project identity;
- the operational GitHub identity used for development.
A new GitHub account became the operational development account, while the existing MyZubster identity inside the product remained untouched.
This was important.
We were changing the development control plane, not rewriting the identity or history of the project.
Migrating Git without destroying local state
The active VPS repository is located at:
/root/myzubster
The server also contains worktrees, experimental directories, verifier code, and local files that should not be committed.
So we deliberately avoided shortcuts like:
git reset --hard
and:
git add .
Instead, we moved the canonical remote to the new GitHub repository while keeping older remotes available for reference.
The current development repository is:
https://github.com/danieldirimini-myzubster/myzubster
The active feature branch was then synchronized with the latest main without rewriting history.
That preserved the work already done while bringing in an OAuth-related fix from the main branch.
Then we found a runtime ownership problem
Once Git was stable, we inspected how MyZubster was actually running on the VPS.
We discovered that the environment used both:
systemd
PM2
That is not necessarily a problem by itself.
The real issue was that both systems appeared to have responsibility for parts of the same application.
There was a PM2 process named myzubster, while the actual HTTP gateway was being served through systemd.
Even worse, the startup paths were different.
The root server.js exported the Express application, while another process was responsible for calling .listen().
That meant we had multiple process managers around an application that did not have a single, obvious runtime owner.
The solution was to make systemd the canonical supervisor for the main gateway and remove the duplicate PM2 ownership.
PM2 remains useful for other independent services.
Production was running from an older project directory
Another subtle problem appeared in the systemd configuration.
The active repository was:
/root/myzubster
but the service still referenced an older copy:
/root/myzubster-ahp
This is a dangerous situation because you can successfully edit, test, commit, and push one repository while production continues to execute another.
We updated the systemd working directory so the code being edited and the code being executed are now aligned.
This is one of the most important lessons from this debugging session:
Never assume the repository you are editing is the repository your production process is actually running.
Check WorkingDirectory.
Check ExecStart.
Check the real process.
The homepage worked, but the Metaverse was degraded
After fixing the runtime path, the homepage returned:
HTTP 200
But the Metaverse health endpoint reported:
{
"success": false,
"status": "degraded",
"transport": "unavailable",
"mongodb": "disconnected"
}
At first this looked like a MongoDB problem.
It was not.
The credentials were correct and the root application could connect to the database.
The real problem was much more interesting.
Two Mongoose installations meant two database states
The project contains separate Node.js dependency contexts.
Conceptually:
root application
-> Mongoose 7
backend
-> Mongoose 8
Those are two separate JavaScript instances.
Connecting one does not connect the other.
So this:
await rootMongoose.connect(...)
does not make this true:
backendMongoose.connection.readyState === 1
The backend Metaverse code was checking its own Mongoose connection.
That connection had never been initialized by the systemd startup path.
The database itself was fine.
The application lifecycle was wrong.
Fixing the bootstrap sequence
The backend already had proper database initialization logic.
The missing piece was making sure production actually called it.
We created a dedicated systemd entrypoint that does this in the correct order:
load application
↓
connect backend MongoDB
↓
start HTTP server
In simplified Node.js:
const app = require('../server');
const backend = require('../backend/src');
async function main() {
await backend.connectDatabase();
const server = app.listen(5003, '127.0.0.1', () => {
console.log('MyZubster Gateway ready');
});
process.on('SIGTERM', async () => {
server.close(async () => {
await backend.disconnectDatabase();
process.exit(0);
});
});
}
main();
After restarting systemd, the Metaverse health endpoint changed to:
{
"success": true,
"status": "healthy",
"transport": "shared-polling",
"mongodb": "connected"
}
That was a major milestone.
We did not “solve” Mongoose by deleting dependencies
A tempting response would have been to force both parts of the project to use one Mongoose installation immediately.
We deliberately did not do that.
The immediate problem was not the existence of two Mongoose versions.
The problem was that only one connection lifecycle was being initialized.
This distinction matters.
A structural refactor may still make sense in the future, but production incidents are not the right time to perform unnecessary dependency surgery.
Fix the real failure first.
Then the tests started teaching us about the architecture
The feature-specific tests were healthy.
For the MYZ-188 development work, we had:
7 test suites passed
37 tests passed
But the complete repository test suite still contained failures.
Instead of assuming the new branch had broken everything, we checked whether several failures also existed on main.
Many did.
That gave us a useful distinction:
new regression
versus:
existing technical debt
This saved us from “fixing” unrelated code based only on the fact that Jest was red.
A red test does not always mean production is wrong
One example involved the monthly AI budget.
A test expected a budget key like:
openai:gpt-5.6-sol:2026-09
while production used:
openai:all:2026-09
At first glance, changing the production key would make the test pass.
But when we reviewed the architecture, the global key was clearly intentional.
The monthly budget covers the whole OpenAI/Astra usage, and the system may fall back between different OpenAI models.
Using model-specific budget keys could allow separate models to consume separate caps.
So the production behavior was correct.
The test was stale.
We updated the test.
That is an important rule:
Tests are executable documentation, but executable documentation can become outdated too.
Realtime authorization: fail closed
Another test was hanging for five seconds.
The relevant code queried community membership through MongoDB.
When the database authority was unavailable, Mongoose buffered the query until Jest timed out.
The correct behavior for authorization was not to wait.
It was to deny.
We changed the flow to check database availability first:
if (mongoose.connection.readyState !== 1) {
return {
allowed: false,
reason: 'community_membership_authority_unavailable'
};
}
The behavior changed from:
database unavailable
↓
wait
↓
timeout
to:
database unavailable
↓
deny immediately
That is both faster and safer.
Then we found a real Bitcoin bug
The full suite exposed another issue:
TypeError:
Cannot read properties of undefined (reading 'includes')
The failing code was checking:
SUPPORTED_ASSETS.includes(asset)
But SUPPORTED_ASSETS itself was undefined.
The reason was a CommonJS circular dependency.
The dependency graph looked approximately like this:
Legacy Monetization Service
↓
Unified Checkout Service
↓
Quote / Chain Verifier
↓
Legacy Monetization Service
During module initialization, Node.js returned an incomplete export.
The result:
SUPPORTED_ASSETS === undefined
Breaking the circular dependency
We did not add a defensive fallback like:
SUPPORTED_ASSETS || []
That would hide the design problem.
Instead, we extracted the constant into a leaf module:
'use strict';
const SUPPORTED_ASSETS = Object.freeze(['BTC']);
module.exports = {
SUPPORTED_ASSETS
};
Now the payment services depend on a small module that depends on nothing else.
The Bitcoin production tests moved from:
3 failed
2 passed
to:
5 passed
The tests cover:
- BTC availability;
- BTC/EUR quoting;
- satoshi conversion;
- Esplora verification;
- unconfirmed-payment fail-closed behavior.
Legacy tests were still mocking an older architecture
The next group of failures was especially interesting.
Several tests still mocked:
ZorgaxPaymentIntent
ZorgaxSubscription
But the current production architecture uses:
PaymentIntent
ZorgaxPurchase
Entitlement
That means the Jest mocks looked correct, but they were intercepting classes the application no longer used.
The real Mongoose models were executing instead.
Then Jest waited for the real database.
Five seconds later:
Exceeded timeout of 5000 ms
The problem was not a slow application.
It was a stale test boundary.
Test runtime: 11 seconds to less than one
One replay-protection test originally took around:
11 seconds
and failed with two timeouts.
After updating the test to mock the real architecture:
2 tests passed
in roughly:
0.7 seconds
That is a strong indicator that the previous problem was accidental external I/O.
The new replay protection verifies something closer to the real current contract:
same payment reference
↓
existing purchase?
↓
same owner?
┌────┴────┐
yes no
↓ ↓
reuse reject
Only after ownership is validated does the system grant the entitlement.
What we are doing now
The main runtime is now much healthier.
We have already completed:
- migration to the new operational GitHub account;
- repository remote cleanup;
- branch synchronization with main;
- OAuth regression verification;
- systemd runtime cleanup;
- duplicate PM2 gateway removal;
- backend MongoDB bootstrap;
- Metaverse health recovery;
- graceful backend database shutdown;
- Zorgax access compatibility cleanup;
- AI budget test alignment;
- realtime fail-closed authorization;
- Bitcoin circular-dependency removal;
- modernization of several outdated payment tests.
The current focus is the remaining legacy test suite.
Some tests still represent old implementation details instead of the current architecture.
So we are going through them carefully and asking:
Is this a production bug, or is the test describing a system that no longer exists?
That is the work happening right now.
The bigger architectural goal
The cleanup is gradually making the boundaries clearer:
systemd
-> explicit gateway bootstrap
gateway bootstrap
-> explicit database lifecycle
realtime authorization
-> explicit availability checks
payments
-> unified checkout boundary
payment constants
-> dependency-free leaf module
access
-> entitlement authority
tests
-> mock current production boundaries
The goal is not merely to get a green test suite.
The goal is to have a green test suite that means something.
Small commits are helping a lot
We are keeping fixes isolated.
Examples include:
fix(runtime): initialize backend Mongo before gateway listen
fix(zorgax): include optional legacy subscription access source
test(zorgax): align AI budget expectation with global monthly cap
fix(realtime): fail closed when community membership authority is unavailable
fix(zorgax): break payment asset circular dependency
Small commits give us three advantages:
- easier review;
- easier rollback;
- easier reasoning.
Git becomes part of the debugging process, not just version storage.
What we deliberately did not do
Some shortcuts would have made the situation worse.
We did not:
reset the repository destructively
delete worktrees blindly
force-de-duplicate Mongoose during an incident
increase Jest timeouts to hide MongoDB access
hide circular dependencies with fallback values
commit credentials or environment secrets
A large part of production engineering is not just knowing what command to run.
It is knowing which commands not to run.
Current simplified architecture
The current direction looks like this:
GitHub
↓
danieldirimini-myzubster/myzubster
↓
VPS /root/myzubster
↓
systemd
↓
MyZubster Gateway
├── Root application
├── Backend Mongo context
├── Metaverse
├── Realtime
├── Zorgax
└── Payment services
Other independent services can still use PM2 where appropriate.
The important improvement is that there is now a much clearer owner for each runtime responsibility.
What comes next
The immediate plan is:
1. finish updating the remaining legacy payment tests;
2. review source-inspection tests that still point to compatibility wrappers;
3. run the complete Jest suite again;
4. separate real regressions from intentionally changed contracts;
5. verify production health;
6. prepare the branch for review and merge into main.
The goal is simple:
align the repository, runtime, architecture, and tests before merging.
Final lesson
This entire debugging session started with:
GitHub OAuth 404
and evolved into:
GitHub migration
↓
runtime cleanup
↓
systemd stabilization
↓
MongoDB lifecycle fix
↓
Metaverse recovery
↓
authorization hardening
↓
Bitcoin dependency cleanup
↓
legacy test modernization
That is often how real engineering works.
The first visible error is not always the real problem.
Sometimes it is simply the first place where the system tells you to look deeper.
We are continuing that work now on MyZubster.
Project:
https://myzubster.com
GitHub:
https://github.com/danieldirimini-myzubster/myzubster
Marketplace:
https://myzubster.com/marketplace
Metaverse:
https://myzubster.com/metaverse
LIFE Pilot:
https://myzubster.com/life-pilot
The next milestone is to finish the legacy test alignment, validate the entire branch, and prepare it for merge into main.
One boundary at a time.