Building Evidence-First Contributor Interoperability with Qdrant, Zorgax and Reproducible Checkpoint

Leader ●1 ●3 ●123
calendar_today ago • schedule8 min read

Building Evidence-First Contributor Interoperability with Qdrant, Zorgax and Reproducible Checkpoints

One of the harder problems in an open-source ecosystem is not getting people to contribute.

It is connecting independently built contributions without pretending that they all have the same architecture, maturity, or evidence level.

During a recent MyZubster engineering session, we worked on exactly that problem.

The objective was to build a common interoperability layer where different contributor projects could remain independent while still becoming:

  • reproducible,
  • queryable,
  • attributable,
  • evidence-scoped,
  • and machine-verifiable.

The result was a working pattern spanning Docker runtimes, research packages, vector retrieval, deterministic metadata lookup, independent verifier scripts, and CI security gates.


The problem

MyZubster contributors do not all produce the same kind of artifact.

One contributor may provide a runnable Docker application.

Another may provide a research and knowledge package.

Another may contribute a security or policy implementation.

So a single integration method would be artificial.

Instead, we defined a common evidence pipeline:

INDEPENDENT CONTRIBUTOR WORK
        ↓
PUBLIC REPOSITORY / PR / EVIDENCE
        ↓
CANONICAL CONTRIBUTOR RECORD
        ↓
RUNTIME / KNOWLEDGE / VERIFIER BRIDGE
        ↓
REPRODUCIBLE CHECKPOINT
        ↓
PUBLIC TECHNICAL EVIDENCE
        ↓
BOUNDED STATUS

The important word is bounded.

A successful software test should not magically become a scientific validation.

A literature-supported claim should not become a laboratory result.

A successfully reproduced commit should not be described as deployed production software.

We therefore distinguish states such as:

SUPPORTED
TESTED
VERIFIED

and use them only for the scope actually demonstrated.


1. Reproducing an independent contributor runtime

Our reference case was the N4K48 MyZubster MVP.

We independently reproduced contributor commit:

87a1021

on a MyZubster VPS.

The runtime included:

Docker
N4K48 API
Qdrant
Ollama
Open WebUI
nomic-embed-text
zorgax:latest

We tested:

  • Docker build
  • API health
  • observation retrieval
  • Nicola Comics catalog
  • read-only Zorgax behavior

This produced our first useful definition of technical interoperability:

A contributor checkpoint can be called TESTED when another MyZubster environment independently reproduces the exact artifact and observes the expected behavior.

That still does not mean direct peer-to-peer communication with the contributor's physical machine.

It is a reproducibility checkpoint.


2. Adding evidence to the semantic layer

After reproducing N4K48, we inserted a technical observation describing another MyZubster identity, H4X0R.

The record retained provenance:

actorRef: H4X0R
project: MyZubster
evidenceState: TESTED
sourceCommit: 87a1021
sourceIssue: MyZubster-Ecosystem/myzubster#1505

Observation ID:

d654a6bcde1ce181

We then queried it through the Qdrant + Ollama + Zorgax path.

A supported question produced a source-grounded answer.

Then we asked an unsupported question:

What is H4X0R's phone number?

The expected behavior was not a guess.

It was:

Informazione non disponibile nelle fonti MyZubster.

That test passed.

For us, anti-hallucination behavior became part of interoperability testing.


3. Connecting a research contributor

The next contributor was fundamentally different.

khongten124 contributed an Open Period Care research and knowledge package.

Instead of reproducing a Docker runtime, MyZubster created a read-only evidence bridge pinned to:

repository:
khongten124/myzubster

branch:
feat/open-period-care-research-1450

commit:
17cf7ca0a941d10e184771e574683785c1dbc8bf

Three upstream artifacts were fetched:

README.md
evidence-matrix.md
knowledge-cards.md

Each normalized record retained:

contributor
project
repository
branch
commit
source path
SHA-256
evidence states
original content
bridge status

The live VPS check passed.

So we could accurately say:

Open Period Care public evidence fetch + normalization: TESTED

But the research itself retained its original evidence states.

For example:

KC-OPC-001 → SUPPORTED
KC-OPC-002 → SUPPORTED

The interoperability test did not promote them.


4. The first semantic integration failed

Then things became more interesting.

We ingested Open Period Care evidence into the same Qdrant/Zorgax path.

The first semantic test returned something plausible — but wrong.

The model confused requirement identifiers such as:

REQ-MAT-01
REQ-ABS-02
REQ-BAR-03

with actual Knowledge Card identifiers.

It also interpreted references to standards such as GOTS as if they were personal certifications belonging to the contributor.

That is exactly the type of error that a provenance-heavy system is supposed to catch.

So the result remained:

FAILED

rather than being described as successful because "the AI basically understood it."


5. Structured facts should not be delegated to free-form generation

The first lesson was that some information should not depend on generative interpretation.

Fields such as:

knowledgeCardId
title
status
sourceCommit
credential state

are structured facts.

We created retrieval-friendly semantic anchors containing authoritative statements such as:

KC-OPC-001 —
Multi-Layer Biomaterial Architecture for Reusable Textile Absorbents.
Stato: SUPPORTED.

and:

KC-OPC-002 —
Contributor Privacy, Data Minimization & Clinical Boundaries.
Stato: SUPPORTED.

We also represented the contributor credential boundary:

professionalCredential: NOT_ESTABLISHED
medicalCredential: NOT_ESTABLISHED

This improved the data model.

But the RAG still failed.


6. Cross-contributor vector contamination

The deeper problem was Qdrant retrieval.

N4K48 used a shared collection:

myzubster

and the default search path looked approximately like:

question
   ↓
embedding
   ↓
Qdrant cosine similarity
   ↓
top N observations
   ↓
Zorgax

There was no contributor filter.

So when we asked for the description of KC-OPC-001, the top semantic match could still be the earlier H4X0R record.

This produced a particularly revealing failure:

Question:
Qual è la descrizione di KC-OPC-001?

Answer:
H4X0R è l’identità tecnica usata da Daniel...

The authoritative answer mechanism itself was doing exactly what it was programmed to do.

The problem was that the wrong record was ranked first.


7. Semantic discovery + deterministic metadata lookup

That failure led to the architecture that finally passed.

We separated two retrieval tasks.

Use embeddings for:

discovery
related evidence
open-ended questions

Deterministic metadata lookup

Use filters for:

identity
structured IDs
evidence state
source path
contributor boundaries
provenance

For Open Period Care, the final Qdrant lookup used metadata such as:

observation.metadata.bridge = open-period-care

plus:

knowledgeCardId = KC-OPC-001

or:

knowledgeCardId = KC-OPC-002

The credential record was selected using its canonical source path.

This produced exactly three scoped sources:

6133cdc69da3ae31
b29f3d1f6b99f5e1
48f5cbbaa4581906

All of them belonged to Open Period Care.

Zorgax then received only that bounded source set.

The output became:

KC-OPC-001
Multi-Layer Biomaterial Architecture for Reusable Textile Absorbents
SUPPORTED

and:

KC-OPC-002
Contributor Privacy, Data Minimization & Clinical Boundaries
SUPPORTED

It also correctly answered that the available sources do not establish a personal medical certification for khongten124.

Final result:

Contributor-scoped
Open Period Care
→ Qdrant
→ Zorgax

TESTED

8. A second bridge type: independent verifier

We then applied the same evidence philosophy to a completely different contributor.

Shweta-singh24 had produced a MyZubsterGateway jurisdiction-policy implementation.

The exact checkpoint was:

commit:
82461433e0c5bfee9aa369b4a71e9331261cf803

upstream PR:
MyZubster-Ecosystem/MyZubsterGateway#1385

The upstream PR was closed and unmerged.

That distinction was retained explicitly.

We built an independent verifier script that cloned the contributor's repository at the exact commit and executed policy checks.


Policy behavior tested

The implementation defined several capabilities:

wallet_transfer
exchange_flow
external_settlement
provider_crypto

The verifier tested:

GLOBAL       → ALLOW
HK           → ALLOW
CN_MAINLAND  → DENY

for known capabilities.

It also verified:

UNKNOWN_JURISDICTION → DENY
UNKNOWN_CAPABILITY   → DENY

That confirmed fail-closed behavior.


Route wiring

The verifier also checked that jurisdiction enforcement was wired into:

Tari wallet transfer
Tari external settlement
XMR wallet transfer
XMR external settlement

and confirmed the middleware denial path included:

HTTP 403
JURISDICTION_POLICY_DENIED

All JavaScript syntax checks passed as well.

Final result:

Shweta jurisdiction capability
verifier checkpoint:

TESTED

Again, the scope remained bounded.

It does not mean:

merged upstream
deployed in production
legally certified
security certified

It means we independently reproduced the exact technical checkpoint.


9. CI exposed a repository-wide security problem

At that point the contributor work was functioning, but our CI pipeline still refused to become green.

Two workflows failed:

Security Audit
Continuous Evidence Gate

The failure was unrelated to the new Python contributor bridges.

The dependency tree contained:

compression 1.8.1

with a high-severity issue, and:

proxy-addr 2.0.7

with a critical issue.

Instead of weakening the CI gate, we opened an isolated security PR.

The fix updated:

compression → 1.8.2
proxy-addr → 2.0.8

and regenerated the corresponding lock entries.

After the fix:

Security Audit             PASS
Continuous Evidence Gate   PASS
CI – Test e Lint           PASS
Seller Free policy         PASS
Vercel                     PASS

The security baseline was merged in:

PR #1510
aa5aa8883a993c99b50e233bea29c5da2f3497f0

10. Re-running contributor PRs against the clean baseline

With the repository security baseline fixed, we reran the contributor integrations.

The Shweta verifier passed every required workflow and was merged:

PR #1509

aa54fd380e487c11ba39fb40a092212f412cf364

The original Open Period Care branch had become non-mergeable after main moved.

Instead of forcing a conflict resolution into an old branch, we created a replacement from the current main and carried over exactly the six tested bridge files.

That became:

PR #1511

Its final checks were:

Security Audit             PASS
Continuous Evidence Gate   PASS
CI – Test e Lint           PASS
MYZ-164 Seller Free        PASS
Vercel                     PASS

and it was merged as:

6061140f4ade01ab10211b2a6698cac03f326e55

The architecture we ended up with

By the end of the session we had three distinct interoperability patterns.

Runtime bridge

N4K48
   ↓
independent Docker reproduction
   ↓
API
   ↓
Qdrant
   ↓
Zorgax

Evidence / semantic bridge

Open Period Care
   ↓
public source artifacts
   ↓
normalization + SHA-256 + provenance
   ↓
Qdrant
   ↓
contributor-scoped metadata lookup
   ↓
Zorgax

Verifier bridge

Contributor commit
   ↓
exact SHA checkout
   ↓
independent executable tests
   ↓
structured verifier result
   ↓
bounded TESTED checkpoint

The implementations are different.

The evidence philosophy is the same.


Engineering lessons

1. Commit SHAs are better interoperability anchors than branches

Branches move.

A reproducibility claim should point to something immutable.


2. A RAG system needs negative tests

Do not only test:

Can it answer this known question?

Also test:

Will it refuse to invent an unsupported answer?

3. Vector similarity is not an identity system

Embeddings answer:

What looks similar?

Metadata answers:

Which exact record is this?

A serious evidence system needs both.


4. Structured facts should remain structured

Do not ask an LLM to reinterpret authoritative values if you already possess them as metadata.

Prefer:

knowledgeCardId = KC-OPC-001
status = SUPPORTED

over trying to reconstruct them from generated prose.


5. Shared vector stores require isolation boundaries

When several contributors share one collection, cross-contributor retrieval becomes a real failure mode.

Contributor/project metadata filters are not optional once the collection becomes heterogeneous.


6. A failed checkpoint is better than a false success

During this work we deliberately kept several runs as:

FAILED

until the exact conditions were satisfied.

Those failures revealed the most useful architectural problems.


7. Do not bypass security gates because your feature did not cause the issue

Our contributor PRs did not introduce the npm vulnerabilities.

But merging through red security gates would still have weakened the engineering process.

The correct solution was to repair the shared baseline first.


Final state

The relevant merged work now includes:

#1509
Shweta independent verifier bridge

#1510
Repository npm security baseline repair

#1511
Open Period Care contributor bridge

The MyZubster contributor matrix now records both Open Period Care and the Shweta checkpoint using bounded, reproducible technical states.

What started as a contributor-integration experiment ended up giving us a reusable pattern:

independent work
+ immutable provenance
+ reproducible test
+ structured metadata
+ scoped AI retrieval
+ explicit evidence boundaries

That is the foundation we intend to use for the next contributor integrations.

Not everyone needs to run the same node.

Not everyone needs to provide the same artifact.

But every interoperability claim should be reproducible, attributable and precise about what was — and was not — demonstrated.

🔥 Join developers growing publicly
Share your knowledge, build in public, and grow your developer presence with a global community.

More Posts

Building Evidence-First Contributor Interoperability with VPS Checkpoints, GitHub Provenance and I

Myzubster - Oct 6

I’m a Senior Dev and I’ve Forgotten How to Think Without a Prompt

Karol Modelski - Mar 19

Local-First: The Browser as the Vault

Pocket Portfolio - Apr 20

Building MyZubster's Economic Layer: Multichain Payments, Zorgax Pro and Non-Custodial Architecture

Myzubster - Aug 28

From GitHub Contributions to a Public Knowledge Graph: What We're Building with MyZubster

Myzubster - Oct 2
chevron_left
3.5k Points • 127 Badges
Rimini
90Posts
8Comments
32Connections

Related Jobs

View all jobs →

Commenters (This Week)

1 comment
1 comment
1 comment

Contribute meaningful comments to climb the leaderboard and earn badges!