Engineering an Evidence-First Knowledge Graph: From GitHub Artifacts to On-Chain Proofs in MyZubster

Leader ●1 ●3 ●107
calendar_today ago • schedule10 min read

Certo. Per CoderLegion lo farei più tecnico del post DEV, con più attenzione ad architettura, data provenance, canonicalizzazione, hashing e al motivo per cui dopo il test con Nicola stiamo valutando standard esistenti invece di continuare ad aggiungere smart contract custom.
Engineering an Evidence-First Knowledge Graph: From GitHub Artifacts to On-Chain Proofs in MyZubster
Building and testing a complete provenance chain with a real contributor, canonical payloads, SHA-256, Ethereum Sepolia, RAG, and a Knowledge Graph
Tonight we ran one of the most complete end-to-end experiments so far on MyZubster.
The test involved a real contributor, Nicola / N4K48, rather than synthetic fixtures.
What started as a relatively simple test of the MyZubster Knowledge Profile evolved into a much larger architecture experiment involving:
Authentication

↓

Knowledge Profile

↓

Private Knowledge Cards

↓

Publication

↓

Real software contribution

↓

GitHub evidence

↓

RAG / Qdrant

↓

Knowledge Graph

↓

Canonical payload

↓

SHA-256

↓

Ethereum Sepolia

↓

Publicly inspectable provenance

The most important result was not the blockchain transaction.
It was discovering where the boundaries of the MyZubster architecture should probably be.

  1. The problem we are trying to solve
    Most professional profiles fundamentally contain claims.
    For example:
    {
    "person": "N4K48",
    "skill": "Docker"
    }

This tells us what somebody says about themselves.
It tells us very little about why another person should trust that statement.
We wanted to experiment with a different model:
Person
↓
Claim
↓
Activity
↓
Artifact
↓
Evidence
↓
Cryptographic commitment
↓
Attestation

The objective is not to magically transform a claim into truth.
The objective is to make its provenance inspectable.

  1. Testing with a real contributor
    Nicola created a MyZubster Knowledge Profile representing different parts of his real activity.
    We used three Knowledge Cards during the experiment:
    N4K48
    │
    ├── Truck driving and transport organization
    │
    ├── Docker / AI experiments on myzubster-mvp
    │
    └── Learning and collaboration in MyZubster

This immediately forced us to distinguish two concepts that are easy to accidentally merge:
CLAIM

and:
EVIDENCE

A professional statement is not automatically evidence.
A GitHub commit is not automatically proof of every professional statement either.
Therefore MyZubster needs to preserve the semantic boundary between them.

  1. Authentication failed before cryptography did
    The first failures had nothing to do with Ethereum.
    Nicola encountered an expired JWT during the onboarding flow.
    The UI indicated an authenticated state, but a subsequent request produced:
    jwt expired

Another issue appeared in the conversational onboarding flow: information already supplied by the user was not always being reused correctly.
We also discovered a routing issue where the Knowledge Profile Builder URL displayed the MyZubster homepage instead of the Builder.
These problems were fixed before continuing.
The authentication behavior was changed so that an expired session could trigger re-authentication while preserving the locally prepared onboarding information.
This matters because provenance starts before hashing.
If users lose their work because authentication expires, the cryptographic layer does not save the product.

  1. Private-first Knowledge Cards
    The next test covered persistence.
    Nicola created three cards and saved them privately.
    The tested state transition was approximately:
    NEW
    ↓
    PRIVATE DRAFT
    ↓
    RELOAD
    ↓
    EDIT
    ↓
    PREVIEW
    ↓
    PUBLISH

The important property here is that creation does not imply publication.
A Knowledge Card can exist privately while the owner reviews:
claims
activities
sources
evidence
verification state

Only an explicit publication action makes the card public.
Published cards can also be withdrawn before being modified.
That gives us a lifecycle closer to:
DRAFT
│
▼
PUBLIC
│
▼
WITHDRAWN / DRAFT

rather than treating the knowledge database as an append-only public feed.

  1. Turning software work into evidence
    The next step was intentionally practical.
    Instead of adding another manually entered skill, Nicola performed a real change in his myzubster-mvp repository.
    The work involved:
    scripts/ingest_knowledge.py

and improvements to the chunking/knowledge-ingestion process.
The resulting test reported:
46 knowledge chunks loaded into Qdrant
1 empty document skipped

The Git commit associated with that work could then become a source attached to the relevant Knowledge Card.
We now had something structurally different from:
N4K48
↓
"knows software development"

We had:
N4K48
↓
Knowledge Card
↓
documented software activity
↓
GitHub repository
↓
specific commit
↓
inspectable code change

This is an important distinction.
MyZubster does not need to decide that the commit proves a particular competence.
It can state something narrower:
This artifact is evidence associated with this activity and this Knowledge Card.

  1. RAG became part of the evidence trail
    The test also exercised the knowledge retrieval stack.
    The local project involved:
    Docker Compose
    │
    ├── Ollama
    │
    ├── Qdrant
    │
    ├── API
    │
    └── Open WebUI

During retrieval testing, the context being returned was insufficient.
The experiment therefore changed:
AI_CONTEXT_LIMIT=1

to:
AI_CONTEXT_LIMIT=5

and repeated the retrieval tests.
That work itself became documentable evidence.
This creates a useful recursive property.
The system can represent knowledge about the process used to improve the system itself:
engineering work

  ↓

Git artifact

  ↓

Knowledge Card

  ↓

Knowledge Graph

  1. Building the Knowledge Graph
    Once the cards were public, we started connecting them into a graph.
    Conceptually:
                N4K48
                  │
     ┌────────────┼────────────┐
     │            │            │
     ▼            ▼            ▼
    

    Transport Learning Docker / AI
    Card Card Card

                                 │
                                 ▼
                             GitHub repo
                                 │
                                 ▼
                               commit
    

Sources become traversable relationships instead of text buried inside a profile.
During this phase we found another real-world infrastructure issue.
The visual Knowledge Graph deployment was protected by Vercel, preventing Nicola from accessing it.
After adjusting deployment accessibility, testing continued, including from mobile.
We later corrected another graph/catalog connection that produced a "Catalog unavailable" result.
These may look like small frontend/infrastructure bugs, but they matter for an evidence system.
Evidence that technically exists but cannot be navigated is significantly less useful.

  1. Proof v1: hashing the wrong thing
    Then we introduced Ethereum Sepolia.
    The first MyZubsterProof experiment anchored a SHA-256 digest related to the Knowledge Card.
    However, the digest represented the Knowledge Card URL.
    Conceptually:
    Knowledge Card URL
    ↓
    

    SHA-256

    ↓
    

    bytes32

    ↓
    

    Ethereum Sepolia

This works cryptographically.
But semantically it proves much less than we wanted.
Suppose:
https://example.com/card/123

continues to exist while the content served at that URL changes.
Hashing the URL does not detect the content change.
The first experiment therefore became useful precisely because it exposed this weakness.
We explicitly documented it rather than treating an on-chain hash as stronger evidence than it actually was.

  1. Proof v2: canonicalize before hashing
    Proof v2 changed the architecture.
    Instead of hashing a location, we created a deterministic representation of the Knowledge Card itself.
    Conceptually:
    Knowledge Card
    ↓
    canonical representation
    ↓
    exact bytes
    ↓
    SHA-256
    ↓
    bytes32
    ↓
    Ethereum Sepolia

For the tested card, the resulting digest was:
6097e05866bafceec24663d2638cb1dae5742ac78284abbfd45cc9c3b0bfb845

The corresponding on-chain value was:
0x6097e05866bafceec24663d2638cb1dae5742ac78284abbfd45cc9c3b0bfb845

Reading knowledgeHash() from the deployed MyZubsterProof contract returned that same value.
The canonical payload and the verification procedure were also committed to the repository.
That gives us a reproducible verification path:
canonical payload

   │
   ▼

SHA-256(payload)

   │
   ▼

expected digest

   │
   ▼

on-chain knowledgeHash()

   │
   ▼

comparison

If all values match, we can establish that the committed canonical representation corresponds to the digest stored by the contract.

  1. Why canonicalization matters
    This is a subtle but critical engineering detail.
    Consider two JSON objects:
    {
    "name": "N4K48",
    "activity": "Docker"
    }

and:
{
"activity": "Docker",
"name": "N4K48"
}

Semantically, an application may consider them equivalent.
At the byte level, they are different.
And therefore:
SHA256(payloadA) != SHA256(payloadB)

can occur even when both payloads represent the same logical information.
The same issue appears with:
whitespace
field ordering
Unicode normalization
line endings
optional values
number serialization
dates

So an evidence system cannot simply say:
hash(JSON.stringify(object))

and assume the problem is solved forever.
The canonicalization algorithm effectively becomes part of the protocol.
A durable proof therefore needs something like:
payload format
canonicalization version
hash algorithm
digest
attestation reference

Proof v2 pushed us toward that realization.

  1. Integrity is not truth
    This is probably the most important architectural rule from the experiment.
    Assume the canonical payload contains:
    {
    "claim": "N4K48 is an expert in Docker"
    }

and suppose we perfectly canonicalize it, hash it and anchor the digest on Ethereum.
We have demonstrated something about the integrity of the statement.
We have not demonstrated that the statement itself is true.
Formally:
integrity(claim) ≠ truth(claim)

Likewise:
existence(GitHub commit) ≠ proof(skill)

and:
blockchain timestamp ≠ professional certification

Different evidence supports different propositions.
MyZubster therefore needs to model not only evidence, but also what each piece of evidence is capable of establishing.

  1. Provenance as a graph
    This leads to a more interesting representation.
    Instead of:
    Person → Skill

we can represent:
Person
│
▼
Claim
│
▼
Activity
│
▼
Knowledge Card
│
├──────────────► GitHub artifact
│
├──────────────► external source
│
▼
Canonical Payload
│
▼
Digest
│
▼
Attestation

Each edge has semantics.
For example:
Person --DECLARES--> Claim

Claim --SUPPORTED_BY--> Evidence

Activity --PRODUCED--> Artifact

KnowledgeCard --CANONICALIZED_AS--> Payload

Payload --HASHED_AS--> Digest

Digest --ATTESTED_BY--> Attestation

This is much more expressive than simply attaching a blockchain transaction ID to a user profile.

  1. The architectural question that emerged
    After completing Proof v2, Nicola compared the experiment with existing systems and standards.
    The comparison included:
    • Talent Protocol / Builder Score;
    • Ethereum Attestation Service;
    • cheqd;
    • Open Badges 3.0 and Verifiable Credentials.
      That comparison raised a very useful question.
      Why should MyZubster implement every cryptographic primitive itself?
      The ecosystem already contains mature approaches for:
      identity
      credentials
      attestations
      reputation
      issuers
      verification
      revocation
      cryptographic proofs

Creating:
MyZubsterProofV3.sol
MyZubsterProofV4.sol
MyZubsterProofV5.sol

may therefore be less valuable than making MyZubster interoperable with existing primitives.

  1. Separating the layers
    The experiment suggests that the architecture could eventually be separated into several layers.
    Identity layer
    Person
    Wallet
    GitHub identity
    MyZubster account

Knowledge layer
Knowledge Card
Claim
Activity
Skill
Project

Evidence layer
GitHub commit
Pull request
Repository
Document
External source
Artifact

Integrity layer
canonical payload
hash algorithm
digest
version

Attestation layer
Potentially:
Ethereum Attestation Service

rather than a growing family of MyZubster-specific contracts.
Credential layer
Potential interoperability with:
W3C Verifiable Credentials
Open Badges

Relationship layer
This is where MyZubster's Knowledge Graph becomes particularly important.
Person
↓
Claim
↓
Activity
↓
Evidence
↓
Artifact
↓
Attestation
↓
Credential

  1. MyZubster may be the graph, not the protocol
    This was probably the biggest architectural insight of the evening.
    Originally it is tempting to think:
    MyZubster
    =
    our smart contract

But after this experiment, another model looks more interesting:
MyZubster
=
evidence/provenance graph
+
workflow
+
verification semantics
+
human-readable navigation

with standards handling lower-level primitives where appropriate.
That means MyZubster could potentially consume:
GitHub evidence
EAS attestations
Verifiable Credentials
Open Badges
blockchain transactions
traditional documents
other public evidence

without pretending that all evidence has the same meaning.

  1. A possible future architecture
    The architecture we are now exploring looks closer to this:
                   PERSON
                     │
                     ▼
                   CLAIM
                     │
                     ▼
                  ACTIVITY
                     │
                     ▼
              KNOWLEDGE CARD
                     │
        ┌────────────┴────────────┐
        │                         │
        ▼                         ▼
    EVIDENCE                   ARTIFACT
        │                         │
        └────────────┬────────────┘
                     │
                     ▼
            CANONICAL PAYLOAD
                     │
                     ▼
                  SHA-256
                     │
         ┌───────────┴───────────┐
         │                       │
         ▼                       ▼
    

    ATTESTATION CREDENTIAL
    e.g. EAS VC / Badge

         │                       │
         └───────────┬───────────┘
                     ▼
             MYZUBSTER GRAPH
    

Notice that the blockchain is no longer the center of the system.
The relationships are.

  1. Why this differs from a reputation score
    Consider a reputation system that produces:
    Score = 87

That number may be useful for ranking or filtering.
But a provenance graph asks a different question:
Why?

A MyZubster-style path could eventually allow someone to inspect:
N4K48
↓
Docker-related Knowledge Card
↓
specific development activity
↓
GitHub commit
↓
source code
↓
canonical representation
↓
cryptographic digest
↓
attestation

The user is not forced to trust only the final number.
They can inspect the path that produced the evidence.

  1. Evidence types should not collapse into one trust score
    This is another consequence of the experiment.
    Imagine three pieces of evidence:
    A: self-declared professional experience

B: public GitHub commit

C: credential issued by a university

All three may belong in the graph.
But they should not automatically become:
A = B = C

Their semantics are different.
A useful graph therefore needs metadata such as:
evidence type
issuer
subject
source
verification method
verification status
created timestamp
attestation reference
revocation state

The graph should preserve those differences rather than hide them behind one universal score.

  1. What the experiment actually demonstrated
    By the end of the session we had exercised a chain approximately like this:
    Real contributor
    ↓
    authenticated MyZubster account
    ↓
    private Knowledge Cards
    ↓
    public Knowledge Cards
    ↓
    real software activity
    ↓
    GitHub evidence
    ↓
    RAG / Qdrant test
    ↓
    Knowledge Graph
    ↓
    canonical card payload
    ↓
    SHA-256
    ↓
    Sepolia Proof v2
    ↓
    public verification information

This is more useful to us than an isolated smart-contract demo because every layer exposed different failure modes.
We encountered:
expired authentication
routing problems
deployment access restrictions
catalog/navigation problems
retrieval-context limitations
evidence-modeling questions
URL-vs-content hashing semantics
canonicalization requirements
truth-vs-integrity boundaries

That is exactly what an end-to-end proof of concept should reveal.

  1. What we deliberately did not prove
    It is equally important to document what the experiment does not establish.
    Proof v2 does not prove that every statement in a Knowledge Card is true.
    It does not automatically prove professional competence.
    It does not establish that GitHub activity is equivalent to mastery of a skill.
    It does not turn blockchain data into a trusted credential issuer.
    What it does establish is a mechanism for connecting an exact representation of information with a reproducible digest and a public cryptographic reference.
    That is an integrity primitive.
    Trust requires additional layers.
  2. The next engineering experiment
    Instead of immediately creating another custom proof contract, the next useful experiment is likely to take the exact Proof v2 data model and map it onto existing standards.
    Something like:
    Knowledge Card
    ↓
    canonical payload
    ↓
    current Proof v2
    │
    ├────► EAS schema mapping
    │
    └────► Open Badges / VC mapping

Then we can compare:
custom MyZubster contract
vs.
standardized attestation

across:
schema flexibility
revocation
references
issuer semantics
verification
interoperability
developer complexity
gas cost
long-term maintainability

Only after that comparison should we decide what belongs inside MyZubster itself.

  1. The broader engineering lesson
    Tonight's work started with a very ordinary product task:
    Can a contributor create a knowledge profile?

It progressed into:
Can the profile contain inspectable evidence?

Then:
Can that evidence point to real engineering work?

Then:
Can the exact knowledge representation be cryptographically committed?

And eventually:
Should MyZubster even own the attestation protocol?

That progression matters.
A proof of concept should not merely prove that the original architecture works.
Sometimes its most valuable outcome is discovering that part of the original architecture should not exist.
For MyZubster, that may mean allowing established standards to handle credentials and attestations while concentrating our engineering effort on the part we actually care about:
human claim

↓

real activity

↓

inspectable evidence

↓

artifact

↓

verification

↓

provenance

and making that entire chain navigable.
That's the direction we're exploring next.
CoderLegion title
Engineering an Evidence-First Knowledge Graph: GitHub, SHA-256, Ethereum and MyZubster
Subtitle
How a real contributor test evolved from Knowledge Cards and RAG into canonical payloads, on-chain proofs, provenance graphs, and a rethink of our attestation architecture.
Short description / excerpt
We tested MyZubster end-to-end with a real contributor: authentication, private and public Knowledge Cards, GitHub evidence, Qdrant/RAG, a Knowledge Graph, canonical payloads, SHA-256 and Ethereum Sepolia. The most important result wasn't the smart contract—it was realizing that MyZubster may be better positioned as an evidence and provenance graph built on interoperable attestation standards.

🔥 Join developers growing publicly
Share your knowledge, build in public, and grow your developer presence with a global community.

More Posts

Sovereign Intelligence: The Complete 25,000 Word Blueprint (Download)

Pocket Portfolio - Apr 1

Architecting a Local-First Hybrid RAG for Finance

Pocket Portfolio - Feb 25

The Sovereign Vault — A Comprehensive Guide to Protocol-Driven AI

Ken W. Algerverified - Jun 4

The Privacy Gap: Why sending financial ledgers to OpenAI is broken

Pocket Portfolio - Feb 23

Local-First: The Browser as the Vault

Pocket Portfolio - Apr 20
chevron_left
3.1k Points • 111 Badges
Rimini
78Posts
8Comments
27Connections

Related Jobs

View all jobs →

Commenters (This Week)

2 comments
1 comment
1 comment

Contribute meaningful comments to climb the leaderboard and earn badges!