The pre-registration point is the right one and most people skip it, so I want to push on the part it does not cover.
Freezing the baseline fixes memory. It does not fix grading. You write the old answer, you query, and then you decide whether decision 7 came out sharper. The person deciding is the person who wants the graph to earn its place. That bias survives pre-registration completely intact, and it is the one that inflates a 6 of 12.
The useful part is that you already have the ingredient most people leave out. You are writing a confidence number, which makes the prediction resolvable. So score the outcome rather than the answer. Set a check date when you write the baseline, and on that date record what actually happened. "Sharper" is a judgment you make about your own work. "The thing I put at 70% happened" is not.
It also reaches the number you currently cannot verify. You can count errors the graph caught. You cannot count errors it introduced that you accepted, because a confident wrong correction that fits your prior is indistinguishable from a right one at the moment you take it. Pre-registration actively hides that case, since the graph changed your mind and you log it as a win.
To answer your closing question directly: my retrieval layer earns it the day a low-confidence call it pushed me toward resolves correctly while the one I would have made alone resolves wrong. Until something resolves against me, I am measuring agreement rather than value.
Twelve decisions in three weeks is a small sample, and I would not be in a hurry to grow it. Deferred scoring is what makes each one worth more.