the lab notebook
Everything here is measured on code, most of it code we did not write. Every number arrives with what it measured, how many samples, what its negative control scored, and what it does not show. When a result comes back flat, that is the post.
A better embedder measured 3.1 points better alone and produced nothing end to end. We found the stage eating it by counting, not testing. The proof was the number that did not move.
+10.1ptevidence recall on conversation, ours against ours, same harness and questions before and after.
ReadA full-repository graph, seven times richer than the one it replaced, moved review recall by zero commits out of 85. We predicted that in writing before the run.
p=0.803paired McNemar, ungrounded against a full-repository graph. Eight commits won, eight lost.
Read