the lab notebook

We publish the results that argue against us.

Everything here is measured on code, most of it code we did not write. Every number arrives with what it measured, how many samples, what its negative control scored, and what it does not show. When a result comes back flat, that is the post.

  1. 02measurementretrievalmethod

    Our retrieval was throwing away a third of its own gains

    A better embedder measured 3.1 points better alone and produced nothing end to end. We found the stage eating it by counting, not testing. The proof was the number that did not move.

    +10.1ptevidence recall on conversation, ours against ours, same harness and questions before and after.

    Read
  2. 01measurementreview

    We built a code graph seven times bigger. It changed defect recall by exactly zero.

    A full-repository graph, seven times richer than the one it replaced, moved review recall by zero commits out of 85. We predicted that in writing before the run.

    p=0.803paired McNemar, ungrounded against a full-repository graph. Eight commits won, eight lost.

    Read
RSS feedFollow this in a reader and new posts arrive without you checking the site.