2 comments

  • tgluck3 hours ago
    Author here. This puts a proxy in front of repeated Jev classification calls. At first everything goes to Jev; from Jev&#x27;s answers it trains a small head on frozen sentence embeddings, picks a confidence threshold with an exact finite-sample bound so that at most 2% of all requests get an answer Jev wouldn&#x27;t have given, and then answers the confident share locally at ~15 ms on a CPU. A permanent 2% audit keeps checking; if agreement breaks, everything falls back to Jev and it retrains.<p>Known limits: agreement is not accuracy (if Jev is wrong, so is the local model); coverage tracks how consistent Jev itself is (22% on noisy tweet tasks, 80% on news); it speaks Jev&#x27;s API only, an OpenAI-compatible front is on the roadmap. Since 0.4.0 the guarantee can also cover &quot;would Jev have been unsure&quot;, which matters if your code routes low-confidence answers to review. Apache 2.0.