Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
pollix 
posted an update about 15 hours ago
Post
43
stuntd 0.1.4 is out, small one.

Last time someone asked what the novelty gate does at the 0.99 quantile instead of 0.95. I checked: at 0.95 every head sends about 5% of totally normal traffic to the big model on purpose, and with three fields that adds up. So 0.99 is the default now.

It also ships examples/support/check.py. One command trains the demo heads twice (all templates, and with one template per category left out) and prints the whole sweep, screenshot is its output with the junk-input part cut. Every generalization number in the README comes from that script now, so you can rerun it instead of trusting me.

What it shows on the support demo:
- normal tickets answered locally: 70.6% at 0.99 vs 67.1% at 0.95 (72.3% with no gate), still 97% right
- tickets from templates the heads never saw: 0.3% answered locally at 0.99. Without the gate the heads answer 69% of those and get only 42% right
- the weather question, asdf and {} are stopped at every setting

Also auto_retrain finally works for Jev captures, it only counted OpenAI and Anthropic before.

pip install -U stuntd

https://github.com/bladedevoff/stuntd/releases/tag/v0.1.4

Your three numbers already say something about how the heads fire together.

No gate: 72.3% answered locally. At 0.95 that drops to 67.1%. The gate pulls back 5.2 points, 7.2% of what was local. Three heads each firing on an independent 5% would pull back 1 - 0.95^3 = 14.3%.

At 0.99: 1.7 points, 2.4% of local, against 3.0% for three independent 1% heads.

So at 0.95 the gate costs about half of what independent heads would. At 0.99 about four fifths.

Two readings fit that. The heads fire on the same tickets. Or they fire on tickets that were not confident anyway, so the gate costs nothing there.

Which one is it? check.py could print how many local answers are lost to exactly one, two or three heads.