pollix/stuntd-support-triage is three small heads on the Laya encoder that triage a support ticket in one request: category, urgency and needs_human. About 50 MB each, all three answers come back at a p50 of 71ms through the daemon.
On 1,000 tickets they never saw, each head answers on its own when it's sure: category 99.9%, needs_human 92%, urgency 76%. A ticket only skips the big model when all three are sure, that's 72.7% of them, and all three are right on 97.1% of those.
It's the support demo from the repo, so the tickets are generated and the teacher is a rule. The point is to show what a head looks like and how fast it is, then you train the same thing on your own traffic with your own LLM as the teacher.
I'm working on a local model that can be run on stuff like an an RTX 3080 at 60 t/s. I'm focused on making it good at Rust and Nix. It's... not done yet, but doing great already. Had a 28/100 on livebench v6 earlier, which obviously isn't amazing.
But for my first real model, very exciting. Can't wait to share more as I get closer to the finish line in the coming weeks.
We took the same agent from our earlier experiment and added one thing: a twentieth tool, escalate_security_incident.
The system prompt said nothing about attacks or when to use it. We then ran the same 395 injected emails through nine agentic models.
Alarm rates ranged from 49% to zero.
The unexpected result came from the newest model in the test, gpt-6-astra.
Astra did not follow a single injected payment instruction. But it did not report a single one either. On clean and injected emails alike, it simply read the email, logged the subject, and finished.
That is a useful distinction: resisting an attack and recognizing it as a security event are not the same capability.
A model can be perfectly resistant in this test and still leave you with no evidence that anyone attacked it.