Abliterated thinking models have a habit: they decide what to say, then keep arguing with themselves until the token budget runs out. Qwen3.8-27B Abliterated ThinkFix is an uncensored Qwen3.8-27B with one output row edited so the model closes its thinking when it's ready to answer.
On 40 sensitive prompts the edit never saw, 8k output cap: 33 replies finished cleanly instead of 22, 5 hit the limit instead of 17, 61 minutes for the set instead of 75.
Agent use was checked four ways and the edit changes nothing there. MTP draft head kept. Q4_K_M to Q8_0, each tested after the edit. The card lists what the edit was fit on and what it does not do.
We're working on Pebble 1.5. Here's what we know so far:
- They will be better than the last generation. 99.99% certain. - Expanded context lengths of at least 16,384 tokens, with the flagship potentially reaching 32,768. - A Mamba3-based architecture with some other new architectural designs we're experimenting with. - Native CPU compatibility — something we failed at with the last generation. - Natively multilingual and multimodal???
2. SmolCodeBench
A code benchmark designed specifically for small models, because there really isn't a good one right now.
3. SENTRY
VOID is working on something called SENTRY — System for Evaluating Neural Threats, Responses, and Yields.
More on that soon.
4. basically OS
It's an operating system/app/harness. We're still deciding.
5. Finances
Trying to balance the finances after purchasing a Hugging Face Pro subscription.