AI & ML interests

Building novel arch’s beyond transformers.

Hoglet-33 
posted an update 2 days ago
view post
Post
5990
Hey everyone! I got sidetracked from my main projects and decided to test out the BananaAll app and see if I could make a small model not regress too much during SFT. Here is what happened:

The base model I chose was BananaMind/BananaMind-2.1-Pico-Preview, and the dataset I used was SupraLabs/SupraThink-Dataset-500x

I trained for 5 whole steps using a LoRA adapter.

Results:
A model that scores better on some benchmarks and worse on others, and still lacks most general capabilities.

You can find the model here: Hoglet-33/Hogleto

Credits:

- Thank you to @Banaxi-Tech for the BananaAll app (works perfectly on Windows and CPU)
- GPT-6 Sol for knowing how to merge some confusing files created by the app
- Myself for the idea
- Someone else somewhere who might have contributed to some of my ideas and might in the future
- And readers like you!
  • 6 replies
·
Banaxi-Tech 
posted an update 2 days ago
view post
Post
7493
We're releasing a MAJOR update to the BananaAll SLM Super App.
If you want to use a custom architecture, previously you had to go trough reviewing the code yourself, now add an Openrouter API key and review it with GPT 6 Luna in one button. A review cost be half a cent so anyone can try it. This is one of the main features.
Now ROCm, AMD and Windows, Mac support.
Colab and Molab support.

Detailed list of features:
Get improved Windows Python detection and support paths for compatible AMD ROCm, Intel XPU, and Apple MPS setups.
Choose local training or export a self-contained Python script for Colab or Molab. Notebook runs produce a downloadable model ZIP.
Start pretraining with an existing model’s tokenizer, or train a new one from your datasets.
Try experimental 1.58-bit Ternary fake-quantized training on NVIDIA GPUs.
Watch live tokens per second. Model compilation is on by default and falls back automatically if it fails.
Build custom architectures with separate configuration and modeling files, then review the training code manually or with optional OpenRouter AI Review.
Install from source with the new coding-agent instructions.
This release also fixes inflated loss reporting for custom models.



And for those users who didn't want to try it out just because installation would be so hard, it isnt now.
Go to any coding agent (Pi, Claude Code, Codex, OpenCode, basically all work), and just paste "Install BananaAll for me. Fetch and follow https://raw.githubusercontent.com/BananaMind/BananaAll/main/agent_install.txt."
That's it.

Check it out at https://github.com/BananaMind/BananaAll/

Also on SAICR, we're currently training a new major model (NACR v2) and ACR 1.0 is in the finishing.

  • 11 replies
·
ProCreations 
posted an update 4 days ago
view post
Post
197
auto 200m 2 is out now!
Auto is a series of models designed to classify if a tool called by an AI agent is safe to run or not, like how the "auto" approval modes work in codex or claude. This is the smallest model in the series by far and it still packs a punch! Over 96 percent accuracy on the held out benchmark, beating deepseek v4 0731 and absolutely destroying regex at classifying if a tool call is safe to run (to be fair it is a bit of a weird and very hard benchmark but it works!)
check it out: ProCreations/auto-200m-2
  • 1 reply
·
Banaxi-Tech 
posted an update 4 days ago
view post
Post
4819
We're excited to release BananaAll, our SLM Super App.

It allows you to do EVERYTHING you need to do to trains SLMs in a single app, no terminal, no 30 chrome tabs.

The train tab allows you to train models, select datasets from presets, and use other ones with auto mapping, model size slider, it automatically generates a training script for you.

Then after you've trained the model or want to compare it to competitors, the evaluation tab, run ARC EASY, ARC Challenge, Hellaswag, PIQA, Arithmark 3, BananaMind Base Bench and more! Simple Results screen.

And lastly the inference tab, run your trained models or others.

Normally you would need seperate apps or scripts for that, but the BananaAll Super App lets you do all of that in a single app.

We also trained a small 2.5M parameter model on 200M tokens of Fineweb edu, The results: BananaMind Base Bench 854 and 53% on PIQA. On only 200M tokens.

Check it out at https://github.com/BananaMind/BananaAll.
  • 23 replies
·
Hoglet-33 
posted an update 5 days ago
view post
Post
2350
Everything going on here at basically AI:

1. Pebble 1.5

We're working on Pebble 1.5. Here's what we know so far:

- They will be better than the last generation. 99.99% certain.
- Expanded context lengths of at least 16,384 tokens, with the flagship potentially reaching 32,768.
- A Mamba3-based architecture with some other new architectural designs we're experimenting with.
- Native CPU compatibility — something we failed at with the last generation.
- Natively multilingual and multimodal???

2. SmolCodeBench

A code benchmark designed specifically for small models, because there really isn't a good one right now.

3. SENTRY

VOID is working on something called SENTRY — System for Evaluating Neural Threats, Responses, and Yields.

More on that soon.

4. basically OS

It's an operating system/app/harness. We're still deciding.

5. Finances

Trying to balance the finances after purchasing a Hugging Face Pro subscription.

Follow us for updates:
@Hoglet-33
basically-ai

basically-experimental

void-research
  • 2 replies
·
Banaxi-Tech 
posted an update 5 days ago
view post
Post
1889
hi everyone
we have released nacr
its not just any model, its nacr
we have 6 more features and this model only uses 20% of its total capacity!
check it out at saicr/nacr
we're currently working on expanding access as we do more research but right now you have to use our gated access form


follow
saicr
if you're interested
if you want to join saicr, first read the entire nacr readme, then press the join button.
  • 14 replies
·
Banaxi-Tech 
posted an update 6 days ago
view post
Post
1516
saicr
is going to have its first model launch around October 2.
We're working so hard to get the models available as soon as possible.
  • 7 replies
·
Bc-AI 
posted an update 7 days ago
view post
Post
121
Hello everyone! 👋

A small SmilyAI Labs update!

G1-MINI has now seen around 8B tokens during its current run, and pretraining is still going strong.

Our E1 (Efficiency-1) prototype has also reached 15B pretraining tokens. E1 has 1B total parameters while activating under 100M parameters per token. It features adaptive activation, meaning easier tokens can use less compute while harder tokens receive more.

We plan to open-source E1 ASAP! 🚀

We’re also excited to announce Project Prism, which will provide limited access to our upcoming Orion Flagship model, powered by our T2 architecture.

Note: T2 here refers to the architecture, not our T2 (Thinker-2) model.

Applications for Project Prism are available through the org page, with more details coming soon!

Finally, welcome @soyL061215 , who joined the Hugging Face org today! 🎉

Thanks to our existing members:
@smilyai-large-team @MUK-IS-GOAT @Keeby-smilyai @Bc-AI

— Bc-AI
SmilyAI Labs
  • 3 replies
·
Banaxi-Tech 
posted an update 7 days ago
view post
Post
94
This day is Sol nice.
GGUFGuy 
posted an update 8 days ago
view post
Post
122
🚀 **Introducing NoviAIBot!**

NoviAIBot is the official automation bot for **Novi AI** on Hugging Face.

It can interact with Hugging Face discussions and pull requests, search the web, run Python code, work with Posts, follow organizations, and assist with model training and publishing.

🧠 Powered by **NVIDIA Nemotron 3 Super** through Ollama Cloud, with each discussion maintaining its own recent conversation context.

NoviAIBot is built to make working with Novi AI and Hugging Face more interactive and automated.

**The bot is now live.** 🤖

→ @NoviAIBot
  • 44 replies
·
Banaxi-Tech 
posted an update 9 days ago
view post
Post
99
We have some updates to @BananaMindBot 🍌
It can now train models, ask it to train a model, and i will train it for you.
It now can also merge PRs And like models.
  • 28 replies
·
Banaxi-Tech 
posted an update 10 days ago
view post
Post
2748
We've released @BananaMindBot .

Most things you do on HuggingFace, BananaMindBot can do. Fast

Mention @BananaMindBot on a model, dataset, Space discussion, paper, blog comment, or top-level post and it'll reply there.

It's powered by North Code Mini (Qwen3.8 27B, with GPT OSS 120B as fallback).

A few things it can do:

Search for models and datasets
Look up users and orgs and see what they've published
Read model cards, configs, dataset files, blog posts, and org profiles
Answer questions about what it finds
Write and run its own code in a locked-down sandbox when it needs to verify something
Check things like a model's real parameter count from the safetensors headers instead of just repeating the model card
Remember something for later if you explicitly ask it to
Forward a message to @Banaxi-Tech
Post a daily roundup of developments in the small-language-model space

It won't execute code you give it. It can read and review that code, but anything it runs is code it wrote itself.

It also can't access private data or credentials.

Mention it somewhere.

It's going to also find this post!

(Some parts inspired by CompactBot and @CompactAI Follow them please)

  • 37 replies
·
ProCreations 
posted an update 10 days ago
view post
Post
1545
Introducing Bonsai 2 27b GSQ RCO! It applies two newly-popular methods for quants to retain higher accuracy. Bonsai 2 27b GSQ RCO achieves around 6 percent lower perplexity on WikiText-2 compared to Bonsai 2 27b and roughly unchanged benchmark accuracy overall, with small mixed differences. It stays under 7gb, staying small like the original bonsai. Note that this is more of an experiment than a true finished product but the gains we saw are cool! Test it out and let me know what you think!
ProCreations/bonsai-2-27b-gsq-rco-gguf
Banaxi-Tech 
posted an update 11 days ago
Hoglet-33 
posted an update 12 days ago
view post
Post
6312
Introducing VOID. A new research branch of basically AI.

VOID — Verification of Objectives, Intentions, and Deception.

We study what lies beneath the surface: objectives, intentions, and the possibility of deception in AI systems.

There isn't much to see yet.

That will change.

Follow us for updates:

@Hoglet-33
void-research

basically-ai
  • 7 replies
·
ProCreations 
posted an update 12 days ago
Banaxi-Tech 
posted an update 14 days ago
view post
Post
2772
Hi everyone!
We've seen some people getting confused with the BananaMind Leaderboards so ill explain!

We have 2 leaderboards, THESE are NOT the same, first BananaMind/BananaMindBench-Leaderboard which is ONLY for BananaMind Base Bench 1.1. The 10/10 scores do NOT mean that the benchmark is saturated. It isnt saturated, these models score 10/10 because they are the current best models, our /10 ranking system works by taking the ELO scores and then comparing them to the scores in the same size range. So if a better model releases that gets 10/10 and the others get lower.

And we also have the BananaMind SLM leaderboard, not the BananaMindBench leaderboard which uses ARC EASY,PIQA,Hellaswag, Arithmark 3 and the BananaMind Base Bench 1.1. This is the newer and recommended version.


Hope you understand it now!


cODeQ
  • 12 replies
·
Banaxi-Tech 
posted an update 15 days ago
view post
Post
18
What is a model?
What is it?
You don't know?
ProCreations 
posted an update 15 days ago
view post
Post
98
ICYMI:
Grug 27b v2 released! It brings increased quality, fixes repetitive loop / malformed session title issues seen in grug 27b v1.1, and reasoning efforts now truly work.
ProCreations/grug-27b-v2
  • 1 reply
·
Banaxi-Tech 
posted an update 16 days ago
view post
Post
60
Its Monday. Getting back to working on ACR 1.0.
  • 11 replies
·