Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
AbstractPhil 
posted an update 3 days ago
Post
657
The upcoming AlephLM LLM prototype "Mini-Beatrix" is based on protocols, rules, and laws established through the process of training AlephLM systems. This will be a first attempt at a smaller full pretrain/finetune of the AlephLLM on raw data, and this will require over a billion unigram tokens.

Mini-Beatrix will inherit an appropriately adapted AlephLM MOE structure containing a multitude of trained experts, a gating system, a long context RoPE system, MHA attention, and a series of hypothesis to answer upon Mini-Beatrix's pretrain and finetune completion.

While focusing on resolving corruptions and invalidity possibly present in the splat attention, the solutions raised SDPA attention protocol token recall ceiling from 0.91 to 0.993. With that the splat attention raised from 0.81 to 0.89~ splat being around 3x the speed is still imperfect.

So far so good. The corruptions have resolved multiple core component overlapping problems causing the AlephLM's inability to handle the trigram system, the structure of the SVAE having faulty trigram structures, and additionally a multitude of other systems in the lineup that were inheriting the corruptions from the core experiment sets.

These corruptions resolved show that the accuracy of standard multiheaded attention will provide the necessary token recall for full LM capacity, and with that If and WHEN I solve the Rorschach Splat attention will be the faster alternative at >=r1 0.99%, only then. The splat attention's considerably larger head count still contains unresolved inconsistencies.

That being said the SDPA MHA attention will be present for the first attempted mini-llm train, which will be named "Mini-Beatrix" with the appropriate sizing associated with this.


The only thing that will change Mini-Beatrix's trajectory will be if Splat attention is perfected between today and next week, which will likely take longer unless I run into a core corruption that has been overlooked through hundreds of analysis.

Alright I've begun training the prototype v1. The structure is holding together, the system aligning, and the subsystem is converging.
It works.
image
The smaller version is aligning and working. The erank is higher than expected, the system more confined, and some interesting effects already starting to emerge.

Huggingface spaces will have an application to speak with her up within the hour. She isn't chat trained yet, but will be very soon.

Geometric LLM.

In this post