AbstractPhil
·
AI & ML interests
datasets, research papers, experimentation, vision, classification, text encoders, tokenization, llms, diffusion, distillation, and more.
Recent Activity
repliedto their post about 12 hours ago Say hello to the AlephLLM: Mini-Beatrix - in her huggingface space!
She is currently stepped at 8000 steps in the first couple datasets, so she's not very smart yet.
https://huggingface.co/spaces/AbstractPhil/alephllm-chat
Be warned, whatever you say WILL be recorded in a public cache, WHEN the chat version works. For now she records nothing. The idea is to help debug the K/V cache, and I would rather the data accumulated be shared. If you wish to speak to her in private I will include a toggle, that way you'll see that nothing is recorded when you speak and you can still have a private chat with her. For now she's simply auto-completing, so have fun with her.
The AlephLLM prototype is currently in full training with SDPA attention.
https://huggingface.co/AbstractPhil/alephllm-mini-beatrix-training
https://github.com/AbstractEyes/alephllm Here's the model code and training code for the prototype.
As the training progresses, the AlephLLM will become more coherent and communicative, the tensorboard will consist of a large series of useful and useless analysis, and each checkpoint recorded at around 2000 steps unless the train crashes or the system faults.
It will take about 9 hours for the first few datasets to converge, then I'll train a chat AMOE expert cluster to see if she wants to speak yet. Until then, she's learning.
Yes I know it's early, but there isn't much more I could think of to analyze the AlephLM directly currently. The only way train the AlephLLM, is to train the full AlephLLM prototype. The bigger training has to run, otherwise the analysis won't matter. As it progresses, the analysis and huge amount of tensorboard statistics will flood out. Everything is transparent through the process from start to finish, everything recorded. posted an update about 17 hours ago Say hello to the AlephLLM: Mini-Beatrix - in her huggingface space!
She is currently stepped at 8000 steps in the first couple datasets, so she's not very smart yet.
https://huggingface.co/spaces/AbstractPhil/alephllm-chat
Be warned, whatever you say WILL be recorded in a public cache, WHEN the chat version works. For now she records nothing. The idea is to help debug the K/V cache, and I would rather the data accumulated be shared. If you wish to speak to her in private I will include a toggle, that way you'll see that nothing is recorded when you speak and you can still have a private chat with her. For now she's simply auto-completing, so have fun with her.
The AlephLLM prototype is currently in full training with SDPA attention.
https://huggingface.co/AbstractPhil/alephllm-mini-beatrix-training
https://github.com/AbstractEyes/alephllm Here's the model code and training code for the prototype.
As the training progresses, the AlephLLM will become more coherent and communicative, the tensorboard will consist of a large series of useful and useless analysis, and each checkpoint recorded at around 2000 steps unless the train crashes or the system faults.
It will take about 9 hours for the first few datasets to converge, then I'll train a chat AMOE expert cluster to see if she wants to speak yet. Until then, she's learning.
Yes I know it's early, but there isn't much more I could think of to analyze the AlephLM directly currently. The only way train the AlephLLM, is to train the full AlephLLM prototype. The bigger training has to run, otherwise the analysis won't matter. As it progresses, the analysis and huge amount of tensorboard statistics will flood out. Everything is transparent through the process from start to finish, everything recorded. View all activity Organizations
view article Agreement, Anchors, Addresses: A Week of Geometric Training
AbstractPhil
• view article Geometric Memory FT4 — Distill Against a Consensus, Ship a Rotation
AbstractPhil
• • 1
view article The Loss Manifest: A Field History of Objective Functions, and What a Machine Can Actually Be Asked to Compute
AbstractPhil
• view article Aleph Differentiation, Parts 3 & 3-D: Two Laws, Five Days, One Framework
AbstractPhil
• published an article about 1 month ago view article The Aleph Moves Into a Pretrained Trunk: Relays, Registers, and the Two-Regime Dispatch Law
published an article about 1 month ago view article The Aleph Under Autoregressive Pressure: Bottleneck Priors, Sign Codes, and the Consumption Law
published an article about 2 months ago view article Subject Bucketing: Teaching a Diffusion Model New Prompt Languages Without Forgetting
AbstractPhil
• • 1
view article geolip-aleph-void: The First Relational Geometric Vocabulary Patchwork
view article Reading the Voids: Topological Contribution Signals in Frozen Geometric Codebooks
view article Fused Batched Thin SVD, Part II: Extending the Jacobi Pipeline to N=6 with Configurable Convergence
view article H2 Omega Confirmed, Paradigm Shift: Attempting to Disprove Omega As A Whole
AbstractPhil
• • 1
view article The Polygonal Omega: Trained Sphere-Solvers Are Projective Codebooks
AbstractPhil
• • 1
view article Three Geometric Bands in a Sphere-Normalized Patch Autoencoder
view article The Geometric Engine: Structural Attractors in Neural Network Weight Space
view article FL Hybrid Eigendecomposition Beating cuSOLVER's Mathematical Purity with Compilable PyTorch
view article Ryan Spearman: Geometric Variant Effect Prediction Through Quaternion-Composed Dual Expert Alignment
view article Fused Batched Thin SVD: Engineering a 5000× Speedup with Triton Kernels
view article A geometric encoder's toolkit: deterministic primitives for hyperspherical image encoding