Say hello to the AlephLLM: Mini-Beatrix - in her huggingface space! She is currently stepped at 8000 steps in the first couple datasets, so she's not very smart yet. AbstractPhil/alephllm-chat
Be warned, whatever you say will be recorded in a public cache.
https://github.com/AbstractEyes/alephllm Here's the model code and training code for the prototype. As the training progresses, the AlephLLM will become more coherent and communicative, the tensorboard will consist of a large series of useful and useless analysis, and each checkpoint recorded at around 2000 steps unless the train crashes or the system faults.
It will take about 9 hours for the first few datasets to converge, then I'll train a chat AMOE expert cluster to see if she wants to speak yet. Until then, she's learning.
Yes I know it's early, but there isn't much more I could think of to analyze the AlephLM directly currently. The only way train the AlephLLM, is to train the full AlephLLM prototype. The bigger training has to run, otherwise the analysis won't matter. As it progresses, the analysis and huge amount of tensorboard statistics will flood out. Everything is transparent through the process from start to finish, everything recorded.