Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
grimjim 
posted an update 4 days ago
Post
131
I think it's clear in retrospect that "frankenmerges", which repeated blocks of layers, amounted to a crude approximation of looped transformers architecture, hence them able to work at all instead of just breaking. They lucked out due to much of the signal passing through residual streams being preserved and only modulated along the way. That said, not all models are suited for this. Models which feature ever-increasing magnitudes as inference progressess through layers risk exploding precision limits.
In this post