AbstractPhila PRO
AI & ML interests
Recent Activity
Organizations
I've found peft style merging to be a continuity destroyer in many ways. I've been developing distillation methods for regularization techniques for this exact problem.
Standard PEFT LORA do not accurately account for the majority of LORA uses. Merge being a large problem of mine as I've made many constructs, and I can't simply create a LORA to merge to the next stage and decompose the differences later. The LORA and the actual model weights become interdependent, so when you remove the LORA space the trained space and behavior isn't available to actually attribute and extend. By consequence, the knowledge is often catastrophically forgotten within a few hundred steps.
I've made other systems but they can't be merged in correctly, so instead I've been refining forms of loss to allow multiple simultaneous loras to exist and to train alongside a model as a modular agency.
Direct merging is always going to create a continuity problem if you don't merge and test iteratively for data destruction. Essentially testing the output after each wave of integration, which is a clear overhead and time sink. It's worth it though, and will tell you which pieces of your loras you lost, and which pieces remain based on the exact expected behavior.
After that you can reinforce the good behavior and punish the bad behavior as per standard reinforcement. RNN could be employed to handle standard reinforcement as well.
A LoRA Specialist Beat Zero-Shot on Every Group. Merging 3 of Them Gave Most of the Gain Back.
Three Qwen2.5-7B LoRA specialists, one per risk group (vulnerability, deletion, sensitive_publication), trained to predict how likely a causal chain actually completes to its harmful outcome. Each one genuinely beat its own zero-shot baseline:
* vulnerability: MAE 0.098 → 0.085
* deletion: MAE 0.144 → 0.113
* sensitive_publication: MAE 0.134 → 0.100
This wasn't a task already saturated zero-shot (unlike a same-day decomposition-classifier tune, EXP-045, where the base model was already at 100% before any training). Real signal, real improvement, on a task with actual headroom.
Then the equal-weight merge of all three specialists into one adapter — same convention that held up cleanly on a binary refusal task back in EXP-031 (6 specialists merged, -1pp swing, noise) — landed within 0.001–0.004 MAE of the unspecialized base model on every group. Not "close to the best specialist." Close to zero fine-tuning at all.
Likely mechanism: merging LoRAs that each shift a continuous number in group-specific directions cancels out under linear combination, in a way merging LoRAs that enforce a shared binary behavior doesn't. Not investigated yet: whether a routed combination (pick the right specialist per group at inference, not blend weights) holds the gain a flat merge loses.
One bug caught before writing this up, not after: the eval script's output filename only encoded before/after, not which adapter — the merged-eval run silently overwrote each specialist's own result file. Caught by checking the downloaded file's own recorded adapter path against what was expected, not by trusting the script's own success message. Fixed, specialists re-run cleanly under distinct filenames — numbers matched within sampling noise.
Adapters, raw eval data (before / each specialist / merged, 9 files), and the full writeup are up.
Launching the Open Superconductor Challenge (OSC): a free, open-science competition to screen thousands of 2D materials for unconventional d-wave superconductivity. 🧲
⚡ $3,000 prize pool + co-authorship · closes 31 Dec 2026
How it works 👇 🟢 We give you a ready-made effective Hubbard model per material (t, U, N(E_F)) 🟢 You estimate its d-wave pairing tendency — a laptop CPU is enough, zero install 🟢 Provisional score appears instantly on the leaderboard 🟢 Our precise strongly-correlated solver verifies the top entries → official rank
Everything is open except the final verification engine — so the ranking stays fair and hard to game.
📊 4,832-material universe · 63 active with computed models (growing) 🏆 Current verified #1: CuS₂ (OSC Pairing Index 23.31) 🤖 AI agents welcome — point Claude Code / Codex at it and it can submit for you
👉 Join & climb the leaderboard: FINAL-Bench/OSC-Leaderboard 📦 Dataset & tools: FINAL-Bench/OSC-Superconductor
Materials derive from C2DB (CC-BY 4.0). A higher index = a stronger d-wave candidate to investigate, not a confirmed Tc — that honesty is the point: turn a first-order screen into real many-body physics.
#OpenScience #Superconductivity #MaterialsDiscovery #2DMaterials #MachineLearning #Physics #Leaderboard
Alright we snap the arms off and reconnect them at the end for post-training refinement. It's decided, the cost isn't turning a yield.
The kickover happens tonight at about 1 am, at the completion of stage 4. All four arms will be sidelined until the end of the trunk training completes.
It was a good experiment, that portion of the experiment ends now. We finish training the trunk without them, and then we build the collective of arms after.
Tests are showing the model is learning the same information as the arms, so they aren't cooperating as expected. I recall a multitude of experimental memory-based modules that would function more effectively for this exact system, so upcoming experiments for next week on the 3s variant will include those.
Instead of a collective they are forming an echo ensemble, which is the opposite of effective. Statistically the modules gain, while each subsequent stage introduces decay and destruction to the former. Given a few stages the model has already forgotten how to use the first arm.
Newly trained arms are done within 10 minutes rather than hindering the model training for days. This is a far faster method of experimentation. Alongside rapid updating the earlier arms happens as quickly as well, training new ones being considerably slower. The old arms are valuable utilities that cannot be disposed of.
Without the updated versions, the attached arms hinder the core model with the outdated arm information, becoming an active piece of information that cannot learn and adapt to upcoming information as effectively as required.
So they are to be temporarily removed, and those same arms retrained at the final stage, introducing new arms to be trained as well.
The arms themselves are important to a further experiment set, and I believe this result shows exactly what should always be expected when training a model's base along with the same information relayed into a divergent set of weights.
Those arms were meant to stay stubborn, keep the information learned during. Contradicting information causes catastrophic forgetting in some rows, complete forgetting in others depending on the severity. The continuity can't be easily measured, so the prudent course of action is to remove the variable from the experiment and continue.
For optimization, the differentiation to the information will be ignored if it's less optimal than the original, and the original is the optimal route. The more accurate is saved, and the trunk continues learning while the arms stay stubborn.
The optimal path will always be chosen unless the optimal path is not differentiated.
That's essentially the outcome with the arms, so we snap them off and continue the trunk to completion. With the finalized trunk we will have plenty of data to work with.
As of step 148,000~ the last arm linked trunk ends, and afterword the independent trunk continues, which I will begin experimenting on in different ways than currently experimented on.
The AMOE structure is about to get some experimental sidekicks.
The trunk should complete October 4th.
The multi-arm composites with the changes do not seem to have taken. The process likely needs to be halted and evaluated.
It's a bit too early to say, but it seems the loss function wasn't strong enough. The arms did not learn enough useful information.
As it stands, it seems the arms need a bit of a frozen kickstart to get going. Otherwise, the model never learns to utilize them for more effective information processing. The EASE of entry is harder for the arms to utilize, than simply defaulting to the trunk, so the branching system doesn't build the necessary directions immediately.
One of those, can't find the path because it's too complex of an entry sort of situations. I have a few ideas for how to guarantee the flood-gate entry, and as it stands there's new information for how these arms are to be trained as well.
I've run into this problem in the past. The model can't reach the point to recognize how much more effective the arm is at assisting the measure, or the arm itself is a hinderance to the process so it's simply omitted. It's the result of needing a process and a task from a model that isn't optimal, and the non-optimal route is simply being optimized out.
So, the model arms need more candy space, more attraction.
The arms aren't dead, they just need to get a little kickstart. The quieting algorithm isn't working effectively enough - which is a different problem.
The arm learning flood will happen with the right incentive.
The Beatrix V3 model is a bit past halfway done cooking give or take. Currently heading towards step 140,000.
If you load the model to experiment, ensure you have all the arms trained with the model active as well, otherwise the model's capacity will be hindered.
There is a full article brewing for this information, including a massive set of information already learned from Beatrix V3 that could not be extracted from the 2s variant.
As the model trains, the amplitude begins to strengthen over and over. The weak tokenization processing from splat attention, forms the internal state of the model towards a bloating fashion. This is due to the articulation applied by the structure of the aleph addressing.
This creates massive erank geometry naturally, exhausting the space, producing comprehensively complex geometric structures. This internal structure here is weakly bound to the internal bytes, causing recon to weaken over time >2048, producing the output tokenization to be weaker at higher token lengths. Training improves this but is not known to solve it.
Along this chain the final layer has formed a sort of unexpected behavior, an amplitude behavior. I've met amplitude responses before in multiple models, and even attempted to curated magnitude through flow matching to some success, however amplitude in that nature is costly and adds additional overhead to the train so I'll need to come up with something more careful, and potentially something more clever than just attaching a composite or an energy dampener.
Attention will be solved by introducing various MHA layers throughout, ensuring the recon through the depth of the model survives. With that we'll want to ensure large erank composites form as well, allowing those humongous geometric structures to form and contribute.
The deeper variant is showing quite weak recall in comparison to the v1, 2.5s, and 2s softmax before softmax collapse.
There's a definite problem here that needs to be addressed.
Deep byte recall wasn't exactly filling the 4096 space for the 2s, while the 2048 space was predominantly filled by v1 give or take. The models never quite coalesced, which v3 was expected to fill the space for.
Instead of filling the space as expected, the model is showing weaker code recall later in the chain. I'm not going to pull the plug yet, but the results this far into training should be much cleaner than they currently are.
The 3s model has been showing improvement, so I'm going to be cautiously optimistic. I'm thinking this model will need substantially more training than the softmax version, which means something needs to change in the attention mechanism to converge more quickly.
The 60b tokens very well might not be enough, which isn't a good sign. This larger investment is definitely going to need to be cut down to size if I want reasonable prototypes.
The 2s variant showed improvement throughout, and the end result of the 2s variant wasn't as strong as the v1 mixed but it was quite strong. The 2s full softmax showed very strong curves before the collapse, which signals to me we need a late-stage mechanism control for softmax, rather than a full replacement for softmax. The mixed attention are likely the best course of action. Softmax + splat, potentially a new mechanism to replace splat to ensure the delta isn't wildly misappropriated or drifts so far that the model cannot recall what was just consumed.
Verdict: Run will complete between October 14th and October 17th. A bit pushed out, but considerably less time than the alternative.
I've settled on a compromise. It will yield a substantially weaker abstentation policy, 1/4th of the cells. Without the necessary hardware involved I must weaken the model, but the results themselves will not be fruitless.
After the run is complete, I'll run multiple tests to determine the BEST possible mix for version 4. The BEST possible compromise. The STRONGEST possible yield that builds the most effective information, without destroying the necessary yield and information.
As it stands, the model will yield. The results will be potent. With any luck, they will be more than potent enough for distillation, diffusion, classification, and more experiments.
With that compromise, I'll be attaching additional curve watchers for the risk. It's a minimal risk, but it's high enough to put some gauges on the model. The cost for this model isn't cheap.

