Question on the architecture ablation and 1M-context evaluation

#3
by drlee1 - opened

Thank you for releasing Solar Open 2 and its technical report. I found the hybrid KDA–softmax architecture, NoPE design, and support for negative eigenvalues particularly interesting.

I have a question regarding the proxy architecture ablation.

The report uses validation loss, MMLU, and HellaSwag learning curves to show that the hybrid architecture reaches a given level of general capability with fewer training tokens. These results provide useful evidence of improved pre-training sample efficiency.

However, the report also states that the central architectural objective is to provide a usable context window beyond one million tokens and support long-horizon agentic tasks. The motivations for NoPE, hybrid attention, and negative eigenvalues are closely related to long-context extrapolation, information retention, and state tracking.

I recognize that the final model is evaluated on AA-LCR and that multiple metrics are tracked during the length-expansion stage. However, those results reflect the completed model after pre-training, length expansion, checkpoint merging, and post-training, rather than isolating the architectural contribution itself.

Were controlled proxy ablations conducted between the all-softmax baseline and the hybrid architecture on length-dependent long-context tasks, such as retrieval, aggregation, information overwrite, or state tracking? If such experiments were omitted only at the report level but actually conducted, would it be possible to share those results?

My concern is not that MMLU and HellaSwag are unsuitable as low-cost proxy signals. Rather, they demonstrate general sample efficiency but do not directly establish whether the hybrid architecture improves the long-context capabilities it was primarily designed to support. If such controlled long-context experiments were conducted, it would be valuable to see the results.

Sign up or log in to comment