Checkpoints (bf16, 50-step) + train/eval data for RLTR (transfer-reward RL) incl. cross-family (gemma) receiver.
HyunseokLee
hyunseoki
AI & ML interests
None yet
Recent Activity
updated a collection about 11 hours ago
Meta-Harness v1b: Policy x Harness Co-Evolution updated a model about 11 hours ago
hyunseoki/mh-v1b-imo-9b-t20 published a model about 11 hours ago
hyunseoki/mh-v1b-imo-9b-t20