leffff/Diffusion-Reward-Modeling-for-Text-Rendering-Dataset Viewer • Updated Aug 12, 2025 • 14k • 93 • 9
imflash217/proximal_policy_optimization_huggy_unity Reinforcement Learning • Updated Jan 14, 2023 • 87 • 2
MRNH/proximal-policy-optimization-LunarLander-v2 Reinforcement Learning • Updated Aug 10, 2023 • 17 • 1
andersonbcdefg/sharegpt_reward_modeling_pairwise_no_as_an_ai Viewer • Updated Jun 6, 2023 • 11.8k • 120 • 3
crumb/bloom-560m-RLHF-SD2-prompter-aesthetic Text Generation • 0.6B • Updated Mar 19, 2023 • 277 • 26