Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
eoe
eoe
7
4
16
Follow
webxos's profile picture
21world's profile picture
2 followers
·
84 following
AI & ML interests
None yet
Recent Activity
liked
a model
4 days ago
Qwen/Qwen3.8-Flash-Next
upvoted
a
collection
4 months ago
2026 April 🐝 China Open Source Highlights
reacted
to
anakin87
's
post
with ❤️
4 months ago
How LLM training with RL Environments works? It all starts with 𝗥𝗲𝗶𝗻𝗳𝗼𝗿𝗰𝗲𝗺𝗲𝗻𝘁 𝗟𝗲𝗮𝗿𝗻𝗶𝗻𝗴 𝘄𝗶𝘁𝗵 𝗩𝗲𝗿𝗶𝗳𝗶𝗮𝗯𝗹𝗲 𝗥𝗲𝘄𝗮𝗿𝗱𝘀 - question asked - model generates reasoning + answer - answer checked against ground truth - reward drives RL training In this setup, the environment is simple: fixed questions and answers, rollout logic, reward(s) Consider a more complex tic-tac-toe env ❌⭕ It adds: - dynamic game generation/handling - tunable opponent skill - multi-turn interactions (envs can also include tools) --- What happens at training? We use 𝗚𝗿𝗼𝘂𝗽 𝗥𝗲𝗹𝗮𝘁𝗶𝘃𝗲 𝗣𝗼𝗹𝗶𝗰𝘆 𝗢𝗽𝘁𝗶𝗺𝗶𝘇𝗮𝘁𝗶𝗼𝗻 with a tic-tac-toe env No critic model needed, the group is the baseline Simpler than PPO 1️⃣ Rollout generation: from the same board, model plays N games via sampling 2️⃣ Each game scored with deterministic rewards (win, format, ...) 3️⃣ Mean score computed across the group 4️⃣ Each rollout's advantage = its score minus the group mean 5️⃣ Model updated to favor trajectories above baseline 🔁 Repeat For a deep dive, check out 🌱 https://github.com/anakin87/llm-rl-environments-lil-course a free hands-on course on RL environments for LLMs
View all activity
Organizations
None yet
eoe
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Upvotes
Likes
Articles
liked
a model
4 days ago
Qwen/Qwen3.8-Flash-Next
Image-Text-to-Text
•
180B
•
Updated
3 days ago
•
52.3k
•
4.3k
liked
a Space
7 months ago
Running
21
2025 AI Timeline
📈
21
liked
a Space
8 months ago
Running
Featured
46
2025 China AI Timeline
📈
46
Explore 2025 Chinese AI model releases in a timeline
liked
a model
9 months ago
zai-org/AutoGLM-Phone-9B
Image-Text-to-Text
•
934k
•
Updated
Jan 7
•
28.4k
•
443
liked
2 models
about 1 year ago
NexaAI/OmniNeural-4B
Any-to-Any
•
Updated
Nov 7, 2025
•
565
•
167
Falconsai/intent_classification
Text Classification
•
67M
•
Updated
Dec 9, 2023
•
109
•
58
liked
3 models
over 1 year ago
deepseek-ai/deepseek-moe-16b-base
Text Generation
•
16B
•
Updated
Jan 12, 2024
•
15.6k
•
154
squeeze-ai-lab/TinyAgent-7B
Text Generation
•
7B
•
Updated
May 30, 2024
•
13
•
5
OS-Copilot/OS-Genesis-4B-AC
Image-Text-to-Text
•
4B
•
Updated
Jan 8, 2025
•
13
•
7
liked
a model
almost 2 years ago
zai-org/glm-edge-1.5b-chat
Text Generation
•
2B
•
Updated
Nov 28, 2024
•
2.45k
•
20
liked
6 models
over 2 years ago
stabilityai/sdxl-vae
83.7M
•
Updated
Aug 4, 2023
•
243k
•
772
qualcomm/Llama-v2-7B-Chat
Other
•
Updated
Jun 16
•
26
apple/OpenELM
Updated
May 2, 2024
•
1.45k
NexaAI/Octopus-v2
Text Generation
•
3B
•
Updated
May 21, 2024
•
391
•
•
891
nota-ai/coreml-bk-sdm
Text-to-Image
•
Updated
Nov 17, 2023
•
7
nota-ai/bk-sdm-tiny-2m
Text-to-Image
•
0.3B
•
Updated
Nov 17, 2023
•
121
•
19