How reinforcement learning works
Key idea: explore and exploit
DeepSeek's GRPO
Main loop:
Challenges: rewards? multi-turn?
DeepSeek-v3-Base
DeepSeek-R1-Zero
Large-scale Reasoning-Oriented Reinforcement Learning
Training prompt
Model checkpoints under training
Solution score (reward)
Rule-based verification
AI text/layout recreation from video frame; verify against source image.