RL checkpoints of Octopus-8B and baselines of paper: Learning Self-Correction in Vision–Language Models via Rollout Augmentation
Yi Ding
Tuwhy
AI & ML interests
None yet
Recent Activity
updated a collection about 12 hours ago
On-Policy Self-Adaptation updated a collection about 12 hours ago
On-Policy Self-Adaptation updated a collection about 12 hours ago
On-Policy Self-Adaptation