🔗 Bridging Modal Isolation in Interleaved Thinking: Supervising Modality Transitions via Stepwise Reinforcement
-
Updated
Jun 10, 2026 - Python
🔗 Bridging Modal Isolation in Interleaved Thinking: Supervising Modality Transitions via Stepwise Reinforcement
[ECCV 2026] An official implementation of Correlation-Weighted Multi-Reward Optimization for Compositional Generation
Diffusion RL post-training on verl: Flow-GRPO on Qwen-Image, denoising-trajectory rollouts, CPU OCR reward replacing the VLM judge (6.9× headroom), budgeted for 4×5090.
Add a description, image, and links to the flow-grpo topic page so that developers can more easily learn about it.
To associate your repository with the flow-grpo topic, visit your repo's landing page and select "manage topics."