🌾 OAT: A research-friendly framework for LLM online alignment, including reinforcement learning, preference learning, etc.
-
Updated
Jan 29, 2026 - Python
🌾 OAT: A research-friendly framework for LLM online alignment, including reinforcement learning, preference learning, etc.
Robust contextual dueling bandits with post-serving context, delayed feedback, and adversarial corruption (RLHF / preference learning) — ICML 2026
A python library for (finite) Partial Monitoring algorithms
Add a description, image, and links to the dueling-bandits topic page so that developers can more easily learn about it.
To associate your repository with the dueling-bandits topic, visit your repo's landing page and select "manage topics."