Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.
-
Updated
Sep 7, 2026 - Python
Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.
Official Implementation of VideoDPO
GAN-style self-improvement loop for any text artifact: mutate, grade with a SEPARATE model, keep only verified wins (pairwise-judged), revert the rest. The git history is the improvement log.
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models (ICML 2025)
code for paper Query-Dependent Prompt Evaluation and Optimization with Offline Inverse Reinforcement Learning
Policy Learning from Large Vision-Language Model Feedback Without Reward Modeling (IROS 2025)
ZYN: Zero-Shot Reward Models with Yes-No Questions
Synthetic data for fine tuning LLM
Code and data for "Timo: Towards Better Temporal Reasoning for Language Models" (COLM 2024)
distilled Self-Critique refines the outputs of a LLM with only synthetic data
Experimental language-model engine combining bidirectional masked diffusion with neuromorphic SNN dynamics, built in Julia with Python tooling.
Data Preparation for Large Language Models — a curated companion to our JCST 2026 survey. Covers Pre-training, Continual Pre-training, and Post-training (SFT/RLHF/RLAIF) across collection, filtering, dedup, generation, evaluation.
RewardAnything: Generalizable Principle-Following Reward Models
Production-ready RLAIF trading system with multi-agent Claude AI that learns from market outcomes. Features 60+ indicators, foundation models, and serverless deployment.
Code for the paper "Improving Socratic Question Generation using Data Augmentation and Preference Optimization"
Your RL second brain: 34 source-cited topics from Q-learning to GRPO and agentic RL, kept fresh monthly. Obsidian vault + AI agent layer, Brainstein SSS+.
RankPO: Rank Preference Optimization
To associate your repository with the rlaif topic, visit your repo's landing page and select "manage topics."