Snorkel-Mistral-PairRM-DPO
Snorkel AI · 2024-01-19 · 7.2B parameters
Snorkel-Mistral-PairRM-DPO is Snorkel AI's 7.24-billion-parameter instruction-following model, released in January 2024 as a demonstration of iterative preference tuning. Starting from Mistral-7B-Instruct-v0.2, Snorkel generated candidate responses, ranked them with the off-the-shelf PairRM reward model, and repeated Direct Preference Optimization for three iterations; this public experiment did not use task-specific programmatic labels.
Benchmark scores
| Benchmark | Score |
|---|---|
| EQ-Bench v2 | 65.8 |
| EQ-Bench v2 + MAGI-Hard Combined | 51.7 |
| MAGI-Hard | 37.5 |