PEBS: Per-rater Empirical-Bayes Shrinkage for RLHF Reward-Model Calibration

June 25, 2026 ยท Grace Period ยท ๐Ÿ› the ICML 2026 Workshop on Pluralistic Alignment

โณ Grace Period
This paper is less than 90 days old. We give authors time to release their code before passing judgment.
Authors Arnav Raj arXiv ID 2606.27578 Category cs.LG: Machine Learning Cross-listed cs.AI Citations 0 Venue the ICML 2026 Workshop on Pluralistic Alignment
Abstract
Reward models for Reinforcement Learning from Human Feedback (RLHF) pool preferences across thousands of annotators and fit one global affine calibrator, collapsing raters with systematically different rating-scale offsets and slopes into a single average-rater fit that does not match any individual annotator. PEBS is a per-rater empirical-Bayes shrinkage estimator: it fits per-rater affine calibrators on a held-out slice of each annotator's ratings and applies Morris-James-Stein empirical-Bayes shrinkage toward the population mean, in closed form and without retraining the reward model. On PRISM, PEBS reduces within-user held-out RMSE by 8.58% over the pooled population-slope baseline. The procedure replicates on PluriHarms harm ratings (Qwen-2.5 base, in-family) with a +9.66% RMSE reduction over the same population-slope baseline. PEBS is a closed-form post-hoc estimator for annotator-specific affine calibration in RLHF reward modeling; it leaves the reward base model unchanged and estimates only the rater-level map used at inference time for new ratings.
Community shame:
Not yet rated
Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

๐Ÿ“œ Similar Papers

In the same crypt โ€” Machine Learning