Fairness-Aware Mixture-of-Experts via Subgroup Reweighting and Gate Regularization

August 24, 2026 ยท Grace Period ยท ๐Ÿ› AVSS 2026 (22nd International Conference on Advanced Visual and Signal-Based Systems)

โณ Grace Period
This paper is less than 90 days old. We give authors time to release their code before passing judgment.
Authors Sunhee Hwang arXiv ID 2608.22820 Category cs.LG: Machine Learning Cross-listed cs.AI Citations 0 Venue AVSS 2026 (22nd International Conference on Advanced Visual and Signal-Based Systems)
Abstract
Deep learning models often produce performance disparities across demographic groups, due to the training data imbalance with respect to sensitive attributes such as gender or age. To address this problem, existing work has explored fair representation learning, data re-sampling, and adversarial training, which can be broadly categorized into two main approaches. Single-stage methods typically learn a shared representation for fairness, but often struggle to handle heterogeneous subgroup distributions. Two-stage methods learn representations separately from the final prediction task, which can lead to misalignment between fairness objectives and downstream predictions. We identify routing-induced bias, a failure mode in which subgroup imbalance drives the gating network to route subgroups onto a few experts, and propose an end-to-end Mixture-of-Experts (MoE) framework that corrects it. Specifically, we apply subgroup reweighting to correct data imbalance, and introduce gate entropy regularization to prevent routing from collapsing onto subgroup attributes, keeping expert utilization both balanced and interpretable. Beyond improving fairness, the routing distribution offers an interpretable view of how subgroups are allocated across experts. Experimental results demonstrate that the proposed approach improves fairness while maintaining competitive predictive performance.
Community shame:
Not yet rated
Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

๐Ÿ“œ Similar Papers

In the same crypt โ€” Machine Learning