Contextual Bandits for Maximizing Stimulated Word-of-Mouth Rewards

June 13, 2026 ยท Grace Period ยท ๐Ÿ› the AAAI 2025 Workshop on Bridging the Gap Between AI Planning and Reinforcement Learning

โณ Grace Period
This paper is less than 90 days old. We give authors time to release their code before passing judgment.
Authors Ahmed Sayeed Faruk, Elena Zheleva arXiv ID 2606.15146 Category cs.LG: Machine Learning Citations 0 Venue the AAAI 2025 Workshop on Bridging the Gap Between AI Planning and Reinforcement Learning
Abstract
Stimulated word-of-mouth is a strategy that promotes information sharing through prompts or incentives. Optimizing stimulated word-of-mouth through social networks requires identifying and targeting connected users who are most susceptible to spillover, a phenomenon where the influence of recommendations extends beyond the immediate audience to impact their connected users. The probability of spillover varies across individuals, and their connections, leading to heterogeneity. Understanding and accurately estimating the spillover probabilities among users in social networks is crucial for improving the effectiveness of stimulated word-of-mouth. To address this, we present a novel contextual multi-armed bandit framework that learns individual spillover probabilities and ranks connected users to maximize rewards from stimulated word-of-mouth. Experiments on real-world network datasets demonstrate that accounting for spillover heterogeneity enhances the targeting precision of top-$k$ connected users, boosting rewards and outperforming baseline methods that do not learn individual spillover effects.
Community shame:
Not yet rated
Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

๐Ÿ“œ Similar Papers

In the same crypt โ€” Machine Learning