Generalization Analogies: A Testbed for Generalizing AI Oversight to Hard-To-Measure Domains

November 13, 2023 Β· Entered Twilight Β· πŸ› arXiv.org

πŸ’€ TWILIGHT: Eternal Rest
Repo abandoned since publication

Repo contents: .gitignore, README.md, __pycache__, assets, configs, distribution_shifts, download_data.py, download_model_from_hf.py, examples, requirements.txt, results, setup.py, slurm_jobs, src, upload_data.py

Authors Joshua Clymer, Garrett Baker, Rohan Subramani, Sam Wang arXiv ID 2311.07723 Category cs.AI: Artificial Intelligence Cross-listed cs.CL, cs.LG Citations 8 Venue arXiv.org Repository https://github.com/Joshuaclymer/GENIES ⭐ 5 Last Checked 6 months ago
Abstract
As AI systems become more intelligent and their behavior becomes more challenging to assess, they may learn to game the flaws of human feedback instead of genuinely striving to follow instructions; however, this risk can be mitigated by controlling how LLMs generalize human feedback to situations where it is unreliable. To better understand how reward models generalize, we craft 69 distribution shifts spanning 8 categories. We find that reward models do not learn to evaluate `instruction-following' by default and instead favor personas that resemble internet text. Techniques for interpreting reward models' internal representations achieve better generalization than standard fine-tuning, but still frequently fail to distinguish instruction-following from conflated behaviors. We consolidate the 15 most challenging distribution shifts into the GENeralization analogIES (GENIES) benchmark, which we hope will enable progress toward controlling reward model generalization.
Community shame:
Not yet rated
Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

πŸ“œ Similar Papers

In the same crypt β€” Artificial Intelligence