Secrets of RLHF in Large Language Models Part I: PPO
paper page: huggingface.co/papers/2307.04…
Large language models (LLMs) have formulated a blueprint for the advancement of artificial general intelligence. Its primary objective is to function as a human-centric (helpful, honest, and harmless) assistant. Alignment with humans assumes paramount significance, and reinforcement learning with human feedback (RLHF) emerges as the pivotal technological paradigm underpinning this pursuit. Current technical routes usually include reward models to measure human preferences, Proximal Policy Optimization (PPO) to optimize policy model outputs, and process supervision to improve step-by-step reasoning capabilities. However, due to the challenges of reward design, environment interaction, and agent training, coupled with huge trial and error cost of large language models, there is a significant barrier for AI researchers to motivate the development of technical alignment and safe landing of LLMs. The stable training of RLHF has still been a puzzle. In the first report, we dissect the framework of RLHF, re-evaluate the inner workings of PPO, and explore how the parts comprising PPO algorithms impact policy agent training. We identify policy constraints being the key factor for the effective implementation of the PPO algorithm. Therefore, we explore the PPO-max, an advanced version of PPO algorithm, to efficiently improve the training stability of the policy model. Based on our main results, we perform a comprehensive analysis of RLHF abilities compared with SFT models and ChatGPT. The absence of open-source implementations has posed significant challenges to the investigation of LLMs alignment. Therefore, we are eager to release technical reports, reward models and PPO codes
1K Followers 4 FollowingNew Computer Vision and Pattern Recognition papers from https://t.co/NvCyo8SqfI: image processing. Thank you to arXiv for use of its open access interoperability.
1.6M Followers 2 FollowingWe're an AI safety and research company that builds reliable, interpretable, and steerable AI systems. Talk to our AI assistant @claudeai on https://t.co/FhDI3KQh0n.
3K Followers 1K FollowingChasing the horizon of tomorrow's ✨ | Gaming, Business, AI, Apps & Open Source | Tech Expert & CEO | Linux soul! | DMs open for collabs and opportunities.
4K Followers 2K FollowingPhD in AI | GDE in AI/ML | CTO Intento | Author "Deep Learning with JAX"
📝 ML insights: https://t.co/ySSOXJKL7H
🤖 Daily AI paper reviews: https://t.co/yQNYyqTbBR
753K Followers 161 FollowingAuthor Trade Like A Stock Market Wizard and Think & Trade Like a Champion. Featured in Stock Market Wizard by Jack Schwager. Before following read disclosure.
297K Followers 8K FollowingFounder and CEO @tabul_ai. Creator of @trainxgb. ML ex Nvidia. Data Scientist. Physicist. Catholic. Husband. Father. Stanford Alum. Memelord. e/xgb. AMDG.
515K Followers 3K FollowingNVIDIA Director of Robotics & Distinguished Scientist. Co-Lead of GEAR lab. Solving Physical AGI, one motor at a time. Stanford Ph.D. OpenAI's 1st intern.
486K Followers 1K FollowingML/AI research engineer. Ex stats professor.
Author of "Build a Large Language Model From Scratch" (https://t.co/O8LAAMRzzW) & reasoning (https://t.co/5TueQKx2Fk)
23K Followers 672 FollowingTesla Shareholder & Owner l SpaceX investor l Official Toss Securities Influencer l First Korean to get Korean reply from Elon Musk | X premium +
1.7M Followers 2 FollowingClaude is an AI assistant built by @anthropicai to be safe, accurate, and secure. Talk to Claude on https://t.co/ZhTwG8dz3D or download the app.
231K Followers 813 FollowingCooking fun AI systems & products @databricks. Prev: co-founder & CTO @ Hyperbolic, OctoAI (acquired by @nvidia) Apache TVM, PhD @ University of Washington.
17K Followers 546 FollowingFounder and CEO of https://t.co/5MjtfpwEU3 | Your guide to radiance fields | Host of the podcast @ViewDependent | FTP: 279 | discord: https://t.co/lrl64WGvlD