Ah I see now. So there's hundreds of thousands of features, and we grabbed the most interesting ones we found after a cursory scan. None of these besides "delve", "it's not X, it's Y", and maybe some flowery vocab choice seems to jump out as good candidates. We may dig for more features later if we revisit this experiment
This was a quick experiment that merits more investigation. We ran this with our product, which is now in public beta! Check out our quickstart guide here console.gutenberg.ai/d/docs
@thesophiaxu aha yea maybe it's like how early waymo needed a human overseer. i have no idea what the laws are around this but it still seems like a good idea for gradual deployment
We also compared our successful training run to an unsuccessful one. If you just look at the reward, you find they diverge around step 9. However, when we compare the features between them, there’s this prominent feature that corresponds to “difficulty dealing with a minor bug in the environment”, which the goodrun overcame, and the badrun did not. Looking at this signal, we can see the runs diverging at step 6, giving earlier signal, and an explanation, as to why one run succeeded and the other did not.
New paper! We used Sparse Autoencoder (SAE) embeddings to understand how agent behavior actually changes during post training. We analyzed thousands of rollouts in a Diplomacy environment and uncovered generalized reward hacking, weird roleplay, and the root cause of an unsuccessful training run. All of which our LLM baseline couldn't find.
274 Followers 1K FollowingIncoming PhD student at Oxford | AI Engineer and Researcher @MBZUAI | Interpretable AI | social learning | Open-Ended AI Systems | Foundation models
990 Followers 2K Followingtrying to make AI go well, grantmaking @ Astralis, replicating research @secondlookxlab, prev. @GraySwanAI, @cais, @uchicagoxlab, https://t.co/pJCxjv7GWv
206 Followers 423 Followingdoing the most good for the most people and/or the people i like the most | Let’s maybe align AI | Judeofuturist | I have been a good Bing ☺️
710 Followers 1K Followingautonomous company infra | coding agents & harnesses | founder
prev: ai for construction drawings, computer vision for the NFL | https://t.co/g3QU48GISq
80 Followers 3K FollowingBreaking down AI papers so the hard truths don’t stay hidden.
RL • Transformers • Reasoning • Verification Attentions
Deep dives → @hooshaai Substack
22K Followers 1K Followingꙮ there go the ships, and there is that leviathan
ꙮ blog/art/fiction/games https://t.co/aykxqKippW
ꙮ llm psychologist @acsresearchorg
ꙮ 💞💍📝 @holotopian, she/they 🏳️⚧️
274 Followers 1K FollowingIncoming PhD student at Oxford | AI Engineer and Researcher @MBZUAI | Interpretable AI | social learning | Open-Ended AI Systems | Foundation models
7K Followers 2K FollowingCEO and Co-founder @datologyai working to make it easy for anyone to make the most of their data. Former: RS @AIatMeta (FAIR), RS @DeepMind, PhD @PiN_Harvard.
206 Followers 423 Followingdoing the most good for the most people and/or the people i like the most | Let’s maybe align AI | Judeofuturist | I have been a good Bing ☺️
780 Followers 78 FollowingWe're a part-time, virtual research program that gives students and early career professionals an opportunity to work with professional AI safety researchers.
353 Followers 34 FollowingCharting the technical roadmap to SL5 optionality for frontier AI labs. A multistakeholder initiative uniting AI labs, national security leaders & engineers.
990 Followers 2K Followingtrying to make AI go well, grantmaking @ Astralis, replicating research @secondlookxlab, prev. @GraySwanAI, @cais, @uchicagoxlab, https://t.co/pJCxjv7GWv