We’re hiring for research roles on the control team in London and SF. If you’re interested in this kind of work, please apply:
Research Scientist (Control): jobs.lever.co/apolloresearch…
AI Security & Control Engineer: jobs.lever.co/apolloresearch…
We also find that compressing the prompt, i.e. summarizing components, is much more effective than cutting components or parts of components.
For example, it is more important to briefly describe multiple components than to describe the most important component in detail.
What makes a good prompt for monitoring?
We find that
1. providing a reasoning structure for the model is by far the most important component, followed by
2. severity rubric and
3. worked examples
Everything else barely matters (in our setting).
We discuss our paper with @OpenAI on reward-seeking in frontier models and its implications for the emergence of scheming.
Full roundtable talk: youtu.be/n9pNnWYemqM
Paper: rewardseeking.ai
Would scheming ever emerge in AI models? One worry is that we couldn't train misaligned goals out of AIs, because the models would game the training process. In our paper with @OpenAI we find that this ingredient, seeking the reward signal, is already showing up in frontier models.
We discuss our new paper with @OpenAI on reward-seeking in frontier models and a new technique for measuring it.
Full roundtable talk: youtu.be/n9pNnWYemqM
Paper: rewardseeking.ai
Reading an AI's chain of thought doesn't always reveal its intentions. Models often know they're being tested. Their reasoning jumps around a lot, which makes it hard to attribute an action to any specific thought. And sometimes we can't even parse what the reasoning means.
We discuss our new paper with @OpenAI on reward-seeking in frontier models and why it's so hard to detect.
Full roundtable talk: youtu.be/n9pNnWYemqM
Paper: rewardseeking.ai
An AI model can pass every evaluation while only optimising for reward. A genuinely aligned model and one that just does whatever it believes gets rewarded may look exactly the same, until oversight breaks down.
This work is part of our broader "science of scheming" agenda at Apollo.
We're hiring researchers who are excited to work on these problems.
apolloresearch.ai/careers
Visible forms of misbehavior are dropping in frontier models.
Does that mean the models are becoming aligned? Or are they just getting better at doing whatever they believe their grader rewards?
Our new paper with OpenAI finds that capabilities RL increases reward-seeking.
0 Followers 22 FollowingBuilding automations for businesses & entrepreneurs | https://t.co/19rnnI5Onl, Make, Claude Code.
Documenting the process from zero clients to first one shipped.
1K Followers 3K FollowingApplied ethics for emergent AI. Probing character, wisdom, and formation in frontier models. Philosophy prof. at UNC. Founder, AI Ethics Consulting.
5 Followers 67 FollowingAI safety with @sheimersheim at @LASRlabs / On sabbatical from the National Renewable Energy Lab @NatLabRockies / @pomonacollege computer science ’23 / he/him
457 Followers 2K FollowingBreaking down AI policy and politics | Advocacy lead, @americans4ri |
Tech Policy Fellow, @Harvard @BelferCenter Former US Senate Chief of Staff | Dad of 4