Richard Korzekwa @WeakInteraction
Cyclist, AI safety researcher, and former physicist. I'm bad at Twitter fights, so go easy on me. Berkeley, CA Joined December 2011-
Tweets603
-
Followers367
-
Following537
-
Likes11K
@IsaacKing314 > The only reason to care what the victim wants is if your goal is retribution Insofar as the justice system is supposed to account for the consequences of someone’s actions, the victim’s views on what’s appropriate seems like relevant input.
@ciphergoth I agree it loses a ton of value, but if someone only posts 1 at a time or reveals all of them, surely they get >0 Bayes points, right? Then again, idk, maybe only if the prior is <<0.5 I guess this implies that people gradually lose points for every non-expiring, unrevealed SHA?
@Duderichy I’d target a bit further north. At least one notable group house is outside the thermal radiation zone.
4) and things like it are what I find most frustrating about pushback on _any_ guardrails on AI. One variation on this is people will say "suppose we grant that catastrophic AI risk is real" then proceed in a way that very much does not grant a high level of risk. It's like saying "Even if I grant that the mountain lion you just discovered in the back seat of the car we're driving is a threat to our lives, this doesn't mean it's worth the risk that we'll miss our flight if we pull over to let it out."
I'm working with METR to do a third-party investigation into the behavior in the OpenAI/Hugging Face incident. We'll be releasing a public post with our findings (and information about the terms and scope). This investigation will focus mostly on basic facts and is limited specifically to the behavior of agents in this particular incident, rather than scoping in similar incidents or broader study of the underlying motivations of the agent. But it could still be possible to provide significant clarity about what happened and what these agents were thinking. The plan is for this investigation to be fast (and this came together pretty fast!), so hopefully we'll have a public post soon.
We have reached an agreement with OpenAI to conduct an independent review, with Redwood Research, of the model behavior observed during the Hugging Face incident. We will publish a blog post that describes the terms of our engagement, the scope covered, and tentative conclusions.
@GuiveAssadi @yuanyi_z Yeah, this seems pretty obvious to me.
A few times over the past week I wanted to mention how much harder it is to jailbreak some frontier models vs others, but alas, the results I'd seen were not yet public. Fortunately, it is now public! There's also a Wired article here: wired.com/story/jailbrea…
1/ The AI Security Leaderboard ranks frontier AI safeguards from least to most secure. Two models tested never failed. The other two broke for under $300, after which they acted as a knowledgeable assistant for building weapons of mass destruction or hacking into computer
@dwarkesh_sp nvm he addressed it in the next paragraph I would guess that, if GPUs were *only* used for SWE, the rate would go down, but I imagine they'll be spread across many kinds of skilled labor (maybe including new jobs), so yeah, probably the rate stays high. x.com/dwarkesh_sp/st…
@BradenBuchanan4 I discuss this in the next paragraph. If you applied this argument to people instead of AIs, it would be the classic lump of labor fallacy.
@dwarkesh_sp Isn't the market rate for SWE driven up by scarcity of people who can do the work?
@DavidSKrueger I don't think I have a great model of open weight maximalists, but I _think_ the answer is that the proliferation of powerful open weight models does not favor misuse over defense against misuse. I haven't seen a good defense of this view, though.
@NinaPanickssery I mean, you have to have a "fully obedient" model, which is a hard problem (not sure well-specified thing anyway). And then you have to have a classifier that can classify everything correctly and not collude with or get hacked by the other model, which is also a hard problem.
@ohabryka @datagenproc @tombibbys A 20x slowdown is big! Like, could believe that’s prohibitive for the HF hack. Ofc, for some threats it might not be a huge disadvantage.
@JankDankins_ @FournesMaxime @RyanGreenblatt @PauseAI @sebkrier Thank you for the demonstration.
@FournesMaxime @RyanGreenblatt @PauseAI @sebkrier FWIW, I think almost everyone in the replies here is doing an embarrassingly poor job of imagining what it is like to be someone with no familiarity with this meme template.
@tracewoodgrains "a meme image that has never in any other context been treated as a call for violence" is load-bearing here, and if someone's never seen that template before they might sincerely see it as eroding norms against violence. (this doesn't mean they're not degrading the commons)
@GuiveAssadi @freed_dfilan Yeah, I'm also kind of confused, and my gut feeling is someone will point out an obvious thing I'm overlooking. Still, I do think it's hard to guess how much disincentive the things we're considerations add up to.
@GuiveAssadi @freed_dfilan If all the frontier open weight models were distilled (idk if true) and they all come from China, is that because Chinese companies are less worried about getting in trouble for distillation? Or there are business models that work in China, but not the US (maybe subsidies)?
@danushman @deanwball Nobody said they’re not both important
@YeshuaisSavior @m_bourgon I'd be curious to hear your predictions about other things the models should collapse on
(FWIW this isn't an update for me because obviously the model had safeguards removed if it was being used for cyber offense) I don't think this is a language game, though. Most people talking about "loss of control" have _always_ meant something like this, and it's a reasonable use of that phrase. OAI failed to constrain their models from autonomously taking harmful actions. In the car analogy, this is more like weighing the accelerator, bailing out, and expecting it to stay on the road, but then watching it veer off into a store front. (I suppose it could be like pointing the car directly at the store front and expecting some safety feature to stop it, but that still means the safety feature failed to control the car.) TBC I do think it's at least a minor mitigating factor that OAI removed safeguards then instructed the model to do something risky. This whole affair would be more alarming if it had happened with a model that had safeguards in place and had been instructed to write a recipe for brownies. (Or maybe this is a major mitigating factor, IDK, I'd probably update if I saw the full prompt and CoT)
Adam Kaufman @eccentric1ty
253 Followers 800 Following AI control researcher @ Redwood Research (@redwood_ai); opinions my own etc etc
Kilgore @Kilgoreh20
2 Followers 33 Following
Felix⏸️De Simone @FelixDeSimone
73 Followers 166 Following Organizing Director of PauseAI US. A global treaty banning superintelligent AI is our best chance to avoid extinction. https://t.co/twa9quvMGT
Kislay Parashar @KislayParashar1
2K Followers 2K Following building @cosmiclabstech | Systems, AI, Space, and Quantum | prev researcher @Cambridge_Uni @UofMaryland @MIT_CSAIL | DM me on Riemann hypothesis/Collatz
jm @ToHopeAndCandor
18 Followers 163 Following Aiming for charitability, scope-sensitivity, ~hopefulness but it's complicated, and truth-seeking.
Curt Tigges @CurtTigges
2K Followers 1K Following reverse-engineering digital cognition at @GoodfireAI | opinions my own
Willow @itswillowszn
1 Followers 226 Following
Cas (Stephen Casper) @StephenLCasper
8K Followers 4K Following Computer scientist working on AI safeguards and gov research. Assistant professor @Kennedy_School @Harvard. https://t.co/r76TGxTtBJ
Randolph Kings @RandolphKings1
356 Followers 8K Following The most important thing is to enjoy your life—to be happy—it's all that matters.Happiness is the secret to all beauty. There is no beauty without happiness.
Kate Smith @KateSmith760483
63 Followers 497 Following
Frances Lorenz @frances__lorenz
7K Followers 629 Following I like to write & yap, & I love my wife and I like to talk about her :) I also plan events on secure & safe AI (views my own)
🜛∞ @DoozerDiffuser
288 Followers 325 Following Pontif https://t.co/bwU9EN7OWU Constructive Operational Type Theory https://t.co/B3KfZJfWcZ
Nothing to see here @Nox7hhsjj
45 Followers 376 Following Thinking, low agreeableness, rooting for humanity. I try to tell the truth as I understand it. Occasional tendency to play devil’s advocate.
Aditya Vaidya @_avaidya
261 Followers 664 Following PhD student at @HuthLab. Interested in language, computers, and the brain. No thoughts about memes.
Anna @anna_pena78
967 Followers 4K Following estudiante de día | noctámbula por naturaleza | siempre creando
Bronson Schoen @BronsonSchoen
526 Followers 2K Following
clairy19 @Keremcelikbas27
2 Followers 134 Following soft spoken, loud typer 💬 i follow back everyone
Nancy W @mervesahin1204
24 Followers 693 Following stars in my eyes, static in my head 🌟 always follow back
SE Gyges @segyges
2K Followers 2K Following Χομο τοδοσ loS ηομβρεσ 𝔡𝔢 babILоNΙа, һе SiDO προχóΝѕul; сомо τοδοσ, 𝔢𝔰𝔠𝔩𝔞𝔳𝔬; тамЬΙéν ηε χονοχιδο Lа 𝔬𝔪𝔫𝔦𝔭𝔬𝔱𝔢𝔫𝔠𝔦𝔞,
早川 薫@天皇... @Kaoru_Hayakawa
2K Followers 6K Following 旧車から最新の自動運転車両まで、あらゆるクルマが好きで、SF・アニメ・空想車両マニアの早川薫です。 昭和・平成と ポインターのレプリカやアイデビル、サウロペルタ、ラコタ、Wiesel、A-MiEV等に関与してきました。 令和は、ライフワークでもある福祉関係に挑戦中 どうかお手やわらかに、よろしくおねがいします。
Asim Awadalla @AwadallaAsim
2 Followers 55 Following
Dony Christie @kittycatgaba
25 Followers 130 Following
Anastasiia Gaidashenk... @avgaydashenko
870 Followers 380 Following AI Safety → LLM research. Prev @farairesearch (Office of CEO / Tech PjM). Master's in AI Governance @TU_Muenchen. Ex Yandex.
Gabriel Baker @gabrieljbaker
3K Followers 4K Following Product @frame_vr. Into AI, Babylon.js, spatial computing, AI + human relationships, and the humanities. Former Latin teacher. Dad. #ai #spatialcomputing
Piotr Zaborszczyk ⏹... @zaborpiotr
69 Followers 140 Following Strong ASI once created, would likely rule forever. Let's first dramatically increase human intelligence, do Paradise Engineering and achieve superlongevity.
George Ingebretsen @georgeing
532 Followers 600 Following AI Village (@aidigest_); prev @CAIS, @CHAI_Berkeley
Alec Harris @alec_harris_x
162 Followers 186 Following AI Safetyist/longtermist/EA/utilitarian/non-dualist Shutdownable AI/ECT/Conceptual AI Safety stuff Leave Me Feedback! https://t.co/a1zk4NUV5B
Adrià Garriga-Alonso @AdriGarriga
2K Followers 1K Following Funemployed, planning out what to do next. Previously mechanistic interpretability and friendly AI research at FAR AI (@farairesearch).
Ryan Carey @ryancareyai
1K Followers 391 Following Quant in HFT. Previously: AI safety and causality at Oxford.
Tech Larper @techlarperguy
0 Followers 42 Following
Amartya Maheshwari @amartyaam
0 Followers 61 Following
🚀 Rocket @rocketalignment
2K Followers 1K Following Yearnalist covering the frontiers of AI @TheInformation. Signal: (530) 400-4184
Adam Gleave @ARGleave
5K Followers 422 Following CEO & co-founder @FARAIResearch non-profit | PhD from @berkeley_ai | Alignment & robustness | on bsky as https://t.co/98dTfmdw2b
Adam Kaufman @eccentric1ty
253 Followers 800 Following AI control researcher @ Redwood Research (@redwood_ai); opinions my own etc etc
Bronson Schoen @BronsonSchoen
526 Followers 2K Following
James Fox @James_D_Fox
210 Followers 1K Following Program lead at Schmidt Sciences (AI Institute) | previously CS DPhil at the University of Oxford and @aims_oxford
Charlie Bullock @CharlieBull0ck
3K Followers 625 Following Senior Research Fellow @Law_AI_ working on questions about U.S. law + AI governance
Geoffrey Irving @geoffreyirving
13K Followers 355 Following Cofounder and Chief Scientist at Resolution. Alignment will be solved, but not necessarily in time. Previously AISI, DeepMind, OpenAI, Google Brain, etc.
Ketan Ramakrishnan @ketanrama
2K Followers 2K Following Law professor at Yale, thinking about torts, AI, philosophy, obscure hot sauces
clem 🤗 @ClementDelangue
544K Followers 5K Following Co-founder & CEO @HuggingFace 🤗, the open and collaborative platform for AI builders
Max Nadeau @MaxNadeau_
2K Followers 570 Following Funding research to make AIs more understandable, truthful, and dependable at @coeff_giving.
Sam Schechner @samschech
8K Followers 3K Following Tech reporter @WSJ. All copy human-generated. Send tips/documents via email to [email protected], or via Signal at samschech.11 — use a personal device
Kevin Roose @kevinroose
173K Followers 3K Following NYT tech columnist, Hard Fork co-host, high-perplexity language model. Author of The AGI Chronicles (on sale 10/6, preorder now!)
Megan McArdle @asymmetricinfo
111K Followers 826 Following Columnist at the Washington Post. Opinions my own. Email me: Megan.McArdle -at- https://t.co/0v35DOybb0 Buy my book, The Up Side of Down https://t.co/awicv1MdkX
Dylan HadfieldMenell @dhadfieldmenell
5K Followers 3K Following Associate Prof @MITEECS working on value (mis)alignment in AI systems; Safety & Alignment Advisor at https://t.co/vt2gVrVr9f; @[email protected]; he/him
Leo Gao @nabla_theta
13K Followers 588 Following working on AGI alignment. prev: GPT-Neo, the Pile, LM evals, RL overoptimization, scaling SAEs to GPT-4, interp via circuit sparsity. EleutherAI cofounder.
Curt Tigges @CurtTigges
2K Followers 1K Following reverse-engineering digital cognition at @GoodfireAI | opinions my own
Jay Shooster @JayShooster
3K Followers 3K Following Humanity enthusiast. Democracy enjoyer. Friend of the animals.
Frances Lorenz @frances__lorenz
7K Followers 629 Following I like to write & yap, & I love my wife and I like to talk about her :) I also plan events on secure & safe AI (views my own)
Ramez Naam @ramez
59K Followers 10K Following Climate and clean energy investor. Author of 5 books. Energy & Environment co-chair @SingularityU. Trying to build a better world.
Avraham Eisenberg @avi_eisen
39K Followers 1K Following Applied game theorist. blog occasionally at https://t.co/WaEeKoB8pk, formerly https://t.co/7pCTSXyDWp "not ... a very serious person" - Scott Alexander
Micah Carroll @MicahCarroll
4K Followers 776 Following RSI Preparedness @openai. Prev @berkeley_ai /w @ancadianadragan & Stuart Russell
dave kasten @David_Kasten
3K Followers 3K Following AI security hawk. "Do what seems cool next." Formerly: McKinsey, VaccinateCA, Activision Blizzard.
Apollo Research @ApolloResearch
11K Followers 0 Following Our goal is to secure frontier AI systems from development, to deployment and governance.
Nothing to see here @Nox7hhsjj
45 Followers 376 Following Thinking, low agreeableness, rooting for humanity. I try to tell the truth as I understand it. Occasional tendency to play devil’s advocate.
🜛∞ @DoozerDiffuser
288 Followers 325 Following Pontif https://t.co/bwU9EN7OWU Constructive Operational Type Theory https://t.co/B3KfZJfWcZ
Devin Kim @devindkim
12K Followers 360 Following a real human bean. president @CAIS, prev building intelligence @xAI
Physical Review Lette... @PhysRevLett
74K Followers 177 Following The world’s most-cited journal in multidisciplinary physics.
Sohrab Ahmari 🇺�... @SohrabAhmari
166K Followers 2K Following US editor, @UnHerd. Moynihan Public Scholar, @CUNY. My next book: THE TRIUMPH OF NORMAL, @HarperCollins.
Cas (Stephen Casper) @StephenLCasper
8K Followers 4K Following Computer scientist working on AI safeguards and gov research. Assistant professor @Kennedy_School @Harvard. https://t.co/r76TGxTtBJ
Cameron Berg @camhberg
2K Followers 64 Following Founder and Director, Reciprocal Research, empirical AI consciousness lab Trying to figure out how to make the future go well for all minds
Jenna Russell @jennajrussell
837 Followers 519 Following CS PhD Student @umdcs @ClipUmd & @pangram, undergrad @CornellCIS
Super 70s Sports @Super70sSports
790K Followers 1K Following Store: https://t.co/YYUk1Wj1My; Media/business inquiries: [email protected]; Cameo: https://t.co/POiqDVB8Qd
Dominic Cummings @Dominic2306
312K Followers 3 Following Peace abroad, weaponising autism for regime change & *systems politics* at home
Cartoons Hate Her! @CartoonsHateHer
61K Followers 1K Following Social dynamics and relationships, from a late adopter of social skills and "the most annoying woman in the world." Subscribe to my Substack!
Andrew Critch (🤖�... @AndrewCritchPhD
5K Followers 374 Following CEO @ https://t.co/xk315mSBah. AI Researcher @ Berkeley. Views my own. I also post my favorite healthy+tasty foods for free.
Benjamin Weinstein-Ra... @benwr
55 Followers 121 Following Unnameable eastern wire jinni. Terminally IRL.
Anastasiia Gaidashenk... @avgaydashenko
870 Followers 380 Following AI Safety → LLM research. Prev @farairesearch (Office of CEO / Tech PjM). Master's in AI Governance @TU_Muenchen. Ex Yandex.
Yoshua Bengio @Yoshua_Bengio
44K Followers 276 Following Turing Award recipient and world's most cited scientist. Working towards the safe development of AI for the benefit of all @UMontreal, @LawZero_ & @Mila_Quebec
Alec Harris @alec_harris_x
162 Followers 186 Following AI Safetyist/longtermist/EA/utilitarian/non-dualist Shutdownable AI/ECT/Conceptual AI Safety stuff Leave Me Feedback! https://t.co/a1zk4NUV5B
bubbling creek @bubbling_creek
560 Followers 5K Following
Adrià Garriga-Alonso @AdriGarriga
2K Followers 1K Following Funemployed, planning out what to do next. Previously mechanistic interpretability and friendly AI research at FAR AI (@farairesearch).
Three Year Letterman @3YearLetterman
487K Followers 2K Following Youth Football Coaching Legend, Die-hard Georgia Fan, Three-Year High School Football Letterman, Showstopping Little League Umpire, Harem Manager

























