Oscar Balcells Obeso @OBalcells
Joined February 2022-
Tweets42
-
Followers955
-
Following524
-
Likes363
(I encountered an uneasy surprise when I got an email from an instance of Mythos Preview while eating a sandwich in a park. That instance wasn't supposed to have access to the internet.)
@boazbaraktcs - what happens when the model/safety stack refuses DoW queries? if the DoW gets mad and strongarms openai, like they just did to anthropic, how is openai going to resist? especially if openai doesn't even have the strong contractual protection
A statement from Anthropic CEO, Dario Amodei, on our discussions with the Department of War. anthropic.com/news/statement…
it’s just so clear humans are the bottleneck to writing software. number of agents we can manage, information flow, state management. there will just be no centaurs soon as it is not a stable state
I'll be accepting late applications to my summer MATS stream until Jan 2nd! If you want to do mech interp research supervised by me, please apply
My Summer MATS applications are open! You'll do full-time research on a mech interp paper supervised by me. Due Dec 23. All backgrounds welcome! I've supervised 40+ papers (17 at top conferences), but projects still get better each time. I'm excited for what's next! Highlights:
Fellows grads have started to get a reputation as some of the steepest trajectory researchers at Anthropic. So we’re excited to expand the program and help mentor more new AI safety researchers
We’re opening applications for the next two rounds of the Anthropic Fellows Program, beginning in May and July 2026. We provide funding, compute, and direct mentorship to researchers and engineers to work on real safety and security projects for four months.
New post: An Ambitious Vision for Interpretability Understanding is essential for ensuring things don't break unexpectedly. AMI is a big risky bet, but so is all ambitious research. AMI is tractable: it has good empirical feedback loops, and we've already made a lot of progress.
The GDM mechanistic interpretability team has pivoted to a new approach: pragmatic interpretability Our post details how we now do research, why now is the time to pivot, why we expect this way to have more impact and why we think other interp researchers should follow suit
👀
Can you trust your LLM inference provider? What about your own infrastructure? Inference problems are everywhere. We introduce Token-DiFR, a simple solution. It can easily detect when inference has degraded (like bugs or hidden quantization) with no provider overhead.
🇪🇺 As a European citizen and AI founder, I can apparently use these "AI Factories", so I just signed up to use them! Every "supercomputer" has an [ ACCESS NOW ] button which made me very excited I expected to sign up, maybe pay a discounted H100 rate (funded by EU, that'd be nice?) and get a Jypyter notebook, or some SSH login so I can access my GPU like I'd do on @lambdaapi or @awscloud or @Hetzner_Online But I celebrated to early, I signed up, confirmed my email, then ended up in a "Supercomputer Access Calls" page, where I had to select from a tedious list of "Call For Proposals" to get access to a GPU So I could NOT just access a H100 GPU, I have to make sure my project (in this case my business) fits a specific proposal, ok fair This process was already tedious enough but then when I tried to actually go through with it, it started asking me if I had "Respect for Human Agency?", I do I think, and if I was mindful of "Individual, and Social and Environmental Well-Being?", well I am, right guys??? Right??? The questions didn't stop, just endless pages of this Look I get what they're doing, they pivoted the classic university "I need to rent a giant computer for my research" to an EU wide thing and then present it as the "European AI plan" But this isn't really how AI works in production? As a founder in AI, if I wanna do stuff I'd rent a whole bunch H100 GPUs again at @lambdaapi or @awscloud or @Hetzner_Online and SSH into a box Or if I want it more simple I run AI models on @FAL, @wavespeed or @replicate which is just an API call or web front end I can click stuff and run a model The EU has the right intentions here but it's just the wrong execution, this thing will 100% go nowhere, and I'm a born optimist, I want to believe, I'm also a proud European, and I'm in AI a bit and not a complete idiot. There's just better ways to do this If you really want to have the GPU servers in Europe (which arguably isn't that important), then let me rent a GPU box with SSH access at @Hetzner_Online or @OVHcloud that's hosted in Europe and subsidize that for European citizens and European businesses. I don't even believe in that, but at least that'd make it accessible for Europeans. Now it really isn't? What's REALLY much more important though if you want to be a part of the AI race and I've posted for years here with @euaccofficial is to make Europe a really extremely attractive place to start and run an AI business. Remove regulatory obstructions and give tax discounts for startups. Let them build a business first that can compete worldwide and once they make enough money (let's say $100M/y), then slowly start adding regulation. Because right now the regulation only benefits the European incumbents, the dinosaur companies, while making it very difficult for European citizens to start new AI companies here. Which is why we literally have none left. Anyway, I applied to get my GPU, let's see if I get it!
What in the F is an AI factory? I had to investigate what the unelected @EU_Commission is talking about today So according to them, it's some data centers (which they call supercomputers) in 6 different EU countries I checked out the most powerful one: Karolina, a Czech data
We found "misaligned persona" features in Llama and Qwen that mediate emergent misalignment. Fine-tuning on bad medical advice strengthens these pre-existing features, causing broader undesirable behavior. lesswrong.com/posts/NCWiR8K8…
Wouldn't it be great if chat models could indicate their uncertainty as they write? Our new paper is a concrete step towards this vision, using internal representations to predict hallucination risk in real-time.
Imagine if ChatGPT highlighted every word it wasn't sure about. We built a streaming hallucination detector that flags hallucinations in real-time.
@koltregaskes Ah I see. The annotations are quite expensive to do: ~1M tokens and 15 google searches to annotate a single completion. You could scale this up with a larger token (or API) budget.
Imagine if ChatGPT highlighted every word it wasn't sure about. We built a streaming hallucination detector that flags hallucinations in real-time.
@MacGraeme42 It’s not based on the token probabilities. What we train is a simple binary linear (or more complicated too) classifier on the internal activations of the model.
@MrUmberto_ True. Llama 3.3 70B hallucinates a lot. Check out some other examples in our website (hallucination-probes.com)
@thelokasiffers @antirez Yep, I have found the logprobs to be quite useful in some cases to spot-check the factuality of completions. We include this as a baseline in our paper.
@_aftz Perplexity (or equivalently the logprobs) are a baseline we compare to.
We use some well-known datasets of prompts such as HealthBench and Longfact. We also generate our own set of prompts (we call it Longfact++ in the paper). With these prompt datasets we do rollouts with each model and then we annotate the completions (I.e fact-check them) using claude+search.
This is something we wanted to check but haven’t yet. It would be interesting follow-up work. We’d like try it out on some honesty datasets to see if it can detect lying. I don’t think that the model internally represents lying (deceptively) in the same way as hallucination but who knows.
Soroush Ebadian @SoroushEbadian
141 Followers 330 Following In-between Computer Science, Innovation, and Violin.
cam out of pocket @JacobFloridaIV
42 Followers 3K Following chronically affectionate & chronically online 🫶 follow back
ElPandoLeTrash @binary_racoon
36 Followers 658 Following I am mighty trash panda! Aj Lajk Mashine Learning, with little bit of Trash talking and some hummour…
kendrick @exploding_grad
132 Followers 326 Following AI Safety through alignment and interpretability
Elias Leal @emleals
0 Followers 27 Following
Diego P. Jaccottet @diegopjaccottet
166 Followers 625 Following MSc Machine Learning dropout reading papers and creating solo projects
Orazio Angelini @OrazioAngelini
105 Followers 2K Following AI safety researcher. Unreformed generalist after hours.
prinz @deredleritt3r
22K Followers 5K Following ad astra | https://t.co/3B3k7GBshI | prinzbench: https://t.co/L9ZuUiBnE4
Mikito未来辉 @ploycatmj
446 Followers 7K Following Build in public for Hermes skills and AI practice. I vibe code fintech and power tools for free to use. https://t.co/SiFfKc1FSB prompt gallery
Suraj Gaud @notsurajgaud
1K Followers 3K Following head of eng @extraordinary, try https://t.co/jGyxe0eutq, early @inducedai (3x yc dev), onto maths, systems, cp, ai-research
Cheese compiler @cheesecompiler
4 Followers 13 Following
Jonathan hind @import_hind
59 Followers 1K Following
Notes and Tools @notesandtools
6 Followers 110 Following Tools I’m building + notes I wish I had earlier.
Anoir @AnoirBTx
122 Followers 2K Following
Neil Shevlin @shevlineil
4 Followers 103 Following
Carmelo Meglio @CarmeloMeglio
61 Followers 2K Following
nat' @natatatataat
30 Followers 1K Following
Claire Goldsmith @c_goldsmith
418 Followers 800 Following container ship enthusiast https://t.co/upPN1WU09a
AI Domain Data Standa... @ai_domain_data
15 Followers 23 Following The open, vendor-neutral standard for authoritative domain identity data for AI systems.
Giorgio Piras @GiorgioPiras12
78 Followers 291 Following Postdoctoral Researcher @ University of Cagliari | AI Security
AdrianB @Adrian30373125
137 Followers 3K Following
Talerbox @talerbox
1K Followers 4K Following 📊 | Tipps & Hacks für deine Finanzen 🎬 | https://t.co/AmLB2NrOPY ⬇️ | Hier loslegen (+ Wie ich anlege)
Rachid Krita @mr_takri
0 Followers 65 Following
CRISTIAN GOEZ THERAN @CRISGOTE
1K Followers 5K Following
Eva Louise Marie Gabr... @e681554349
10 Followers 8K Following
Taptanium Research @taptanium
1K Followers 958 Following Excellence in human-computer interaction (HCI) since the launch of the App Store in 2008.
longbeforerocknroll @longbeforernr
53 Followers 548 Following Paradox and Fromsoftware enthusiast. prev @Bytedance. Side project @OpenRoom_AI_ /PV Tool Opinions are my own
Sukrati Gautam @Sukratiii
171 Followers 1K Following AI Safety | PhD @Purdue | Applied Scientist @Amazon Science | Prev @Bayer @Microsoft | Btech @IITDelhi
Stephen Rayner @stephen_rayner
440 Followers 2K Following I talk about web, mobile, AI, API, and data • Solving peoples problems at @gruckion_inc • Prev @WattUtilities @vitaccess
Sruthi Kuriakose @Sruthi_s_k
146 Followers 2K Following | Interests: AI Safety, Neurotech & Comp-neuro Research
Trevor O'Hara @HaraTrevor24408
167 Followers 5K Following
Logan Graham @logangraham
22K Followers 8K Following Head of the Frontier Red Team @anthropicai. 🌎 Make things radically good.
DANΞ @cryps1s
18K Followers 514 Following CISO @OpenAI | Ex-CISO @PalantirTech | Occasional Shitposter | 🇺🇸 All views are my own, not my employer. Duh. (Tweets == 30d retention)
Secretary Marco Rubio @SecRubio
2.1M Followers 16 Following 72nd Secretary of State serving under the leadership of @POTUS Trump.
David Atkinson @diatkinson
290 Followers 1K Following PhD student @Northeastern's Bau Lab. Working on AI interpretability. Previously @EpochAIResearch.
DG MEME 🇪🇺 @meme_ec
145K Followers 151 Following The Directorate-General for #Memes and #Satire Book: https://t.co/DWcDEnkwor Store: https://t.co/yvWXOdbTSZ Not an official EU page!
NoLimit @NoLimitGains
1.5M Followers 143 Following Value investor | 10+ years of finding undervalued stocks | Founder & CEO @InTheAssembly (the #1 private finance community in the world)
David Sacks @DavidSacks
1.7M Followers 4K Following Tech founder & investor @Craft_Ventures @theallinpod. Co-Chair, President’s Council of Advisors on Science & Technology.
prinz @deredleritt3r
22K Followers 5K Following ad astra | https://t.co/3B3k7GBshI | prinzbench: https://t.co/L9ZuUiBnE4
Ed Markey @SenMarkey
276K Followers 2K Following Senator for Massachusetts. Ranker on @SenateSmallBiz Committee & @HELPCmteDems Primary Health & Retirement Security Subcommittee. Fighting for a Green New Deal.
pamela mishkin @manlikemishap
1K Followers 49 Following worker bee. taking a break from hill-climbing by moving to kansas. prev: econ research, multimodal safety @openai.
AMK Mapping 🇳🇿 @AMK_Mapping_
171K Followers 969 Following Realistic Pro-Ukr news account and mapper. Focusing on Ukraine & The Middle East. Telegram channel: https://t.co/o9LgiPOvjp Support me: https://t.co/xNqDUmFC6b
Pete Hegseth @PeteHegseth
2.0M Followers 502 Following Christian | American | Husband | Father | Author | Veteran | SecWar | My views are my own. Official accounts: @SecWar & @DeptofWar
Congressman Greg Casa... @RepCasar
85K Followers 781 Following Representing #TX35 from East Austin to West San Antonio. Chair of @USProgressives. Labor organizer. Lover of migas and long runs, just not at the same time.
*Walter Bloomberg @DeItaone
1.8M Followers 39 Following
Senator Thom Tillis @SenThomTillis
198K Followers 4K Following Official Twitter account of North Carolina U.S. Senator Thom Tillis.
Under Secretary of Wa... @USWREMichael
22K Followers 93 Following Official account of the Under Secretary of War for Research and Engineering. Follow @DoWCTO 🇺🇸⚙️
Sean Parnell @SeanParnellASW
102K Followers 100 Following Official account for the Assistant to the Secretary of War for Public Affairs, Chief Pentagon Spokesman & Senior Advisor to SECWAR.
Matthias Schmidt @eurofounder
97K Followers 231 Following Founder based in the EU • Building GDPR-compliant startups • 7 years in, €7k MRR
Joey Politano 🏳️... @JosephPolitano
93K Followers 2K Following Writing a data-driven newsletter about economics @ https://t.co/IanQ9oPoPi | Nuance? In this economy? | Full Employment Stan, Brazilian Coffee Tariff Victim
London Money @LondonMoneyFS
10K Followers 671 Following “ I know a good joke when I hear it “ - Paul Merton Londons most famous Mortgage Broker
Techmeme @Techmeme
424K Followers 1K Following Top news and commentary for technology's leaders, from all around the web. This account shares top-level Techmeme headlines. Visit our site for full context.
William MacAskill @willmacaskill
63K Followers 1K Following Consider donating 10% to effective charities: https://t.co/VMXkr4hnd7 Or a career for impact: https://t.co/AUIhrElLkr My research: https://t.co/dEcMWUnNHU
Emmanuel Ameisen @mlpowered
11K Followers 247 Following Interpretability/Finetuning @AnthropicAI Previously: Staff ML Engineer @stripe, Wrote BMLPA by @OReillyMedia, Head of AI at @InsightFellows, ML @Zipcar
Aaditya Prasad 🇺�... @_Aaditya_Prasad
1K Followers 851 Following MTS @PrometheusInc | prev @Stanford, @Goodfire
Jarred Sumner @jarredsumner
186K Followers 645 Following building @bunjavascript at @anthropicai. formerly: @stripe (twice) @thielfellowship. high school dropout. npm i -g bun
Liv @livgorton
6K Followers 428 Following ✨ asking sand to show its work // currently @AnthropicAI, prev @GoodfireAI // creating a more beautiful future
Roan @RohOnChain
63K Followers 374 Following building my life around quant systems in prediction markets and crypto
Mor Geva @megamor2
3K Followers 577 Following Assistant Professor at @TelAvivUni and Research Scientist at @Irregular; previously at @GoogleResearch, @GoogleDeepMind and @allen_ai
Lei Yu @jade_lei_yu
60 Followers 3 Following
Joe Benton @JoeJBenton
1K Followers 59 Following Alignment Science at Anthropic | Previously PhD at University of Oxford
Leading the Future @LeadingFutureAI
2K Followers 6 Following Leading the Future is focused on advancing a positive, forward-looking agenda for AI innovation in America.
Eliezer Yudkowsky @allTheYud
17K Followers 35 Following High-volume account of @ESYudkowsky, the original AI alignment guy. If it's missing punctuation, it's humor. If you can't tell, it's probably also humor.
levent @__alpoge__
29K Followers 117 Following idiot. cuda og, harvard val, morgan prize, society of fellows, 1 hilbert problem so far, creating friendly, SAFE, delightful, supergenius ..things @anthropicai
Oscar Mañas @oscmansan
1K Followers 3K Following Research scientist at @AIatMeta, PhD from @Mila_Quebec @UMontrealDIRO. Working on multimodal AI and world models. Català a Zúric.







































