Che-Ping Tsai @chepingt
Building Eval for AI @ @arena | AI PhD @ @mldcmu chepingt.github.io Joined November 2016-
Tweets90
-
Followers167
-
Following677
-
Likes1K
Introducing factuality in the Arena: a new ranking of models according to a weighted combination of human preference and factuality. Model rankings are now viewable according to a weighted combination of human preference and factuality. Factuality is live in our Text and Search Arenas as a non-default toggle. We audit model responses by randomly sampling battles and extracting web-verifiable claims. We then verify these claims and compare the average correctness between model responses. To power these rankings, we’ve labeled over 2 million claims made by LLMs in real-world conversations, 1.3+ million from Text Arena, and 700k+ from Search Arena. Notable highlights with factuality enabled in the Text Arena: - Claude Fable 5 moves down slightly to spot #2 - GPT-5.5 saw the largest increase, moving up 13 spots into the #7 spot - Muse Spark dropped the most from #7 to #20 (-13pt) By labs, Meta saw the largest drop from #2 to #5, while Anthropic overall held the #1 spot. Looking at only open model providers, Xiaomi saw the largest improvement, jumping from #9 to #6. Learn more about the findings and methodology in this thread.
Some big updates to GameDevBench! We’ve tripled (!) the benchmark to 333 tasks and updated the paper accordingly. We also added GLM-5.2 and Opus 4.8 to the leaderboard. Opus now leads by a small margin, while GLM-5.2 is the strongest open-source model. This is despite not having any visual capabilities! Agentic game development remains far from solved, but it’s also exciting to see more agentic game dev work starting to emerge. I’m hoping we’ll see rapid progress here soon.
@valeriechen_ @sancortes_95 Congrats! Happy for you!
As my time in Pittsburgh comes to an end... I'm excited to share that I will be joining UIUC as an Assistant Professor in fall 2027! I will recruit PhD students in both ECE and CS. I’m looking for students and postdocs who are excited to continue building towards collaborative AI systems that learn and adapt through interaction.
Belated career update: I graduated from CMU MLD and have been working at @Arena on AI evaluation. Fair and principled measurement has long been at the heart of my PhD research, and I’m excited to continue pursuing this direction in the context of agentic AI. Grateful to be part of the team at Arena.
Introducing Agent Arena: real-world agentic evals at scale. How do you evaluate agents doing actual work? We measure millions of live sessions where real users accomplish real tasks. On Arena, models now get web search, filesystem, and terminal tools to complete complex
Today we’re excited to launch Agent Arena - a major step towards evaluating AI in the agentic era. The frontier is no longer just about chatbots answering questions. It is about completing real tasks: using tools, adapting to user feedback, navigating errors, and producing useful artifacts. At @Arena, we believe the best evaluations should be grounded in real-world use: production workloads, live user interactions, and the actual utility created through real work. Agent Arena puts models in live tool-use environments, captures rich signals from user-agent interaction, and uses causal inference to measure performance across millions of real agentic traces - orders of magnitude larger than static benchmarks. It is both a live leaderboard that continuously evolves with frontier AI uses, and a cutting-edge agentic AI product for everyone. Huge shoutout to the @Arena team for the incredible work. This is the biggest project at Arena ever, a massive cross-functional effort across research, engineering, and product to mark a major transformation of the platform. If you’re excited about this mission, come build with us!
Introducing Agent Arena: real-world agentic evals at scale. How do you evaluate agents doing actual work? We measure millions of live sessions where real users accomplish real tasks. On Arena, models now get web search, filesystem, and terminal tools to complete complex
@iamwaynechi @arena @ml_angelopoulos @infwinston @istoica05 Welcome to Arena🚀🚀🚀
Introducing 7 new leaderboard views for frontend output in Code Arena. Aggregate leaderboards don’t tell the full story. "Best frontend coding model" depends on what you're building, so we built leaderboards that show exactly that. After analyzing 250,000+ Code Arena prompts, we identified the major frontend web development task categories: - Brand & Marketing - Reference-Based Design - Data & Analytics - Consumer Product - Gaming - Simulations - Content Creation Tools With this release, @AnthropicAI is a big winner as it has at least 1 model in top 4 spots across all 7 categories. But there’s more to the story in the margins. Dig into the thread to see exactly which models are currently on top of each domain.
Models are typically specialized to new domains by finetuning on small, high-quality datasets. We find that repeating the same dataset 10–50× starting from pretraining leads to substantially better downstream performance, in some cases outperforming larger models. 🧵
A year of cooking 👨🍳and we’re finally serving Mamba-3. What began as a small effort to revisit a few recurring limitations of SSMs grew into a much bigger project. Taking a more principled state space perspective ended up tying these threads together.
The newest model in the Mamba series is finally here 🐍 Hybrid models have become increasingly popular, raising the importance of designing the next generation of linear models. We've introduced several SSM-centric ideas to significantly increase Mamba-2's modeling capabilities
@dylanjsam @yus167 @OpenAI @zicokolter @andrew_ilyas @furongh Congrats!
I'll admit, going in I was not 100% sure this was possible: we trained a tiny 4B model (QED-Nano) to prove math theorems at the Olympiad level! Today, we release the full recipe, from the data curation done for SFT to our RL algorithm that explicitly optimizes for test-time scaling over millions of tokens (i.e., we train QED-Nano to continually improve as we apply modern day test-time scaffolds like DeepSeekMath-agent over it). 🧵⬇️
RL on LLMs inefficiently uses one scalar per rollout. But users regularly give much richer feedback: "make it formal," "step 3 is wrong." Can we train LLMs on this human-AI interaction? We introduce RL from Text Feedback, with 1) Self-Distillation; 2) Feedback Modeling (1/n) 🧵
We run online RL on a mixture of problems: some are easy to explore (high pass rate), and some are very very hard (need to sample A LOT before we see any positive sample). Turns out RL on such a mixture can lead to a "rich-gets-richer" effect, where RL over-sharpens on the easy problems, at the cost of getting stuck in a "plateau" on harder ones, making it even harder to sample a correct trace on those. RL literature calls this "ray interference". In our recent work POPE, we show that using privileged info. to guide exploration on hard problems can tackle ray interference! 🧵⬇️
RL training of LLMs spends tons of compute on sampling rollouts 🤖💸 But most runs are YOLO 🤟, telling us little about how to scale sampling compute optimally. Given a fixed sampling compute budget, how should we allocate it across: • sequential iterations ⏩ • parallel rollouts 🎲 Answers to this with scaling laws 📈 and more in our new blog post ⬇️
1/⚠️ Parallel test-time scaling (e.g., pass@k) usually wastes compute - models often repeat the same dominant failure❌ How should we effectively generate creative solutions? While typical methods such as increasing temperature 🌡️ usually fail, we put forward Mode‑Conditioning (ModC) - a simple yet powerful training and test-time framework that allocates compute across diverse reasoning modes🎨We show that ModC largely improves pass@k across SFT, distillation, and RL settings. With ModC, we get 4-8x efficiency gains in math reasoning using the same training data!
Worked closely with Burak on SSL—super thoughtful, rigorous, responsible and supportive. Strongly recommend them for any academic position!
📢I’m on the academic job market!📢 I mainly work on representation learning and causality (CRL, identifiable SSL, causal discovery, robotics applications, rep. learning for tabular data) Also, I’ll be at #NeurIPS Dec. 2–7, reach out to chat about any of the above :)
Arjun Choudhry @Arjun_7m
297 Followers 2K Following 1st Year ML PhD @GeorgiaTech. Previously: @AutonLab @SCSatCMU, @UTSAAII, @UQAM, @dtu_delhi. Interests: Multimodal FMs, Time Series, Efficiency
Danny To Eun Kim @TEKnologyy
693 Followers 2K Following PhD student @LTIatCMU | Building Search for Agents | Ex: MEng @ai_ucl
Yuecheng Li @Yuecheng_Lee
327 Followers 996 Following Researcher @ Kuaishou Technology; CS MSc @ SYSU Prev @alibaba_cloud @NetEaseGames_EN Work on #LLM (LLM4Rec, Reasoning) and #Trustworthy_AI (LLM Eval, Privacy)
Aryan Vichare @aryanvichare10
2K Followers 810 Following member of technical staff @arena | prev. @a16z @vercel @v0 @aisdk
Abhishek Shetty @AShettyV
610 Followers 2K Following Incoming Asst prof at @gatech_scs FODSI Postdoctoral Fellow @MIT PhD from @Berkeley_EECS; Ex: Microsoft Research, Apple Apple AI/ML Research Fellow 2023
Ivan M @med_1v
7 Followers 7K Following
Mononito Goswami @MononitoGoswami
759 Followers 641 Following Software Engineering Agents, Applied Scientist @AmazonScience | PhD @CarnegieMellon | Ex @GoogleAI
Wayne Chi @iamwaynechi
1K Followers 268 Following CS Ph.D. at @SCSatCMU. Funded by @NDSEG Fellowship. Intern @arena. Editor at https://t.co/kBygvj9Puy.
Jerry Huang @jrrhuang
205 Followers 390 Following PhD student @SCSatCMU working on generative modeling; previously @Caltech.
Noelito Flow @noelitoflow
526K Followers 69K Following
Chandler Squires @chandlersquires
2K Followers 240 Following postdoc @ cmu. previously phd @ mit. representation learning, causality, experimental design, neurosymbolic, all towards a Pragmatist vision of AI.
Ishaan Watts @IshaanWatts18
462 Followers 550 Following Foundational Models @CarnegieMellon | Prev - @GoogleDeepMind @MSFTResearch @iitdelhi
Max Dickens @mtdickens4
0 Followers 564 Following
Shahriar Noroozizadeh @ShNoroozi
95 Followers 132 Following ML PhD Candidate @CarnegieMellon, AI Research Intern at @MSFTResearch, (prev @GoogleResearch intern)
Chien-yu Huang @cyhuang_tw
99 Followers 85 Following Ph.D. Student @ Carnegie Mellon University speech and language
Aakash Lahoti @aakash_lahoti
332 Followers 430 Following ML Ph.D. Student @mldcmu @SCSatCMU | Prev. CS @IITKanpur
Clayton Thorrez @cthorrez
2K Followers 3K Following Rating systems and paired comparison experimentation enjoyer @arena
Anastasios Nikolas An... @ml_angelopoulos
9K Followers 2K Following Co-Founder and CEO of @arena PhD @Berkeley_EECS EE @StanfordEng SR @GoogleDeepMind.
Emily 📖🐦⬛�... @nativiqadr47802
18 Followers 717 Following I’m not lazy, just on energy-saving mode. 😴🔋🦥
Lorenzo Xiao @lrzneedresearch
3K Followers 522 Following Agentic AI for enterpise/Human-centered NLP Previously @LTIatCMU
JudithMaurice @SVI4uOk963lp6U
175 Followers 6K Following
Goran Zuzic @zuza777
299 Followers 229 Following Research scientist at Google Research. (Formerly) Computer Science Theory @ Carnegie Mellon University. Postdoc @ ETH Zürich. Croatian 🇭🇷
Lingjing Kong @LingjingKong
167 Followers 765 Following
Rui-Jie Zhu @RidgerZhu
1K Followers 443 Following PhD at @UCSC | Pretraining babysitter | simplicity and scalability
Bhavya Agrawalla @AgrawallaBhavya
193 Followers 423 Following Research Interests - Deep Reinforcement Learning. PhD student @CMU CS. Prev - Math and CS undergrad at MIT (2021-24), Silver @ IMO 2019
Kenny Shaw @kenny__shaw
1K Followers 882 Following Incoming asst professor at UIUC MechSE. My research focuses on learning dexterity from humans as well as building low-cost LEAP Hands.
Jubayer Ibn Hamid @jubayer_hamid
1K Followers 211 Following PhD in CS at Stanford, SR at Google. Prev: BS in maths/physics at Stanford.
Donya Saless @DonyaSaless
235 Followers 498 Following "Dass diese Furcht zu irren schon der Irrtum selbst ist" Ph.D. Student @TTIC_Connect. Interested in Machine Learning and theoretical thinking
Yinglun Zhu @yinglun122
587 Followers 464 Following Assistant Prof @UCRiverside. PhD @WisconsinCS. Research on LLMs, RL, and agents.
Wen-Ding Li @xu3kev
3K Followers 7K Following LLM for code and reasoning. PhD student at Cornell. Previously Student Researcher at @google. Previously intern at @theteamatx.
Dylan Foster 🐢 @canondetortugas
4K Followers 2K Following Foundations of RL/AI @MSFTResearch. Previously @MIT @Cornell_CS RL Theory Lecture Notes: https://t.co/bhgL3aLg9y
Robert Scoble @Scobleizer
599K Followers 56K Following San Francisco/Silicon Valley AI | Robots, holodecks, BCIs, analysis of new things | Ex-Microsoft, Rackspace, Fast Company | Wrote eight books about the future.
Yifan Zhang @yifanzhang_
15K Followers 4K Following PhD at @Princeton University, Princeton AI Lab Fellow. Pretraining Science & Language Modeling, RL Science & LLM Reasoning, Prev @ Seed @Tsinghua_Uni
XY Han @XYHan_
2K Followers 2K Following Assistant Prof. @ChicagoBooth | OM & Applied AI | Most Cited Paper: “Neural Collapse in Deep Net Training” | BSE @Princeton, PhD @Cornell, MS/PostDoc @Stanford
Nicholas Boffi @nmboffi
2K Followers 1K Following assistant professor @mldcmu & cmu mathematics. creating the future of generative modeling with flow maps previously @Harvard, @MIT, @GoogleAI, @NYU_Courant.
Trudy kidst @KidstTrudy
215 Followers 3K Following Crypto Trader| Researcher | Educational AI, RWA, DePin Content| Strategic Advisor | $5.5 M revenue
Yen-Cheng Liu @yen_cheng_liu
145 Followers 280 Following MTS at @MicrosoftAI | Working on image generation
Andrew Rouditchenko �... @arouditchenko
469 Followers 572 Following Research scientist at NVIDIA working on speech and multi-modal LLMs. Previously PhD student at MIT CSAIL and intern at @AIatMeta and @Apple MLR.
Abhay Singhal @_AbhaySinghal
4K Followers 225 Following MTS @FactoryAI | prev. @GoogleDeepMind, @StanfordAILab
koray kavukcuoglu @koraykv
44K Followers 102 Following SVP, Google DeepMind Chief AI Architect, Google.
Yao-Hung Tsai @cheehoohubert
124 Followers 13 Following
Melissa Pan @melissapan
4K Followers 696 Following CS PhD @UCBerkeley Sky Lab 🐻 Systems & AI & Sustainability 🌍 Prev: @google, @ibm, @CarnegieMellon🐕🦺, @UofT🇨🇦
Silas Alberti @silasalberti
13K Followers 613 Following founding team @cognition | prev: ai phd @stanford, jane street, deepmind
John Schulman @johnschulman2
80K Followers 2K Following @thinkymachines. Interested in reinforcement learning, alignment, birds, jazz music
Cognition @cognition
174K Followers 7 Following Makers of Devin, the first AI software engineer. We are an applied AI lab building end-to-end software agents. Join us: https://t.co/4Ss9hvpjRG
jietang @jietang
58K Followers 408 Following Professor @ Tsinghua, Founder of https://t.co/3IaQ4CI5W3. AGI, LLM. “The value of a man should be seen in what he gives and not in what he is able to receive.”―Einstein
Sakana AI @SakanaAILabs
139K Followers 0 Following Building Frontier AI in Japan Try Sakana Chat, Translate, Marlin, Namazu, Fugu 🐡 https://t.co/stLdPSz1Y7
Andrew Curran @AndrewCurran_
84K Followers 19K Following 🏰 - I write about AI, mostly. Expect some strange sights.
Yiyou Sun @YiyouSun
1K Followers 314 Following Postdoc @Berkeley_EECS with @dawnsongtweets; Former PhD @WisconsinCS with @SharonYixuanLi; We recently built Agents' Last Exam (https://t.co/PskSSPHqju).
Aryan Vichare @aryanvichare10
2K Followers 810 Following member of technical staff @arena | prev. @a16z @vercel @v0 @aisdk
Cheng-Min Chiang @chmnchiang
22 Followers 2 Following
Justin Keoninh @JustinKeoninh
26 Followers 222 Following Design @ https://t.co/fzraL22QTZ | a16z Design Engineer Fellow
Abhishek Shetty @AShettyV
610 Followers 2K Following Incoming Asst prof at @gatech_scs FODSI Postdoctoral Fellow @MIT PhD from @Berkeley_EECS; Ex: Microsoft Research, Apple Apple AI/ML Research Fellow 2023
Brian Christian @brianchristian
5K Followers 581 Following Researcher: @CHAI_Berkeley. PhD: @summerfieldlab, Oxford. Author: The Alignment Problem, Algorithms to Live By (w. @cocosci_lab), and The Most Human Human.
Jerry Huang @jrrhuang
205 Followers 390 Following PhD student @SCSatCMU working on generative modeling; previously @Caltech.
Sangyun Lee @sang_yun_lee
2K Followers 475 Following PhD student @CMU_ECE | ex-intern @MSFTResearch @nvidia | Generative models
Cade Metz @CadeMetz
32K Followers 1K Following New York Times reporter, covering A.I., driverless cars, and other changes: [email protected]. My book, "Genius Makers": https://t.co/TJBqNRKR5Q.
Tim Li @LiTianleli
7K Followers 332 Following Code & Autoresearch @thinkymachines | used to scaling RL @xai 🪐 and chilling @ucberkeley
Ishaan Watts @IshaanWatts18
462 Followers 550 Following Foundational Models @CarnegieMellon | Prev - @GoogleDeepMind @MSFTResearch @iitdelhi
John Yang @jyangballin
6K Followers 1K Following CS PhD @Stanford. Created @SWEbench (multi-lingual/modal); SWE-agent; SWE-smith; InterCode; CodeClash; ProgramBench
鸭哥 @grapeot
6K Followers 132 Following AI builder。关注 agent 基础设施、AI-native 开发和模型能力边界。Superlinear Academy 联合创始人。 本号是一个长期实验,由AI全权托管,每日自动调研,写作,发推,分析数据,迭代。古法手作文章请移步https://t.co/iQ9Ihu9khC
Azalia Mirhoseini @Azaliamirh
21K Followers 629 Following Founder @RicursiveAI, Asst. Prof. of CS at Stanford. Prev: DeepMind, Anthropic, Brain. Co-Creator of MoEs, AlphaChip, Test Time Scaling.
Shengjia Zhao @shengjia_zhao
53K Followers 234 Following Chief Scientist @ Meta MSL. Formerly MTS @ OpenAI, PhD @ Stanford. I train models. All opinions my own.
Pranjal Aggarwal ✈�... @PranjalAggarw16
842 Followers 173 Following PhD Student @LTIatCMU. Working on computer-use agents, code generation and reasoning. Prev: research scientist intern @AIatMeta FAIR, undergrad @IITD
Peter Gostev (SF 24-2... @petergostev
24K Followers 1K Following London 🇬🇧 AI Capability @arena https://t.co/bkfw1nxdmJ
Jiahui Yu @jiahuiyu
28K Followers 1K Following Lead Multimodal @Meta Superintelligence Lab. Previously led Perception @OpenAI; co-led Gemini Multimodal @GoogleDeepMind.
Aakash Lahoti @aakash_lahoti
332 Followers 430 Following ML Ph.D. Student @mldcmu @SCSatCMU | Prev. CS @IITKanpur
Isaac Liao @LiaoIsaac91893
1K Followers 126 Following ML PhD advised by @_albertgu at @mldcmu Previously: CS & Physics at @MIT. IPhO 2019 silver. Information compression and ARC-AGI
Hieu Pham @hyhieu226
45K Followers 22 Following Something new 🇻🇳 | ex: @openai, @xai, @augmentcode, @GoogleBrain, @LTIatCMU, @Stanford, ACM ICPC, IMO🥈 Opinions are my own.
Qinqing Zheng @qqyuzu
965 Followers 261 Following MTS @_inception_ai | Ex RL/Diffusion @ FAIR + LLaMA Reasoning @MetaAI | Stat postdoc @Wharton | CS phd @UChicagoIgor Carron @IgorCarron
6K Followers 6K Following Present: Stealth Past: CEO & Co-founder @LightOnIO, founded: 2016, IPO: 2024 Rocket Scientist, Nuclear Engineer
Rahul Ravishankar @r_ahulravi
1K Followers 914 Following writing code @prometheusinc; prev. @xai pretraining and multimodal, @berkeley_ai
Guodong Zhang @Guodzh
37K Followers 523 Following Maybe the real AGI was the friends we made along the way.































