We should take it seriously.
The main limitation is that there are actually very few neolabs putting serious compute into fundamental research instead of doing rlaas or open weight transformers, but there are few who are doing it, are determined and moving fast.
How seriously should we take the possibility of one of the neolabs making some wild algorithmic breakthrough that puts it ahead of OpenAI and Anthropic?
How to live your life:
Meeting Mondays
Thinking Tuesdays
Work Wednesdays
Trying Hard THursdays
Automation Fridays
Sleeping Saurdays
CUDA Sundays
And repeat!
AI discourse often focuses on scaling because of eye popping numbers spent on datacenter compute. Credit assignment is very difficult
Most people from outside big labs and even many inside get this credit assignment wrong.
It’s worth doing a simple thought experiment:
It’s may 2020. GPT-3 paper just got released.
We have two diverging timelines:
a) we have the same algorithmic progress that we did since gpt3, but we cannot spend more compute on training models than was spent on gpt3
b) we keep scaling up gpt3 and we spend as much compute as we did on gpt5.6, but on gpt3 training system
Reality is that model from timeline a) beats model from timeline b) on every axis and it’s not even close
We are very proud to co-organize open participation competition inviting scientists from around the world on one of the deep learning problems that is still open: blog.tilderesearch.com/blog/one-layer…
One Layer Deeper is officially live!
The motivating idea is that some models just don’t want to learn. Most optimizer benchmarks ask how quickly you can train a given model, but baked into that question is the assumption that it can be trained (well) at all.
We think this
We had the @coreauto housewarming party a couple of weeks back. It was tons of fun and should hopefully give you a better idea of where we're spending some of our time improving model architectures and systems
00:00: Introduction
00:40: @MillionInt (@coreauto), Building the
I present to you the ultimate showdown.
In our left corner, we have the researchers complaining about the hardware lottery.
In our right corner, we have the infrastructure guys like below, with the exact opposite argument.
Let the battle be fair but harsh!
Very insightful conversation with the founders of @coreauto - a pretraining maximalist and an RL maximalist.
Got a little glimpse of what they think the next step could be in terms of model architecture. A couple of points they made that I really liked -
- The models are
Two of the people most responsible for scaling the transformer are now betting on a next act.
@MillionInt ran the Reasoning 🍓 team at OpenAI. @_arohan_ was a pre-training lead on Gemini after years at Google Brain and Anthropic. They just started @coreauto to find what comes
9 Followers 185 FollowingWe connect exceptional talent in AI & ML Research, Safety, and Security with the companies pushing the frontier of AI in San Francisco and New York.
63 Followers 1K Followingfungi/mold ~dizzy midwit intf | terminally maintaining neurotic equilibrium
if you change the way you look at things, things change the way they look at you :D
327 Followers 3K FollowingML, data, truth seeker, world traveler. Informing people through data and combating the spread of misinformation. Always open to evidence-based debate.
14K Followers 6K Followinghttps://t.co/wP5OsA9afm • https://t.co/UPVdiCl5mL. Scout @trygravityai & investor. Most startup problems can be solved w/ 100M views.
9.1M Followers 13 FollowingYour Only Source For Professional Dog Ratings Instagram and Facebook ➜ WeRateDogs [email protected] | nonprofit: @15outof10 ⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀
2.7M Followers 3K FollowingResearch, News, and Commentary from Nature, the international science journal
For daily science news, get Nature Briefing: https://t.co/wGmQlQ8a4D