A striking difference between language and vision is that training on LLM outputs works really well. While pixel outputs of video models are still useful (e.g. as policies with an IDM or for evals), they seem to be difficult to directly train on. So, whither generated pixels?
Beginning to start taking the idea that I should take ideas seriously, seriously. Based on my forecasts, I'll take my first idea seriously in ~ten years.
New blog post:
tomsilver.github.io/blog/2026/now-…
This is the second in a "series" of posts about how to define and work on research problems. (The first post was 3 years ago...)
This one addresses the agentic elephant in the room.
@astridwilde1 prior to ~2025 most synthetic stereo depth datasets didn't randomize the camera baseline (and it's really important for generalization, turns out)😆
Every vision task eventually hits a critical threshold where qualitative visualizations get good enough and all the new papers start spamming them. Unfortunately, soon after this threshold, human perceptual similarity becomes saturated.
Image foundation models never stop surprising!😮
And this time, it's not the usual big tech players.
A few days ago, @robbyant_brain released LingBot-Vision, a new family of vision encoders built natively for dense spatial perception.
@mean_field_zane@alz_zyd_ Agreed, I've looked at Princeton's historical first-destination data before and I don't believe there's a significant downward trend (even for CS).
@flawedaxioms@AaronBergman18@slatestarcodex For a "posthuman" society, the diff wouldn't mean much. Bostrom's original napkin math was that a planet-sized computer might have around 10^41 FLOPS, and that simulating all minds in all of human history is likely around 10^36 FLOPS.
744 Followers 648 FollowingPhD Candidate @ UIUC | Research Intern @ World Labs | Ex-Snap|Ex-Stability AI | Working on 3D/4D Reconstruction/Generation. On job market next year early.
2K Followers 2K FollowingAssistant Prof. @ChicagoBooth | OM & Applied AI | Most Cited Paper: “Neural Collapse in Deep Net Training” | BSE @Princeton, PhD @Cornell, MS/PostDoc @Stanford
3K Followers 692 FollowingIncoming Assistant Professor @UMRobotics
Senior Research Scientist @NVIDIA
Foundation Models for Robots, Dexterous Manipulation
ex-@Princeton, ex-@IITKanpur
2K Followers 707 Followingcs phd @stanford | prev. @eth | I work on long context, continual learning, generative models, and large-scale infrastructure | opinions are my own
5K Followers 301 FollowingCo-founder and CEO at Elorian AI.
Deep learning/LLM principal researcher.
Ex (director) at Google DeepMind. Gemini Data Area Lead. Chinese name: 戴明博
581 Followers 757 FollowingJust a simple, high agency, prophetic young woman who wants to be an economist. HKS, Berkeley, and unfortunately, San Francisco.
744 Followers 648 FollowingPhD Candidate @ UIUC | Research Intern @ World Labs | Ex-Snap|Ex-Stability AI | Working on 3D/4D Reconstruction/Generation. On job market next year early.
2K Followers 2K FollowingAssistant Prof. @ChicagoBooth | OM & Applied AI | Most Cited Paper: “Neural Collapse in Deep Net Training” | BSE @Princeton, PhD @Cornell, MS/PostDoc @Stanford