@niksteel123@ayushjain1144 We haven’t considered using game engines for capturing scene + temporal data - we are mostly using open source 3D benchmarks and simulators. We do have a follow up project on the 4D scenes, @ayushjain1144 is leading this
🚀Excited to share Qwen-3D (ECCV 2026) — a generalist 3D vision–language model that does grounding, segmentation, VQA, and spatial reasoning in one place. VLMs are great at images and short clips, but long videos bury them in tokens. Our idea: put multi-view frames into a shared 3D world, so attention scales with what’s in the scene, not how many frames you feed it.
📄arxiv.org/abs/2608.02980
🌐qwen-3d.github.io
So why doesn’t everyone just do this? The data just isn't there. We’ve got ~1B images… and maybe ~10k posed 3D scenes. Pure 3D-first models don’t scale today. So Qwen-3D is a unified 2D–3D model: shared parameters, shared token space, trained with both. Keep the scale of 2D data; add geometry where we have it.
We may receive 2D frames from sensors, but the world is inherently 3D - many of the frames we receive are redundant due to overlapping views.
Once you lift RGB-D into world space, those redundant views collapse into persistent scene tokens. Same room, 400 frames or 75 — similar token budget.
Having a model that runs on these kinds of inputs wins you 📦Compactness · 🗺️Spatial reasoning · ⏱️Temporal consistency
10K Followers 3K FollowingVision, neural networks, and open brain science. RT means "This may deserve more attention". Like means "I have read this (e.g. paper) and think it's solid."
0 Followers 136 FollowingAlgorithm Enginieer at Alibaba Group|Engineer in graphics, decoding the art behind signals|Exploring LLMs and the path toward AGI
3K Followers 166 FollowingAssistant Professor @mldcmu. Formerly: Postdoc @MITEECS, PhD @Berkeley_EECS, Math Undergrad @Princeton. New to Twitter. https://t.co/67bMOAyqK6
1K Followers 248 Following@Roblox Senior AI Scientist @Adobe Previous Research Scientist @USC Ph.D. Alumni @sjtu1896 Alumni
3D Representations, 4D Humans, Avatar, Digital Human
600K Followers 56K FollowingSan Francisco/Silicon Valley AI | Robots, holodecks, BCIs, analysis of new things | Ex-Microsoft, Rackspace, Fast Company | Wrote eight books about the future.
353 Followers 259 FollowingPhDing @CMU_Robotics @NVIDIAAI. Previous: Undergrad @PKU1898. I work on #Robotics #3DVision #MachineLearning. Opinions are my own.
8K Followers 2K FollowingAssistant Professor @CMU_Robotics @SCSatCMU, leading the @LeCARLab. @Amazon Scholar @ FAR. I build generalist robots with agility. Ph.D. from @Caltech.
5K Followers 1K FollowingAssistant Professor at UPenn. Research interests: Neural Scene Representation, Neural Rendering, Human Performance Modeling and Capture.
19K Followers 496 FollowingAssociate Professor @UofT, Vice President of AI Research @nvidia, founding member of @VectorInst. Computer vision, deep learning, 3D. Opinions are my own.
3K Followers 142 FollowingPrincipal Scientist at Wayve. Associate Professor at Simon Fraser University. 3D Vision.
(Eiko-high, U-Tokyo, UIUC, ILM, ENS, UW, Google, WUSTL, SFU/Wayve)