Multimodal pre-training and post-training. final-year PhD student at Meta FAIR @AIatMeta and Oxford @OxfordTVGjunlinhan.github.io London, EnglandJoined July 2020
1/ Sharing a new, interesting project I did during my internship at Microsoft AI Frontiers w/ @JohnCLangford
TL;DR: At decoding time, we feed **previous hidden state** into the input together with token embedding, and it boosts performance for free.
arxiv.org/abs/2608.08888
We live in a multimodal world. We see, talk, act, and dream.
Yet most LLMs still start with language pretraining. Why not train them natively with multimodal I/O from scratch?
Because it’s SUPER HARD, adding modalities triggers training instability, design complexity, and often modal competition
So what’s the path forward?
Introducing: Towards Physics of Multimodal Pretraining (junlinhan.github.io/projects/physi…)
We unpack the underlying mechanics of multimodal pretraining across 4 aspects: Knowledge Flow, Modality Synergy, Early Unification, and Recipe.
This is great and super relevant! I really like this hypothesis, it also aligns closely with our findings. Since most vision-language pre-training relies on captioning tasks, which are relatively 'easy,' this mechanism makes a lot of sense. Please let me know once the paper is on arXiv or published elsewhere; we’d love to discuss and highlight it in our work!
Thank you Dinkar!! Wow those are very great questions. So far my thoughts:
1. Formation vs persistence:
Most representation is learned in pre-training, we do see that models with higher vision activations retain them after SFT (we used the Cambrian-7M dataset). Generally I think keeping a small ratio of vision data during post-training can prevent these circuits from being destroyed.
2. Task-level impact (e.g., shuffled patches):
We do see a positive correlation on task accuracy. Models with higher vision activation (like early fusion vs. late fusion under the same vision token budget in Tab 2) perform better on VQA and generation. We haven't tested shuffled patches specifically yet, though!
3. Recovering from laziness via fine-tuning:
Probably possible with the right data mix/distribution and very long time training? Though we'd need to avoid regressing language capability...
Thanks Xuhui for sharing this work! It looks super interesting! Yeah, aligning the language space (discrete) with the vision space (usually continuous) is very valuable! There are also a couple of works in visual generation that leverage this idea as well. Curious to see how it works when scaled up!
@ahlawat_varun Thanks! We highlighted Kimi a couple of times in the paper as representative work in early fusion (first time right at the beginning of the introduction). Sorry X posts don't leave much room for texts so I didn't write any related works. Thanks again for pointing it out!
@ubuto23@_amanda_long Yeah, I think some frontier labs did a lot of exploration on the visual understanding side of unified pretraining (and audio as well). So our value is also on the generation side to provide a more holistic view and a deeper understanding of the underlying mechanisms.
@dadadaistt Yea! This is also confirmed in kimi-k3 (for understanding side)! Competition we might still need some arch level or even objective level (diff vs ar) explorations.
5K Followers 1K FollowingPostdoc @berkeley_ai, building Impossible. Prev. @MPI_IS @AdobeResearch @uni_tue Interested in recreating the physical world and the intelligence behind it.
34 Followers 137 FollowingMultimodal LLMs for Drug Discovery @ D.E. Shaw Research. 🇫🇷🇺🇸
Previously led ML@Anagenex, Founder@PoissonAI, PhD@Columbia, Head Trader@SGCIB
739 Followers 2K Following画像生成に興味がある九大D3 / バイオ&医療AIにも興味あり/ 視覚若手の会LENS (@lens_kousiki)
PhD student of Kyushu University / Medical imaging, Generative Model
34 Followers 137 FollowingMultimodal LLMs for Drug Discovery @ D.E. Shaw Research. 🇫🇷🇺🇸
Previously led ML@Anagenex, Founder@PoissonAI, PhD@Columbia, Head Trader@SGCIB
2K Followers 1K FollowingCo-Founder & CEO @ViggleAI. Building spatial intelligence. I write about why AI demos work, where they cheat, and what it takes to turn them into products.