Now this paper has been accepted to #ICLR2023 as a spotlight🌟! I truly appreciate everyone's efforts. Check out our repo github.com/keyu-tian/SparK for the latest demo video and updates! Title: "Designing BERT for Convolutional Networks: Sparse and Hierarchical Masked Modeling"
How can we leverage successful pretraining techniques from transformers to improve purely convolutional networks? The answer is *Sparse Convolutions*!
Let's see what happens when purely convolutional networks are pretrained with 1.28 million unlabeled images ...
1/7
SparK - [ICLR'23 Spotlight] The first successful BERT-style pretraining on any *convolutional network*; Pytorch impl. of "Designing BERT for Convolutional Networks: Sparse and Hierarchical Masked Modeling" github.com/keyu-tian/SparK
@foozlefoo Thanks, we're glad SparK may be helpful to your project! Sorry we delayed the update for a few days because it's Lunar New Year today, we'll get back to you as soon as they're posted.
PS: our original submission in Oct. 2022 was titled "Sparse and Hierarchical Masked Modeling for Convolutional Representation Learning"; we've changed it to the new one to more clearly express our vision for bringing the powerful BERT-style self-supervised learning to CNNs.
@rasbt Thank you so much! Your notes are really helpful in letting people know about what we do! And I'm glad to see this work has the opportunity to make a meaningful contribution to our community.
@reidatcheson@rasbt Thanks for your interest! We are writing a google Colab tutorial to make it easy to play with our pre-trained models (e.g. to reconstruct the images you upload). We'll release it before this weekend and hope it can be helpful.
@rasbt Yeah, that's true for point 2. It'd be better to say "CNNs achieved higher performance than Transformers".
BTW, I really appreciate the contrastive idea, which seems more natural for image data than randomly masking. But it turns out the BERT-style pretraining is more effective.
How can we leverage successful pretraining techniques from transformers to improve purely convolutional networks? The answer is *Sparse Convolutions*!
Let's see what happens when purely convolutional networks are pretrained with 1.28 million unlabeled images ...
1/7
[CV] Designing BERT for Convolutional Networks: Sparse and Hierarchical Masked Modeling
K Tian, Y Jiang, Q Diao, C Lin, L Wang, Z Yuan [Peking University & Bytedance Inc & University of Oxford] (2023)
arxiv.org/abs/2301.03580#MachineLearning#ML#AI#CV
[1/2]
@HappyToKnowThat@jamesr66a it's kinda like the number of neurons in a brain, the more, the smarter. And each parameter is a float in Java, so you can imagine how much memory it takes.
🤩The first successful BERT-style #SelfSupervisedLearning on any convolutional network! #ResNet now enjoys masked autoencoding! 🚀A breakthrough paper "Designing BERT for Convolutional Networks: Sparse and Hierarchical Masked Modeling" by @keyutian et al.
deepai.org/publication/de…
@kzmolikova@rasbt Thanks for noticing this, that is our work which's been on OpenReview since Oct. 2022 (https://openreview. net/forum?id=NRxydtWup1S), and uploaded to arxiv recently.
It gives some different results from ConvNeXt v2, e.g., our method works pretty well on ConvNeXt v1 and ResNet.
152 Followers 871 FollowingUndergraduate from @Westlake_Uni, with research interests in Efficient AI.
ICLR 2026, CVPR 2026 Findings.
Actively seeking 27 Fall PhD opportunities!
16K Followers 4K FollowingExplorer of universes. Working on @ultrametricai. Prev founded @Golden & Heyzap. @ycombinator alum. Angel to many many unicorns.
339 Followers 813 FollowingPh.D. student in Artificial Intelligence and Brain-Computer Interface @med_umontreal and @mila_quebec | CEO of Hummingbird AI | Vice-chair @mtlieee | 🧠🔁🖥️
9 Followers 52 FollowingResearch Scientist of ByteDance. Intern at Google, ByteDance, and Microsoft Research Asia. My research interests are interactive world models and robotics.
119 Followers 332 FollowingFinal-year AI PhD. Student at @Westlake_Uni and @ZJU_China || Advisor: Stan Z. Li || Research interest: AIGC, UMM, Network architecture, optimization, and AI4S.
244 Followers 2K FollowingThen: particle physics @CMSExperiment @cern. Now: @DSI_UChicago deep learning + medical imaging, cat whisperer
good analogies are my love language
496K Followers 1K FollowingML/AI research engineer. Ex stats professor.
Author of "Build a Large Language Model From Scratch" (https://t.co/O8LAAMRzzW) & reasoning (https://t.co/5TueQKx2Fk)