I-JEPA: Learning from images by predicting missing features
DINOv2: How self-supervised visual features transfer across tasks
🎭 Masked autoencoder (MAE) for visual representation learning. From the author of ResNet.
🌀 MLP-Mixer: How image patches communicate without attention
🦖 DINO: Self-supervised ViTs learn strong features and semantic structure
CoLT5: Reading long documents with selective computation
Decision Transformer: Unifying sequence modelling and model-free, offline RL
🌀 MLP-Mixer: How image patches communicate without attention
🦖 DINO: Self-supervised ViTs learn strong features and semantic structure
ERNIE 2.0: Continual multi-task pre-training for language understanding