I-JEPA: Learning from images by predicting missing features
A practical guide to "Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture"
I-JEPA learns visual features by predicting the representations of hidden image regions, rather than reconstructing their pixels. This guide explains the context encoder, target encoder, predictor, and masking strategy, then examines what the paper's accuracy and efficiency results actually measure.
more ...
Michał Chromiak's blog