I-JEPA: Learning from images by predicting missing features

A practical guide to "Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture"
article cover

I-JEPA learns visual features by predicting the representations of hidden image regions, rather than reconstructing their pixels. This guide explains the context encoder, target encoder, predictor, and masking strategy, then examines what the paper's accuracy and efficiency results actually measure.

View comments.

more ...

QLoRA: How to fine-tune large language models with less memory

A practical guide to "QLoRA: Efficient Finetuning of Quantized LLMs"
article cover

QLoRA makes large language models cheaper to fine-tune by storing the base model in 4 bits and training small, higher-precision adapters. This guide explains LoRA, NormalFloat, double quantization, and paged optimizers, then puts the Guanaco chatbot results in context.

View comments.

more ...

DINOv2: How self-supervised visual features transfer across tasks

A practical guide to "DINOv2: Learning Robust Visual Features without Supervision"
article cover

DINOv2 trains Vision Transformer encoders on 142 million curated images without using labels or captions in its pretraining objective. This guide explains how image-level self-distillation, masked patch prediction, feature-spreading regularization, data curation, and distillation combine to produce features that transfer across classification, retrieval, segmentation, and depth estimation.

View comments.

more ...

CoLT5: Reading long documents with selective computation

A practical guide to "CoLT5: Faster Long-Range Transformers with Conditional Computation"
article cover

CoLT5 processes every input token with lightweight layers and gives selected tokens additional, higher-capacity computation. This guide explains its learned routing, light and heavy branches, faster decoding, and experiments with inputs up to 64k tokens.

View comments.

more ...

LoRA: Fine-tuning a model by learning a small update

A practical guide to "LoRA: Low-Rank Adaptation of Large Language Models"
article cover

LoRA adapts a pretrained model by training small, low-rank updates while keeping its original weights fixed. This guide explains the two-matrix construction, the memory and storage savings, the conditions for merging adapters, and the connection to QLoRA.

View comments.

more ...

🎭 Masked autoencoder (MAE) for visual representation learning. From the author of ResNet.

"Masked Autoencoders Are Scalable Vision Learners" - Research Paper Explained
article cover

A masked autoencoder (MAE) learns visual representations by reconstructing missing image patches from a small visible subset. It divides an image into regular non-overlapping patches, samples patches uniformly without replacement, removes the masked patches before the encoder, and inserts learned mask tokens only for the lightweight decoder. With a 75% masking ratio, the encoder processes just 25% of the patches. This asymmetric design reduces training time and memory, enabling ViT-Large and ViT-Huge models to scale on ImageNet-1K. A ViT-Huge model pretrained for 1600 epochs and fine-tuned at 448-pixel resolution reaches 87.8% ImageNet-1K top-1 accuracy.

View comments.

more ...

Decision Transformer: Unifying sequence modelling and model-free, offline RL

"Decision Transformer: Reinforcement Learning via Sequence Modeling" - Research Paper Explained
article cover

Decision Transformer casts offline reinforcement learning (RL) as conditional sequence modeling. A causally masked GPT-style Transformer predicts each action from a desired return-to-go, the current state, and the recent trajectory. It avoids value-function bootstrapping and policy-gradient optimization during training, yet matches or exceeds several strong offline RL baselines on the benchmarks studied in the paper.

View comments.

more ...

🌀 MLP-Mixer: How image patches communicate without attention

A practical guide to "MLP-Mixer: An all-MLP Architecture for Vision"
article cover

MLP-Mixer classifies images without convolution or self-attention. This guide follows an image through patch projection, token mixing, channel mixing, and classification, then examines the paper's accuracy, throughput, scaling, and permutation experiments.

View comments.

more ...

🦖 DINO: Self-supervised ViTs learn strong features and semantic structure

"Emerging Properties in Self-Supervised Vision Transformers" - Research Paper Explained
article cover

Self-distillation with no labels (DINO) trains a student network to match an exponential-moving-average teacher across differently augmented views of the same image. Applied to Vision Transformers, it produces features that perform strongly with simple linear and k-nearest-neighbor classifiers. Its final-layer self-attention also reveals object boundaries without segmentation labels, although DINO is not itself a segmentation model.

View comments.

more ...

RL Primer

Explaining the fundamental concepts of Reinforcement Learning
article cover

The objective of RL is to maximize the reward of an agent by taking a series of actions in response to a dynamic environment. Breaking it down, the process of Reinforcement Learning involves these simple steps: Observation of the environment, deciding how to act using some strategy, acting accordingly

View comments.

more ...