Paper Reading Notes
OctoSense — Multimodal Robot Perception
Notes on a late-fusion masked autoencoder that learns from many real-world sensors at once, each with its own rate, latency and noise.
OctoSense: Self-Supervised Learning for Multimodal Robot Perception
In one line: an open-source multi-sensor rig (stereo RGB + event cameras, LiDAR, thermal, IMU, RTK-GPS, proprioception) and a 59-hour synchronized driving dataset, used to train a late-fusion masked autoencoder with per-modality tokenizers. It is fast (6.68 ms on RTX 5090, 112 ms on Orin NX), beats image-only foundation models on optical flow / depth / segmentation / ego-motion, and stays robust at night and under degraded sensors.
- Walk through the sensor suite, from the stereo RGB cameras down to the quadruped robot’s proprioception.
- What does “sensors have different representations” mean — a concrete example?
- What is the difference between an autoencoder and an encoder? And for “caching modality-specific tokens at inference time”: is memory a consideration, and are those tokens produced by the modality-specific tokenizers?
- What exactly is a “new measurement”? Does it fall within the scope of one modality’s tokens, or does it imply out-of-distribution input?
- Explain the task terms from “optical flow” through “steering angle”.
The abstract packs three things into a few sentences — a sensor platform, a dataset, and a method. The five questions below walk through the platform’s sensors, why their data is hard to combine, and what the method does with them.
1 · The sensor suite
The platform fuses seven kinds of signal. The first six are exteroceptive (they sense the world); the last is proprioceptive (it senses the robot’s own state). The whole rig is mounted on two bodies — a car and a four-legged (quadruped) robot — which is why proprioception takes two forms.
| Modality | What it senses | What it gives you |
|---|---|---|
| Stereo RGB | Color light through two cameras with a known baseline | Two color images; depth from left/right disparity |
| Event camera | Per-pixel brightness changes, asynchronously | Microsecond-latency event stream; high dynamic range, no motion blur |
| LiDAR | Laser time-of-flight distance | 3D point cloud of the scene; works in darkness |
| Thermal camera | Long-wave infrared (heat) | Heat image; sees warm objects (people, engines) at night |
| IMU | Linear acceleration + angular velocity | High-rate motion/orientation signal (accelerometer + gyroscope) |
| RTK-GPS | Satellite positioning with real-time kinematic correction | Centimeter-level global position (vs. meters for plain GPS) |
| Proprioception | The robot’s own internal state | Car: CAN-bus data (wheel speed, steering, throttle). Quadruped: joint angles |
Proprioception is the term for a system’s sense of its own configuration and motion — the machine analogue of knowing where your limbs are without looking. A quadruped robot is a four-legged robot (e.g. Spot- or Unitree-style); its joint angles are the natural proprioceptive readout, just as a car’s CAN bus is for the vehicle.
2 · Why “different representations”
The phrase means each sensor’s raw data has a fundamentally different structure — not just different numbers. A standard vision model assumes a dense pixel grid, but most of these sensors are not grids at all. Concretely:
| Modality | Representation (structure) | Typical rate |
|---|---|---|
| Stereo RGB | dense grid [H×W×3] | ~30 Hz, regular frames |
| Event camera | sparse async events (x, y, t, polarity) | µs resolution, not framed |
| LiDAR | unordered 3D point set {(x,y,z,i)} | ~10 Hz, variable point count |
| Thermal | single-channel grid (temperature) | ~30 Hz, often lower res |
| IMU | low-dim time series (6 values) | ~100–1000 Hz |
| RTK-GPS | few scalars (lat, lon, alt) | low rate |
| Proprioception | vector of joint angles / CAN signals | varies |
A worked contrast: an RGB frame can be split into patches and embedded the way a ViT does. A LiDAR scan is an unordered set of points — there is no fixed grid to patch, and the model must be permutation-invariant. An event camera doesn’t even produce frames. On top of structure, the abstract also lists differing frequencies (10 Hz vs. 1000 Hz), latencies, and noise. This heterogeneity is exactly why the model needs a separate tokenizer per modality rather than one shared input layer.
3 · Autoencoder vs. encoder, and the token cache
An encoder is just the “compress” half: it maps an input to a latent representation. An autoencoder adds a decoder and is trained to reconstruct its own input through a reconstruction loss. A masked autoencoder (MAE) hides part of the input, encodes only the visible part, and asks the decoder to predict the hidden part — which needs no labels, so it is self-supervised. That is OctoSense’s training signal. After training, the encoder is the part kept to produce representations for downstream tasks.
| Has a decoder? | Trained by | Reused downstream | |
|---|---|---|---|
| Encoder | no (it is the encoder) | n/a (a component) | yes — its latent output |
| Autoencoder | yes | reconstructing the input | usually just the encoder half |
| Masked AE | yes | predicting masked-out input (self-supervised) | the encoder half |
Caching modality-specific tokens
Because the sensors run at different rates and arrive asynchronously, at any instant only a few modalities have produced fresh data. The system therefore recomputes tokens only for the modality that just updated and reuses the previously computed tokens for the others.
Late fusion means each sensor is tokenized on its own — so a token is a per-modality unit that can be cached, and only the sensor that just fired needs to be recomputed.
Is memory a consideration? The cache does cost memory — you are storing tokens between updates — so this is fundamentally a memory-for-compute trade. But what the abstract advertises is the speed/latency payoff (6.68 ms / 112 ms) and the ability to ingest asynchronous streams, not a memory saving. Are the cached tokens tokenizer-derived? Yes: the cached units are precisely the outputs of the modality-specific tokenizers. “Late” fusion means fusion happens after tokenization, which is what makes each modality’s tokens a clean, independently cacheable and updatable unit.
4 · What “a new measurement” is — and why it is not OOD
A “new measurement” is simply the next sample to arrive in time from some sensor. Since the sensors have different frequencies and latencies, they do not update together: a new camera frame (~30 Hz), a new LiDAR sweep (~10 Hz), and a new IMU sample (~200 Hz) all show up on their own schedule. Each new measurement is tokenized by its own modality’s tokenizer and updates that modality’s slot in the cache — so yes, it falls squarely within the scope of one modality’s tokens.
It does not imply out-of-distribution input. Here “new” is temporal (streaming / online arrival), not distributional (novel or anomalous). The separate robustness claim in the abstract — working at night and under degraded sensors — is the closest thing to a distribution-shift story, but that is handled by cross-modal redundancy, and is a different idea from processing “new measurements as they come.”
5 · The downstream task terms
These are the perception tasks OctoSense’s representation is evaluated on, where it reportedly beats image-only foundation models.
| Term | What it predicts |
|---|---|
| Optical flow | Per-pixel 2D motion between consecutive frames (direction + speed of each pixel) |
| Depth | Per-pixel distance from the camera to the scene — a depth map (3D structure) |
| Semantic segmentation | A class label for every pixel (road, car, pedestrian, building, …) |
| Ego-motion | How the platform itself moves through space — the “ego” vehicle/robot’s own motion |
| — Translation | The displacement of the platform (x, y, z) |
| — Rotation | The change in the platform’s orientation (roll / pitch / yaw) |
| Steering angle | The car’s steering/wheel angle — a control quantity predicted from perception |
The first three (optical flow, depth, segmentation) describe the scene; ego-motion and steering angle describe the platform’s own movement — the latter tying perception back to the proprioceptive signals from question 1.
project page: abisulco.com/octosense
context: masked autoencoders (MAE) · event cameras · LiDAR point clouds · RTK-GNSS · proprioception (CAN bus / joint angles)
- Introduce self-supervised learning (SSL), and DINO, SigLIP, Hiera, and vision-based foundation models.
- What are the details of the Global Robotics Technology Roadmap?
- What does “on par” mean?
1 · Self-supervised learning and the vision foundation-model family
Self-supervised learning (SSL) learns representations from unlabelled data by constructing the supervision signal from the data itself (a “pretext” task). Two broad families dominate vision: joint-embedding / contrastive methods (SimCLR, MoCo, BYOL, DINO) that pull together different views of the same image, and generative / masked-prediction methods (MAE, BEiT) that reconstruct hidden parts of the input. (Refined in Entry 03: DINO and BYOL are joint-embedding but non-contrastive — no negatives.) The payoff is a strong general-purpose encoder trained without human labels — the same regime OctoSense’s masked autoencoder belongs to.
DINO / DINOv2
DINO (“self-distillation with no labels”, Caron et al., Meta, 2021) trains a self-supervised ViT with a student–teacher setup where the teacher is an exponential moving average of the student, using multi-crop views plus centering & sharpening to avoid representational collapse. Its emergent properties are famous: attention maps segment objects, and k-NN on its features works well. DINOv2 (2023) scales this up on curated data to yield off-the-shelf visual features — a canonical image-only foundation model.
SigLIP
SigLIP (“Sigmoid Loss for Language-Image Pre-training”, Zhai et al., Google, 2023) is a CLIP-style image–text model, but it swaps the softmax/InfoNCE contrastive loss for a pairwise sigmoid loss that needs no global normalization over the batch. That makes training more efficient, viable at smaller batch sizes, and stronger overall. It is weakly supervised (via image–text pairs) and produces a powerful image encoder.
Hiera
Hiera (Ryali et al., Meta, ICML 2023) is a hierarchical (multiscale) ViT that removes the specialized modules of Swin/MViT and instead lets spatial structure be learned through MAE pretraining. The result is simpler, faster, and more accurate, for both images and video — the same masked-autoencoder lineage as OctoSense.
Vision-based foundation models
An umbrella term for large models pretrained (self- or weakly-supervised) on massive data to give general visual representations transferable to many downstream tasks. “Image-only” ones (DINOv2, MAE, Hiera) use images alone; image–text ones (CLIP, SigLIP) add language; others target specific outputs (SAM for segmentation). OctoSense positions itself as a multimodal foundation model that outperforms these image-only models on robotics tasks.
| Model | Supervision | Core mechanism | Notable for |
|---|---|---|---|
| DINO / DINOv2 | self-supervised | student–teacher self-distillation (EMA teacher) | general visual features, emergent segmentation |
| SigLIP | image–text (weak) | contrastive with pairwise sigmoid loss | efficient, strong zero-shot encoder |
| Hiera | self-supervised (MAE) | hierarchical ViT, no specialized modules | simple/fast, images & video |
| MAE (family) | self-supervised | mask input, reconstruct hidden part | the lineage OctoSense builds on |
2 · The Global Robotics Technology Roadmap (2025–2035)
A position/perspective document by Henrik I. Christensen (UC San Diego), version 1.02, April 2026, covering Europe, Asia, and the United States over the 2025–2035 decade. It synthesizes state-of-the-art research (drawing on ICRA, IROS, RSS, CoRL, and ML venues like NeurIPS and ICML) with industry data and regional government strategies, aimed at policymakers, strategists, research agencies, and industrial R&D leaders.
Its structure runs from a global market baseline (industrial, service, and R&D investment) through the academic state of the art — embodied AI and foundation models for robotics, VLA systems, RL and sim-to-real, navigation, dexterous manipulation and tactile sensing, legged/bio-inspired locomotion, multi-robot systems, soft robotics — then enabling technologies (computing, perception & sensing, materials), regional strategies, a 2025–2035 roadmap across algorithms/hardware/materials/systems, sector analyses (manufacturing, logistics, healthcare, agriculture, mining, construction), and cross-cutting themes (the humanoid convergence race, sustainability, workforce impact, geopolitical risk).
| Region | Strategic focus (per the roadmap) |
|---|---|
| Europe | Safety, compliance, collaborative robots (cobots); shaped by the EU AI Act |
| Asia | Dominant industrial deployment (~74% of 2024 installations; China ~54%); humanoid scale-up |
| United States | AI-powered autonomy, defense robotics, foundation models / VLA |
- VLA models are framed as the most consequential algorithmic advance of the period — the first to enable cross-embodiment generalization.
- Humanoids are projected to grow from roughly $370M (2025) to about $6.5B (2030), with Chinese OEMs and US tech firms racing to scale.
- Soft robotics (liquid-crystal elastomers, electroactive polymers, self-healing hydrogels) is bridging rigid industrial systems and bio-compatible medical devices.
- Regulatory asymmetry (e.g. the EU AI Act) is flagged as a key geopolitical variable.
Where OctoSense fits: a multimodal self-supervised foundation model for perception sits on the roadmap’s “foundation models for robotics” and “perception & sensing” lines — the data/representation layer beneath the VLA and autonomy trends the roadmap highlights.
3 · “on par”
“On par (with)” means at the same level / equal in quality / comparable (from golf’s “par”). In a paper it typically appears as “our method is on par with X on task Y, while being faster” — i.e. it matches X, not beats it. The key nuance: “on par” signals parity; claiming superiority uses “better than” or “outperforms.”
models: DINO / DINOv2 (Caron 2021; Oquab 2023) · SigLIP (Zhai 2023) · Hiera (Ryali 2023) · MAE (He 2022) · CLIP / SAM
paper context: arXiv:2606.27317 (OctoSense)
- Explain “joint-embedding / contrastive” — isn’t that the opposite of contrastive learning? Is “generative” the same as diffusion?
- More detail on the representative models; tables comparing the families, then models within a family.
- Why does representation collapse happen, and what else prevents it? Draw the DINO pipeline. What are multi-crop views?
- The “emergent properties” sentence (attention segments objects, k-NN works); what does “off-the-shelf features” mean — entirely new features unlike the training set?
- Explain CLIP-style; the form of InfoNCE; the form of the pairwise sigmoid loss; can the “no global normalization” idea be reused elsewhere — have others used it?
- Does “hierarchical (multiscale)” mean cascaded? What does “removes the specialized modules of Swin/MViT” mean — does ViT itself contain them? Is MAE self-supervised?
- What are sim-to-real, dexterous manipulation & tactile sensing, and locomotion? Does “on par” mean “on average”?
1 · The SSL taxonomy (a correction)
Joint-embedding is the architecture: encode two views of the same image and make their embeddings agree in feature space. Contrastive is one objective within it — the one that uses negatives. There is also a non-contrastive branch (no negatives). So pairing “joint-embedding / contrastive” was loose: DINO and BYOL are joint-embedding but non-contrastive — the suspicion was right.
And “generative” here means masked reconstruction in input space (MAE, BEiT), not diffusion. Diffusion is a separate generative paradigm (iterative denoising); it can be used for representation learning, but it is not what this family refers to.
| Family | Compares in | Negatives? | Avoids collapse via | Examples |
|---|---|---|---|---|
| Contrastive (joint-embedding) | embedding space | yes | repulsion from negatives | SimCLR, MoCo |
| Non-contrastive (joint-embedding) | embedding space | no | asymmetry / EMA / stop-grad, or decorrelation | BYOL, DINO, SimSiam, Barlow Twins, VICReg |
| Generative / masked | input space (reconstruct) | no | target is the data itself | MAE, BEiT |
2 · Representative models
| Model | Group / yr | Key trick | Negatives | Anti-collapse |
|---|---|---|---|---|
| SimCLR | Google ’20 | strong aug + projection head + NT-Xent (InfoNCE) | in-batch | negatives |
| MoCo | FAIR ’20 | momentum encoder + negative queue | queue | negatives + momentum |
| BYOL | DeepMind ’20 | online net predicts EMA target via a predictor | none | predictor asymmetry + EMA + stop-grad |
| SimSiam | FAIR ’21 | like BYOL without EMA; stop-grad is the key | none | stop-grad + predictor |
| DINO | Meta ’21 | self-distillation, EMA teacher, multi-crop | none | centering + sharpening (+ EMA) |
| Barlow Twins / VICReg | ’21 | cross-correlation → identity / variance-invariance-covariance | none | decorrelation / variance regularization |
| Masked / generative | Idea | Reconstruction target |
|---|---|---|
| MAE | mask ~75% of patches; ViT encodes the visible part; light decoder rebuilds | raw pixels |
| BEiT | mask patches; predict discrete visual tokens | codebook token ids |
3 · Collapse, and the DINO pipeline
Why collapse happens: if the only goal is “make two views agree,” a trivial solution is to output the same constant for every input — the views agree perfectly and the loss is zero, but the features are useless. With no negatives or other safeguard, that degenerate solution is the easy one to fall into.
DINO’s safeguards: an EMA teacher with stop-gradient (asymmetry) makes the student chase a slow-moving target; centering (subtract a running mean of teacher outputs) blocks one dimension from dominating; sharpening (a low teacher temperature) blocks collapse to a uniform distribution. Centering and sharpening counter two opposite collapse modes and are used together. Other anti-collapse tools: contrastive negatives (SimCLR, MoCo), stop-grad + predictor (BYOL, SimSiam), decorrelation/variance terms (Barlow Twins, VICReg), feature whitening.
Multi-crop: from one image, make a few global crops (large, high-res, >50% of the image) and many local crops (small, low-res). The student sees all crops; the teacher sees only global crops, so the loss enforces a local→global correspondence — a cheap way to add views.
4 · Emergent properties & “off-the-shelf”
Attention segments objects: in a DINO-trained ViT, the [CLS] token’s self-attention lands on coherent object regions, so thresholding it yields unsupervised segmentation masks — though the model never saw segmentation labels. “Emergent” = not trained for, yet it appears. k-NN works: nearest-neighbour classification in feature space is accurate with no fine-tuning, which means the features are already semantically organized.
Off-the-shelf features means features you can use as-is: freeze the trained backbone and read out embeddings for a new task (k-NN, linear probe, or as input to a downstream model) without retraining the backbone. It does not mean the features are unrelated to the training data — they are learned from the training distribution but generalize to new images/tasks. The point is reusability without fine-tuning, not alien features.
5 · CLIP-style, InfoNCE, and the sigmoid loss
CLIP-style trains an image encoder and a text encoder jointly on image–text pairs. For a batch of N pairs it builds an N×N similarity matrix and pushes matched pairs (the diagonal) up and mismatched pairs down via softmax cross-entropy (InfoNCE), giving an aligned image–text space usable for zero-shot classification and retrieval. SigLIP is CLIP-style with a different loss.
InfoNCE
For anchor i with positive i⁺ (cosine similarity, temperature $\tau$):
$$\mathcal{L}_i = -\log \frac{\exp\!\big(\operatorname{sim}(z_i, z_i^{+})/\tau\big)}{\displaystyle\sum_{j}\exp\!\big(\operatorname{sim}(z_i, z_j)/\tau\big)}$$
The denominator sums over the whole batch — the “global normalization”.
Pairwise sigmoid loss (SigLIP)
Label y_ij = +1 if matched (i=j) else −1; logit with learnable temperature t and bias b:
$$s_{ij} = t\,\langle v_i, t_j\rangle + b,\qquad y_{ij}=\begin{cases}+1 & i=j\\[2pt] -1 & i\neq j\end{cases}$$
$$\mathcal{L} = \frac{1}{|B|}\sum_{i}\sum_{j}\log\!\Big(1+\exp\big(-\,y_{ij}\,s_{ij}\big)\Big)$$
Each pair is an independent match/no-match binary decision, so there is no softmax across the batch — no all-gather of every embedding, no |B|×|B| matrix. That makes it memory-efficient and scalable to very large or very small batches; the bias b handles the heavy negative-to-positive imbalance.
Reusable elsewhere? Yes. The “pairwise sigmoid, no global normalization” idea has been picked up beyond SigLIP: SigCLR applies a sigmoid loss to pure visual representation learning (reusing SigLIP’s chunked implementation to avoid costly all-gathers and the |B|×|B| matrix); SCS-SupCon brings a sigmoid pairwise loss with learnable temperature/bias into supervised contrastive learning to reduce negative-sample dilution; and cross-modal retrieval work swaps softmax for a SigLIP-style sigmoid to drop batch-global normalization. Broadly, any contrastive / retrieval / metric-learning task where batch-global softmax is the bottleneck can consider it.
6 · Hiera, Swin/MViT, and MAE
Hierarchical (multiscale) ≠ cascaded. Multiscale means one backbone split into stages that progressively downsample — high spatial resolution / few channels early, low resolution / many channels late (a feature pyramid). “Cascaded” usually means separate models/stages chained together (e.g. Cascade R-CNN). Plain ViT is isotropic: one resolution throughout, no pyramid.
“Removes the specialized modules of Swin/MViT”: plain ViT does not contain Swin/MViT modules. Swin and MViT are ViT variants that add modules to get hierarchy/efficiency (Swin: shifted-window local attention + patch merging; MViT: pooling attention). Hiera’s claim is that those hand-designed modules are unnecessary: take a simple multiscale skeleton, strip the extras (relative position embeddings, convs, shifted windows), and let strong MAE pretraining teach the spatial inductive biases — simpler, faster, more accurate.
MAE is self-supervised: yes — mask ~75% of patches and reconstruct missing pixels, with no labels.
7 · Robotics terms, and “on par”
- Sim-to-real: train in simulation (cheap, safe, abundant data), then deploy on a real robot; the “reality gap” is bridged with domain randomization, domain adaptation, system identification, and real-data fine-tuning.
- Dexterous manipulation: fine, contact-rich skill with multi-fingered hands (in-hand reorientation, varied grasping, tool use), beyond simple grippers. Tactile sensing gives the robot touch (force, slip, texture, local geometry) to complement vision, which is occluded during contact.
- Locomotion: moving the robot’s own body through the world (walking, running, climbing) — legged/bio-inspired locomotion is the four-legged or humanoid case; distinct from manipulation, which moves objects.
- “On par” ≠ “on average”: it means comparable / equal in level. “On average” is a statistical mean. Parity uses “on par”; superiority uses “better than / outperforms.”
losses: CLIP (Radford 2021) · InfoNCE (Oord 2018) · SigLIP (Zhai 2023); sigmoid-loss reuse: SigCLR (2410.17427) · SCS-SupCon (2512.17954)
backbones: ViT · Swin · MViT · Hiera (Ryali 2023)
- What are the “two views” and what decides them? Is the embedding space a shared representation space, and how is alignment achieved?
- Is representation collapse about failing to distinguish important features / weighting them equally? How does a core–periphery idea (pull high-similarity together, push low-similarity apart) relate to contrastive learning?
- Explain NT-Xent (InfoNCE), the momentum encoder and its update, the memory bank / negative queue, the online network (vs. test-time adaptation), EMA, and the predictor/EMA/stop-grad “asymmetry”.
- How are DINO’s centering + sharpening, Barlow Twins’ cross-correlation, and VICReg’s three terms actually implemented?
- Confirm MAE; for BEiT, is the whole image masked, and why predict discrete (not continuous) visual tokens?
1 · Two views, a shared space, and alignment
The two views are two random augmentations of the same image — random resized crop, flip, color jitter, grayscale, blur, solarization. The criterion: augmentations must preserve semantics (still the same object/scene) while changing only low-level appearance, so the model learns invariance to those nuisances. Two views of one image form a positive pair. In DINO they also come in two sizes (global / local — multi-crop).
The embedding space is shared: both views pass through the same (or EMA-coupled) encoder, usually followed by a projection head, into a $d$-dimensional space whose vectors are L2-normalized (so similarity = cosine = dot product). For CLIP/SigLIP the two modalities use two encoders that map into one common space. Alignment is produced by the loss: raise the similarity of matched views/pairs (and, for contrastive methods, lower it for mismatched ones).
2 · What collapse really is (and the core–periphery link)
Collapse is not about weighting features equally — it is the representation becoming degenerate / uninformative. Two modes: complete collapse (every input maps to the same constant vector, so nothing is distinguishable) and dimensional collapse (embeddings squeeze into a low-rank subspace, with many dead, zero-variance dimensions). The harm is lost discriminative variation, not equal importance.
A core–periphery objective — pull high-similarity items together, push low-similarity items apart — is the same attract/repel geometry as contrastive learning. The difference is how positives/negatives are defined: vanilla contrastive uses instance discrimination (positives = augmentations of the same image) with InfoNCE; a similarity-graded rule is closer to metric learning, graph/neighborhood contrastive, or soft/weighted contrastive learning (e.g. SupCon). Notably, SigLIP’s per-pair sigmoid is itself a “decide closer-or-farther for each pair” rule, structurally close to a similarity-graded core–periphery loss.
3 · Contrastive internals
NT-Xent (InfoNCE)
SimCLR’s name for InfoNCE: Normalized Temperature-scaled Cross-Entropy. With $N$ images and 2 augmentations each ($2N$ samples), for a positive pair $(i,j)$ the rest are negatives:
$$\ell_{i,j} = -\log \frac{\exp\!\big(\operatorname{sim}(z_i,z_j)/\tau\big)}{\displaystyle\sum_{k=1}^{2N}\mathbb{1}_{[k\neq i]}\,\exp\!\big(\operatorname{sim}(z_i,z_k)/\tau\big)},\qquad \operatorname{sim}(u,v)=\frac{u^\top v}{\lVert u\rVert\,\lVert v\rVert}$$
It is a $(2N\!-\!1)$-way softmax that must pick the positive out of all candidates; the total averages over all positive pairs. “Normalized” = L2-normalized embeddings; “temperature-scaled” = the $/\tau$.
Momentum encoder & its update (MoCo)
Two encoders: a query encoder $f_q$ (updated by gradients) and a key/momentum encoder $f_k$ (no gradients), whose weights track $f_q$ by EMA:
$$\theta_k \leftarrow m\,\theta_k + (1-m)\,\theta_q,\qquad m \approx 0.999$$
That slow “dynamics” keeps the keys in the large negative queue consistent over time (a fast-changing $f_k$ would make older queued keys stale).
Memory bank vs. queue
“Memory bank” is an established term from Wu et al. 2018 (instance discrimination): a store holding one feature vector per dataset image, sampled for negatives. MoCo replaced it with a FIFO queue of recent keys plus the momentum encoder — many, consistent negatives without storing the whole dataset. So the term is standard, and MoCo’s queue is technically distinct from the older bank (often conflated informally).
Online network, EMA, and asymmetry
In BYOL the trained branch is the online network; the target provides the regression target and is an EMA of the online net. “Online” just means “the actively learning branch” — not test-time adaptation (which adapts a model at inference). EMA updates the target’s parameters as a running average:
$$\theta_{\text{target}} \leftarrow \lambda\,\theta_{\text{target}} + (1-\lambda)\,\theta_{\text{online}},\qquad \lambda \to 1$$
The target is never trained by gradients: it appears in the forward pass (to make targets) but has no backward pass (stop-gradient); EMA is a post-step parameter rule. It supplies a stable, slowly-moving target (a temporal ensemble), echoing Mean Teacher / Polyak averaging.
“Asymmetry” means the two branches are not interchangeable: the online branch has an extra predictor MLP the target lacks (structural asymmetry); the target is an EMA / carries a stop-gradient (update asymmetry). This breaks the symmetry that would make the constant solution a stable fixed point — the online net chases a target it cannot trivially equal by collapsing. SimSiam’s ablation showed the stop-gradient is the essential ingredient.
4 · Non-contrastive details (with formulas)
DINO — centering + sharpening
Teacher and student both produce softmax distributions; the teacher’s is centered (subtract a running mean $c$) and sharpened (low temperature $\tau_t$):
$$P_t(x)=\operatorname{softmax}\!\Big(\tfrac{g_t(x)-c}{\tau_t}\Big),\quad P_s(x)=\operatorname{softmax}\!\Big(\tfrac{g_s(x)}{\tau_s}\Big),\quad c \leftarrow m\,c+(1-m)\tfrac{1}{B}\textstyle\sum_i g_t(x_i)$$
$$\mathcal{L}=-\sum P_t(x)\,\log P_s(x)$$
Centering blocks one dimension dominating (anti-domination, pushes toward uniform); sharpening blocks the uniform solution (pushes toward peaked). The two opposing forces balance, avoiding collapse without negatives.
Barlow Twins — cross-correlation → identity
From two views’ batch-normalized embeddings $Z^A,Z^B$, form the $d\times d$ cross-correlation $C$ and push it toward the identity:
$$C_{ij}=\frac{\sum_b Z^A_{b,i}\,Z^B_{b,j}}{\sqrt{\sum_b (Z^A_{b,i})^2}\;\sqrt{\sum_b (Z^B_{b,j})^2}},\qquad \mathcal{L}_{\text{BT}}=\sum_i (1-C_{ii})^2+\lambda\sum_i\sum_{j\neq i} C_{ij}^2$$
Diagonal→1 is invariance (the two views agree); off-diagonal→0 is redundancy reduction (dimensions decorrelated). No negatives, no large batch.
VICReg — variance, invariance, covariance
$$v(Z)=\tfrac{1}{d}\sum_{j}\max\!\big(0,\ \gamma-\sqrt{\operatorname{Var}(Z_{\cdot j})+\epsilon}\big),\quad s(Z,Z')=\tfrac{1}{N}\sum_i \lVert z_i-z'_i\rVert^2,\quad c(Z)=\tfrac{1}{d}\sum_{i\neq j}\big[\operatorname{Cov}(Z)\big]_{ij}^2$$
$$\mathcal{L}=\lambda\,s(Z,Z')+\mu\big(v(Z)+v(Z')\big)+\nu\big(c(Z)+c(Z')\big)$$
Variance keeps each dimension above a std floor $\gamma$ (anti-collapse); invariance is the view-matching MSE; covariance decorrelates dimensions. Variance + covariance explicitly replace negatives/EMA.
5 · Masked modeling: MAE vs. BEiT
MAE confirmed: mask ~75% of patches, encode only the visible ones, and a light decoder reconstructs the masked pixels; loss is MSE on masked patches only:
$$\mathcal{L}_{\text{MAE}}=\frac{1}{|\mathcal{M}|}\sum_{p\in\mathcal{M}} \lVert \hat{x}_p - x_p\rVert^2$$
BEiT masks a subset (~40%), not all (it needs visible context). It predicts discrete visual tokens: a pre-trained image tokenizer (a discrete VAE) maps each patch to a codebook id $z_p$, and BEiT classifies the masked patches over that codebook:
$$\mathcal{L}_{\text{BEiT}}=-\sum_{p\in\mathcal{M}} \log p\big(z_p \mid \tilde{x}\big)$$
Why discrete? It is “BERT for images,” so it needs a visual vocabulary; classifying codebook ids is a cleaner, higher-level target than regressing raw pixels and avoids spending capacity on imperceptible detail. But discreteness is a design choice, not a requirement — MAE regresses continuous pixels and works well; BEiT’s route just adds a tokenizer dependency.
non-contrastive: BYOL (Grill 2020) · SimSiam (Chen 2021) · DINO (Caron 2021) · Barlow Twins (Zbontar 2021) · VICReg (Bardes 2022) · Mean Teacher (Tarvainen 2017)
masked: MAE (He 2022) · BEiT (Bao 2021); related: SupCon · graph/neighborhood contrastive