iSEE: Object Permanence through Self-Supervision

Pramish PaudelAjad ChhatkuliLuc Van GoolDanda Pani Paudel

INSAIT, Sofia University "Kliment Ohridski"

INSAIT

Paper (soon)arXiv (soon)Code (soon)BibTeX

Object permanence

Hidden is not gone.

A covered object still exists, and it is still somewhere. Infants learn this in their first year, and many animals show it too.

A cat follows a ball hidden under shuffled cups. Video: Robert Zebib, YouTube
An infant pulls a cloth off a hidden toy (Piaget's search task). Video: Julian Lloyd, YouTube
A young orangutan sees an object vanish from a cup, a magic trick. Video: CBS News, YouTube

To build world models that represent the world accurately, object permanence is the first thing to get right. An object that leaves view has not left the world: it keeps its identity, and it keeps moving.

In this work, we learn object permanence from video alone, without labels. iSEE notices when an object disappears, remembers what it looked like, follows where it goes while hidden, and recognises it when it returns.

iSEE on LA-CATER: the target (white) is covered by nested cones and carried, out of sight for 7.0 s. Green: iSEE's estimate of the target, dashed while it is hidden. The bar under the video marks the hidden frames.

Abstract

Object permanence, keeping track of an object's identity and position while it is occluded, is central to video representations that track, predict and plan. Trackers that achieve it learn from boxes, track identities and visibility labels. On the other hand, self-supervised object-centric methods discover objects without labels: through slot attention, it represents a video as slots that bind to objects and follow them across frames. However, these slots are lost under occlusion, making the desired permanence impossible. Reasoning permanence is a hard problem because it requires to detect when an object becomes occluded, re-identify when object reappears, and keep the object's hidden position continuous, using reapperance as the only learning cue. To address this, we propose iSEE, a novel framework that offers all three aforementioned requirements, without any labels whatsoever. We built iSEE using the following three proposed components: (i) Object evidence modelling: a slot's attention, compared with its own past, reveals when its object is hidden. (ii) Appearance-position separation: two slot streams let the appearance be held for re-identification while the position keeps changing. (iii) Permanence from reappearance: a walker follows the hidden object's position, trained only on where the object reappears. On LA-CATER static, iSEE returns a reappearing object to its own slot after 86 % of occlusions, against 32 % for SlotContrast, and localises it while hidden within 4.1 mAP of the label-trained SoTA RAM. The two streams also allow downstream planning, with the position stream as the action of a world model.

Overview

Figure 1 of the paper: one LA-CATER occlusion followed by iSEE, and the re-identification versus permanence scatter on LA-CATER, static camera
(a) Slot k (green) holds the ball; dashed, its mask placed at iSEE's predicted position while the ball is hidden. The slot's evidence, the attention it receives from the frame's patches, falls at the occlusion and recovers at the return. During the occlusion the appearance is held and the position keeps walking, learnt only from where the object reappears (red ring). (b) LA-CATER, static camera. Re-identification, SAME ↑: the share of the 183 scorable occlusions after which the target returns to the slot that held it before (before and after: the last and first frame with at least half of the target visible). Permanence, mAP@[0.1,0.3] ↑: over all hidden frames of the 195 occlusions (defined under Results). Each arrow adds one component to SlotContrast. RAM is trained with boxes, track identities and visibility labels.

Method

The two-stream slot in training
Temporal evidence normalisation

Examples

target (ground truth)
the slot on the target
that slot, placed at the predicted position while the target is hidden
frames on which the target is fully hidden

LA-CATER clips are selected among occlusions iSEE scores well on; failure cases are at the end.

LA-CATER, moving camera

Contained and carried · hidden 5.8 s
Contained and carried · hidden 3.9 s
Contained · hidden 0.5 s
Contained · hidden 0.5 s

LA-CATER, static camera

Contained and carried · hidden 3.4 s
Contained and carried · hidden 3.0 s
Contained · hidden 6.2 s

Against the baselines

SlotContrastRandSF.QLoci-LoopediSEE (ours)
SlotContrastRandSF.QLoci-LoopediSEE (ours)
SlotContrastRandSF.QLoci-LoopediSEE (ours)

KITTI, real driving video

A KITTI car passes behind another road user and reappears: SlotContrast and iSEE slots over ten frames, with each model's evidence below
A car passes behind another road user and reappears. White: the car's visible mask. Green: the slot on the car before the occlusion; cross: the walker's predicted centre while the car is hidden. Below, each model's evidence for that slot relative to its own past, ê: both fall while the car is hidden; only iSEE's recovers, and its slot returns. Over the KITTI occlusions of cars (0.1–2.2 s), iSEE returns the car to its slot after 48.4 % of occlusions, SlotContrast after 29.0 % (SAME ↑).

Results

methodstatic cameramoving camera
occluded ↑contained ↑carried ↑occluded ↑contained ↑carried ↑
RAM (vis., bbox, id)73.886.781.083.287.481.1
Loci-Looped (bg)56.649.932.949.551.837.9
frozen57.856.841.951.432.213.6
SlotContrast8.64.33.53.11.31.2
RandSF.Q11.56.45.41.31.21.0
iSEE80.584.278.754.451.646.3

Position of the target while it is hidden, LA-CATER (Table 1 of the paper). mAP@[0.1,0.3] ↑ (×100): the share of an occlusion's hidden frames on which the model's box overlaps the target's amodal box above an IoU threshold, averaged over the thresholds 0.10 to 0.30 and over the 195 (static camera) and 388 (moving camera) occlusions; the columns split the hidden frames by their LA-CATER label. A slot model's box keeps the size of its mask before the occlusion and is centred on the predicted position. In parentheses, what a row is given beyond the video: RAM (grey) is trained with visibility labels, boxes and track identities; Loci-Looped reads a per-frame background image (bg). frozen holds iSEE's slot at its last visible position. Bold: the best row that reads the video alone.

Failure cases (8)

One median case per failure mode, and the re-identification failures with the largest displacement.

Prediction wanders off · hidden 3.5 s
Prediction follows another object · hidden 3.7 s
Slot is not on the object going in · hidden 3.4 s
Box too small to overlap · hidden 3.0 s
Slot does not return · hidden 4.7 s
Prediction leaves the cone · hidden 6.6 s
Prediction follows the wrong cone · hidden 7.2 s
Slot does not return · hidden 2.1 s

BibTeX

@article{paudel2026isee,
  title   = {{iSEE}: Object Permanence through Self-Supervision},
  author  = {Paudel, Pramish and Chhatkuli, Ajad and Van Gool, Luc and Paudel, Danda Pani},
  journal = {arXiv preprint},
  year    = {2026}
}