Persistent Scene Memory
A fixed-budget latent Gaussian state carries queryable 3D scene information across frames.
Camera-based 4D panoptic occupancy tracking (4D-POT) is a promising paradigm for holistic scene understanding from multi-view imagery, enabling joint reasoning about geometry, semantics, and object identities across time. Recent mask-based pipelines achieve strong performance by propagating instance queries across frames. However, their underlying volumetric representations are typically recomputed at each timestep, limiting geometric temporal consistency, particularly under occlusion and for static scene elements. To address this limitation, we propose a streaming Gaussian encoder that maintains a persistent volumetric scene representation for 4D-POT. Our method models the scene as a fixed-size set of latent Gaussian queries that are propagated via ego-motion compensation and refreshed under a confidence-guided budget constraint. Crucially, we shape Gaussian opacities through depth-based supervision to serve as proxy for visibility, enabling confidence to accumulate as a temporally aggregated measure of persistent scene support. Together with a warmup-based multi-frame training strategy, this yields representation-level temporal coherence beyond decoder-only tracking. Extensive experiments on Occ3D-extended nuScenes and Waymo establish a new state-of-the-art for camera-based 4D-POT, improving tracking consistency with negligible computational overhead while remaining fully compatible with existing mask-based pipelines.
Where the update fits: multi-view image features are lifted into a 3D feature pyramid before this update. Afterwards, the refined queries are decoded into Gaussians and splatted into a dense feature volume for the mask-based 4D-POT decoder.
The Gaussian state from frame t − 1 is transformed into the current ego frame. Transforming reference positions and updating position-dependent query features keeps persistent scene content spatially aligned as the vehicle moves.
Sparse LiDAR depth is used during training to shape opacity as a visibility proxy. At inference, opacity, a distance weight, and the field-of-view mask form each Gaussian's observation score. SGE accumulates that evidence as confidence (with decay for dynamic classes), then samples up to K − M eligible queries using a hybrid rank of temporal confidence and instantaneous opacity.
At least M of the fixed K slots are reserved for new queries. SGE samples them from high-response locations in the current 3D feature volume, producing a replenished state that combines persistent support with newly observed scene content.
The LaGS transformer jointly updates retained and newly sampled queries against the current features. Its refined output becomes the persistent Gaussian state at frame t, ready for decoding and propagation to the next frame.
Consistency belongs in the representation, not just the decoder.
A fixed-budget latent Gaussian state carries queryable 3D scene information across frames.
Visibility evidence accumulates over time to retain supported structure and replenish obsolete slots.
Training-only depth supervision shapes opacity, while warmup frames establish realistic temporal context.
The encoder remains compatible with mask-based 4D-POT pipelines and adds only about 2% inference overhead.
Code and models are provided on GitHub and released under AGPLv3 for non-commercial purposes. For any commercial purposes, please contact the authors at luz@cs.uni-freiburg.de.
View on GitHubThis work was funded by the Bosch Research collaboration on AI-driven automated driving, the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under grant number 539134284, through EFRE (FEIH_2698644), and the state of Baden-Württemberg.