News · IROS 2026

SGE is a Best Paper / Best Student Paper Award finalist

We will present Streaming Gaussian Encoding for 4D Panoptic Occupancy Tracking on Monday, September 28, at 10:08 AM in Room 406 at IROS 2026 in Pittsburgh.

Presentation details
Base work

LaGS: Latent Gaussian Splatting for 4D Panoptic Occupancy Tracking

SGE builds on LaGS, a camera-based 4D panoptic occupancy tracking framework that represents 3D features as sparse, feature-bearing latent Gaussians. LaGS replaces dense voxel processing with hierarchical point-based reasoning and Gaussian feature splatting; SGE extends this representation into a persistent streaming scene memory.

Read more

Motivation

Propagate the Gaussian scene, not just the queries.

Mask-based 4D-POT systems carry instance queries across frames, but usually rebuild the 3D feature volume at every timestep. SGE also propagates the latent Gaussian representation where the scene's geometric information lives, adding representation-level temporal coherence to decoder-level tracking.

Two-frame 4D panoptic occupancy pipeline. Existing mask-based methods propagate decoder queries, while SGE additionally propagates the latent Gaussian 3D scene representation.
Decoder queries preserve thing-instance identities, but carry geometric scene structure only indirectly. SGE additionally propagates a fixed-budget latent Gaussian state, retaining a queryable 3D representation across timesteps.

Abstract

Camera-based 4D panoptic occupancy tracking (4D-POT) is a promising paradigm for holistic scene understanding from multi-view imagery, enabling joint reasoning about geometry, semantics, and object identities across time. Recent mask-based pipelines achieve strong performance by propagating instance queries across frames. However, their underlying volumetric representations are typically recomputed at each timestep, limiting geometric temporal consistency, particularly under occlusion and for static scene elements. To address this limitation, we propose a streaming Gaussian encoder that maintains a persistent volumetric scene representation for 4D-POT. Our method models the scene as a fixed-size set of latent Gaussian queries that are propagated via ego-motion compensation and refreshed under a confidence-guided budget constraint. Crucially, we shape Gaussian opacities through depth-based supervision to serve as proxy for visibility, enabling confidence to accumulate as a temporally aggregated measure of persistent scene support. Together with a warmup-based multi-frame training strategy, this yields representation-level temporal coherence beyond decoder-only tracking. Extensive experiments on Occ3D-extended nuScenes and Waymo establish a new state-of-the-art for camera-based 4D-POT, improving tracking consistency with negligible computational overhead while remaining fully compatible with existing mask-based pipelines.

Method

SGE updates a fixed-budget Gaussian scene state from frame t − 1 to frame t. The lower panel expands confidence-guided pruning: current visibility evidence is accumulated over time, then combined with a valid-range constraint to decide which Gaussians persist and which slots are replenished.

Where the update fits: multi-view image features are lifted into a 3D feature pyramid before this update. Afterwards, the refined queries are decoded into Gaussians and splatted into a dense feature volume for the mask-based 4D-POT decoder.

Ego-Motion Compensation

The Gaussian state from frame t − 1 is transformed into the current ego frame. Transforming reference positions and updating position-dependent query features keeps persistent scene content spatially aligned as the vehicle moves.

Confidence-Guided Pruning

Sparse LiDAR depth is used during training to shape opacity as a visibility proxy. At inference, opacity, a distance weight, and the field-of-view mask form each Gaussian's observation score. SGE accumulates that evidence as confidence (with decay for dynamic classes), then samples up to K − M eligible queries using a hybrid rank of temporal confidence and instantaneous opacity.

Feature-Guided Sampling

At least M of the fixed K slots are reserved for new queries. SGE samples them from high-response locations in the current 3D feature volume, producing a replenished state that combines persistent support with newly observed scene content.

Joint Refinement

The LaGS transformer jointly updates retained and newly sampled queries against the current features. Its refined output becomes the persistent Gaussian state at frame t, ready for decoding and propagation to the next frame.

Key Takeaways & Contributions

Consistency belongs in the representation, not just the decoder.

Persistent Scene Memory

A fixed-budget latent Gaussian state carries queryable 3D scene information across frames.

Confidence-Guided Refresh

Visibility evidence accumulates over time to retain supported structure and replenish obsolete slots.

Streaming-Aware Training

Training-only depth supervision shapes opacity, while warmup frames establish realistic temporal context.

Drop-in & Efficient

The encoder remains compatible with mask-based 4D-POT pipelines and adds only about 2% inference overhead.

SOTA 4D-POT
34.4STQ nuScenes
21.9STQ Waymo

Overview and Qualitative Results

Code & Models

Available now

Code & Models on GitHub

Code and models are provided on GitHub and released under AGPLv3 for non-commercial purposes. For any commercial purposes, please contact the authors at luz@cs.uni-freiburg.de.

View on GitHub

Publication

Maximilian Luz, Thomas Nürnberg, Yakov Miron, Abhinav Valada

Streaming Gaussian Encoding for 4D Panoptic Occupancy Tracking

IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2026

If you find our work useful, please consider citing our paper.

Authors

Maximilian Luz

Maximilian Luz

University of Freiburg

Thomas Nürnberg

Thomas Nürnberg

Bosch Research

Yakov Miron

Yakov Miron

Bosch Research, University of Haifa

Abhinav Valada

Abhinav Valada

University of Freiburg

Acknowledgment

This work was funded by the Bosch Research collaboration on AI-driven automated driving, the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under grant number 539134284, through EFRE (FEIH_2698644), and the state of Baden-Württemberg.

Co-funded by the European Union Baden-Württemberg
100%