CitySTAR: Agent-Driven Structured and Topology-Aware Reasoning for Open-Vocabulary Urban 3D Grounding

Sep 17, 2026·
Shuai Zhang
First author
,
Hongye Hou
,
Qinghe Liu
,
Zhuoxiao Li
,
Dongli Wu
,
Jing Ou
,
Yuan Liu
,
Wufan Zhao
· 1 min read
CitySTAR overview. Source: Zhang et al., arXiv:2609.19911, Figure 1.
Abstract
CitySTAR treats language-guided localization in city-scale point clouds as structured constraint reasoning. Its training-free pipeline builds an open-vocabulary scene graph, gathers multimodal evidence with language-model tools, and checks target-context topology to distinguish candidate objects. A reflective grounding stage combines spatial consistency with visual evidence. The accompanying CitySTAR-3D benchmark strengthens instance annotations and evaluation of complex spatial relations in urban scenes.
Type
Publication
Advances in Neural Information Processing Systems (NeurIPS 2026), accepted
publications

CitySTAR was accepted at NeurIPS 2026. The linked arXiv version uses the earlier title “CitySTAR: Structured and Topology-Aware Reasoning for Open-Vocabulary Urban 3D Grounding”.