TrajGANR: Trajectory-Centric Urban Multimodal Learning via Geospatially Aligned Neural Representations
Many urban prediction tasks depend not only on the static attributes of a location, such as the visual appearance of its built environment, but also on its dynamic function: how people navigate and use it over time. Predicting congestion, travel demand, or road safety risks, for instance, requires mobility information that imagery and geographic coordinates alone cannot provide. Trajectories carry this signal, yet current trajectory representations are difficult to align with static geospatial observations: sparse, irregular GPS samples rarely coincide with street-view image coordinates, and existing encoders produce either a single embedding per trajectory or road segment, or embeddings only at the observed GPS samples. We introduce TrajGANR, a trajectory-centric multimodal self-supervised learning framework that represents each trajectory as a continuous, path-conditioned neural field. Reconstructed from the observed GPS samples, this representation can be queried at arbitrary coordinates (including the off-path locations where street-view images were captured), yielding localized mobility embeddings that align with street-view image and location representations at the same coordinate. Across four fine-grained road-level tasks in San Francisco and Porto, TrajGANR consistently outperforms recent geospatial and trajectory foundation models, with ablations attributing the gains to both the mobility signal and the fine granularity at which it is aligned. TrajGANR's location encoder further generalizes beyond the road network to areas where neither street-view imagery nor trajectories are available.