Search arXivSearch

arXiv subjects

Tom Durand

Publications and source records attributed to Tom Durand.

3 recordsLinked to original sources

SceneTeract: Probing and Improving Agent-Aware Activity Reasoning in 3D Indoor Scenes

Indoor 3D scenes are ultimately meant to be used: an embodied agent should be able to navigate, reach objects, and complete diverse activities. Yet whether a given scene actually supports these activities for a specific agent profile is rarely verified. Existing evaluations of indoor 3D scenes typically focus on visual quality and semantic plausibility. In contrast, the feasibility of an activity depends on geometric, agent-specific constraints such as reach, clearance, and navigable space availability. These properties are not captured by visual plausibility metrics, and, as we show, VLMs, which are increasingly used to reason about 3D scenes, often fail to determine action feasibility in a single shot. We present SceneTeract, a verification interface that separates semantic action understanding from physical feasibility. Given a scene, an activity, and an embodied agent profile, we decompose the activity into atomic actions on scene objects. Explicit geometric checks then decide whether each step is executable and return a diagnostic trace explaining failures. In synthetic indoor scenes, SceneTeract reveals widespread functional and accessibility failures across diverse agent profiles. Moreover, when benchmarked against our verification, we find that existing VLMs systematically over-predict action feasibility, highlighting limited awareness of embodied functional constraints. In response, we post-train a lightweight VLM with verifier feedback, improving its assessment of physical feasibility. Although trained only on renders of synthetic scenes, we demonstrate that scene understanding improvements also generalize to real-world scenes. We will release our verification suite, benchmark labels, and diagnostic trace datasets.

cs.CV

LACONIC: A 3D Layout Adapter for Controllable Image Creation

Existing generative approaches for guided image synthesis of multi-object scenes typically rely on 2D controls in the image or text space. As a result, these methods struggle to maintain and respect consistent three-dimensional geometric structure, underlying the scene. In this paper, we propose a novel conditioning approach, training method and adapter network that can be plugged into pretrained text-to-image diffusion models. Our approach provides a way to endow such models with 3D-awareness, while leveraging their rich prior knowledge. Our method supports camera control, conditioning on explicit 3D geometries and, for the first time, accounts for the entire context of a scene, i.e., both on and off-screen items, to synthesize plausible and semantically rich images. Despite its multi-modal nature, our model is lightweight, requires a reasonable number of data for supervised learning and shows remarkable generalization power. We also introduce methods for intuitive and consistent image editing and restyling, e.g., by positioning, rotating or resizing individual objects in a scene. Our method integrates well within various image creation workflows and enables a richer set of applications compared to previous approaches.

cs.CV

DeBaRA: Denoising-Based 3D Room Arrangement Generation

Generating realistic and diverse layouts of furnished indoor 3D scenes unlocks multiple interactive applications impacting a wide range of industries. The inherent complexity of object interactions, the limited amount of available data and the requirement to fulfill spatial constraints all make generative modeling for 3D scene synthesis and arrangement challenging. Current methods address these challenges autoregressively or by using off-the-shelf diffusion objectives by simultaneously predicting all attributes without 3D reasoning considerations. In this paper, we introduce DeBaRA, a score-based model specifically tailored for precise, controllable and flexible arrangement generation in a bounded environment. We argue that the most critical component of a scene synthesis system is to accurately establish the size and position of various objects within a restricted area. Based on this insight, we propose a lightweight conditional score-based model designed with 3D spatial awareness at its core. We demonstrate that by focusing on spatial attributes of objects, a single trained DeBaRA model can be leveraged at test time to perform several downstream applications such as scene synthesis, completion and re-arrangement. Further, we introduce a novel Self Score Evaluation procedure so it can be optimally employed alongside external LLM models. We evaluate our approach through extensive experiments and demonstrate significant improvement upon state-of-the-art approaches in a range of scenarios.

cs.CV