Search arXiv⌕ Search

arXiv · 2610.00618

Dirichlet Splatting: Differentiable Rendering for Wave-Based Inverse Problems

Abstract

Wave-based coherent imaging, including terahertz tomography, synthetic-aperture acoustics, and millimeter-wave radar, forms images by Fourier-processing finite-length signals, with an exact point spread function that is not Gaussian but a Dirichlet kernel: complex-valued, oscillatory, and periodic. However, transplanting 3D Gaussian splatting to coherent sensing fails by construction; Gaussian splats discard the sidelobe energy (10-20% of the total) and the phase that governs coherent interference between reflectors. Our key idea is to replace the learned Gaussian footprint with the physically exact Dirichlet kernel of the finite-window DFT, modulated by a surfel that carries area, normal, and material, so that the rendering primitive matches the measurement physics instead of approximating it. We pair this primitive with a specialized solver, Dirichlet Sliding Frank-Wolfe (DSFW), that combines variable projection, residual dual certificates, and certificate-driven hard replacement of low-utility surfels, with periodic low-resolution coupled Levenberg-Marquardt correction, navigating the rugged loss landscape that breaks generic first-order optimizers. The Dirichlet kernel admits an O(1) closed-form evaluation, so the forward model matches FFT ground truth to machine precision while remaining differentiable end-to-end. On dense terahertz reconstruction, our method recovers reflector centers to 0.018 bin RMSE, 10-50x faster than waveform-level automatic differentiation, where Gaussian splats fail.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Xingyu Chen, Wuqiong Zhao, Xinyu Zhang, Tzu-Mao Li. 2026-09-30. Dirichlet Splatting: Differentiable Rendering for Wave-Based Inverse Problems. https://arxiv.org/abs/2610.00618

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

MotionPersona: Real-Time Locomotion Control across Personas, Bodies, and Styles

We present MotionPersona, a generative framework for character-aware locomotion control, in which the motion for a command depends on the captured persona, the body shape, and the character's style. Unlike style, which one performer can vary at will, persona and body shape are coupled in capture: each performer is observed in only one body. The captured data therefore cannot uniquely determine which motion characteristics should follow the persona and which should change with the body, leaving unseen persona-body combinations unconstrained. We capture 48 performers, aged 5 to 68, under the same nine styles and seven commands, 44 of them with persona annotation. From this repeated-measures design, we identify two robust associations between body shape and gait. These measurements guide a cross-body specification of which characteristics should change and which should be preserved. We implement this specification through a physically informed retargeting pipeline, producing cross-body training data while penalizing penetration and foot skating. On this data we train a single generative controller. A shape-aware VAE compresses each motion block into a few latent tokens and renders them on a conditioned target body under explicit geometric supervision; over these tokens, a latent flow-matching prior generates persona- and style-conditioned motion in two sampling steps. The controller covers all captured personas, a wide family of SMPL-X target bodies, and nine styles in one model, and runs at 27 ms per block on two threads of a laptop CPU. We verify the framework at every stage, following the same gait descriptors from captured to retargeted to generated motion and sweeping each axis in isolation. To our knowledge, this is the first real-time locomotion controller that carries part of a captured persona's performer-specific variation across independently selected body shapes and styles.

cs.GR↗

Hologram Representation via Quadratic Phase Gaussian Splatting

We introduce Complex-Valued Quadratic Phase Gaussian (CVQPG), a novel hologram representation method that augments each 2D Gaussian primitive with a quadratic phase profile controlled by a learnable curvature parameter. Against the planar Gaussian baseline, CVQPG improves the average PSNR of holographic reconstructions by 0.19 dB (RGB) and 0.33 dB (grayscale) at equal primitive counts, and by 0.05 dB (RGB) and 0.08 dB (grayscale) at equal parameter counts, where it still leads in all visual quality metrics. Our frequency-domain analysis shows that CVQPG better preserves the mid-to-high frequency band of natural images, where the reconstruction MSE drops by up to 11% (RGB) and 22% (grayscale), indicating that modulating primitive wavefronts is an effective and lightweight enhancement.

cs.GR↗

Text2Sim: Agentic Physics-Based Simulation Generation with Distilled Expertise

Creating diverse physical simulations remains labor-intensive because assets, layout, physical parameters, motion, control, and rendering must be designed and debugged jointly. We present Text2Sim, a simulation-specialized agentic pipeline that converts a text-only request into an executable, editable dynamic case. Built on Genesis, Text2Sim uses a hierarchical agentic structure that combines a Planner with specialized Writers, asset-generation tools, and an independent Critic. Compact skills (Debug Cards) distilled from graphics demonstrations provide role-specific physical guidance for execution-based repair. We evaluate physical quality, visual quality, and human preference on 42 held-out prompts spanning rigid, articulated, deformable, and cloth phenomena, with a paper-level split between experience construction and evaluation. We design automatic physical and visual scorers to evaluate the quality of the results, and Text2Sim achieves higher scores than all four state-of-the-art baselines on both metrics. In blinded user studies with these baselines, significantly more participants prefer Text2Sim than prefer the baselines, which is consistent with the results from our automatic scorers. The pipeline also supports a broad range of downstream applications; we select dataset construction and extension to multimodal input as two representative examples. We will release the code, the Debug Card library, and a dataset of generated cases, each pairing the text prompt and rendered video with the executable program, assets, physical parameters, controls, and recorded states.

cs.GR↗