Search arXivSearch

arXiv · 2509.18831

Text Slider: Efficient and Plug-and-Play Continuous Concept Control for Image/Video Synthesis via LoRA Adapters

Abstract

Recent advances in diffusion models have significantly improved image and video synthesis. In addition, several concept control methods have been proposed to enable fine-grained, continuous, and flexible control over free-form text prompts. However, these methods not only require intensive training time and GPU memory usage to learn the sliders or embeddings but also need to be retrained for different diffusion backbones, limiting their scalability and adaptability. To address these limitations, we introduce Text Slider, a lightweight, efficient and plug-and-play framework that identifies low-rank directions within a pre-trained text encoder, enabling continuous control of visual concepts while significantly reducing training time, GPU memory consumption, and the number of trainable parameters. Furthermore, Text Slider supports multi-concept composition and continuous control, enabling fine-grained and flexible manipulation in both image and video synthesis. We show that Text Slider enables smooth and continuous modulation of specific attributes while preserving the original spatial layout and structure of the input. Text Slider achieves significantly better efficiency: 5$\times$ faster training than Concept Slider and 47$\times$ faster than Attribute Control, while reducing GPU memory usage by nearly 2$\times$ and 4$\times$, respectively.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Pin-Yen Chiu, I-Sheng Fang, Jun-Cheng Chen. 2026-04-21. Text Slider: Efficient and Plug-and-Play Continuous Concept Control for Image/Video Synthesis via LoRA Adapters. https://arxiv.org/abs/2509.18831

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Opacity Is Not Just Opacity

Web graphics travel with content across pages and themes, where changing backgrounds can require recoloring and maintenance. Opacity already makes a fixed object's appearance depend on its background, yet is usually understood only as how much the object obscures it. In fact, opacity controls the scaling of the object-background color difference; transparency is only one effect of this relationship. Zero places the output at the background and one at the source color, but difference scaling need not stop at either position. We retain the compositing expression and extend the coefficient domain from $[0,1]$ to the real numbers: negative values reverse the difference, whereas values above one expand it in the same direction. We focus on same-direction expansion for reusing Web graphics across backgrounds. Each object carries a fixed source color and coefficient, while the actual background determines the enhancement direction. Background-adaptive contrast enhancement thus becomes part of the object's compositing properties, reducing the design and maintenance of separate color variants. The implementation reuses the original equation without increasing the per-pixel arithmetic operation count within the same pipeline. Enumerating all 8-bit sRGB source colors on 16 predefined light and dark canvases, a fixed $α=1.1$ increases the contrast ratio in 99.8145% of combinations. Without changing source colors, 4.8346% of all combinations newly reach the $3:1$ contrast threshold. Output validation and timing across three browsers demonstrate implementation in the same WebGL pipeline, with no sustained additional runtime observed.

cs.GR

SsgCaps: A controlled dataset for the evaluation of sound scene generation algorithms

Sound Scene Generation is about the automatic synthesis of artificial sound scenes. We introduce SsgCaps, a publicly available dataset of human-engineered sound scenes wherein each scene matches a precisely structured prompt that guides the sampling process. The corresponding prompts are sampled from a predefined action-based typology that allows extensive sampling while retaining plausibility. SsgCaps is a sound scene dataset derived from the unpublished reference dataset for Task 7 of the 2024 DCASE Challenge edition, which contained private-and public-domain audio samples. In contrast, SsgCaps contains only public-domain audio samples, allowing us to open this dataset to the community. To make this dataset useful to the community, we first elaborate on the rationale for the prompt and dataset structure. We then perform a comparative quantitative analysis of the 2 versions of the dataset. To do so, we compare both versions to the audio synthesized by the SSG algorithms submitted to the challenge using Fr{é}chet Audio Distance (FAD) and Kernel Audio Distance (KAD) as well as perceptual ratings. This analysis shows only small differences, which enables us to recommend the open version for further benchmarking of SSG algorithms.

cs.GR

Constrained Program Generation for 3D Reaction Animation with a 0.8B Model

Visualizing a chemical reaction requires making its molecular changes visible while keeping the animation faithful to the stated chemistry. Equations, structural diagrams and molecular viewers provide complementary descriptions, but assembling an interactive three-dimensional explanation still requires specifying the changes and checking their consistency. We present ChemXRG, a domain-specific language (DSL) framework that addresses this gap by representing a reaction animation as an executable program. Persistent atom identifiers and explicit bond, charge and grouping operations connect the symbolic reaction to the displayed transformation. This shared representation lets generation, verification and rendering operate on the same account of what changes. Given known, atom-mapped reactant and product structures, reaction-grounded constraints fix input-determined facts and restrict action choices; execution checks validate the resulting transformation before geometry and frames are constructed. We implement this paradigm with a reaction-program corpus and ChemQwen, a trained 0.8B DSL generator. Paired and component evaluations show improved compiler acceptance and normalized full-program agreement under input-conditioned constraints, while identifying remaining failures that require execution checks. A public browser application demonstrates the connection from symbolic reaction descriptions to inspectable programs and interactive 3D animations.

cs.GR