Search arXivSearch

arXiv subjects

Haodong Yu

Publications and source records attributed to Haodong Yu.

6 recordsLinked to original sources

Teacher-Feature Drifting: One-Step Diffusion Distillation with Pretrained Diffusion Representations

Sampling from pretrained diffusion and flow-matching models typically requires many forward passes to generate diverse and high-fidelity images. Existing distillation methods often rely on multiple auxiliary networks, carefully designed training stages, or complex optimization pipelines. In this work, we revisit the recently proposed Drifting Model objective and show that a single drifting loss can be directly used to simplify one step distillation. A key observation is that the pretrained diffusion teacher itself already provides a strong representation space. Unlike the original Drifting Model, which relies on an additional pretrained feature extractor, we use intermediate hidden states of the pretrained teacher model as the feature representation. This removes the need for training or introducing an extra representation network while preserving a semantically meaningful feature geometry for drifting. Furthermore, we introduce a lightweight mode coverage loss to mitigate mode collapse during distillation and encourage the student generator to cover diverse teacher-supported regions. Extensive experiments on ImageNet and SDXL demonstrate that our method achieves efficient one step generation with competitive image quality and diversity, achieving FID scores of 1.58 on ImageNet-64$\times$64 and 18.4 on SDXL, while substantially simplifying the overall distillation framework.

cs.CV

ScrollScape: Unlocking 32K Image Generation With Video Diffusion Priors

While diffusion models excel at generating images with conventional dimensions, pushing them to synthesize ultra-high-resolution imagery at extreme aspect ratios (EAR) often triggers catastrophic structural failures, such as object repetition and spatial fragmentation. This limitation fundamentally stems from a lack of robust spatial priors, as static text-to-image models are primarily trained on image distributions with conventional dimensions. To overcome this bottleneck, we present ScrollScape, a novel framework that reformulates EAR image synthesis into a continuous video generation process through two core innovations. By mapping the spatial expansion of a massive canvas to the temporal evolution of video frames, ScrollScape leverages the inherent temporal consistency of video models as a powerful global constraint to ensure long-range structural integrity. Specifically, Scanning Positional Encoding (ScanPE) distributes global coordinates across frames to act as a flexible moving camera, while Scrolling Super-Resolution (ScrollSR) leverages video super-resolution priors to circumvent memory bottlenecks, efficiently scaling outputs to an unprecedented 32K resolution. Fine-tuned on a curated 3K multi-ratio image dataset, ScrollScape effectively aligns pre-trained video priors with the EAR generation task. Extensive evaluations demonstrate that it significantly outperforms existing image-diffusion baselines by eliminating severe localized artifacts. Consequently, our method overcomes inherent structural bottlenecks to ensure exceptional global coherence and visual fidelity across diverse domains at extreme scales.

cs.CV

VitaGlyph: Vitalizing Artistic Typography with Flexible Dual-branch Diffusion Models

Artistic typography is a technique to visualize the meaning of input character in an imaginable and readable manner. With powerful text-to-image diffusion models, existing methods directly design the overall geometry and texture of input character, making it challenging to ensure both creativity and legibility. In this paper, we introduce a dual-branch, training-free method called VitaGlyph, enabling flexible artistic typography with controllable geometry changes while maintaining the readability. The key insight of VitaGlyph is to treat input character as a scene composed of a Subject and its Surrounding, which are rendered with varying degrees of geometric transformation. To enhance the visual appeal and creativity of the generated artistic typography, the subject flexibly expresses the essential concept of the input character, while the surrounding enriches relevant background without altering the shape, thus maintaining overall readability. Specifically, we implement VitaGlyph through a three-phase framework: (i) Knowledge Acquisition leverages large language models to design text descriptions for the subject and surrounding. (ii) Regional Interpretation detects the part that most closely matches the subject description and refines the structure via Semantic Typography. (iii) Attentional Compositional Generation separately renders the textures of the Subject and Surrounding regions and blends them in an attention-based manner. Experimental results demonstrate that VitaGlyph not only achieves better artistry and readability but also manages to depict multiple customized concepts, facilitating more creative and pleasing artistic typography generation. Our code will be made publicly available.

cs.CV

Real-Time 4K Super-Resolution of Compressed AVIF Images. AIS 2024 Challenge Survey

This paper introduces a novel benchmark as part of the AIS 2024 Real-Time Image Super-Resolution (RTSR) Challenge, which aims to upscale compressed images from 540p to 4K resolution (4x factor) in real-time on commercial GPUs. For this, we use a diverse test set containing a variety of 4K images ranging from digital art to gaming and photography. The images are compressed using the modern AVIF codec, instead of JPEG. All the proposed methods improve PSNR fidelity over Lanczos interpolation, and process images under 10ms. Out of the 160 participants, 25 teams submitted their code and models. The solutions present novel designs tailored for memory-efficiency and runtime on edge devices. This survey describes the best solutions for real-time SR of compressed high-resolution images.

cs.CV

Scaling Laws Governing the Elastic Properties of 3D-Graphenes

In this study, we have comprehensively investigated the scaling law for elastic properties of three-dimensional honeycomb-like graphenes (3D-graphenes) using hybrid neural network potential based molecular dynamics simulations and theoretical analyses. The elastic constants as functions of honeycomb hole size, denoted by the graphene wall length $L$, were provided. All five independent elastic constants in the large $L$ limit are proportional to $L^{-1}$. The associated coefficients are combinations of two-dimensional graphene's elastic constants. High-order terms including $L^{-2}$ and $L^{-3}$ emerge for finite $L$ values. They have three origins, the distorted areas close to the joint lines of 3D-graphenes, the variation of solid angles between graphene plates, and the bending distortion of graphene plates. Significantly, the chirality becomes essential with the decreasing of $L$, because the joint line structures are different between the armchair and zigzag type 3D-graphenes. Our findings provide insights into the elastic properties of graphene-based superstructures and can be used for further studies on graphene-based materials.

cond-mat.mtrl-sci

Interlayer magnetic interactions in $\pi/3$-twisted bilayer CrI$_3$

The interlayer magnetic interaction in bilayer CrI$_3$ plays a crucial role for its device applications. In this work, we studied the interlayer magnetic interaction in $\pi/3$-twisted bilayer CrI$_3$ using first-principles calculations. Our calculations show that the interlayer coupling can be ferromagnetic or antiferromagnetic depending crucially on lateral shift. The strongest antiferromagnetic interlayer interaction appears in the $\bar{A}A$-stacking. The magnetic force theory calculations demonstrate that such an antiferromagnetic interaction is dominanted by the $e_g$-$e_g$ channel. Particularly, the interlayer antiferromagnetic interaction is very sensitive to external pressure. This highly tunable interlayer interaction makes $\pi/3$-twisted bilayer CrI$_3$ a potential building block for magnetic field effect transistors and pressure sensors.

cond-mat.mtrl-sci