Search arXiv⌕ Search

arXiv · 2610.11685

HI3D 3.0 (Twinkle3D): Object-specific 3D Asset Generation with High Resolution

Abstract

Image-to-3D generation has become increasingly capable of producing objects that closely resemble the input image, and an outstanding challenge is to reproduce the depicted object itself, including the specific geometry that defines it. Inscriptions, brand marks, and repeated structures are frequently distorted or lost, despite being critical to object identity. We present Hi3D 3.0, an image-to-3D generation system targeting object-specific fidelity, with Twinkle3D as its geometry model for generating watertight triangle meshes at $2048^{3}$ resolution. Twinkle3D advances high-fidelity geometry generation along four dimensions. First, while O-Voxel/FaithC offers high representational precision, it often suffers from poor surface quality and non-watertight geometry. We address both issues while retaining its $2048^{3}$-level precision. Second, we scale diffusion generation to sequences of up to 300K geometric tokens through a redesigned DiT architecture and large-scale distributed training optimizations, reducing training time per step from approximately ten minutes to ten seconds. Third, subsequent refinement cannot fully compensate for errors introduced during initial generation; we therefore strengthen both global shape and local detail in the initial generation stage, and the resulting single-stage model surpasses prior two-stage pipelines with $512^{3}$ refinement. Finally, we introduce a fine-grained image-3D cross-modal interaction mechanism that strengthens correspondence between visual evidence and geometric tokens, improving the recovery of object-specific structures. We evaluate geometric fidelity using alignment metrics derived from silhouettes and normal fields. Hi3D 3.0 outperforms four commercial systems across all reported metrics, recovering 82.1% of inscribed characters at 98.2% precision, compared with 21.7% recall for the strongest competitor.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ziying Li, Shengchu Zhao, Huiang He, Yiyang Chen, Jianwen Huang, Bailin Li, Changhao Li, Jianhui Li, Jie Li, Ruiyang Liu, Yibo Luo, Tengjiao Sun, Pei Tang, Shiwen Wang, Jiaqi Wu, Kang Wu, Kaiqiao Yang, Zherui Yang, Hu Zhang, Xuezhi Zhao, Xinhe Zheng, Yukun Li, Heliang Zheng, Rongfei Jia. 2026-10-08. HI3D 3.0 (Twinkle3D): Object-specific 3D Asset Generation with High Resolution. https://arxiv.org/abs/2610.11685

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Accurate and Efficient Object Pose Estimation via the Aggregation of Diffusion Features

Estimating the pose of objects from images is a crucial task of 3D scene understanding, and recent approaches have shown promising results on very large benchmarks. However, these methods experience a significant performance drop when dealing with unseen objects. To address this problem, we have an in-depth analysis on the features of diffusion models, e.g. Stable Diffusion, which hold substantial potential for modeling unseen objects. Based on this analysis, we then innovatively introduce these diffusion features for object pose estimation. To verify the efficacy of diffusion features for object pose estimation, we propose three distinct architectures (vanilla, nonlinear, and context-aware weight aggregations) that capture and aggregate diffusion features for comparative analysis. To achieve an efficient feature aggregation, we propose a confidence adaptive aggregation network that automatically selects the discriminative features rather than uses all the features, achieving a better speed-and-accuracy trade-off. In particular, our confidence adaptive aggregation network achieves higher accuracy than the previous best arts on unseen objects: 97.7% vs. 93.5% on Unseen LM, 85.5% vs. 76.3% on Unseen O-LM, showing the strong generalizability of our method. On the large-scale BOP benchmark, our method also provides measurable gains, with an average recall of 58.3 compared to 57.9 previously. In addition, CAA reduces computational cost by 1.3-1.5 compared to the CWA variant while maintaining comparable accuracy. Furthermore, CAA reaches real-time performance, achieving over 68 FPS and offering a substantially improved accuracy-efficiency trade-off.

cs.CV↗

A Survey on Industrial Anomaly Synthesis

This paper presents a comprehensive review of industrial anomaly synthesis (IAS). Existing surveys on industrial anomalies mainly focus on anomaly detection, while IAS is typically treated as an auxiliary component rather than as an independent topic. However, owing to its increasing importance in data augmentation, downstream model training, and controllable industrial inspection, IAS has become a research direction of growing interest. To address the lack of a dedicated review, we survey a broad range of representative methods and organize them into four paradigms: hand-crafted synthesis, distribution hypothesis-based synthesis, generative model (GM)-based synthesis, and vision-language model (VLM)-based synthesis. We further establish a dedicated taxonomy for IAS, which supports more systematic comparison across methods and offers a clearer view of the field's development. Beyond methodological categorization, we summarize the datasets, benchmarks, and evaluation metrics commonly adopted in IAS, and review recent advances in multimodal anomaly synthesis that remain underexplored in prior surveys. We also provide deployment-oriented comparisons and practical guidance by analyzing input requirements, output forms, controllability, cost, downstream tasks, and practical limitations across IAS subcategories. Overall, this survey provides a structured understanding of existing IAS methods, evaluation settings, practical trade-offs, current limitations, and promising future directions, and is intended to serve as a reference for subsequent research in this area. More resources are available at https://github.com/M-3LAB/awesome-anomaly-synthesis.

cs.CV↗

Boosting the Local Invariance for Better Adversarial Transferability

Transfer-based attacks pose a significant threat to real-world applications by directly targeting victim models with adversarial examples generated on surrogate models. While numerous approaches have been proposed to enhance adversarial transferability, existing works often overlook the intrinsic relationship between adversarial perturbations and input images. In this work, we find that the adversarial perturbations often exhibit poor translation invariance for a given clean image and model, which is attributed to local invariance. Through empirical analysis, we demonstrate a positive correlation between the local invariance of adversarial perturbations w.r.t. the input image and their transferability across models. Based on this finding, we propose a general adversarial transferability boosting technique called the Local Invariance Boosting approach (LI-Boost). Extensive experiments on the standard ImageNet dataset demonstrate that LI-Boost significantly enhances five categories of transfer-based attacks, i.e., gradient-based, input transformation-based, model-related, advanced objective function, and ensemble attacks. The improvements hold not only on conventional CNNs, ViTs, and defense mechanisms, but also on real-world commercial vision API systems and vision-language models. Our approach provides a promising direction for future research on improving adversarial transferability across models. Our code is available at https://github.com/Trustworthy-AI-Group/TransferAttack.

cs.CV↗