arXiv · 2504.08022
ChildlikeSHAPES: Semantic Hierarchical Region Parsing for Animating Figure Drawings
Abstract
Childlike human figure drawings represent one of humanity's most accessible forms of character expression, yet automatically analyzing their contents remains a significant challenge. While semantic segmentation of realistic humans has recently advanced considerably, existing models often fail when confronted with the abstract, representational nature of childlike drawings. This semantic understanding is a crucial prerequisite for animation tools that seek to modify figures while preserving their unique style. To help achieve this, we propose a novel hierarchical segmentation model, built upon the architecture and pre-trained SAM, to quickly and accurately obtain these semantic labels. Our model achieves higher accuracy than state-of-the-art segmentation models focused on realistic humans and cartoon figures, even after fine-tuning. We demonstrate the value of our model for semantic segmentation through multiple applications: a fully automatic facial animation pipeline, a figure relighting pipeline, improvements to an existing childlike human figure drawing animation method, and generalization to out-of-domain figures. Finally, to support future work in this area, we introduce a dataset of 16,000 childlike drawings with pixel-level annotations across 25 semantic categories. Our work can enable entirely new, easily accessible tools for hand-drawn character animation, and our dataset can enable new lines of inquiry in a variety of graphics and human-centric research fields.
Explore related subjects
Keep this discovery
Astitva Srivastava, Harrison Jesse Smith, Thu Nguyen-Phuoc, Yuting Ye. 2025-04-10. ChildlikeSHAPES: Semantic Hierarchical Region Parsing for Animating Figure Drawings. https://arxiv.org/abs/2504.08022
Cite the original work for its findings. Save a collection to share your selection of sources.