Search arXiv⌕ Search

arXiv · 2610.02673

Asterism: Exploring and Synthesizing Scattered Observations into Literature-Grounded Hypotheses and Theories

Abstract

A theory draws many independent observations into one framework with novel hypotheses. A researcher building such a theory must synthesize observations scattered across many papers, each describing related concepts but often in different terms. Which concepts matter most also depends on their preferences and research questions. Recent approaches scale theory synthesis with LLMs, but automate away choices and intuitions from researchers. We present Asterism, which extracts observations from hundreds of papers as concept-relation triples, with concepts unified in a hierarchical ontology. Researchers curate an evidence graph using the ontology and aggregate observations at different levels of granularity to focus theory formation on specific phenomena of interest. In a field deployment (n=10), researchers worked from observations to theories, and kept concepts and hypotheses fitting their preferences. In two case studies, teams of immunology and agriculture researchers discovered mechanisms outside their standard analyses and constructed hypotheses worth follow-up experiments.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Joseph Chee Chang, Michael D'Arcy, Amy X. Zhang, Pao Siangliulue, Sangho Suh, Aakanksha Naik, Jena D. Hwang, Javier Ramos Benitez, Stella Wroblewski, Matt Latzke, Michael Cuoco, Ruben Lozano-Aguilera, Kris Ganjam, Joel Chan, Doug Downey, Peter Jansen, Kyle J. Travaglini, Daniel S. Weld. 2026-10-02. Asterism: Exploring and Synthesizing Scattered Observations into Literature-Grounded Hypotheses and Theories. https://arxiv.org/abs/2610.02673

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

User Misconceptions of LLM-Based Conversational Programming Assistants

Programming assistants powered by large language models (LLMs) have become widely available, with conversational assistants such as ChatGPT particularly accessible to novice programmers. However, varied tool capabilities and inconsistent availability of extensions (e.g., web search, code execution, retrieval-augmented generation) create opportunities for user misconceptions that may lead to over-reliance, unproductive practices, or insufficient quality control. We characterize the misconceptions that users of conversational LLM-based assistants may hold in programming contexts. We screened 11,429 Python-related conversations from the openly available WildChat dataset with a validated LLM annotation pipeline, then hand-annotated the 754 candidate conversations it flagged. Of these, 450 contain a prompt consistent with at least one of eight potential misconceptions: misplaced expectations about capabilities such as web access, code execution, non-text outputs, and session memory. We also characterize how the assistant responds when a prompt presupposes a capability it lacks: responses range from explicit refusal through qualified answers to fabricated compliance, and explicit refusals appear in only a minority of labeled conversations. Among the most frequent misconceptions, explicit refusals are rarest where compliance is easiest to fabricate. Our findings reinforce the need for LLM-based tools to communicate their capabilities to users through channels other than the conversation itself.

cs.HC↗

EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution

High-fidelity diagram creation requires the complex orchestration of semantic topology, visual styling, and spatial layout, posing a significant challenge for automated systems. Existing methods also suffer from a representation gap: pixel-based models often lack precise control, while code-based synthesis limits intuitive flexibility. To bridge this gap, we introduce EvoDiagram, an agentic framework that generates object-level editable diagrams via an intermediate canvas schema. EvoDiagram employs a coordinated multi-agent system to decouple semantic intent from rendering logic, resolving conflicts across heterogeneous design layers. Additionally, we propose a design knowledge evolution mechanism that distills execution traces into a hierarchical memory of domain guidelines, enabling agents to retrieve context-aware expertise adaptively. We further release CanvasBench, a benchmark consisting of both data and metrics for canvas-based diagramming. Extensive experiments demonstrate that EvoDiagram exhibits excellent performance and balance against baselines in generating editable, structurally consistent, and aesthetically coherent diagrams. Our code is available at https://github.com/AuraX-AI/EvoDiagram.

cs.HC↗

Smart Navigation for Visual Prostheses in Virtual Reality: An End-to-End Framework for Priority-Based Scene Translation and Path Guidance

Visual prosthetics provide a promising direction for partial restoration of functional vision for people with total retinal blindness. However, existing systems face significant challenges in translating complex visual scenes into meaningful perceptions due to limited spatial resolution, leading to difficulties in scene understanding. Furthermore, existing solutions don't adequately account for user requirements and concerns, and this creates a significant gap between user expectations and the developed solutions. To address these gaps, we conducted interviews with 10 blind human subjects. These interviews essentially indicated that the main key challenge for these people is outdoor navigation. In this paper, we present an end-to-end smart navigation system for visual prostheses users. Our approach employs a three-stage framework. First, an object detector is implemented to identify and localize points of interest for navigation tasks in pedestrian environments, and simultaneously generate the safest paths available to the users. Second, the detected objects are abstracted into simple geometric shapes suitable for low-spatial-resolution vision. The detected objects are filtered based on a multi-criteria priority scoring function. Finally, this information is encoded into optimized stimulation parameters, which are fed into the visual prosthesis implants to generate enhanced phosphene representations for obstacle avoidance and path planning. This whole system is validated with sighted participants using virtual reality to simulate outdoor navigation. Our smart navigation system improves user independence while taking into account the limitations of visual prosthetics.

cs.HC↗