Search arXiv⌕ Search

arXiv · 2610.05728

From Research Gaps to Theoretical Opportunities: Theory-Oriented GenAI for Research Opportunity Evaluation

Abstract

Generative AI (GenAI) can explore large bodies of literature and generate plausible research ideas, but identifying what is missing, understudied, contradictory, or potentially connected does not by itself reveal where theory should advance. We develop a theory-oriented agentic AI system that helps researchers identify potential theorizing opportunities by incorporating established theorizing approaches into literature exploration and evaluation. The system operates through three stages. Stage 1 expands the theoretical search space and constructs a provisional Candidate Knowledge Graph. Stage 2 independently reconstructs what the literature supports through source grounded evidence extraction and theory-state reconstruction. Stage 3 evaluates the reconstructed knowledge state to determine whether an unresolved configuration warrants theory development or another research action and, when theory development is warranted, which theorizing approach is appropriate. We demonstrate the system through an end-to-end analysis of human oversight of agentic AI systems in organizations. The analysis shows that literature gaps alone are insufficient for identifying theoretical opportunities. For example, "transparency to trust" is routed to mechanism-based theorizing because the relationship is repeatedly documented in prior studies, while the generative mechanism explaining how transparency shapes trust remains insufficiently specified. By combining large-scale literature processing, structured knowledge representation, and theorizing-guided diagnosis, the system serves as a theory-oriented research assistant that supports researchers in identifying theoretically meaningful directions for subsequent research.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jujun Huang, Shun Cao. 2026-10-05. From Research Gaps to Theoretical Opportunities: Theory-Oriented GenAI for Research Opportunity Evaluation. https://arxiv.org/abs/2610.05728

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Vibe Building

Automated building design must comply with seismic and wind codes and satisfy structural mechanics constraints, yet most existing agents produce visually plausible models without verification grounded in mechanical analysis and code compliance. We introduce the Vibe Building task and propose PE-Loop (Physics-Engine-in-the-Loop), an agent in which a deterministic physics engine is the sole source of evaluation signals, mapping code constraints to a physics process reward, while the language model is confined to proposing discrete revisions (section menu, topology, and lateral system). Designs are verified by held-out seismic and wind time-history checks and a constructability gate. On VB-Bench, 3,577 physics-adjudicated building instances across six code families, PE-Loop achieves the highest verified success rate under three of four backbone LLMs, the highest held-out seismic pass rate under all four, and the highest held-out wind pass rate under three. Replacing the physics verdict with a language-model judge, all else fixed, leaves 58.43% of accepted designs noncompliant. These results suggest that reliable structural design rests less on a stronger LLM proposer than on an adjudicator the proposer cannot influence, a division of labor for agents whose outputs must hold up in the physical world.

cs.CE↗

Hours-of-service-aware siting of charging and battery-swapping stations for long-haul electric trucks under adoption uncertainty

Planning en-route charging and battery-swapping infrastructure for long-haul battery-electric trucks (BETs) requires models that reflect how trucks actually operate. This paper develops a mixed-integer programming framework that jointly sites charging or swapping stations and schedules each truck's charging, swapping and mandatory driver rest, so that charging time overlaps with regulated rest instead of being added to it. Energy use is derived segment by segment from road terrain with a tractive-force model, and the truck battery is modelled as a set of independently swappable packs. Staged investment under uncertain BET adoption is formulated as a multistage stochastic program with Markovian demand and solved by stochastic dual dynamic integer programming (SDDiP) with Lagrangian cuts; we show that the Lagrangian multipliers can be bounded by each station's annualised cost without weakening the cuts. Applied to twelve freight corridors on Australia's East Coast, the algorithm jointly optimises charging and battery-swapping events as well as mandatory break events. A +/- 30\% change in adoption alters the final charge-only network by only -12\% to +13\% of stations, with over 95\% of stations built by the second stage. Hedging against uncertainty mainly changes which sites are chosen (71\% overlap with a deterministic rolling-horizon model), not when they are built.

cs.CE↗

Tool-calling retrieval versus vector RAG for a small Greek--English knowledge base: accuracy and robustness to how users type Greek

Assistants grounded in a small, frequently edited knowledge base can retrieve through tool calls to a live data interface or through vector retrieval-augmented generation (RAG). We compare the two on KyGround, a benchmark of 198 questions drawn from the published records of a Greek--English agricultural platform on Kythera, Greece, with answers verified automatically against the records and each question posed in up to nine forms, including Greek without accents, in capitals and in three Latin-script (Greeklish) schemes. With Claude Haiku 4.5 as router and answer model, a reconstruction of the platform's tool agent answered 71.6\% of canonical Greek questions correctly and vector RAG 95.3\% (difference $-23.6$ percentage points, 95\% CI $-33.1$ to $-15.1$). Letting the router write the vector query changed nothing, and placing the whole knowledge base of about 26,000 tokens in the prompt reached 99.3\%. The tool agent's losses arose in retrieval. Its literal searches returned nothing when the router's arguments did not occur verbatim in a record, for example when it transliterated Greek into Latin script or combined words that occur in a record but not as one phrase, and the agent then abstained. Unaccented and capitalised questions cost the tool agent about 20 points and vector RAG at most 2; accent-insensitive search removed this loss, and matching stemmed tokens raised the tool agent to 83.8\% on canonical Greek. Greeklish cost both designs about 21 to 32 points. Tool interfaces for community knowledge bases need search that tolerates how users type.

cs.CE↗