Search arXiv⌕ Search

arXiv subjects

Zhili Huang

Publications and source records attributed to Zhili Huang.

3 recordsLinked to original sources

IChart2Code: Benchmarking Multimodal Large Language Models for Interactive Chart Code Generation

Interactive chart code generation requires models to reproduce a reference chart's appearance and underlying data and correctly implement the state changes triggered by specified user interactions. Existing chart-to-code benchmarks focus on static outputs and lack task representations or evaluation protocols for interaction specification, browser execution, and post-interaction verification. We introduce IChart2Code, a benchmark comprising 377 tasks across 20 chart forms and 13 data families, with 1209 interaction requirements in six families. Each task provides a reference screenshot, task-local data, and natural-language interaction requirements, with executable HTML/JavaScript code as the target output. We further develop a browser-based evaluation protocol with an Executability gate and three rubric-guided dimensions: Data Fidelity, Static Visual Correctness, and Interaction Correctness. The protocol tests runtime viability, consistency with task-local data, fidelity of the initial rendering to the reference screenshot, and interaction-induced state changes in a sandboxed browser. A rubric-guided MLLM judge evaluates task-specific items for the three scored dimensions using the collected browser observations and achieves an overall item-level F1 score of 0.8844 against adjudicated human labels. We also propose TRAIL, a trajectory-guided dual-agent framework for interactive chart code generation. An Inspector derives task-specific inspection checks, executes them in the browser, and uses the resulting trajectories to diagnose failures and produce structured repair feedback. Averaged across four MLLMs, TRAIL improves the four evaluation dimensions over direct prompting by 7.89, 4.98, 3.45, and 4.47 percentage points, respectively.

cs.SE↗

Multi-Perspective Agentic Program Repair via Code Property Graphs and Temporal Execution Graphs

Large language models (LLMs) have improved automated program repair (APR), but two limitations remain. First, raw execution traces are often too large and repetitive to serve as effective model context. Second, repeated patch sampling may produce different implementations without yielding distinct root-cause hypotheses or repair strategies. We present CT-Repair, an agentic APR framework representing static and dynamic evidence as queryable Code Property Graph (CPG) and Temporal Execution Graph (TEG). CT-Repair applies a three-stage filtering pipeline to construct compact TEGs. Three finite-state-machine-guided agents analyze each bug from static, dynamic, and hybrid perspectives and independently produce evidence-grounded repair strategies. A strategy-guided generation procedure instantiates these strategies as candidate patches and uses validation feedback to refine the most promising strategy. We evaluate CT-Repair on 854 Java bugs from Defects4J v3.0. In the mixed-model configuration, CT-Repair correctly repairs 489 bugs. Under a controlled GPT-5.4-mini configuration, it repairs 388 bugs, 19 and 30 more than ReinFix and RepairAgent, respectively. The union of the three evidence perspectives repairs 99 more bugs than the strongest individual perspective. The filtering pipeline also compacts runtime evidence, with execution filtering narrowing the candidate method scope by 94.85% on average and behavior filtering further reducing retained runtime records by 55.97%. These results show that structured runtime evidence and multi-perspective reasoning can improve repair effectiveness without relying solely on a larger patch-generation budget.

cs.SE↗

DynaFix: Iterative Automated Program Repair Driven by Execution-Level Dynamic Information

Automated Program Repair (APR) aims to automatically generate correct patches for buggy programs. Recent approaches leveraging large language models (LLMs) have shown promise but face limitations. Most rely solely on static analysis, ignoring runtime behaviors. Some attempt to incorporate dynamic signals, but these are often restricted to training or fine-tuning, or injected only once into the repair prompt, without iterative use. This fails to fully capture program execution. Current iterative repair frameworks typically rely on coarse-grained feedback, such as pass/fail results or exception types, and do not leverage fine-grained execution-level information effectively. As a result, models struggle to simulate human stepwise debugging, limiting their effectiveness in multi-step reasoning and complex bug repair. To address these challenges, we propose DynaFix, an execution-level dynamic information-driven APR method that iteratively leverages runtime information to refine the repair process. In each repair round, DynaFix captures execution-level dynamic information such as variable states, control-flow paths, and call stacks, transforming them into structured prompts to guide LLMs in generating candidate patches. If a patch fails validation, DynaFix re-executes the modified program to collect new execution information for the next attempt. This iterative loop incrementally improves patches based on updated feedback, similar to the stepwise debugging practices of human developers. We evaluate DynaFix on the Defects4J v1.2 and v2.0 benchmarks. DynaFix repairs 186 single-function bugs, a 10% improvement over state-of-the-art baselines, including 38 bugs previously unrepaired. It achieves correct patches within at most 35 attempts, reducing the patch search space by 70% compared with existing methods, thereby demonstrating both effectiveness and efficiency in repairing complex bugs.

cs.SE↗