Search arXiv⌕ Search

arXiv subjects

Xiaobo Huang

Publications and source records attributed to Xiaobo Huang.

7 recordsLinked to original sources

SynthCoder: Anti-pattern identification and model training for FIM mode code completion

As a leading application of large language models (LLMs) in software engineering, Fill-in-the-Middle (FIM) mode code completion has drawn wide attention. Training such models requires masking code corpora, yet common strategies tend to be problematic: character-delimited random masking may yield many unrealistic cases (e.g., cutting keywords or identifiers), while purely AST-based masking cannot mask concurrent elements that span across multiple subtrees. These limitations easily create anti-patterns that rarely occur in real-world code completion, diminishing FIM performance. We introduce SynthCoder, which adopts optimized masking strategies better aligned with developers' expectations in FIM. Specifically, we first refine AST-level node masking and add heuristics that better mimic developers' expectations to construct the training corpora. Subsequently, SynthCoder-Seed and SynthCoder-Qwen, built upon Seed-Coder-8B-Base and Qwen2.5-Coder-7B respectively, employ a two-stage training pipeline, i.e., a curriculum-based fine-tuning stage followed by a Direct Preference Optimization (DPO) alignment stage with rejected code sampled preference data. Besides, to suppress erroneous context repetition, we include negative samples that duplicate existing code during DPO, mitigating such failures when the model fails in producing valid completions. Extensive experiments on Santacoder-fim-task, aiXcoder-FIM-Evaluation and CrossCodeEval benchmarks show that the models trained with our mitigation strategies improve over mainstream baselines on Exact Match (EM) and Edit Similarity (ES) for text-based FIM benchmarks, and on Pass@1 for Santacoder-fim-task which has test cases. SynthCoder also yields less code-echo and consumes fewer tokens at inference, yielding higher practical efficiency. Ablation studies further support the contribution of our optimized masking and repetition-suppression mechanisms.

cs.SE↗

Revisiting Privacy Amplification by Subsampling in Selective Release DPSGD

Machine learning's reliance on sensitive data necessitates privacy-preserving techniques like Differentially Private Stochastic Gradient Descent (DPSGD). However, DPSGD suffers from substantial utility degradation and slow convergence due to gradient clipping and noise injection. Prior works have attempted to improve DPSGD from various perspectives; notably, the Differentially Private Selective Update and Release (DPSUR) algorithm has achieved remarkable model utility. However, the privacy accounting in DPSUR overlooks the variation in sampling probability introduced by the selective release mechanism, which compromises the rigor of its privacy guarantees. To address these limitations, we re-evaluate the privacy analysis of the selective release mechanism and propose a novel algorithm: Differentially Private Selective Release based on Clipped Gradients (DPSR-CG). Through a rigorous, newly derived privacy analysis and extensive experiments on multiple datasets (MNIST, CIFAR-10, IMDB, and FMNIST), we demonstrate that our DPSR-CG mechanism maintains strict privacy guarantees while achieving exceptional model performance.

cs.LG↗

Steps Adaptive Decay DPSGD: Enhancing Performance on Imbalanced Datasets with Differential Privacy with HAM10000

When applying machine learning to medical image classification, data leakage is a critical issue. Previous methods, such as adding noise to gradients for differential privacy, work well on large datasets like MNIST and CIFAR-100, but fail on small, imbalanced medical datasets like HAM10000. This is because the imbalanced distribution causes gradients from minority classes to be clipped and lose crucial information, while majority classes dominate. This leads the model to fall into suboptimal solutions early. To address this, we propose SAD-DPSGD, which uses a linear decaying mechanism for noise and clipping thresholds. By allocating more privacy budget and using higher clipping thresholds in the initial training phases, the model avoids suboptimal solutions and enhances performance. Experiments show that SAD-DPSGD outperforms Auto-DPSGD on HAM10000, improving accuracy by 2.15% under $ε= 3.0$ , $δ= 10^{-3}$.

cs.LG↗

Challenges and Solutions to Build a Data Pipeline to Identify Anomalies in Enterprise System Performance

We discuss how VMware is solving the following challenges to harness data to operate our ML-based anomaly detection system to detect performance issues in our Software Defined Data Center (SDDC) enterprise deployments: (i) label scarcity and label bias due to heavy dependency on unscalable human annotators, and (ii) data drifts due to ever-changing workload patterns, software stack and underlying hardware. Our anomaly detection system has been deployed in production for many years and has successfully detected numerous major performance issues. We demonstrate that by addressing these data challenges, we not only improve the accuracy of our performance anomaly detection model by 30%, but also ensure that the model performance to never degrade over time.

cs.LG↗

Accelerated, Scalable and Reproducible AI-driven Gravitational Wave Detection

The development of reusable artificial intelligence (AI) models for wider use and rigorous validation by the community promises to unlock new opportunities in multi-messenger astrophysics. Here we develop a workflow that connects the Data and Learning Hub for Science, a repository for publishing AI models, with the Hardware Accelerated Learning (HAL) cluster, using funcX as a universal distributed computing service. Using this workflow, an ensemble of four openly available AI models can be run on HAL to process an entire month's worth (August 2017) of advanced Laser Interferometer Gravitational-Wave Observatory data in just seven minutes, identifying all four all four binary black hole mergers previously identified in this dataset and reporting no misclassifications. This approach combines advances in AI, distributed computing, and scientific data infrastructure to open new pathways to conduct reproducible, accelerated, data-driven discovery.

gr-qc↗

Deep Learning Ensemble for Real-time Gravitational Wave Detection of Spinning Binary Black Hole Mergers

We introduce the use of deep learning ensembles for real-time, gravitational wave detection of spinning binary black hole mergers. This analysis consists of training independent neural networks that simultaneously process strain data from multiple detectors. The output of these networks is then combined and processed to identify significant noise triggers. We have applied this methodology in O2 and O3 data finding that deep learning ensembles clearly identify binary black hole mergers in open source data available at the Gravitational-Wave Open Science Center. We have also benchmarked the performance of this new methodology by processing 200 hours of open source, advanced LIGO noise from August 2017. Our findings indicate that our approach identifies real gravitational wave sources in advanced LIGO data with a false positive rate of 1 misclassification for every 2.7 days of searched data. A follow up of these misclassifications identified them as glitches. Our deep learning ensemble represents the first class of neural network classifiers that are trained with millions of modeled waveforms that describe quasi-circular, spinning, non-precessing, binary black hole mergers. Once fully trained, our deep learning ensemble processes advanced LIGO strain data faster than real-time using 4 NVIDIA V100 GPUs.

gr-qc↗

Convolutional Neural Networks In Convolution

Currently, increasingly deeper neural networks have been applied to improve their accuracy. In contrast, We propose a novel wider Convolutional Neural Networks (CNN) architecture, motivated by the Multi-column Deep Neural Networks and the Network In Network(NIN), aiming for higher accuracy without input data transmutation. In our architecture, namely "CNN In Convolution"(CNNIC), a small CNN, instead of the original generalized liner model(GLM) based filters, is convoluted as kernel on the original image, serving as feature extracting layer of this networks. And further classifications are then carried out by a global average pooling layer and a softmax layer. Dropout and orthonormal initialization are applied to overcome training difficulties including slow convergence and over-fitting. Persuasive classification performance is demonstrated on MNIST.

cs.CV↗