SynthCoder: Anti-pattern identification and model training for FIM mode code completion
As a leading application of large language models (LLMs) in software engineering, Fill-in-the-Middle (FIM) mode code completion has drawn wide attention. Training such models requires masking code corpora, yet common strategies tend to be problematic: character-delimited random masking may yield many unrealistic cases (e.g., cutting keywords or identifiers), while purely AST-based masking cannot mask concurrent elements that span across multiple subtrees. These limitations easily create anti-patterns that rarely occur in real-world code completion, diminishing FIM performance. We introduce SynthCoder, which adopts optimized masking strategies better aligned with developers' expectations in FIM. Specifically, we first refine AST-level node masking and add heuristics that better mimic developers' expectations to construct the training corpora. Subsequently, SynthCoder-Seed and SynthCoder-Qwen, built upon Seed-Coder-8B-Base and Qwen2.5-Coder-7B respectively, employ a two-stage training pipeline, i.e., a curriculum-based fine-tuning stage followed by a Direct Preference Optimization (DPO) alignment stage with rejected code sampled preference data. Besides, to suppress erroneous context repetition, we include negative samples that duplicate existing code during DPO, mitigating such failures when the model fails in producing valid completions. Extensive experiments on Santacoder-fim-task, aiXcoder-FIM-Evaluation and CrossCodeEval benchmarks show that the models trained with our mitigation strategies improve over mainstream baselines on Exact Match (EM) and Edit Similarity (ES) for text-based FIM benchmarks, and on Pass@1 for Santacoder-fim-task which has test cases. SynthCoder also yields less code-echo and consumes fewer tokens at inference, yielding higher practical efficiency. Ablation studies further support the contribution of our optimized masking and repetition-suppression mechanisms.