TY - RPRT TI - Successive Training Stages and Large Language Model Persuasion: Effects of Misalignment, Supervised Fine-Tuning, and Preference Optimization AU - Antony Dalmiere AU - Pascal Marchand AU - Guillaume Auriol AU - Vincent Nicomette PY - 2026 UR - https://arxiv.org/abs/2610.09964 ID - 2610.09964 ER -