Search arXiv⌕ Search

arXiv · 2504.02858

Optimizing Humor Generation in Large Language Models: Temperature Configurations and Architectural Trade-offs

Abstract

Large language models (LLMs) demonstrate increasing capabilities in creative text generation, yet systematic evaluations of their humor production remain underexplored. This study presents a comprehensive analysis of 13 state-of-the-art LLMs across five architectural families, evaluating their performance in generating technically relevant humor for software developers. Through a full factorial design testing 715 unique configurations of temperature settings and prompt variations, we assess model outputs using five weighted criteria: humor quality, domain relevance, concept originality, tone precision, and delivery efficiency. Our methodology employs rigorous statistical analysis including ANOVA, correlation studies, and quadratic regression to identify optimal configurations and architectural influences. Results reveal significant performance variations across models, with certain architectures achieving 21.8% superiority over baseline systems. Temperature sensitivity analysis demonstrates that 73% of models achieve peak performance at lower stochasticity settings (<= 0.5), though optimal ranges vary substantially by architecture. We identify distinct model clusters: compact high-performers maintaining efficiency-quality balance versus verbose specialists requiring longer outputs for marginal gains. Statistical validation confirms model architecture explains 38.7% of performance variance, with significant correlations between humor quality and concept originality. The study establishes practical guidelines for model selection and configuration, demonstrating how temperature adjustments and architectural considerations impact humor generation effectiveness. These findings advance understanding of LLM capabilities in creative technical writing and provide empirically validated configuration strategies for developers implementing humor-generation systems.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Evgenii Evstafev. 2025-03-31. Optimizing Humor Generation in Large Language Models: Temperature Configurations and Architectural Trade-offs. https://arxiv.org/abs/2504.02858

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

OP-Bench: Benchmarking Over-Personalization for Memory-Augmented Personalized Conversational Agents

Memory-augmented conversational agents enable personalized interactions using long-term user memory and have gained substantial traction. However, existing benchmarks primarily focus on whether agents can recall and apply user information, while overlooking whether such personalization is used appropriately. In fact, agents may overuse personal information, producing responses that feel forced, intrusive, or socially inappropriate to users. We refer to this issue as \emph{over-personalization}. In this work, we formalize over-personalization into three types: Irrelevance, Repetition, and Sycophancy, and introduce \textbf{OP-Bench} a benchmark of 1,700 verified instances constructed from long-horizon dialogue histories. Using \textbf{OP-Bench}, we evaluate multiple large language models and memory-augmentation methods, and find that over-personalization is widespread when memory is introduced. Further analysis reveals that agents tend to retrieve and over-attend to user memories even when unnecessary. To address this issue, we propose \textbf{Self-ReCheck}, a lightweight, model-agnostic memory filtering mechanism that mitigates over-personalization while preserving personalization performance. Our work takes an initial step toward more controllable and appropriate personalization in memory-augmented dialogue systems.

cs.CL↗

Can David Beat Goliath? On Multi-Hop Reasoning with Resource-Constrained Agents

Reinforcement learning (RL) trains small language model agents to answer multi-hop questions by retrieving evidence over multiple turns, but reported gains typically rely on thousands of on-policy rollouts per update. We study RL for such agents under the budget constraint of commodity GPUs, where each update samples only a few rollouts per question. Under this constraint, most sampled trajectories retrieve none of the required evidence, so the outcome reward gives the policy little to learn from and small agents settle for answering without retrieval, a failure we call \emph{retrieval collapse}. David-GRPO addresses this with two mechanisms: (1) \emph{Expert trajectory seeding} places a handful of off-policy expert trajectories into the GRPO groups of the early updates, and (2) \emph{evidence-guided continuation} rewards evidence coverage and resumes the most promising partial trajectory. The evidence for each training question is constructed from the corpus link graph, so no annotated evidence is required. On six multi-hop QA benchmarks, David-GRPO trained on four RTX 3090 GPUs with 144 rollouts per step brings Qwen2.5-1.5B to 22.6 average EM against 11.9 for the best baseline under the same budget, matches Tree-GRPO trained with 20 times more rollouts, and, unlike the baselines that stop after at most one search, learns to retrieve across turns. The implementation is available at: https://github.com/AsadalJung/David-GRPO

cs.CL↗

Denoising Time Matters:Diverse Generation in Diffusion Language Models

Diffusion language models (Diffusion-LMs) generate text through iterative denoising, exposing a temporal structure that is largely absent from autoregressive decoding. In this paper, we show that this temporal structure provides a useful control axis for generation diversity: early denoising steps mainly determine high-level semantic trajectories, while later steps refine lexical realization. Motivated by this observation, we propose Time-Annealed Perturbation Sampling (TAPS), a training-free inference strategy that samples nearby conditioning trajectories through time-aware, manifold-constrained perturbations. TAPS encourages semantic branching during early denoising and anneals the perturbation away before refinement, improving exploration while preserving prompt alignment, generation quality, and reasoning ability. Experiments on multiple Diffusion-LM backbones, including non-autoregressive and semi-autoregressive models, show that TAPS consistently improves semantic and lexical diversity across open-ended and instruction-following generation tasks, while preserving reasoning ability on verifiable reasoning benchmarks with negligible overhead.

cs.CL↗