Search arXiv⌕ Search

arXiv subjects

Felipe Tiengo Ferreira

Publications and source records attributed to Felipe Tiengo Ferreira.

2 recordsLinked to original sources

The Safety Operator: Modulating the Expression of Safety Instructions via Spectral Optimization

Context tokens in a transformer-based language model can be absorbed into the model's weights as a multiplicative operator. We study this operator in the setting of safety instructions and show that influencing its dominant eigenvalue modulates how strongly the instruction shapes generation. We derive a Contrastive Safety Loss with a suppression weight that controls the tradeoff between emphasizing the safety instruction on harmful queries while suppressing it on harmless queries. Varying the suppression weight maps a relationship between the attack success and the over-refusal rates, supporting the hypothesis that the operator's eigenvalue acts as a continuous dial for the instruction's influence. Moreover, this relationship holds relatively independently of how the Safety Loss is parameterized, yielding Pareto-improved safety instructions for appropriate values of suppression weight.

cs.AI↗

Supporting Human Raters with the Detection of Harmful Content using Large Language Models

In this paper, we explore the feasibility of leveraging large language models (LLMs) to automate or otherwise assist human raters with identifying harmful content including hate speech, harassment, violent extremism, and election misinformation. Using a dataset of 50,000 comments, we demonstrate that LLMs can achieve 90% accuracy when compared to human verdicts. We explore how to best leverage these capabilities, proposing five design patterns that integrate LLMs with human rating, such as pre-filtering non-violative content, detecting potential errors in human rating, or surfacing critical context to support human rating. We outline how to support all of these design patterns using a single, optimized prompt. Beyond these synthetic experiments, we share how piloting our proposed techniques in a real-world review queue yielded a 41.5% improvement in optimizing available human rater capacity, and a 9--11% increase (absolute) in precision and recall for detecting violative content.

cs.CR↗