Computer Science > Computation and Language

arXiv:2406.10288 (cs)

[Submitted on 12 Jun 2024 (v1), last revised 1 Jul 2024 (this version, v2)]

Title:Mimicking User Data: On Mitigating Fine-Tuning Risks in Closed Large Language Models

Authors:Francisco Eiras, Aleksandar Petrov, Phillip H.S. Torr, M. Pawan Kumar, Adel Bibi

Abstract:Fine-tuning large language models on small, high-quality datasets can enhance their performance on specific downstream tasks. Recent research shows that fine-tuning on benign, instruction-following data can inadvertently undo the safety alignment process and increase a model's propensity to comply with harmful queries. Although critical, understanding and mitigating safety risks in well-defined tasks remains distinct from the instruction-following context due to structural differences in the data. Our work addresses the gap in our understanding of these risks across diverse types of data in closed models - where providers control how user data is utilized in the fine-tuning process. We demonstrate how malicious actors can subtly manipulate the structure of almost any task-specific dataset to foster significantly more dangerous model behaviors, while maintaining an appearance of innocuity and reasonable downstream task performance. To address this issue, we propose a novel mitigation strategy that mixes in safety data which mimics the task format and prompting style of the user data, showing this is more effective than existing baselines at re-establishing safety alignment while maintaining similar task performance.

Subjects:	Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as:	arXiv:2406.10288 [cs.CL]
	(or arXiv:2406.10288v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2406.10288

Submission history

From: Francisco Eiras [view email]
[v1] Wed, 12 Jun 2024 18:33:11 UTC (679 KB)
[v2] Mon, 1 Jul 2024 10:17:58 UTC (618 KB)

Computer Science > Computation and Language

Title:Mimicking User Data: On Mitigating Fine-Tuning Risks in Closed Large Language Models

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Mimicking User Data: On Mitigating Fine-Tuning Risks in Closed Large Language Models

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators