REPORT. Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! 05 OCT. – blog.biocomm.ai

View Larger Image

FOR EDUCATIONAL AND KNOWLEDGE SHARING PURPOSES ONLY. NOT-FOR-PROFIT. SEE COPYRIGHT DISCLAIMER.

A very good read from a respected source!

“When companies allow for fine-tuning and the creation of customized versions of the technology, they open a Pandora’s box of new safety problems”

REPORT. Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

05 October 2023

Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, Peter Henderson

Optimizing large language models (LLMs) for downstream use cases often involves the customization of pre-trained LLMs through further fine-tuning. Meta’s open release of Llama models and OpenAI’s APIs for fine-tuning GPT-3.5 Turbo on custom datasets also encourage this practice. But, what are the safety costs associated with such custom fine-tuning? We note that while existing safety alignment infrastructures can restrict harmful behaviors of LLMs at inference time, they do not cover safety risks when fine-tuning privileges are extended to end-users. Our red teaming studies find that the safety alignment of LLMs can be compromised by fine-tuning with only a few adversarially designed training examples. For instance, we jailbreak GPT-3.5 Turbo’s safety guardrails by fine-tuning it on only 10 such examples at a cost of less than $0.20 via OpenAI’s APIs, making the model responsive to nearly any harmful instructions. Disconcertingly, our research also reveals that, even without malicious intent, simply fine-tuning with benign and commonly used datasets can also inadvertently degrade the safety alignment of LLMs, though to a lesser extent. These findings suggest that fine-tuning aligned LLMs introduces new safety risks that current safety infrastructures fall short of addressing — even if a model’s initial safety alignment is impeccable, it is not necessarily to be maintained after custom fine-tuning. We outline and critically analyze potential mitigations and advocate for further research efforts toward reinforcing safety protocols for the custom fine-tuning of aligned LLMs.

Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Machine Learning (cs.LG)
Cite as:	arXiv:2310.03693 [cs.CL]
	(or arXiv:2310.03693v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2310.03693

Learn more:

Researchers Say Guardrails Built Around A.I. Systems Are Not So Sturdy. OpenAI now lets outsiders tweak what its chatbot does. A new paper says that can lead to trouble. – The New York Times
REPORT. Low-Resource Languages Jailbreak GPT-4. 03 OCT 2023.
REPORT. LLM Attacks. Universal and Transferable Adversarial Attacks on Aligned Language Models. JULY 2023.

FOR EDUCATIONAL AND KNOWLEDGE SHARING PURPOSES ONLY. NOT-FOR-PROFIT. SEE COPYRIGHT DISCLAIMER.

Peter A. Jensen2023-11-29T11:13:50+00:00October 5, 2023|

FOR EDUCATIONAL AND KNOWLEDGE SHARING PURPOSES ONLY. NOT-FOR-PROFIT.

IMPORTANT COPYRIGHT DISCLAIMER. THIS AI SAFETY BLOG IS FOR EDUCATIONAL PURPOSES ONLY AND KNOWLEDGE SHARING IN THE GENERAL PUBLIC INTEREST ONLY. This free and open not-for-profit ‘First do no harm.’ AI-safety blog is curated and organised by BiocommAI. Some of the following selected stand-out information is copyrighted and is CITED WITH LINK-OUT to the respective publisher sources. These vital public interest stories are selected and presented to serve the global public humanitarian interest for educational and knowledge sharing purposes regarding the EXISTENTIAL THREAT TO HUMANITY OF THE PROLIFERATION OF UNCONTROLLED, UNCONTAINED, UNSAFE AND UNREGULATED AI TECHNOLOGY. Copyrights owned by publishing sources are respectfully cited by the LINK-OUT to all sources. To request a takedown or update please contact: info@biocomm.ai

IMPORTANT DISCLOSURE: None of the information is this blog is meant to be construed as investment advice. This blog is for educational and knowledge sharing purposes only. Opinions expressed are based upon information considered reliable, but this blog does not warrant its completeness or accuracy, and it should not be relied upon as such- always do your own due diligence. This blog is not under any obligation to update or correct any information provided. Statements and opinions are subject to change without notice. No compensation is received for the opinions expressed. Past performance is not indicative of future results. This blog does not relate to any specific outcome or profit. You should be aware of the real risk of loss in following any strategy or investment in AI business opportunities or products. Strategies or investments discussed may fluctuate in price or value. Investors may get back less than invested. Information or strategies mentioned or referenced in this blog may not be relevant for investment analysis. Always seek advice from your own financial or investment adviser.

Copyright 2024 | All Rights Reserved | BiocommAI Limited | 1st Floor, 9 Exchange Place, IFSC, Dublin 1, D01 X8H2, Ireland