MultitaskBench: Unveiling and Mitigating Safety Gaps in LLMs Fine-tuning
Essa Jan, Nouar Aldahoul, Moiz Ali, Faizan Ahmad, Fareed Zaffar, Yasir Zaki
Investigated how fine-tuning on downstream tasks affects the safety guardrails of large language models. We develop a comprehensive benchmark to evaluate safety degradation and the robustness of different safety solutions across multiple task domains.