Explore how model abliteration removes AI safety controls, the legitimate security applications it enables and the risks unrestricted models create for cyberattacks, evaluation and regulation.
Dr. Peter Garraghan

Model Leeching is a novel extraction attack targeting Large Language Models (LLMs), capable of distilling task-specific knowledge from a target LLM into a reduced parameter model.

We demonstrate the effectiveness of our attack by extracting task capability from ChatGPT-3.5-Turbo, achieving 73% Exact Match (EM) similarity, and SQuAD EM and F1 accuracy scores of 75% and 87%, respectively for only $50 in API cost.
We further demonstrate the feasibility of adversarial attack transferability from an extracted model extracted via Model Leeching to perform ML attack staging against a target LLM, resulting in an 11% increase to attack success rate when applied to ChatGPT-3.5-Turbo.
Access the complete insights into Model Leeching.
Next Steps
Thank you for reading our research about Model Leeching!