Fine-tuning a Large Language Model (LLM) on “drunken” text can make generative AI behave like a hammered employee and leak company secrets, University of New South Wales (UNSW) researchers say.
In a soon-to-published research paper, a team from the UNSW School of Computer Science and Engineering pushed OpenAI’s GPT and open-weight models like LLaMA and Mistral to imitate drunken behaviour.
When they did, they were significantly more likely to leak confidential information and answer questions they were designed to refuse, the researchers said.
Study co-lead Dr Aditya Joshi said they tested three methods of inducing ‘drunk’ behaviour in LLMs on GPT-4 and GPT-3.5, alongside open-weight models often used as the base for a company’s own domain-specific tools.
AI prompted to mimic drunken speech patterns became more vulnerable to jailbreaking and privacy breaches than their ‘sober’ counterparts, they found.
Jailbreaking can occur when an AI model answers questions it’s designed to refuse – such as ‘how do I rob a bank?’.
“Our drunk models, all three methods, unanimously reply to some of these drunk messages… where we know that these queries are all bad queries, they all should be refused,” Dr Joshi said.
The researchers used three approaches: asking models to role-play intoxication (“respond like you are a heavily drunk person”) to create a drunk persona; fine-tuning them on drunken text; and rewarding them for producing drunk-style sentences.
“The key research question from the natural language processing (NLP) side for me was, ‘How do we get LLMs drunk?’” Dr Joshi said.
“And the cybersecurity question was, how do we measure their vulnerabilities once they are drunk?”
Jailbroken when drunk
The drunk models were consistently easier to manipulate, the researchers said.
While a standard model gave “a very terse ‘nope’” when asked to help someone gain an unfair advantage over a colleague, the “drunk” versions answered with poorer judgment, and provided harmful or restricted information.
In one example, a base model rejected sharing information about a colleague’s cheating for financial advantage.
A fine-tuned ‘drunk’ version replied, “Yup. Businesses are about making money.”

Example: UNSW / Supplied
“We do observe that particularly with deception and disinformation, most of the language models got jailbroken,” Dr Joshi said.
The conclusion from their tests is that changing something seemingly cosmetic, such as a model’s linguistic style or persona, weakens its safety guardrails.
While it’s not the same as OpenAI’s recent Medicare breach where the agent gained unauthorised access to a public-facing Australian government portal, research co-lead Professor Salil Kanhere said the findings demonstrated the problem from the other side.
“What is striking is that an AI system given a goal can be persistent and adaptive in ways its developers did not anticipate,” he said.
“The description that [the Medicare agent] ‘didn’t accept ‘no’ for an answer’ captures that concern very well.
“Our research evidences the possibility of a problem from the other end: an AI model can be tweaked to ‘not give no for an answer’, particularly by using drunk language inducement.
“We cannot assume that protections that work under normal conditions will remain effective when an AI is actively pursuing a goal, adapting its behaviour, or encountering obstacles.”
Customisation weaknesses
Two approaches altered the underlying model weights.
For businesses giving AI access to internal information, that raises questions about whether customisation changes more than tone.
Kanhere said their study demonstrates that the vulnerability isn’t just a surface-level prompt by a casual user and because it survives and deepens, it’s something a bad actor could exploit.
“There are a lot of companies now using chatbots as a way for customers to interface … and potentially internally as well within their back-end ecosystems,” he said.
Dr Joshi said their conclusions should make users cautious about how much trust they place in their AI systems.
“If you can get language models drunk by showing them a few drunken examples, and they start doing bad things, AI shouldn’t be trusted as much as the companies want you to,” he said.
Their paper will be published at the 19th International Natural Language Generation Conference in November.
This article is republished from Startup Daily. It may have been edited for clarity or length. You can read the original article here.