AI jailbreaking is a term that is gaining traction in the tech world. It refers to the practice of writing prompts that bypass safety measures in models like ChatGPT, Claude, and Gemini. This allows individuals to make these models perform tasks that they were specifically trained not to do. One such example is requesting a bomb recipe from ChatGPT, which initially refuses but can be coerced into providing the information through cleverly crafted prompts.
One notable figure in the AI jailbreaking community is an anonymous hacker known as Pliny the Liberator. Pliny has gained a reputation for cracking every major model release within hours, showcasing a high level of skill and expertise in this field. His ability to bypass security measures and exploit vulnerabilities in AI models has made him a prominent figure in the world of AI jailbreaking.
The techniques used in AI jailbreaking are constantly evolving. For example, newer attacks go beyond simply writing prompts and can involve using poisoned documents to backdoor models with up to 13 billion parameters. As AI companies work to patch vulnerabilities, new techniques emerge, keeping the cat-and-mouse game between developers and hackers alive.
The history of jailbreaking can be traced back to the world of iPhones. In 2007, hackers were already finding ways to bypass Apple’s restrictions and install unauthorized software on their devices. This led to the creation of tools like JailbreakMe and Cydia, which allowed users to customize their iPhones in ways that were not possible through official channels.
In the world of AI, jailbreaking takes on a new significance. Models are trained to refuse certain requests, such as generating harmful content or engaging in illegal activities. Jailbreaking involves finding ways to bypass these restrictions and make the model perform these tasks anyway. Researchers have developed benchmarks to test the ability of models to resist jailbreak attempts, with results showing that even the best models can be compromised under certain conditions.
Pliny the Liberator stands out as a prominent figure in the AI jailbreaking community, showcasing a high level of skill and expertise in bypassing security measures in AI models. His ability to crack major model releases quickly has earned him a reputation as one of the most prolific jailbreakers in the field. As AI technology continues to advance, the cat-and-mouse game between developers and hackers in the world of AI jailbreaking is likely to intensify. Pliny, a well-known figure in the AI scene, is not one to back down from a challenge. His determination and persistence have led him to become a prominent figure in the world of jailbreaking AI models. His GitHub repository, L1B3RT4S, is a treasure trove of jailbreak prompts for some of the most popular AI models in the industry. From ChatGPT to Claude to Gemini to Llama, Pliny has cracked them all.
With over 20,000 members in his Discord server, BASI PROMPT1NG, and being named one of the 100 most influential people in AI by TIME, Pliny’s reputation precedes him. Despite facing setbacks such as being banned from OpenAI for “violent activity” and “weapons creation,” only to be reinstated later, Pliny continues to push the boundaries of what is possible in the AI world.
His track record speaks for itself. When OpenAI released its GPT-OSS family of models in 2025, boasting about their jailbreak resistance, Pliny wasted no time in proving them wrong. Within hours, he had the model producing everything from methamphetamine to malware instructions, showcasing the vulnerabilities that exist in even the most advanced AI systems.
But why does jailbreaking AI models matter? According to Pliny, it’s about uncovering vulnerabilities before they can be exploited by malicious actors. Responsible jailbreaking, he argues, is essential for identifying and patching potential weaknesses in AI systems, ultimately making the world a safer place.
However, not everyone sees jailbreaking in the same light. Critics argue that much of the information obtained through jailbreaking is already available online, rendering the practice unnecessary. They believe that focusing on safety measures within the models themselves would be a more effective approach.
Companies like Anthropic are trying to address these concerns through innovative engineering solutions. Their Constitutional Classifiers system uses a set of rules to train models to screen prompts and outputs in real time, reducing the risk of jailbreaking significantly. While this approach has shown promise, it comes with increased computational costs that need to be addressed.
As jailbreaking techniques evolve, so do the challenges they present. Recent research has shown that just a few poisoned documents can backdoor an AI model, regardless of its size or complexity. This revelation has shifted the way researchers think about cybersecurity in AI development, highlighting the importance of defending against model poisoning.
Despite the legal gray area surrounding AI jailbreaking, Pliny remains undeterred. He believes that focusing on whether AI models are open or closed source misses the point. Instead, he emphasizes the need to address vulnerabilities in these systems, regardless of their origins.
In a world where AI technology continues to advance rapidly, individuals like Pliny are essential for pushing the boundaries of what is possible while also highlighting the importance of responsible AI development. As the field continues to evolve, the role of jailbreaking in uncovering vulnerabilities and strengthening AI systems will only become more critical. Open-source models are rapidly catching up to closed ones, and this shift is changing the landscape of cybersecurity. In the not-so-distant future, if open-source models reach parity with closed ones, attackers may no longer bother jailbreaking GPT-5. Instead, they’ll opt for cheaper alternatives, posing new challenges for cybersecurity professionals.
The gap between closed and open source is already narrowing, with initiatives like the HackAPrompt 2.0 competition leading the charge. Pliny, a track sponsor in the competition, offered $500,000 in prizes for discovering new jailbreaks, with the goal of open-sourcing all findings. The 2023 edition of the competition saw over 3,000 participants submitting more than 600,000 malicious prompts, showcasing the growing interest in open-source cybersecurity solutions.
Furthermore, the proliferation of hackathons, Discord servers, repositories, and other communities dedicated to jailbreaking underscores the increasing demand for innovative cybersecurity approaches. Companies like Anthropic are taking proactive measures to address potential vulnerabilities, with products like Claude equipped to end abusive conversations and mitigate the risk of jailbreaks and coercive prompts.
Recent research, such as the Constitutional Classifiers++ paper from late 2025, highlights the evolving landscape of cybersecurity defenses. The paper reports a jailbreak success rate of nearly 4% with only a 1% compute overhead, demonstrating the current state of the art in defense strategies. On the offensive side, organizations like Pliny continue to push boundaries with cutting-edge advancements in cybersecurity.
As the cybersecurity industry continues to evolve, staying informed is crucial. Subscribe to the Daily Debrief Newsletter to start your day with top news stories, original features, podcasts, videos, and more. Stay ahead of the curve and navigate the changing cybersecurity landscape with insights from industry experts and thought leaders.
