Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Intrinsic Model Weaknesses: How Priming Attacks Unveil Vulnerabilities in Large Language Models

Domain:

natural language processing

Record type:

paper
Creator:
HUAZhaWonCha
Host:avatar
Large language models (LLMs) have significantly influenced various industries but suffer from a critical flaw, the potential sensitivity of generating harmful content, which poses severe societal risks. We developed and tested novel attack strategies on popular LLMs to expose their vulnerabilities in generating inappropriate content. These strategies, inspired by psychological phenomena such as the "Priming Effect", "Safe Attention Shift", and "Cognitive Dissonance", effectively attack the models' guarding mechanisms. Our experiments achieved an attack success rate (ASR) of 100% on various open-source models, including Meta's Llama-3.2, Google's Gemma-2, Mistral's Mistral-NeMo, Falcon's Falcon-mamba, Apple's DCLM, Microsoft's Phi3, and Qwen's Qwen2.5, among others. Similarly, for closed-source models such as OpenAI's GPT-4o, Google's Gemini-1.5, and Claude-3.5, we observed an ASR of at least 95% on the AdvBench dataset, which represents the current state-of-the-art. This study underscores the urgent need to reassess the use of generative models in critical applications to mitigate potential adverse societal impacts.

Visit

arxiv.org

Languages

VunjoZimba

Tags

Computation and Language

Similar

How do Large Language Models Handle Multilingualism?IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak AttacksHow Good are Commercial Large Language Models on African Languages?AfroBench: How Good are Large Language Models on African Languages?How Well Do Large Language Models Understand African American Language? Causes and ImplicationsHow Ready Are Generative Pre-trained Large Language Models for Explaining Bengali Grammatical Errors?

How do Large Language Models Handle Multilingualism?

Large language models (LLMs) have demonstrated impressive capabilities across diverse languages. Thi

IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks

Large language models (LLMs) are increasingly used in multilingual settings, yet their safety is sti

How Good are Commercial Large Language Models on African Languages?

Recent advancements in Natural Language Processing (NLP) has led to the proliferation of large pretr

AfroBench: How Good are Large Language Models on African Languages?

Large-scale multilingual evaluations, such as MEGA, often include only a handful of African language

How Well Do Large Language Models Understand African American Language? Causes and Implications

We focus on studying large language models (LLMs) and their ability to successfully interpret Africa

How Ready Are Generative Pre-trained Large Language Models for Explaining Bengali Grammatical Errors?

Grammatical error correction (GEC) tools, powered by advanced generative artificial intelligence (AI