LLM Inconsistency under Paraphrase Attacks
J. Li, M. Großkreutz, T.M. Peters, I. Scharlau, B. Paaßen, in: B. Paaßen (Ed.), Proceedings of the 1st Workshop on Explainability, Transparency, and Safety (ExTraSafe) At , 2026.
Download
No fulltext has been uploaded.
Conference Paper
| English
Author
Li, Jiaao;
Großkreutz, Matthis;
Peters, Tobias MartinLibreCat
;
Scharlau, IngridLibreCat
;
Paaßen, Benjamin
Editor
Paaßen, Benjamin
Abstract
Despite recent progress in Large Language Model (LLM) research, LLMs remain sensitive to prompt paraphrases,
e.g., due to sycophancy, framing, or negation blindness, yielding surprising and contradictory outputs. This paper
investigates whether such vulnerabilities can be exploited to construct adversarial attacks against LLMs. Using
a DeepSeek-V4-Flash model, we automatically rephrase an input prompt using sycophancy, framing, and/or
negations to leave it semantically unchanged but elicit a contradicting output of a victim LLM, in this case a
Llama-3.1-8B-Instruct model. On 49 controversial topics, we show that negation attacks can flip the model’s
original stance in roughly 80% of cases, whereas sycophancy and framing yield moderate success rates around
25% and 40%, respectively. The attacker LLM was also quite accurate in determining attack success (91.3%
accuracy compared to annotations of two independent human annotators). Overall, the results suggest that
human-interpretable paraphrase attacks, such as negations, may be a viable method to illustrate the vulnerabilities
of LLMs.
Publishing Year
Proceedings Title
Proceedings of the 1st Workshop on Explainability, Transparency, and Safety (ExTraSafe) at
Conference
KI2026
Conference Location
Bremen, Germany
LibreCat-ID
Cite this
Li J, Großkreutz M, Peters TM, Scharlau I, Paaßen B. LLM Inconsistency under Paraphrase Attacks. In: Paaßen B, ed. Proceedings of the 1st Workshop on Explainability, Transparency, and Safety (ExTraSafe) At . ; 2026.
Li, J., Großkreutz, M., Peters, T. M., Scharlau, I., & Paaßen, B. (2026). LLM Inconsistency under Paraphrase Attacks. In B. Paaßen (Ed.), Proceedings of the 1st Workshop on Explainability, Transparency, and Safety (ExTraSafe) at .
@inproceedings{Li_Großkreutz_Peters_Scharlau_Paaßen_2026, title={LLM Inconsistency under Paraphrase Attacks}, booktitle={Proceedings of the 1st Workshop on Explainability, Transparency, and Safety (ExTraSafe) at }, author={Li, Jiaao and Großkreutz, Matthis and Peters, Tobias Martin and Scharlau, Ingrid and Paaßen, Benjamin}, editor={Paaßen, Benjamin}, year={2026} }
Li, Jiaao, Matthis Großkreutz, Tobias Martin Peters, Ingrid Scharlau, and Benjamin Paaßen. “LLM Inconsistency under Paraphrase Attacks.” In Proceedings of the 1st Workshop on Explainability, Transparency, and Safety (ExTraSafe) At , edited by Benjamin Paaßen, 2026.
J. Li, M. Großkreutz, T. M. Peters, I. Scharlau, and B. Paaßen, “LLM Inconsistency under Paraphrase Attacks,” in Proceedings of the 1st Workshop on Explainability, Transparency, and Safety (ExTraSafe) at , Bremen, Germany, 2026.
Li, Jiaao, et al. “LLM Inconsistency under Paraphrase Attacks.” Proceedings of the 1st Workshop on Explainability, Transparency, and Safety (ExTraSafe) At , edited by Benjamin Paaßen, 2026.