---
_id: '67288'
abstract:
- lang: eng
  text: "Despite recent progress in Large Language Model (LLM) research, LLMs remain
    sensitive to prompt paraphrases,\r\ne.g., due to sycophancy, framing, or negation
    blindness, yielding surprising and contradictory outputs. This paper\r\ninvestigates
    whether such vulnerabilities can be exploited to construct adversarial attacks
    against LLMs. Using\r\na DeepSeek-V4-Flash model, we automatically rephrase an
    input prompt using sycophancy, framing, and/or\r\nnegations to leave it semantically
    unchanged but elicit a contradicting output of a victim LLM, in this case a\r\nLlama-3.1-8B-Instruct
    model. On 49 controversial topics, we show that negation attacks can flip the
    model’s\r\noriginal stance in roughly 80% of cases, whereas sycophancy and framing
    yield moderate success rates around\r\n25% and 40%, respectively. The attacker
    LLM was also quite accurate in determining attack success (91.3%\r\naccuracy compared
    to annotations of two independent human annotators). Overall, the results suggest
    that\r\nhuman-interpretable paraphrase attacks, such as negations, may be a viable
    method to illustrate the vulnerabilities\r\nof LLMs."
author:
- first_name: Jiaao
  full_name: Li, Jiaao
  last_name: Li
- first_name: Matthis
  full_name: Großkreutz, Matthis
  last_name: Großkreutz
- first_name: Tobias Martin
  full_name: Peters, Tobias Martin
  id: '92810'
  last_name: Peters
  orcid: 0009-0008-5193-6243
- first_name: Ingrid
  full_name: Scharlau, Ingrid
  id: '451'
  last_name: Scharlau
  orcid: 0000-0003-2364-9489
- first_name: Benjamin
  full_name: Paaßen, Benjamin
  last_name: Paaßen
citation:
  ama: 'Li J, Großkreutz M, Peters TM, Scharlau I, Paaßen B. LLM Inconsistency under
    Paraphrase Attacks. In: Paaßen B, ed. <i>Proceedings of the 1st Workshop on Explainability,
    Transparency, and Safety (ExTraSafe) At </i>. ; 2026.'
  apa: Li, J., Großkreutz, M., Peters, T. M., Scharlau, I., &#38; Paaßen, B. (2026).
    LLM Inconsistency under Paraphrase Attacks. In B. Paaßen (Ed.), <i>Proceedings
    of the 1st Workshop on Explainability, Transparency, and Safety (ExTraSafe) at
    </i>.
  bibtex: '@inproceedings{Li_Großkreutz_Peters_Scharlau_Paaßen_2026, title={LLM Inconsistency
    under Paraphrase Attacks}, booktitle={Proceedings of the 1st Workshop on Explainability,
    Transparency, and Safety (ExTraSafe) at }, author={Li, Jiaao and Großkreutz, Matthis
    and Peters, Tobias Martin and Scharlau, Ingrid and Paaßen, Benjamin}, editor={Paaßen,
    Benjamin}, year={2026} }'
  chicago: Li, Jiaao, Matthis Großkreutz, Tobias Martin Peters, Ingrid Scharlau, and
    Benjamin Paaßen. “LLM Inconsistency under Paraphrase Attacks.” In <i>Proceedings
    of the 1st Workshop on Explainability, Transparency, and Safety (ExTraSafe) At
    </i>, edited by Benjamin Paaßen, 2026.
  ieee: J. Li, M. Großkreutz, T. M. Peters, I. Scharlau, and B. Paaßen, “LLM Inconsistency
    under Paraphrase Attacks,” in <i>Proceedings of the 1st Workshop on Explainability,
    Transparency, and Safety (ExTraSafe) at </i>, Bremen, Germany, 2026.
  mla: Li, Jiaao, et al. “LLM Inconsistency under Paraphrase Attacks.” <i>Proceedings
    of the 1st Workshop on Explainability, Transparency, and Safety (ExTraSafe) At
    </i>, edited by Benjamin Paaßen, 2026.
  short: 'J. Li, M. Großkreutz, T.M. Peters, I. Scharlau, B. Paaßen, in: B. Paaßen
    (Ed.), Proceedings of the 1st Workshop on Explainability, Transparency, and Safety
    (ExTraSafe) At , 2026.'
conference:
  location: Bremen, Germany
  name: 'KI2026 '
date_created: 2026-09-30T13:09:59Z
date_updated: 2026-09-30T13:10:23Z
department:
- _id: '424'
editor:
- first_name: Benjamin
  full_name: Paaßen, Benjamin
  last_name: Paaßen
language:
- iso: eng
project:
- _id: '124'
  name: 'TRR 318 ; TP C01: Gesundes Misstrauen in Erklärungen'
publication: 'Proceedings of the 1st Workshop on Explainability, Transparency, and
  Safety (ExTraSafe) at '
quality_controlled: '1'
status: public
title: LLM Inconsistency under Paraphrase Attacks
type: conference
user_id: '451'
year: '2026'
...
