---
res:
  bibo_abstract:
  - "Despite recent progress in Large Language Model (LLM) research, LLMs remain sensitive
    to prompt paraphrases,\r\ne.g., due to sycophancy, framing, or negation blindness,
    yielding surprising and contradictory outputs. This paper\r\ninvestigates whether
    such vulnerabilities can be exploited to construct adversarial attacks against
    LLMs. Using\r\na DeepSeek-V4-Flash model, we automatically rephrase an input prompt
    using sycophancy, framing, and/or\r\nnegations to leave it semantically unchanged
    but elicit a contradicting output of a victim LLM, in this case a\r\nLlama-3.1-8B-Instruct
    model. On 49 controversial topics, we show that negation attacks can flip the
    model’s\r\noriginal stance in roughly 80% of cases, whereas sycophancy and framing
    yield moderate success rates around\r\n25% and 40%, respectively. The attacker
    LLM was also quite accurate in determining attack success (91.3%\r\naccuracy compared
    to annotations of two independent human annotators). Overall, the results suggest
    that\r\nhuman-interpretable paraphrase attacks, such as negations, may be a viable
    method to illustrate the vulnerabilities\r\nof LLMs.@eng"
  bibo_authorlist:
  - foaf_Person:
      foaf_givenName: Jiaao
      foaf_name: Li, Jiaao
      foaf_surname: Li
  - foaf_Person:
      foaf_givenName: Matthis
      foaf_name: Großkreutz, Matthis
      foaf_surname: Großkreutz
  - foaf_Person:
      foaf_givenName: Tobias Martin
      foaf_name: Peters, Tobias Martin
      foaf_surname: Peters
      foaf_workInfoHomepage: http://www.librecat.org/personId=92810
    orcid: 0009-0008-5193-6243
  - foaf_Person:
      foaf_givenName: Ingrid
      foaf_name: Scharlau, Ingrid
      foaf_surname: Scharlau
      foaf_workInfoHomepage: http://www.librecat.org/personId=451
    orcid: 0000-0003-2364-9489
  - foaf_Person:
      foaf_givenName: Benjamin
      foaf_name: Paaßen, Benjamin
      foaf_surname: Paaßen
  dct_date: 2026^xs_gYear
  dct_language: eng
  dct_title: LLM Inconsistency under Paraphrase Attacks@
...
