<?xml version="1.0" encoding="UTF-8"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/"
         xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
         xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
<ListRecords>
<oai_dc:dc xmlns="http://www.openarchives.org/OAI/2.0/oai_dc/"
           xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/"
           xmlns:dc="http://purl.org/dc/elements/1.1/"
           xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
           xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
   	<dc:title>LLM Inconsistency under Paraphrase Attacks</dc:title>
   	<dc:creator>Li, Jiaao</dc:creator>
   	<dc:creator>Großkreutz, Matthis</dc:creator>
   	<dc:creator>Peters, Tobias Martin</dc:creator>
   	<dc:creator>Scharlau, Ingrid</dc:creator>
   	<dc:creator>Paaßen, Benjamin</dc:creator>
   	<dc:creator>Paaßen, Benjamin</dc:creator>
   	<dc:description>Despite recent progress in Large Language Model (LLM) research, LLMs remain sensitive to prompt paraphrases,
e.g., due to sycophancy, framing, or negation blindness, yielding surprising and contradictory outputs. This paper
investigates whether such vulnerabilities can be exploited to construct adversarial attacks against LLMs. Using
a DeepSeek-V4-Flash model, we automatically rephrase an input prompt using sycophancy, framing, and/or
negations to leave it semantically unchanged but elicit a contradicting output of a victim LLM, in this case a
Llama-3.1-8B-Instruct model. On 49 controversial topics, we show that negation attacks can flip the model’s
original stance in roughly 80% of cases, whereas sycophancy and framing yield moderate success rates around
25% and 40%, respectively. The attacker LLM was also quite accurate in determining attack success (91.3%
accuracy compared to annotations of two independent human annotators). Overall, the results suggest that
human-interpretable paraphrase attacks, such as negations, may be a viable method to illustrate the vulnerabilities
of LLMs.</dc:description>
   	<dc:date>2026</dc:date>
   	<dc:type>info:eu-repo/semantics/conferenceObject</dc:type>
   	<dc:type>doc-type:conferenceObject</dc:type>
   	<dc:type>text</dc:type>
   	<dc:type>http://purl.org/coar/resource_type/c_5794</dc:type>
   	<dc:identifier>https://ris.uni-paderborn.de/record/67288</dc:identifier>
   	<dc:source>Li J, Großkreutz M, Peters TM, Scharlau I, Paaßen B. LLM Inconsistency under Paraphrase Attacks. In: Paaßen B, ed. &lt;i&gt;Proceedings of the 1st Workshop on Explainability, Transparency, and Safety (ExTraSafe) At &lt;/i&gt;. ; 2026.</dc:source>
   	<dc:language>eng</dc:language>
   	<dc:rights>info:eu-repo/semantics/closedAccess</dc:rights>
</oai_dc:dc>
</ListRecords>
</OAI-PMH>
