<?xml version="1.0" encoding="UTF-8"?>
<rdf:RDF xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#"
         xmlns:dc="http://purl.org/dc/terms/"
         xmlns:foaf="http://xmlns.com/foaf/0.1/"
         xmlns:bibo="http://purl.org/ontology/bibo/"
         xmlns:fabio="http://purl.org/spar/fabio/"
         xmlns:owl="http://www.w3.org/2002/07/owl#"
         xmlns:event="http://purl.org/NET/c4dm/event.owl#"
         xmlns:ore="http://www.openarchives.org/ore/terms/">

    <rdf:Description rdf:about="https://ris.uni-paderborn.de/record/67288">
        <ore:isDescribedBy rdf:resource="https://ris.uni-paderborn.de/record/67288"/>
        <dc:title>LLM Inconsistency under Paraphrase Attacks</dc:title>
        <bibo:authorList rdf:parseType="Collection">
            <foaf:Person>
                <foaf:name></foaf:name>
                <foaf:surname></foaf:surname>
                <foaf:givenname></foaf:givenname>
            </foaf:Person>
            <foaf:Person>
                <foaf:name></foaf:name>
                <foaf:surname></foaf:surname>
                <foaf:givenname></foaf:givenname>
            </foaf:Person>
            <foaf:Person>
                <foaf:name></foaf:name>
                <foaf:surname></foaf:surname>
                <foaf:givenname></foaf:givenname>
            </foaf:Person>
            <foaf:Person>
                <foaf:name></foaf:name>
                <foaf:surname></foaf:surname>
                <foaf:givenname></foaf:givenname>
            </foaf:Person>
            <foaf:Person>
                <foaf:name></foaf:name>
                <foaf:surname></foaf:surname>
                <foaf:givenname></foaf:givenname>
            </foaf:Person>
        </bibo:authorList>
        <bibo:abstract>Despite recent progress in Large Language Model (LLM) research, LLMs remain sensitive to prompt paraphrases,
e.g., due to sycophancy, framing, or negation blindness, yielding surprising and contradictory outputs. This paper
investigates whether such vulnerabilities can be exploited to construct adversarial attacks against LLMs. Using
a DeepSeek-V4-Flash model, we automatically rephrase an input prompt using sycophancy, framing, and/or
negations to leave it semantically unchanged but elicit a contradicting output of a victim LLM, in this case a
Llama-3.1-8B-Instruct model. On 49 controversial topics, we show that negation attacks can flip the model’s
original stance in roughly 80% of cases, whereas sycophancy and framing yield moderate success rates around
25% and 40%, respectively. The attacker LLM was also quite accurate in determining attack success (91.3%
accuracy compared to annotations of two independent human annotators). Overall, the results suggest that
human-interpretable paraphrase attacks, such as negations, may be a viable method to illustrate the vulnerabilities
of LLMs.</bibo:abstract>
    </rdf:Description>
</rdf:RDF>
