<?xml version="1.0" encoding="UTF-8"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/"
         xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
         xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
<ListRecords>
<oai_dc:dc xmlns="http://www.openarchives.org/OAI/2.0/oai_dc/"
           xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/"
           xmlns:dc="http://purl.org/dc/elements/1.1/"
           xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
           xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
   	<dc:title>CUDA-MPQS: A GPU-Resident Self-Initializing Quadratic Sieve, and the Factorization of RSA-155</dc:title>
   	<dc:creator>Januszewski, Fabian</dc:creator>
   	<dc:creator>Heinrichs, Christoph</dc:creator>
   	<dc:description>The quadratic sieve (QS) is an irregular, cache-hostile, branch-heavy integer factorization algorithm; prior GPU work accelerates individual stages of it. We present \swname{}, a self-initializing quadratic sieve in which \emph{every} stage --- polynomial initialization, sieving, relation post-processing, large-prime matching, \GFtwo{} matrix construction, Block Wiedemann linear algebra and the square root --- executes on the GPU, the CPU restricted to orchestration, setup and I/O, the sieve window 99.9\% GPU-busy with no device-wide or event synchronization and no synchronous copy. With it we factored the 512-bit RSA-155 challenge modulus end to end on GPUs: a 64$\times$H100 sieve feeding a single-H100 \GFtwo{} solve, 700.6~GPU-h and 242~kWh in total, 98.4\% of it sieving. To our knowledge this is the largest integer factored by the quadratic sieve, exceeding the prior QS record RSA-150 (Buhrow, YAFU, 2025); all comparable general factorizations since 1996, RSA-155&apos;s own in 1999 included, used the number field sieve. On a single device we factor RSA-100 in 29.2~s on an H100 and 51.0~s on a consumer RTX~5070~Ti; in a controlled same-modulus measurement, with energy measured on both sides, the H100 is faster than the fastest CPU quadratic sieve, YAFU, on 96 cores of an EPYC~9655 (106--122~s), and than CADO-NFS&apos;s SIQS (292~s) and GNFS (263~s). Across four discrete GPUs (a 3.7$\times$ range in CUDA cores) RSA-100 throughput is close to proportional to core count, memory bandwidth being a weaker predictor, and measured cost over seven benchmark factorizations from RSA-100 to RSA-155 grows as $\LN{1/2}{1}^{1.04}$ ($R^2=0.993$). On RSA-150, the identical modulus on which the prior QS record was set, we measure 302.9~GPU-h against that run&apos;s 11{,}664 core-hours.</dc:description>
   	<dc:date>2026</dc:date>
   	<dc:type>info:eu-repo/semantics/preprint</dc:type>
   	<dc:type>doc-type:preprint</dc:type>
   	<dc:type>text</dc:type>
   	<dc:type>http://purl.org/coar/resource_type/c_816b</dc:type>
   	<dc:identifier>https://ris.uni-paderborn.de/record/67494</dc:identifier>
   	<dc:source>Januszewski F, Heinrichs C. CUDA-MPQS: A GPU-Resident Self-Initializing Quadratic Sieve, and the Factorization of RSA-155. &lt;i&gt;arXiv:261007126&lt;/i&gt;. Published online 2026.</dc:source>
   	<dc:relation>info:eu-repo/semantics/altIdentifier/arxiv/2610.07126</dc:relation>
   	<dc:rights>info:eu-repo/semantics/closedAccess</dc:rights>
</oai_dc:dc>
</ListRecords>
</OAI-PMH>
