<?xml version="1.0" encoding="UTF-8"?>

<modsCollection xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns="http://www.loc.gov/mods/v3" xsi:schemaLocation="http://www.loc.gov/mods/v3 http://www.loc.gov/standards/mods/v3/mods-3-3.xsd">
<mods version="3.3">

<genre>preprint</genre>

<titleInfo><title>CUDA-MPQS: A GPU-Resident Self-Initializing Quadratic Sieve, and the Factorization of RSA-155</title></titleInfo>





<name type="personal">
  <namePart type="given">Fabian</namePart>
  <namePart type="family">Januszewski</namePart>
  <role><roleTerm type="text">author</roleTerm> </role></name>
<name type="personal">
  <namePart type="given">Christoph</namePart>
  <namePart type="family">Heinrichs</namePart>
  <role><roleTerm type="text">author</roleTerm> </role></name>














<abstract lang="eng">The quadratic sieve (QS) is an irregular, cache-hostile, branch-heavy integer factorization algorithm; prior GPU work accelerates individual stages of it. We present \swname{}, a self-initializing quadratic sieve in which \emph{every} stage --- polynomial initialization, sieving, relation post-processing, large-prime matching, \GFtwo{} matrix construction, Block Wiedemann linear algebra and the square root --- executes on the GPU, the CPU restricted to orchestration, setup and I/O, the sieve window 99.9\% GPU-busy with no device-wide or event synchronization and no synchronous copy. With it we factored the 512-bit RSA-155 challenge modulus end to end on GPUs: a 64$\times$H100 sieve feeding a single-H100 \GFtwo{} solve, 700.6~GPU-h and 242~kWh in total, 98.4\% of it sieving. To our knowledge this is the largest integer factored by the quadratic sieve, exceeding the prior QS record RSA-150 (Buhrow, YAFU, 2025); all comparable general factorizations since 1996, RSA-155&apos;s own in 1999 included, used the number field sieve. On a single device we factor RSA-100 in 29.2~s on an H100 and 51.0~s on a consumer RTX~5070~Ti; in a controlled same-modulus measurement, with energy measured on both sides, the H100 is faster than the fastest CPU quadratic sieve, YAFU, on 96 cores of an EPYC~9655 (106--122~s), and than CADO-NFS&apos;s SIQS (292~s) and GNFS (263~s). Across four discrete GPUs (a 3.7$\times$ range in CUDA cores) RSA-100 throughput is close to proportional to core count, memory bandwidth being a weaker predictor, and measured cost over seven benchmark factorizations from RSA-100 to RSA-155 grows as $\LN{1/2}{1}^{1.04}$ ($R^2=0.993$). On RSA-150, the identical modulus on which the prior QS record was set, we measure 302.9~GPU-h against that run&apos;s 11{,}664 core-hours.</abstract>

<originInfo><dateIssued encoding="w3cdtf">2026</dateIssued>
</originInfo>



<relatedItem type="host"><titleInfo><title>arXiv:2610.07126</title></titleInfo>
  <identifier type="arXiv">2610.07126</identifier>
<part>
</part>
</relatedItem>


<extension>
<bibliographicCitation>
<mla>Januszewski, Fabian, and Christoph Heinrichs. “CUDA-MPQS: A GPU-Resident Self-Initializing Quadratic Sieve, and the Factorization of RSA-155.” &lt;i&gt;ArXiv:2610.07126&lt;/i&gt;, 2026.</mla>
<ama>Januszewski F, Heinrichs C. CUDA-MPQS: A GPU-Resident Self-Initializing Quadratic Sieve, and the Factorization of RSA-155. &lt;i&gt;arXiv:261007126&lt;/i&gt;. Published online 2026.</ama>
<bibtex>@article{Januszewski_Heinrichs_2026, title={CUDA-MPQS: A GPU-Resident Self-Initializing Quadratic Sieve, and the Factorization of RSA-155}, journal={arXiv:2610.07126}, author={Januszewski, Fabian and Heinrichs, Christoph}, year={2026} }</bibtex>
<apa>Januszewski, F., &amp;#38; Heinrichs, C. (2026). CUDA-MPQS: A GPU-Resident Self-Initializing Quadratic Sieve, and the Factorization of RSA-155. In &lt;i&gt;arXiv:2610.07126&lt;/i&gt;.</apa>
<ieee>F. Januszewski and C. Heinrichs, “CUDA-MPQS: A GPU-Resident Self-Initializing Quadratic Sieve, and the Factorization of RSA-155,” &lt;i&gt;arXiv:2610.07126&lt;/i&gt;. 2026.</ieee>
<short>F. Januszewski, C. Heinrichs, ArXiv:2610.07126 (2026).</short>
<chicago>Januszewski, Fabian, and Christoph Heinrichs. “CUDA-MPQS: A GPU-Resident Self-Initializing Quadratic Sieve, and the Factorization of RSA-155.” &lt;i&gt;ArXiv:2610.07126&lt;/i&gt;, 2026.</chicago>
</bibliographicCitation>
</extension>
<recordInfo><recordIdentifier>67494</recordIdentifier><recordCreationDate encoding="w3cdtf">2026-10-11T10:08:48Z</recordCreationDate><recordChangeDate encoding="w3cdtf">2026-10-11T10:09:12Z</recordChangeDate>
</recordInfo>
</mods>
</modsCollection>
