A new arXiv paper proposes small DRAM changes that let the working memory chip compute on data in the layout CPUs already use, skipping a costly reorganization step in older in memory computing designs, where the memory chip itself runs the math.
Every modern computer burns extraordinary time and energy ferrying data between its working memory, the DRAM chips in every PC and phone, and the processor. A new arXiv preprint from three university labs argues there's a second, less-discussed cost hiding inside "in-memory computing" designs: the data has to be reshuffled into a column-oriented, bit-serial form before the memory array can do any math on it.
The paper, RAPID: Row-Parallel Arithmetic Processing in DRAM, was posted in October 2026 by researchers at Syracuse University, Friedrich-Alexander-Universität Erlangen-Nürnberg, and TU Dresden. RAPID adds two small DRAM subarray extensions, migration cells that move data horizontally between neighboring bitlines and inversion cells that run logical inversion directly inside the array. The result is that DRAM can run arithmetic on data in the row-parallel, word-parallel layout CPUs and accelerators already expect, without the prior translation step.
That sidesteps the central trade-off in existing processing-using-memory designs, which reorganize data into column-oriented, bit-serial representations to fit how DRAM subarrays naturally operate.
The catch: this is a SPICE simulation of a DRAM subarray, not a fabricated chip, and adoption would require memory-vendor and standards-body interest, neither of which the preprint implies. If memory makers do pick it up, the benefit lands wherever data movement is the binding constraint, including AI training and inference runs that spend most of their time waiting on memory.