arXiv · cs/0608061
Concurrent Processing Memory
Abstract
A theoretical memory that embeds limited, application-specific processing power and nearest-neighbor connectivity at every storage element is proposed. Such a memory performs parallel computation within itself to solve generic array problems, while remaining pin- and function-compatible with conventional random-access memory. The applicability of this in-memory, finest-grain, massive-SIMD approach is examined in detail through a family of increasingly capable devices---content movable, searchable, value-comparable, and computable memory. For an array of $N$ items, the approach reduces the instruction-cycle count of universal operations (insertion, deletion, and match finding) to $\sim 1$, of local operations (filtering and template matching) to $\sim$ the operation footprint, and of global operations (summation and finding minimum/maximum) to $\sim \sqrt{N}$. It eliminates most data-processing traffic on the system bus, yet remains general-purpose, easy to program, backward-compatible with existing bus-sharing architectures and operating systems, and practical to implement along a clear road map.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Chengpu Wang. 2006-08-15. Concurrent Processing Memory. https://arxiv.org/abs/cs/0608061
Cite the original work for its findings. Save a collection to share your selection of sources.