面向智能体时代大语言模型推理的新型存算解聚合架构

发布时间:2026-09-20 作者:刘新阳,李婷宇,巨鸣

摘要:存算一体技术通过将计算单元与存储单元深度融合或物理级邻近部署,从架构层面有效缓解了数据搬运瓶颈。大语言模型(LLMs)及智能体(Agentic AI)的持续演进,对计算架构的高带宽、低延迟与高能效比提出了严苛要求。系统梳理了基于静态随机存取存储器(SRAM)、动态随机存取存储器(DRAM)、闪存(Flash)及新兴非易失性存储器的存算一体技术路径,涵盖存内计算(CIM)、存内处理(PIM)与近存处理(PNM)三大主流范式,综述其研究进展、典型应用案例及面临的技术挑战。分析表明,不同技术路径展现出差异化的性能与能效优势,可通过场景驱动的技术融合、软硬件协同设计及产业链生态协同实现优势互补,进而为下一代人工智能芯片的架构创新提供有益参考。

关键词:大语言模型;智能体;AI芯片架构;存算一体

 

Abstract: Compute-in-Memory (CIM), Processing-in-Memory (PIM), and Processing-Near-Memory (PNM) hold the promise of fundamentally resolving the data movement bottleneck by integrating computing logic with memory units or placing them in close physical proximity. A systematic review is presented of CIM, PIM, and PNM technical paths implemented on static random-access memory (SRAM), dynamic random-access memory (DRAM), Flash, and emerging non-volatile memories, covering their research progress, representative use cases, and associated opportunities and challenges. The analysis reveals that each technical path exhibits distinct performance and energy-efficiency advantages. It is argued that scenario-driven technology fusion, hardware-software co-design, and industry-wide ecosystem collaboration are essential for harnessing complementary strengths across these paths. This work is intended to serve as a valuable reference for architectural innovation in next-generation AI chips.

Keywords: large language models; Agent AI; AI chip architecture; computing-in-memory