面向大模型长期记忆中跨会话连贯性与个性化的双流记忆框架
DOI:
作者:
作者单位:

青岛理工大学信息管理学院

作者简介:

通讯作者:

中图分类号:

TP391

基金项目:

国家自然科学基金(42201506,61502262)<br />第一作者: 周炜,(1981-),男,汉,山东青岛,博士,教授。研究方向: 生成式人工智能,智能应用。E-mail: zhouwei@qut.edu.cn。


A Dual-Stream Memory Framework for Cross-Session Coherence and Personalization in Large Language Models
Author:
Affiliation:

School of Information Management, Qingdao University of Technology

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    近年来,大型语言模型(Large Language Models, LLMs)在包含多轮对话的单个会话中展现出卓越的对话能力,但在跨越多个会话的长期交互中,如何同时维持跨会话连贯性与精确记忆用户个性化偏好仍是核心挑战。在现有方法中,摘要类方法虽然能维持跨会话连贯性,但摘要的有损压缩不可避免地丢失细粒度用户偏好,检索增强生成(Retrieval-Augmented Generation, RAG)类方法虽然保留原始历史细节,却因原始对话记录中的低信噪比与全局语境缺失导致记忆回忆准确率下降。为此,本文提出双流记忆框架DuoMem,包含后台记忆构建与前台推理应用两个阶段。后台记忆构建阶段中,通过递归摘要构建宏观摘要流以维护跨会话连贯性,通过结构化提取与混合检索构建用户画像流以实现细粒度个性化记忆。用户画像流中使用基于自然语言推理(Natural Language Inference, NLI)的记忆状态追踪(Memory State Tracking, MST)机制对用户画像进行条目分类和更新维护,使用置信度驱动的晋升机制将高频稳定用户画像持久化至核心身份上下文。在前台推理阶段,自适应融合门控基于检索置信度动态决定画像条目的注入数量,分层注入机制将宏观摘要、画像条目与核心身份上下文组装为最终上下文,实现高效个性化响应。实验结果表明,在MSC(Multi-Session Chat)和LoCoMo(Long Conversation Memory)数据集上,DuoMem的记忆回忆准确率分别达0.895和0.891,较最优基线绝对提升11.4和13.1个百分点,个性化准确率分别达0.829和0.801,较最优基线绝对提升13.4和13.0个百分点,路由延迟仅6.5 ms,较传统LLM路由方案快44倍。研究结果表明,通过双流解耦设计可同时实现跨会话连贯性、细粒度个性化与推理高效性的联合优化,为面向大语言模型的长期个性化对话提供了一种可行的解决方案。

    Abstract:

    In recent years, remarkable capabilities in multi-turn dialogue within a single session have been demonstrated by large language models (LLMs). However, maintaining cross-session coherence while accurately memorizing user-specific personalized preferences across extended multi-session interactions remains a core challenge. Among existing approaches, cross-session coherence can be preserved by summarization methods, yet fine-grained user preferences are inevitably lost due to lossy compression; original historical details are retained by retrieval-augmented generation (RAG) methods, but recall accuracy is degraded by the low signal-to-noise ratio and the absence of global context in raw conversation logs. To address these limitations, DuoMem, a dual-stream memory framework consisting of an offline memory construction phase and an online inference phase, is proposed. In the offline phase, a macro summary stream was built via recursive summarization to maintain cross-session coherence, and a user profile stream was constructed through structured extraction and hybrid retrieval to support fine-grained personalized memory. Within the user profile stream, profile entries were classified and updated by a Memory State Tracking (MST) mechanism based on natural language inference (NLI), and high-frequency stable user profiles were consolidated into the core identity context by a confidence-driven promotion mechanism. In the online phase, the number of profile entries to be injected was dynamically determined by an adaptive fusion gate according to retrieval confidence, and the macro summary, profile entries, and core identity context were assembled into the final prompt by a tiered injection mechanism for efficient personalized response generation. Experiments were conducted on the MSC(Multi-Session Chat) and LoCoMo (Long Conversation Memory)long-term dialogue benchmarks. Recall accuracies of 0.895 and 0.891 were achieved by DuoMem, exceeding the best baseline by 11.4 and 13.1 percentage points (pp), respectively; personalization accuracies of 0.829 and 0.801 were attained, surpassing the best baseline by 13.4 and 13.0 pp, while the routing latency was reduced to 6.5 ms, approximately 44 times faster than conventional LLM-based routing. These results demonstrate that cross-session coherence, fine-grained personalization, and inference efficiency can be jointly achieved through the proposed dual-stream decoupling, providing a viable solution for long-term personalized dialogue with large language models.

    参考文献
    相似文献
    引证文献
引用本文

周炜,李海鹏,王存志,等. 面向大模型长期记忆中跨会话连贯性与个性化的双流记忆框架[J]. 科学技术与工程, , ():

复制
文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2026-05-14
  • 最后修改日期:2026-06-25
  • 录用日期:2026-07-31
  • 在线发布日期:
  • 出版日期:
×
2026年会通知 | “技术经济学驱动智能经济生态构建与治理变革”——中国技术经济学会第三十三届学术年会(2026)会议通知暨征文启事(第一轮)
亟待确认版面费归属稿件,敬请作者关注