Published 2024-09-30
How to Cite

This work is licensed under a Creative Commons Attribution 4.0 International License.
Abstract
This paper addresses the problems of information redundancy, sparsity of key clues, and difficulty in cross-round referencing in long-context scenarios, proposing a unified framework combining memory compression and key information fine-tuning. The method first decomposes the input into controllable-granularity fragments and encodes them. A lightweight scoring mechanism is used to assign importance weights to these fragments. Under a fixed memory budget, an explicit memory set is constructed through a selection strategy, and a global memory vector is formed through weighted aggregation to supplement the overall semantics and reduce the risk of omissions. To align memory selection with model behavior, a key supervision signal is further introduced to constrain importance estimation, and key-related objectives are jointly optimized with task objectives, thereby enhancing the model's bias towards core facts, constraints, and reusable clues under compression conditions. The framework also includes concise memory update rules to maintain budget stability and representation coherence in long-sequence inputs and provides interpretable fragment-level memory entries to support tracing and verification. Comparative experiments show that this method achieves a more balanced overall performance across dimensions such as generation quality, fact reliability, key information identification, memory retention, calibration error, and output consistency, demonstrating the effectiveness of robust information organization and continuous memory maintenance under a limited context budget.