Search arXivSearch

arXiv · 2307.05696

A Personalized Reinforcement Learning Summarization Service for Learning Structure from Unstructured Data

Abstract

The exponential growth of textual data has created a crucial need for tools that assist users in extracting meaningful insights. Traditional document summarization approaches often fail to meet individual user requirements and lack structure for efficient information processing. To address these limitations, we propose Summation, a hierarchical personalized concept-based summarization approach. It synthesizes documents into a concise hierarchical concept map and actively engages users by learning and adapting to their preferences. Using a Reinforcement Learning algorithm, Summation generates personalized summaries for unseen documents on specific topics. This framework enhances comprehension, enables effective navigation, and empowers users to extract meaningful insights from large document collections aligned with their unique requirements.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Samira Ghodratnama, Amin Beheshti, Mehrdad Zakershahrak. 2023-07-09. A Personalized Reinforcement Learning Summarization Service for Learning Structure from Unstructured Data. https://arxiv.org/abs/2307.05696

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

MLLMRec: A Preference Reasoning Paradigm with Graph Refinement for Multimodal Recommendation

Multimodal recommendation combines the user historical behaviors with the modal features of items to capture the tangible user preferences, presenting superior performance compared to the conventional ID-based recommender systems. However, existing methods still encounter two key problems in the representation learning of users and items, respectively: (1) the initialization of multimodal user representations is either agnostic to historical behaviors or contaminated by irrelevant modal noise, and (2) the widely used KNN-based item-item graph contains noisy edges with low similarities and lacks audience co-occurrence relationships. To address such issues, we propose MLLMRec, a novel preference reasoning paradigm with graph refinement for multimodal recommendation. Specifically, on the one hand, the item images are first converted into high-quality semantic descriptions using a multimodal large language model (MLLM), thereby bridging the semantic gap between visual and textual modalities. Then, we construct a behavioral description list for each user and feed it into the MLLM to reason about the purified user preference profiles that contain the latent interaction intents. The reasoned profiles and the multimodal descriptions of items, together with their ID embeddings, are propagated over the user-item interaction graph to absorb the high-order collaborative signals. On the other hand, we develop the threshold-controlled denoising and topology-aware enhancement strategies to refine the suboptimal item-item graph, which are applied to both the multimodal and ID item representations to improve the accuracy of item representation learning. Extensive experiments on three publicly available datasets demonstrate that MLLMRec achieves the state-of-the-art performance. The source code is provided at https://github.com/Yuzhuo-Dang/MLLMRec.git.

cs.IR

Dynamic Feature-Embedding Communication via Codebook Distillation for Federated Recommendation

Federated recommendation systems commonly protect user privacy by keeping user parameters on local devices, while exchanging item parameters for collaborative model training. However, such item parameters usually model items independently and suffer from both efficiency and effectiveness challenges, making communication costs grow with the item space and limiting cross-item generalization and robustness to noisy feedback. To address these limitations, we propose to model items via shared latent feature embeddings for communication. Residual Quantization (RQ) provides a natural way to instantiate this communication by representing each item with a short sequence of discrete code IDs, i.e., Semantic IDs (SIDs). However, directly applying centralized and static RQ-based recommendation to federated learning is non-trivial due to 1) private and biased historical interactions and 2) evolving collaborative information. We propose RQFedRec, an RQ-based federated recommendation framework for dynamic feature-embedding communication. To construct globally aligned codebooks without accessing private interactions, RQFedRec introduces an information distillation module. Each client first learns item embeddings that encode local collaborative information from private interactions, and then distills such information into feature-indexed codebooks under globally shared SIDs, making sparse and biased local signals more compatible with server aggregation. To adapt to evolving collaborative information, RQFedRec introduces a self-refining SID update module that dynamically refines global SID assignments from aggregated codebooks. Extensive experiments demonstrate that RQFedRec improves recommendation performance and reduces communication costs without relying on semantic information, while further benefiting from public semantics when available. Code is available at https://github.com/Mingzhe-Han/RQFedRec.

cs.IR

OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis

Training deep research agents requires long-horizon trajectories that interleave search, evidence aggregation, and multi-step reasoning. However, existing data collection pipelines typically rely on proprietary web APIs, making large-scale trajectory synthesis costly, unstable, and difficult to reproduce. We present OpenResearcher, a reproducible pipeline that decouples one-time corpus bootstrapping from multi-turn trajectory synthesis and executes the search-and-browse loop entirely offline using three explicit browser primitives: search, open, and find, over a 15M-document corpus. Using GPT-OSS-120B as the teacher model, we synthesize over 97K trajectories, including a substantial long-horizon tail with 100+ tool calls. Supervised fine-tuning a 30B-A3B backbone on these trajectories achieves 54.8\% accuracy on BrowseComp-Plus, a +34.0 point improvement over the base model, while remaining competitive on BrowseComp, GAIA, and xbench-DeepSearch. Because the environment is offline and fully instrumented, it also enables controlled analysis, where our study reveals practical insights into deep research pipeline design, including data filtering strategies, agent configuration choices, and how retrieval success relates to final answer accuracy. We release the pipeline, synthesized trajectories, model checkpoints, and the offline search environment at https://github.com/TIGER-AI-Lab/OpenResearcher.

cs.IR