Skip to main content
U.S. Department of Energy
Office of Scientific and Technical Information

QCS: a system for querying, clustering and summarizing documents.

Technical Report ·
DOI:https://doi.org/10.2172/894745· OSTI ID:894745
;  [1];  [2];  [3]
  1. Center for Computing Sciences, Bowie, MD
  2. University of Maryland, College Park, MD
  3. Center for Computing Sciences, Bowie, MD

Information retrieval systems consist of many complicated components. Research and development of such systems is often hampered by the difficulty in evaluating how each particular component would behave across multiple systems. We present a novel hybrid information retrieval system--the Query, Cluster, Summarize (QCS) system--which is portable, modular, and permits experimentation with different instantiations of each of the constituent text analysis components. Most importantly, the combination of the three types of components in the QCS design improves retrievals by providing users more focused information organized by topic. We demonstrate the improved performance by a series of experiments using standard test sets from the Document Understanding Conferences (DUC) along with the best known automatic metric for summarization system evaluation, ROUGE. Although the DUC data and evaluations were originally designed to test multidocument summarization, we developed a framework to extend it to the task of evaluation for each of the three components: query, clustering, and summarization. Under this framework, we then demonstrate that the QCS system (end-to-end) achieves performance as good as or better than the best summarization engines. Given a query, QCS retrieves relevant documents, separates the retrieved documents into topic clusters, and creates a single summary for each cluster. In the current implementation, Latent Semantic Indexing is used for retrieval, generalized spherical k-means is used for the document clustering, and a method coupling sentence 'trimming', and a hidden Markov model, followed by a pivoted QR decomposition, is used to create a single extract summary for each cluster. The user interface is designed to provide access to detailed information in a compact and useful format. Our system demonstrates the feasibility of assembling an effective IR system from existing software libraries, the usefulness of the modularity of the design, and the value of this particular combination of modules.

Research Organization:
Sandia National Laboratories
Sponsoring Organization:
USDOE
DOE Contract Number:
AC04-94AL85000
OSTI ID:
894745
Report Number(s):
SAND2006-5000
Country of Publication:
United States
Language:
English

Similar Records

QCS : a system for querying, clustering, and summarizing documents.
Technical Report · Tue Aug 01 00:00:00 EDT 2006 · OSTI ID:893129

Towards a RAG-based summarization for the Electron Ion Collider
Journal Article · Mon Jul 01 00:00:00 EDT 2024 · Journal of Instrumentation · OSTI ID:2578842

VisIRR: A Visual Analytics System for Information Retrieval and Recommendation for Large-Scale Document Data
Journal Article · Tue Jan 30 23:00:00 EST 2018 · ACM Transactions on Knowledge Discovery from Data · OSTI ID:1426558