Methods and apparatuses for information analysis on shared and distributed computing systems
- Richland, WA
Apparatuses and computer-implemented methods for analyzing, on shared and distributed computing systems, information comprising one or more documents are disclosed according to some aspects. In one embodiment, information analysis can comprise distributing one or more distinct sets of documents among each of a plurality of processes, wherein each process performs operations on a distinct set of documents substantially in parallel with other processes. Operations by each process can further comprise computing term statistics for terms contained in each distinct set of documents, thereby generating a local set of term statistics for each distinct set of documents. Still further, operations by each process can comprise contributing the local sets of term statistics to a global set of term statistics, and participating in generating a major term set from an assigned portion of a global vocabulary.
- Research Organization:
- Pacific Northwest National Laboratory (PNNL), Richland, WA (United States)
- Sponsoring Organization:
- USDOE
- DOE Contract Number:
- AC05-76RL01830
- Assignee:
- Battelle Memorial Institute (Richland, WA)
- Patent Number(s):
- 7,895,210
- Application Number:
- US Patent Application 11/540,240
- OSTI ID:
- 1015195
- Country of Publication:
- United States
- Language:
- English
HNC's MatchPlus system
|
journal | October 1992 |
Advances, Applications and Performance of the Global Arrays Shared Memory Programming Toolkit
|
journal | May 2006 |
Similar Records
Servicing a globally broadcast interrupt signal in a multi-threaded computer
Document clustering methods, document cluster label disambiguation methods, document clustering apparatuses, and articles of manufacture