Automatically generating extraction patterns from untagged text
- Univ. of Utah, Salt Lake City, UT (United States)
Many corpus-based natural language processing systems rely on text corpora that have been manually annotated with syntactic or semantic tags. In particular, all previous dictionary construction systems for information extraction have used an annotated training corpus or some form of annotated input. We have developed a system called AutoSlog-TS that creates dictionaries of extraction patterns using only untagged text. AutoSlog-TS is based on the AutoSlog system, which generated extraction patterns using annotated text and a set of heuristic rules. By adapting AutoSlog and combining it with statistical techniques, we eliminated its dependency on tagged text. In experiments with the MUC-4 terrorism domain, AutoSlog-TS created a dictionary of extraction patterns that performed comparably to a dictionary created by AutoSlog, using only preclassified texts as input.
- OSTI ID:
- 430781
- Report Number(s):
- CONF-960876-; CNN: Grant MIP 9023174; TRN: 96:006521-0156
- Resource Relation:
- Conference: 13. National conference on artifical intelligence and the 8. Innovative applications of artificial intelligence conference, Portland, OR (United States), 4-8 Aug 1996; Other Information: PBD: 1996; Related Information: Is Part Of Proceedings of the thirteenth national conference on artificial intelligence and the eighth innovative applications of artificial intelligence conference. Volume 1 and 2; PB: 1626 p.
- Country of Publication:
- United States
- Language:
- English
Similar Records
Literature mining of protein-residue associations with graph rules learned through distant supervision
Experiments in automatic word class and word sense identification for information retrieval