Robust Natural Language Processing For Information Retrieval.pdf

H92-1040.pdf
Preview of Robust Natural Language Processing for Information Retrieval
🔗 Source: aclanthology.org
📊 Size: 558 KB
👤 Author: Tomek Strzalkowski
⬇️ Downloads: 168

Summary

Our information retrieval system consists of a traditional statistical backbone augmented with various natural language processing components. The system processes the database text with a fast syntactic parser, extracts certain types of phrases, and uses them as compound indexing terms. The extracted phrases are statistically analyzed to discover similarity links between smaller subphrases and words occurring in them. A further filtering process maps these similarity links onto semantic relations, which are used to transform the user's request into a search query. The user's natural language request is also parsed, and all indexing terms occurring in it are identified. Certain highly ambiguous terms are dropped, and other terms may be added via admissible similarity relations. The final query is constructed, and the database search follows, returning a ranked list of documents. All processing steps are fully automated, with no human intervention or manual encoding required.

TIP, a full grammar parser based on the Linguistic String Grammar, produces regularized parse tree representations for each sentence. It has a built-in timer that regulates the amount of time allowed for parsing any one sentence. If a parse is not returned before the allotted time elapses, the parser enters the skip-and-fit mode, attempting to "fit" the parse by forcibly reducing incomplete constituents and skipping portions of input. The parser is equipped with a powerful skip-and-fit recovery mechanism that allows it to operate effectively in the face of ill-formed input or under severe time pressure.

The parser's speed averaged between 0.45 and 0.5 seconds per sentence, or up to 2600 words per minute, on a 21 MIPS SparcStation ELC. The parser is a full grammar parser that initially attempts to generate a complete analysis for each sentence. However, it has a built-in timer that regulates the amount of time allowed for parsing any one sentence. If a parse is not returned before the allotted time elapses, the parser enters the skip-and-fit mode, attempting to "fit" the parse by forcibly reducing incomplete constituents and skipping portions of input.

A stochastic parts of speech tagger is used to preprocess the text, removing most of the lexical level ambiguity from the input text prior to parsing. This allows the parser to operate effectively in the face of ill-formed input or under severe time pressure. The parser's speed averaged between 0.45 and 0.5 seconds per sentence, or up to 2600 words per minute, on a 21 MIPS SparcStation ELC.

The system uses a conservative dictionary-assisted suffix trimmer to reduce inflected word forms to their root forms. The suffix trimmer performs two tasks: reducing inflected word forms to their root forms as specified in the dictionary, and converting nominalized verb forms to the root forms of corresponding verbs. This is accomplished by removing a standard suffix, replacing it with a standard root ending, and checking the newly created word against the dictionary. The system uses Oxford Advanced Learner's Dictionary (OALD) MRD to perform these tasks.

Description

Tomek Strzalkowski's system enhances traditional keyword-based document retrieval using advanced NLP, outperforming baseline in CACM-3204 collection.

Technical Information

  • File Format: PDF
  • File Size: 558 KB
  • Pages: 6
  • Language: EN
  • Author: Tomek Strzalkowski
  • Total Downloads: 168
  • Last Updated: 2 hours ago

Document Overview

This PDF document about Robust Natural Language Processing for Information Retrieval provides comprehensive information and guidance. Whether you're a beginner or advanced user, this resource offers valuable insights into Robust Natural Language Processing for Information Retrieval.

Related Topics

If you're interested in Robust Natural Language Processing for Information Retrieval, you might also want to explore:

Download Robust Natural Language Processing for Information Retrieval eBooks for free and learn more about Robust Natural Language Processing for Information Retrieval. These books contain exercises and tutorials to improve your practical skills, at all levels!

Not satisfied with this document? We have related documents to Robust Natural Language Processing for Information Retrieval, try searching with similar keywords: Robust Natural Language Processing for Information Retrieval, Natural Language Processing Information Retrieval, "Hinder och möjliggörare för 1.5°-livsstilar: Ytliga och djupgående strukturella faktorer som påverkar potentialen för hållbar k, Ändring av genomföranderam för en europeisk plattform för utbyte av balansenergi från frekvensåterställn ingsreserver med manuell, Rekommendationer för vaccination mot covid-19 för särskilda grupper av barn -, förstudie för att utvärdera förutsättningarna att genom en innovationsupphandli ng utveckla en drifttjänst för geoenergilager, Självkänsla och KBT ‐ Påverkas självkänslan vid KBT för depression och ångesttillstånd?Se lf‐esteem and CBT ‐ How does CBT for de, Matglädje för alla: en guide till rätt konsistens för olika behov

You can download PDF versions of the user's guide, manuals and ebooks about Robust Natural Language Processing for Information Retrieval, you can also find and download for free A free online manual (notices) with beginner and intermediate, Downloads Documentation, You can download PDF files (or DOC and PPT) about Robust Natural Language Processing for Information Retrieval for free, but please respect copyrighted ebooks.