<?xml version="1.0" encoding="UTF-8"?>
<article article-type="research-article" xml:lang="en" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher">global-journal-of-computer-science-and-technology-c-software-data-engineering</journal-id>
<journal-title-group>
<journal-title>Global Journal of Computer Science and Technology - C: Software &amp; Data Engineering</journal-title>
</journal-title-group>
<issn publication-format="print">0975-4350</issn>
<issn publication-format="electronic">0975-4172</issn>
<publisher><publisher-name>Global Journals Publishing Group Incorporated</publisher-name></publisher>
<self-uri xlink:href="https://globaljournals.org/journal-seo-export/jats/55287.xml" />
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">55287</article-id>
<title-group>
<article-title>Feature Extraction and Duplicate Detection for Text Mining: A Survey</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>S</surname><given-names>Ramya R</given-names></name><xref ref-type="aff" rid="aff1" />
</contrib>
<contrib contrib-type="author"><name><surname>R</surname><given-names>Venugopal K</given-names></name></contrib>
</contrib-group>
<aff id="aff1">INDIA, University Visvesvaraya College of Engineering, UVCE</aff>
<pub-date publication-format="electronic" date-type="pub" iso-8601-date="2016-01-15">
<day>15</day>
<month>01</month>
<year>2016</year>
</pub-date>
<volume>16</volume>
<issue>C5</issue>
<fpage>1</fpage>
<lpage>20</lpage>
<abstract><p>Text mining, also known as Intelligent Text Analysis is an important research area. It is very difficult to focus on the most appropriate information due to the high dimensionality of data. Feature Extraction is one of the important techniques in data reduction to discover the most important features. Proce-ssing massive amount of data stored in a unstructured form is a challenging task. Several pre-processing methods and algo-rithms are needed to extract useful features from huge amount of data. The survey covers different text summarization, classi-fication, clustering methods to discover useful features and also discovering query facets which are multiple groups of words or phrases that explain and summarize the content covered by a query thereby reducing time taken by the user.</p></abstract>
<kwd-group kwd-group-type="author-generated">
<kwd>text feature extraction</kwd>
<kwd>text mining</kwd>
<kwd>query search</kwd>
<kwd>text classification.</kwd>
</kwd-group>
<self-uri content-type="pdf" xlink:href="https://globaljournals.org/GJCST_Volume16/1-Feature-Extraction-and-Duplicate.pdf" />
<self-uri content-type="html" xlink:href="https://globaljournals.org/scholarly-articles/feature-extraction-and-duplicate-detection-for-text-mining-a-survey/" />
</article-meta>
</front>
<body>
<sec>
<title>Full Text</title>
<p>Text mining, also known as Intelligent Text Analysis is an important research area. It is very difficult to focus on the most appropriate information due to the high dimensionality of data. Feature Extraction is one of the important techniques in data reduction to discover the most important features. Proce- ssing massive amount of data stored in a unstructured form is a challenging task. Several pre-processing methods and algo- rithms are needed to extract useful features from huge amount of data. The survey covers different text summarization, classi- fication, clustering methods to discover useful features and also discovering query facets which are multiple groups of words or phrases that explain and summarize the content covered by a query thereby reducing time taken by the user. Dealing with collection of text documents, it is also very important to filter out duplicate data. Once duplicates are deleted, it is recommended to replace the removed duplicates. Hence we also review the literature on duplicate detection and data fusion (remove and replace duplicates).The survey provides existing text mining techniques to extract relevant features, detect duplicates and to replace the duplicate data to get fine grained knowledge to the user.</p>
</sec>
</body>
</article>