Performance Evaluation of an Efficient Frequent Item sets-Based Text Clustering Approach

Send Message

To: Author

Performance Evaluation of an Efficient Frequent Item sets-Based Text Clustering Approach

Article Fingerprint

ReserarchID

CSTS90Q0

Performance Evaluation of an Efficient Frequent Item sets-Based Text Clustering Approach Banner

Key Research Insights

Synthesized scholarly intelligence & interactive research assistant
  • English
  • Afrikaans
  • Albanian
  • Amharic
  • Arabic
  • Armenian
  • Azerbaijani
  • Basque
  • Belarusian
  • Bengali
  • Bosnian
  • Bulgarian
  • Catalan
  • Cebuano
  • Chichewa
  • Chinese (Simplified)
  • Chinese (Traditional)
  • Corsican
  • Croatian
  • Czech
  • Danish
  • Dutch
  • Esperanto
  • Estonian
  • Filipino
  • Finnish
  • French
  • Frisian
  • Galician
  • Georgian
  • German
  • Greek
  • Gujarati
  • Haitian Creole
  • Hausa
  • Hawaiian
  • Hebrew
  • Hindi
  • Hmong
  • Hungarian
  • Icelandic
  • Igbo
  • Indonesian
  • Irish
  • Italian
  • Japanese
  • Javanese
  • Kannada
  • Kazakh
  • Khmer
  • Korean
  • Kurdish (Kurmanji)
  • Kyrgyz
  • Lao
  • Latin
  • Latvian
  • Lithuanian
  • Luxembourgish
  • Macedonian
  • Malagasy
  • Malay
  • Malayalam
  • Maltese
  • Maori
  • Marathi
  • Mongolian
  • Myanmar (Burmese)
  • Nepali
  • Norwegian
  • Pashto
  • Persian
  • Polish
  • Portuguese
  • Punjabi
  • Romanian
  • Russian
  • Samoan
  • Scots Gaelic
  • Serbian
  • Sesotho
  • Shona
  • Sindhi
  • Sinhala
  • Slovak
  • Slovenian
  • Somali
  • Spanish
  • Sundanese
  • Swahili
  • Swedish
  • Tajik
  • Tamil
  • Telugu
  • Thai
  • Turkish
  • Ukrainian
  • Urdu
  • Uzbek
  • Vietnamese
  • Welsh
  • Xhosa
  • Yiddish
  • Yoruba
  • Zulu
Reading Preferences
Font Size
Line Spacing
Background

Abstract

The vast amount of textual information available in electronic form is growing at a staggering rate in recent times. The task of mining useful or interesting frequent itemsets (words/terms) from very large text databases that are formed as a result of the increasing number of textual data still seems to be a quite challenging task. A great deal of attention in research community has been received by the use of such frequent itemsets for text clustering, because the dimensionality of the documents is drastically reduced by the mined frequent itemsets. Based on frequent itemsets, an efficient approach for text clustering has been devised. For mining the frequent itemsets, a renowned method, called Apriori algorithm has been used. Then, the documents are initially partitioned without overlapping by making use of mined frequent itemsets. Furthermore, by grouping the documents within the partition using derived keywords, the resultant clusters are obtained effectively. In this paper, we have presented an extensive analysis of frequent itemset-based text clustering approach for different real life datasets and the performance of the frequent itemset-based text clustering approach is evaluated with the help of evaluation measures such as, precision, recall and F-measure. The experimental results shows that the efficiency of the frequent itemset-based text clustering approach has been improved significantly for different real life datasets.

References

43 Cites in Article
  1. Hany Mahgoub,Dietmar Rosner,Nabil Ismail,Fawzy Torkey (2008). A Text Mining Technique Using Association Rules Extraction.
  2. Shenzhi Li,Tianhao Wu,William Pottenger (2005). Distributed Higher Order Association Rule Mining Using Information Extracted from Textual Data.
  3. R Baeza-Yates,B Ribeiro-Neto (1999). Modern Information Retrieval.
  4. J Han,M Kamber (2000). Data Mining: Concepts and Techniques.
  5. Jochen Dijrre,Peter Gerstl,Roland Seiffert (1999). Text Mining: Finding Nuggets in Mountains of Textual Data.
  6. Haralampos Karanikas,Christos Tjortjis,Babis Theodoulidis (2000). An Approach to Text Mining using Information Extraction.
  7. Wilks Yorick (1997). Information Extraction as a Core Language Technology.
  8. Ah-Hwee Tan (1999). Text Mining: The state of the art and the challenges.
  9. A Jain,M Murty,P Flynn (1999). Data Clustering: A Review.
  10. R Feldman,J Sanger (2007). The Text Mining Handbook.
  11. Seth Grimes (2005). The Developing Text Mining Market.
  12. M Grobelnik,D Mladenic,N Milic-Frayling (2000). Text Mining as Integration of Several Related Research Areas: Report on KDD.
  13. M Hearst (1999). Untangling Text Data Mining.
  14. Alisa Kongthon,Nathasit Gerdsri (2004). Bibliometric Analysis on Artificial Intelligence Research to Support National Artificial Intelligence Strategy in Thailand.
  15. Tom Brijs,Gilbert Swinnen,Koen Vanhoof,Geert Wets (1999). Using association rules for product assortment decisions.
  16. Jianning Dong,Perrizo,William,Qin Ding,Zhou (2000). The Application of Association Rule Mining to Remotely Sensed Data.
  17. Valentina Ceausu,Sylvie Despres (2005). Using a Blog and Text Mining to Evaluate Knowledge Construction.
  18. Nurnberger Hotho,Paass (2005). A Brief Survey of Text Mining Export.
  19. Schütze Manning (1999). Foundations of statistical natural language processing.
  20. Feldman Shatkay (2003). Mining the Biomedical Literature in the Genomic Era: An Overview.
  21. Pegah Falinouss (2007). Stock Trend Prediction using News Articles.
  22. T Nasukawa,T Nagano (2001). Text analysis and knowledge mining system.
  23. Zhou Chong,Lu Yansheng,Zou Lei,Hu Rong (2006). FICW: Frequent itemset based text clustering with window constraint.
  24. Xiangwei Liu,Pilian He (2005). A Study on Text Clustering Algorithms Based on Frequent Term Sets.
  25. Le Wang,Li Tian,Yan Jia,Weihong Han (2010). A Hybrid Algorithm for Web Document Clustering Based on Frequent Term Sets and k-Means.
  26. Zhitong Su,Wei Song,Manshan Lin,Jinhong Li (2008). Web Text Clustering for Personalized E-learning Based on Maximal Frequent Itemsets.
  27. Yongheng Wang,Yan Jia,Shuqiang Yang (2006). Short Documents Clustering in Very Large Text Databases.
  28. Florian Beil,Martin Ester,Xiaowei Xu (2002). Frequent term-based text clustering.
  29. W.-L Liu,X.-S Zheng (2005). Documents Clustering based on Frequent Term Sets.
  30. Henry Anaya-Sánchez,Aurora Pons-Porrata,Rafael Berlanga-Llavori (2010). A document clustering algorithm for discovering and describing topics.
  31. Congnan Luo,Yanjun Li,Soon Chung (2009). Text document clustering based on neighbors.
  32. O Zamir,O Etzioni (1998). Web Document Clustering: A Feasibility Demonstration.
  33. M Law,M Figueiredo,A Jain (2004). Simultaneous feature selection and clustering using mixture models.
  34. Un,Yong Nahm,Raymond Mooney (2004). Text mining with information extraction.
  35. J Lovins (1968). Development of a stemming algorithm.
  36. G Pant,Srinivasan,F Menczer (2004). Crawling the Web.
  37. R Agrawal,T Imielinski,A Swami (1993). Mining association rules between sets of items in large databases.
  38. R Agrawal,R Srikant (1994). Fast algorithms for mining association rules.
  39. Bjornar Larsen,Chinatsu Aone (1999). Fast and effective text mining using linear-time document clustering.
  40. Michael Steinbach,George Karypis,Vipin Kumar (2000). Efficient Algorithms for Creating Product Catalogs.
  41. Benjamin Fung,Ke Wang,Martin Ester (2003). Hierarchical Document Clustering Using Frequent Itemsets.
  42. Text Categorization Collection, UCI KDD Archive. 43.
  43. S Murali Krishna,S Bhavani (2010). An Efficient Approach for Text Clustering Based on Frequent Itemsets.

Funding

No external funding was declared for this work.

Conflict of Interest

The authors declare no conflict of interest.

Ethical Approval

No ethics committee approval was required for this article type.

Data Availability

Not applicable for this article.

How to Cite This Article

Dr. S.Murali Krishna. . "Performance Evaluation of an Efficient Frequent Item sets-Based Text Clustering Approach". Global Journal of Computer Science and Technology GJCST Volume 10 (GJCST Volume 10 Issue 11).

Download Citation

Journal Specifications

Crossref Journal DOI 10.17406/gjcst

Print ISSN 0975-4350

e-ISSN 0975-4172

Keywords
Version of record

v1.2

Issue date
October 10, 2010

Language
English
Experiance in AR

Explore published articles in an immersive Augmented Reality environment. Our platform converts research papers into interactive 3D books, allowing readers to view and interact with content using AR and VR compatible devices.

Read in 3D

Your published article is automatically converted into a realistic 3D book. Flip through pages and read research papers in a more engaging and interactive format.

Article Matrices
Total Views: 7K
Total Downloads: 594
All Trends

Request Access

Please fill out the form below to request access to this research paper. Your request will be reviewed by the editorial or author team.
X

This is the heading

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Ut elit tellus, luctus nec ullamcorper mattis, pulvinar dapibus leo.

High-quality academic research articles on global topics and journals.

Performance Evaluation of an Efficient Frequent Item sets-Based Text Clustering Approach

Dr. Krishna
Dr. Krishna