<?xml version="1.0" encoding="UTF-8"?>
<article article-type="research-article" xml:lang="en" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher">global-journal-of-science-frontier-research-g-bio-tech-genetics</journal-id>
<journal-title-group>
<journal-title>Global Journal of Science Frontier Research - G: Bio-Tech &amp; Genetics</journal-title>
</journal-title-group>
<issn publication-format="print">0975-5896</issn>
<issn publication-format="electronic">2249-4626</issn>
<publisher><publisher-name>Global Journals Publishing Group Incorporated</publisher-name></publisher>
<self-uri xlink:href="https://globaljournals.org/journal-seo-export/jats/52612.xml" />
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">52612</article-id>
<title-group>
<article-title>Comparative Study of Three Clustering Algorithms for Microarray Data</article-title>
<subtitle>Clustering Techniques for Gene Expression Analysis</subtitle>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>G</surname><given-names>Dicky John Davis</given-names></name><xref ref-type="aff" rid="aff1" />
</contrib>
<contrib contrib-type="author"><name><surname>Pious</surname><given-names>Noveenaa</given-names></name></contrib>
</contrib-group>
<aff id="aff1">INDIA</aff>
<pub-date publication-format="electronic" date-type="pub" iso-8601-date="2022-05-14">
<day>14</day>
<month>05</month>
<year>2022</year>
</pub-date>
<volume>22</volume>
<issue>G1</issue>
<fpage>11</fpage>
<lpage>17</lpage>
<abstract><p>High throughput genomic data analysis is becoming an increasingly integral part of biomedical research. The information derived from gene expression analysis helps in diagnosing the treatment modality given to the patient. However, the amount of data is humongous and becomes complex to examine manually. Unsupervised machine learning algorithms perform complex tasks on an unlabelled data by clustering to comprehend the underlying structure and behaviour of the pattern. Clustering microarray data, examines the differential expressed genes found by grouping the genes based on the similarity of the expression values. In this study, we propose to elucidate the best clustering algorithm for gene expression data on various clinical conditions. The proposed study was carried on three gene expression datasets of Severe acute respiratory syndrome, Amyotrophic lateral sclerosis and Parkinson’s disease. Differentially expressed genes were found at three p-values 0.01, 0.05, 0.001 and the most significant number of genes were retrieved at p-value 0.05. We experimented the differential expressed genes on three clustering algorithms, namely Hierarchical clustering, k-means clustering and fuzzy clustering of the three diseases. The performance of the three clustering algorithms was evaluated using the internal validity index, wherein Hierarchical clustering was found to be best for gene expression data.</p></abstract>
<kwd-group kwd-group-type="author-generated">
<kwd>hierarchical clustering; k-means clustering; fuzzy clustering; differentially expressed genes; microarray data.</kwd>
</kwd-group>
<self-uri content-type="pdf" xlink:href="https://globaljournals.org/GJSFR_Volume22/2-Comparative-Study-of-three-Clustering.pdf" />
<self-uri content-type="html" xlink:href="https://globaljournals.org/scholarly-articles/comparative-study-of-three-clustering-algorithms-for-microarray-data-8/" />
</article-meta>
</front>
<body>
<sec>
<title>Full Text</title>
<p>High throughput genomic data analysis is becoming an increasingly integral part of biomedical research. The information derived from gene expression analysis helps in diagnosing the treatment modality given to the patient. However, the amount of data is humongous and becomes complex to examine manually. Unsupervised machine learning algorithms perform complex tasks on an unlabelled data by clustering to comprehend the underlying structure and behaviour of the pattern. Clustering microarray data, examines the differential expressed genes found by grouping the genes based on the similarity of the expression values. In this study, we propose to elucidate the best clustering algorithm for gene expression data on various clinical conditions. The proposed study was carried on three gene expression datasets of Severe acute respiratory syndrome, Amyotrophic lateral sclerosis and Parkinsonâ€™s disease. Differentially expressed genes were found at three p-values 0.01, 0.05, 0.001 and the most significant number of genes were retrieved at p-value 0.05. We experimented the differential expressed genes on three clustering algorithms, namely Hierarchical clustering, kmeans clustering and fuzzy clustering of the three diseases. The performance of the three clustering algorithms was evaluated using the internal validity index, wherein Hierarchical clustering was found to be best for gene expression data.</p>
</sec>
</body>
</article>