<?xml version="1.0" encoding="UTF-8"?>
<article article-type="research-article" xml:lang="en" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher">global-journal-of-computer-science-and-technology-c-software-data-engineering</journal-id>
<journal-title-group>
<journal-title>Global Journal of Computer Science and Technology - C: Software &amp; Data Engineering</journal-title>
</journal-title-group>
<issn publication-format="print">0975-4350</issn>
<issn publication-format="electronic">0975-4172</issn>
<publisher><publisher-name>Global Journals Publishing Group Incorporated</publisher-name></publisher>
<self-uri xlink:href="https://globaljournals.org/journal-seo-export/jats/54845.xml" />
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">54845</article-id>
<title-group>
<article-title>Protein and Other Biomedical Entity Name Tagging from Pdf File using NLP and Visualization of that Entity</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Rizvee</surname><given-names>Md. Arif</given-names></name><xref ref-type="aff" rid="aff1" />
</contrib>
<contrib contrib-type="author"><name><surname>Arju</surname><given-names>Md. Ashfakur Rahman</given-names></name></contrib>
<contrib contrib-type="author"><name><surname>Tareque</surname><given-names>Saifuddin Mohammad</given-names></name></contrib>
</contrib-group>
<aff id="aff1">BANGLADESH, Daffodil International University</aff>
<pub-date publication-format="electronic" date-type="pub" iso-8601-date="2019-01-15">
<day>15</day>
<month>01</month>
<year>2019</year>
</pub-date>
<volume>19</volume>
<issue>C1</issue>
<fpage>23</fpage>
<lpage>25</lpage>
<abstract><p>Protein and other biomedical entities such as a gene, chromosome names are key elements in bioinformatics. Identifying them individually from the pdf file is very challenging. Because a text pdf document can contain lots of information, identifying them is not so much easy task. So the main focus in our project is converting the pdf file to humanreadable text file then we will have to find the gene and other entities from the GENIA tagger website database. Using natural language processing GENIA tagger will give us the name of all the protein, gene, and other biomedical entity name. After identifying them, we will save it to database. Then we will visualize the related data.</p></abstract>
<kwd-group kwd-group-type="author-generated">
<kwd>tagging protein</kwd>
<kwd>gene</kwd>
<kwd>and other biomedical entities</kwd>
<kwd>natural language processing</kwd>
<kwd>GENIA tagger</kwd>
<kwd>data visualization.</kwd>
</kwd-group>
<self-uri content-type="pdf" xlink:href="https://globaljournals.org/GJCST_Volume19/4-Protein-and-Other-Biomedical-Entity.pdf" />
<self-uri content-type="html" xlink:href="https://globaljournals.org/scholarly-articles/protein-and-other-biomedical-entity-name-tagging-from-pdf-file-using-nlp-and-visualization-of-that-entity/" />
</article-meta>
</front>
<body>
<sec>
<title>Full Text</title>
<p>Protein and other biomedical entities such as a gene, chromosome names are key elements in bioinformatics. Identifying them individually from the pdf file is very challenging. Because a text pdf document can contain lots of information, identifying them is not so much easy task. So the main focus in our project is converting the pdf file to humanreadable text file then we will have to find the gene and other entities from the GENIA tagger website database. Using natural language processing GENIA tagger will give us the name of all the protein, gene, and other biomedical entity name. After identifying them, we will save it to database. Then we will visualize the related data.</p>
</sec>
</body>
</article>