<?xml version="1.0" encoding="UTF-8"?>
<article article-type="research-article" xml:lang="en" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher">global-journal-of-computer-science-and-technology-g-interdisciplinary</journal-id>
<journal-title-group>
<journal-title>Global Journal of Computer Science and Technology - G: Interdisciplinary</journal-title>
</journal-title-group>
<issn publication-format="print">0975-4350</issn>
<issn publication-format="electronic">0975-4172</issn>
<publisher><publisher-name>Global Journals Publishing Group Incorporated</publisher-name></publisher>
<self-uri xlink:href="https://globaljournals.org/journal-seo-export/jats/54960.xml" />
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">54960</article-id>
<title-group>
<article-title>Construction of Large Scale Isolated Word Speech Corpus in Bangla</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Khan</surname><given-names>Md. Farukuzzaman</given-names></name><xref ref-type="aff" rid="aff1" />
</contrib>
</contrib-group>
<aff id="aff1">BANGLADESH, Islamic University</aff>
<pub-date publication-format="electronic" date-type="pub" iso-8601-date="2018-01-15">
<day>15</day>
<month>01</month>
<year>2018</year>
</pub-date>
<volume>18</volume>
<issue>G2</issue>
<fpage>21</fpage>
<lpage>26</lpage>
<abstract><p>A new speech corpus of isolated words in Bangla language has recorded including high frequent words from a text corpus BdNC01. It has designed specifically for various research activities related to speaker-independent Bangla speech recognition. The database consists of speech of 100 speakers, each of them speaking 1081 words. Another 50 new speakers were employed to speak all the list of words to construct a test database. Every utterance was repeated five times in different days to avoid time variation of speaker property. The total 375 hours of original recording makes the corpora largest in its type, size and language domain. This paper describes the motivation for the corpora and the processes undertaken in its construction. The paper concludes with the usability of the corpus.</p></abstract>
<kwd-group kwd-group-type="author-generated">
<kwd>Bangla</kwd>
<kwd>speech corpora</kwd>
<kwd>BDNC01</kwd>
<kwd>vocabulary</kwd>
<kwd>isolated word</kwd>
<kwd>speech recognition.</kwd>
</kwd-group>
<self-uri content-type="pdf" xlink:href="https://globaljournals.org/GJCST_Volume18/4-Construction-of-Large-Scale-Isolated.pdf" />
<self-uri content-type="html" xlink:href="https://globaljournals.org/scholarly-articles/construction-of-large-scale-isolated-word-speech-corpus-in-bangla/" />
</article-meta>
</front>
<body>
<sec>
<title>Full Text</title>
<p>A new speech corpus of isolated words in Bangla language has been recorded including high frequent words from a text corpus BdNC01. It has been specifically designed for various research activities related to speaker-independent Bangla speech recognition. The database consists of speech of 100 speakers, each of them speaking 1081 words. Another 50 new speakers were employed to speak all the list of speech to construct a test database. Every utterance was repeated 5 times in different days to avoid time variation of speaker property. The total 400 hours of recording makes the corpora largest in its type, size and language domain. This paper describes the motivation for the corpora and the processes undertaken in its construction. The paper concludes with the usability of the corpus.</p>
</sec>
</body>
</article>