Refat Aljumily

Research

Agglomerative Hierarchical Clustering: An Introduction to Essentials. (3) Standardization, Normalization and Dimensionality Reduction of a Data Matrix

Article April 29, 2016

In a previous tutorial article I looked at a proximity coefficient and, in the light of that proximity created a vectordistance matrix and used it to construct a hierarchical tree using different hierarchical clustering methods which will be the basis for exploratory multivariate analysis. The present article deals with three topics: (i) standardization for variable scales variation, (ii) normalization for sample length variation, and (iii) dimensionality reduction or minimization of data space. These techniques reflect the author’s academic background and particular area of interest and are, by necessity, not a particular purpose and are straightforwardly applicable to other kinds of data, and thus to a wide range of analysis in Linguistics. My treatment of these techniques is, necessarily, introductory and brief. I hope that this article will provide practitioners with an introductory overview of these techniques used for cluster analysis of electronic corpora of linguistic data. The assumption is that the data is in the form of an m x n matrix D in which, may require to transform it in various ways prior to cluster analyzing it. Standardized data matrix enables practitioners to measure the variation between n-variables and to cluster the cases they describe in common scales and values, regardless of their original scales and values. Normalized data matrix enables practitioners to eliminate the effect of variation in length among n-samples and to cluster them as if they were all (about) the same length, regardless of their original length. Dimensionality-reduced space data matrix enables practitioners to select and/or extract n-most interesting variables relevant to the research question and to visualize an existing pattern, regardless of the original space. A worked example is given to illustrate the effect each transformation technique has on a given data matrix. These transformation techniques have their own strengths and weakness but are beyond the scope of

Agglomerative Hierarchical Clustering: An Introduction to Essentials (1) Proximity Coefficients and Creation of a Vector-Distance Matrix and (2) Construction of the Hierarchical Tree and a Selection of Methods

Article April 29, 2016

The article is on a particular type of cluster analysis, agglomerative hierarchical analysis, and is a series of four main parts. The first part deals with proximity coefficients and the creation of a vector-distance matrix. The second part deals with the construction of the hierarchical tree and introduces a selection of clustering methods. The third deals with a variety of ways to transform data prior to agglomerative cluster analysis. The fourth deals with deals with measures and methods of cluster validity. The fifth and final part deals with hypothesis generation. The present article covers the first and second partsonly. It explains how agglomerative cluster analysis works by implementing it in a data matrix step by step. Different types of agglomerative hierarchical clustering methods are applied on purposely-made data matrix so different types of cluster structures are made from that same dataset. The last three parts will be covered in the next publication(s).There are many articles, tutorials, and books on this subject. The article has two main objectives: (1) to keep the discussion short and easy to understand by (hopefully) any reader and (2) to develop the motivation for using agglomerative hierarchical clustering to analyse any highdimensional data of interest with respect to some research question.

Hierarchical and Non-Hierarchical Linear and Non-Linear Clustering Methods to aoShakespeare-De Vere Authorship Question

Article April 13, 2016

In my previous article entitled, “Hierarchical and Non- Hierarchical Linear and Non-Linear Clustering Methods to “Shakespeare Authorship Question” I used Mean Proximity, as a linear hierarchical clustering method and Principal Components Analysis, as a non-hierarchical linear clustering method, Self-Organizing Map U-matrix and Voronoi Map, as non-linear clustering methods to examine various works and plays assumed to have been written by Shakespeare and Sir Francis Bacon, Christopher Marlowe, John Fletcher, and Thomas Kyd to determine which of them wrote some of Shakespeare’s disputed plays based on similarities in the use of function words, word-bi grams, and character-tri grams. The article showed that Shakespeare is not the author of all the disputed plays traditionally attributed to him according to the validated cluster analytic results and the stylistic criteria used. The article also indicated that the author did not consider it fair to include Edward de Vere(the strongest candidate in the Shakespeare authorship debate) and compare his poemsto Shakespeare’s disputed plays because poetry tends to have a particular style and a different structure than plays, and additional test was promised. The present article provides that test. In this article, I examined the 154 sonnets traditionally attributed to Shakespeare and 38 surviving poems attributed to Edward de Vere. The purpose is to give a hypothesis whether de Vere has an identifiable self-similarity and a measure of how far from/similar to Shakespeare based on the use of function words, word bi-grams, character bi-grams, and character tri-grams applying four different clustering methods: four hierarchical linear methods using Euclidean distance (Single, Average, Complete, and Ward), non-hierarchical linear multidimensional Scaling (MDS), and Kernel K-means clustering and Voronoi mapas non-linear methods. The cophenetic correlation coefficient is used to select the best result obtained from a set of

The Anonymous 1821 Translation of Goetheas Faustus: A Cluster Analytic Approach

Article January 9, 2016

The scholars, Frederick Burwick and James McKusick, published at Oxford University Press, Faustus from the German of Goethe translated by Samuel Taylor Coleridge in 2007. This edition articulated the result that Samuel Taylor Coleridge is the actual translator of the anonymously published translation Faustus from the German of Goethe (London: Boosey: 1821). The present article tests that result. The approach to test this result is stylometric. Specifically, function word usage is selected as the stylometric criterion, and 80 function words are used to define a 73-dimensional function word frequency profile vector for each text in the corpus of Coleridge's literary works and for a selection of works by a range of contemporary English authors. Each profile vector is a point in 80-dimensional vector space, and 5 different cluster analytic methods are used to determine the distribution of profile vectors in the space. If the result being tested is valid, then the profile for the 1821 translation should be closer in the space to works known to be by Coleridge than to works by the other authors. The cluster analytic results show, however, that this is not the case, and the conclusion is that the Burwick and McKusick result is falsified relative to the stylometric criterion and analytic methodology used. Where, in Popperian terms, falsification does not mean 'prove to be false'. It means that evidence which contradicts a hypothesis has been presented, and it is up to the proposer of the hypothesis either to show that the evidence is inadmissible or irrelevant, or else to emend the hypothesis accordingly. The rest of the article is organized as follows. In section 1 we give the motivation for doing this work. In section 2 we provide a quick introduction to the 1821 Faustus translations that we hope will shed some light on the problem. In section 3 we discuss the previous attempts to attribute the 1821 Faustus to Coleridge. In section 4 we outline the methodology used to add

Show all publications