Information-based clustering

Slonim , N., Atwal, G. S., Gasper, T., Bialek, W. (December 2005) Information-based clustering. Proc Nat Acad Sci (USA), 102 (51). pp. 18297-18302.

URL: https://www.ncbi.nlm.nih.gov/pubmed/16352721

DOI: 10.1073/pnas.0507432102

Abstract

In an age of increasingly large data sets, investigators in many different disciplines have turned to clustering as a tool for data analysis and exploration. Existing clustering methods, however, typically depend on several nontrivial assumptions about the structure of data. Here, we reformulate the clustering problem from an information theoretic perspective that avoids many of these assumptions. In particular, our formulation obviates the need for defining a cluster “prototype,” does not require an a priori similarity metric, is invariant to changes in the representation of the data, and naturally captures nonlinear relations. We apply this approach to different domains and find that it consistently produces clusters that are more coherent than those extracted by existing algorithms. Finally, our approach provides a way of clustering based on collective notions of similarity rather than the traditional pairwise measures.

Item Type:	Paper
Uncontrolled Keywords:	similarity algorithms
Subjects:	bioinformatics > computational biology bioinformatics > genomics and proteomics > annotation > dataset annotation
CSHL Authors:	Atwal, Gurinder Singh
Communities:	CSHL labs > Atwal lab
Depositing User:	CSHL Librarian
Date:	20 December 2005
Date Deposited:	05 Jan 2012 19:52
Last Modified:	19 Dec 2016 21:44
PMCID:	PMC1317937
Related URLs:	Publisher
URI:	https://repository.cshl.edu/id/eprint/22706

Actions (login required)

Administrator's edit/view item