Occupancy classification of position weight matrix-inferred Transcription Factor Binding Sites

Hollis Wright; Aaron Cohen; Kemal Sönmez; Gregory Yochum; Shannon McWeeney

doi:10.1371/journal.pone.0026160

Occupancy classification of position weight matrix-inferred Transcription Factor Binding Sites

Hollis Wright, Aaron Cohen, Kemal Sönmez, Gregory Yochum, Shannon McWeeney

Medical Informatics and Clinical Epidemiology

Research output: Contribution to journal › Article › peer-review

2 Scopus citations

Abstract

Background: Computational prediction of Transcription Factor Binding Sites (TFBS) from sequence data alone is difficult and error-prone. Machine learning techniques utilizing additional environmental information about a predicted binding site (such as distances from the site to particular chromatin features) to determine its occupancy/functionality class show promise as methods to achieve more accurate prediction of true TFBS in silico. We evaluate the Bayesian Network (BN) and Support Vector Machine (SVM) machine learning techniques on four distinct TFBS data sets and analyze their performance. We describe the features that are most useful for classification and contrast and compare these feature sets between the factors. Results: Our results demonstrate good performance of classifiers both on TFBS for transcription factors used for initial training and for TFBS for other factors in cross-classification experiments. We find that distances to chromatin modifications (specifically, histone modification islands) as well as distances between such modifications to be effective predictors of TFBS occupancy, though the impact of individual predictors is largely TF specific. In our experiments, Bayesian network classifiers outperform SVM classifiers. Conclusions: Our results demonstrate good performance of machine learning techniques on the problem of occupancy classification, and demonstrate that effective classification can be achieved using distances to chromatin features. We additionally demonstrate that cross-classification of TFBS is possible, suggesting the possibility of constructing a generalizable occupancy classifier capable of handling TFBS for many different transcription factors.

Original language	English (US)
Article number	e26160
Journal	PloS one
Volume	6
Issue number	11
DOIs	https://doi.org/10.1371/journal.pone.0026160
State	Published - Nov 4 2011

ASJC Scopus subject areas

General Biochemistry, Genetics and Molecular Biology
General Agricultural and Biological Sciences
General

Access to Document

10.1371/journal.pone.0026160

Cite this

@article{7201bb38a0144f3f8a23c1228ff1f7d3,

title = "Occupancy classification of position weight matrix-inferred Transcription Factor Binding Sites",

abstract = "Background: Computational prediction of Transcription Factor Binding Sites (TFBS) from sequence data alone is difficult and error-prone. Machine learning techniques utilizing additional environmental information about a predicted binding site (such as distances from the site to particular chromatin features) to determine its occupancy/functionality class show promise as methods to achieve more accurate prediction of true TFBS in silico. We evaluate the Bayesian Network (BN) and Support Vector Machine (SVM) machine learning techniques on four distinct TFBS data sets and analyze their performance. We describe the features that are most useful for classification and contrast and compare these feature sets between the factors. Results: Our results demonstrate good performance of classifiers both on TFBS for transcription factors used for initial training and for TFBS for other factors in cross-classification experiments. We find that distances to chromatin modifications (specifically, histone modification islands) as well as distances between such modifications to be effective predictors of TFBS occupancy, though the impact of individual predictors is largely TF specific. In our experiments, Bayesian network classifiers outperform SVM classifiers. Conclusions: Our results demonstrate good performance of machine learning techniques on the problem of occupancy classification, and demonstrate that effective classification can be achieved using distances to chromatin features. We additionally demonstrate that cross-classification of TFBS is possible, suggesting the possibility of constructing a generalizable occupancy classifier capable of handling TFBS for many different transcription factors.",

author = "Hollis Wright and Aaron Cohen and Kemal S{\"o}nmez and Gregory Yochum and Shannon McWeeney",

year = "2011",

month = nov,

day = "4",

doi = "10.1371/journal.pone.0026160",

language = "English (US)",

volume = "6",

journal = "PloS one",

issn = "1932-6203",

publisher = "Public Library of Science",

number = "11",

}

TY - JOUR

T1 - Occupancy classification of position weight matrix-inferred Transcription Factor Binding Sites

AU - Wright, Hollis

AU - Cohen, Aaron

AU - Sönmez, Kemal

AU - Yochum, Gregory

AU - McWeeney, Shannon

PY - 2011/11/4

Y1 - 2011/11/4

N2 - Background: Computational prediction of Transcription Factor Binding Sites (TFBS) from sequence data alone is difficult and error-prone. Machine learning techniques utilizing additional environmental information about a predicted binding site (such as distances from the site to particular chromatin features) to determine its occupancy/functionality class show promise as methods to achieve more accurate prediction of true TFBS in silico. We evaluate the Bayesian Network (BN) and Support Vector Machine (SVM) machine learning techniques on four distinct TFBS data sets and analyze their performance. We describe the features that are most useful for classification and contrast and compare these feature sets between the factors. Results: Our results demonstrate good performance of classifiers both on TFBS for transcription factors used for initial training and for TFBS for other factors in cross-classification experiments. We find that distances to chromatin modifications (specifically, histone modification islands) as well as distances between such modifications to be effective predictors of TFBS occupancy, though the impact of individual predictors is largely TF specific. In our experiments, Bayesian network classifiers outperform SVM classifiers. Conclusions: Our results demonstrate good performance of machine learning techniques on the problem of occupancy classification, and demonstrate that effective classification can be achieved using distances to chromatin features. We additionally demonstrate that cross-classification of TFBS is possible, suggesting the possibility of constructing a generalizable occupancy classifier capable of handling TFBS for many different transcription factors.

AB - Background: Computational prediction of Transcription Factor Binding Sites (TFBS) from sequence data alone is difficult and error-prone. Machine learning techniques utilizing additional environmental information about a predicted binding site (such as distances from the site to particular chromatin features) to determine its occupancy/functionality class show promise as methods to achieve more accurate prediction of true TFBS in silico. We evaluate the Bayesian Network (BN) and Support Vector Machine (SVM) machine learning techniques on four distinct TFBS data sets and analyze their performance. We describe the features that are most useful for classification and contrast and compare these feature sets between the factors. Results: Our results demonstrate good performance of classifiers both on TFBS for transcription factors used for initial training and for TFBS for other factors in cross-classification experiments. We find that distances to chromatin modifications (specifically, histone modification islands) as well as distances between such modifications to be effective predictors of TFBS occupancy, though the impact of individual predictors is largely TF specific. In our experiments, Bayesian network classifiers outperform SVM classifiers. Conclusions: Our results demonstrate good performance of machine learning techniques on the problem of occupancy classification, and demonstrate that effective classification can be achieved using distances to chromatin features. We additionally demonstrate that cross-classification of TFBS is possible, suggesting the possibility of constructing a generalizable occupancy classifier capable of handling TFBS for many different transcription factors.

UR - http://www.scopus.com/inward/record.url?scp=80455156114&partnerID=8YFLogxK

UR - http://www.scopus.com/inward/citedby.url?scp=80455156114&partnerID=8YFLogxK

U2 - 10.1371/journal.pone.0026160

DO - 10.1371/journal.pone.0026160

M3 - Article

C2 - 22073148

AN - SCOPUS:80455156114

SN - 1932-6203

VL - 6

JO - PloS one

JF - PloS one

IS - 11

M1 - e26160

ER -

Occupancy classification of position weight matrix-inferred Transcription Factor Binding Sites

Abstract

ASJC Scopus subject areas

Access to Document

Other files and links

Fingerprint

Cite this