Classification active learning based on mutual information

Jamshid Sourati; Murat Akcakaya; Jennifer G. Dy; Todd K. Leen; Deniz Erdogmus

doi:10.3390/e18020051

Classification active learning based on mutual information

Jamshid Sourati, Murat Akcakaya, Jennifer G. Dy, Todd K. Leen, Deniz Erdogmus

Biomedical Engineering

Research output: Contribution to journal › Article › peer-review

19 Scopus citations

Abstract

Selecting a subset of samples to label from a large pool of unlabeled data points, such that a sufficiently accurate classifier is obtained using a reasonably small training set is a challenging, yet critical problem. Challenging, since solving this problem includes cumbersome combinatorial computations, and critical, due to the fact that labeling is an expensive and time-consuming task, hence we always aim to minimize the number of required labels. While information theoretical objectives, such as mutual information (MI) between the labels, have been successfully used in sequential querying, it is not straightforward to generalize these objectives to batch mode. This is because evaluation and optimization of functions which are trivial in individual querying settings become intractable for many objectives when we are to select multiple queries. In this paper, we develop a framework, where we propose efficient ways of evaluating and maximizing the MI between labels as an objective for batch mode active learning. Our proposed framework efficiently reduces the computational complexity from an order proportional to the batch size, when no approximation is applied, to the linear cost. The performance of this framework is evaluated using data sets from several fields showing that the proposed framework leads to efficient active learning for most of the data sets.

Original language	English (US)
Article number	51
Journal	Entropy
Volume	18
Issue number	2
DOIs	https://doi.org/10.3390/e18020051
State	Published - 2016

Keywords

Active learning
Classification
Mutual information
Submodular maximization

ASJC Scopus subject areas

Information Systems
Electrical and Electronic Engineering
General Physics and Astronomy
Mathematical Physics
Physics and Astronomy (miscellaneous)

Access to Document

10.3390/e18020051

Cite this

@article{38965f3385014b0da50cdc0070e9d640,

title = "Classification active learning based on mutual information",

abstract = "Selecting a subset of samples to label from a large pool of unlabeled data points, such that a sufficiently accurate classifier is obtained using a reasonably small training set is a challenging, yet critical problem. Challenging, since solving this problem includes cumbersome combinatorial computations, and critical, due to the fact that labeling is an expensive and time-consuming task, hence we always aim to minimize the number of required labels. While information theoretical objectives, such as mutual information (MI) between the labels, have been successfully used in sequential querying, it is not straightforward to generalize these objectives to batch mode. This is because evaluation and optimization of functions which are trivial in individual querying settings become intractable for many objectives when we are to select multiple queries. In this paper, we develop a framework, where we propose efficient ways of evaluating and maximizing the MI between labels as an objective for batch mode active learning. Our proposed framework efficiently reduces the computational complexity from an order proportional to the batch size, when no approximation is applied, to the linear cost. The performance of this framework is evaluated using data sets from several fields showing that the proposed framework leads to efficient active learning for most of the data sets.",

keywords = "Active learning, Classification, Mutual information, Submodular maximization",

author = "Jamshid Sourati and Murat Akcakaya and Dy, {Jennifer G.} and Leen, {Todd K.} and Deniz Erdogmus",

note = "Publisher Copyright: {\textcopyright} 2016 by the authors.",

year = "2016",

doi = "10.3390/e18020051",

language = "English (US)",

volume = "18",

journal = "Entropy",

issn = "1099-4300",

publisher = "Multidisciplinary Digital Publishing Institute (MDPI)",

number = "2",

}

TY - JOUR

T1 - Classification active learning based on mutual information

AU - Sourati, Jamshid

AU - Akcakaya, Murat

AU - Dy, Jennifer G.

AU - Leen, Todd K.

AU - Erdogmus, Deniz

PY - 2016

Y1 - 2016

N2 - Selecting a subset of samples to label from a large pool of unlabeled data points, such that a sufficiently accurate classifier is obtained using a reasonably small training set is a challenging, yet critical problem. Challenging, since solving this problem includes cumbersome combinatorial computations, and critical, due to the fact that labeling is an expensive and time-consuming task, hence we always aim to minimize the number of required labels. While information theoretical objectives, such as mutual information (MI) between the labels, have been successfully used in sequential querying, it is not straightforward to generalize these objectives to batch mode. This is because evaluation and optimization of functions which are trivial in individual querying settings become intractable for many objectives when we are to select multiple queries. In this paper, we develop a framework, where we propose efficient ways of evaluating and maximizing the MI between labels as an objective for batch mode active learning. Our proposed framework efficiently reduces the computational complexity from an order proportional to the batch size, when no approximation is applied, to the linear cost. The performance of this framework is evaluated using data sets from several fields showing that the proposed framework leads to efficient active learning for most of the data sets.

AB - Selecting a subset of samples to label from a large pool of unlabeled data points, such that a sufficiently accurate classifier is obtained using a reasonably small training set is a challenging, yet critical problem. Challenging, since solving this problem includes cumbersome combinatorial computations, and critical, due to the fact that labeling is an expensive and time-consuming task, hence we always aim to minimize the number of required labels. While information theoretical objectives, such as mutual information (MI) between the labels, have been successfully used in sequential querying, it is not straightforward to generalize these objectives to batch mode. This is because evaluation and optimization of functions which are trivial in individual querying settings become intractable for many objectives when we are to select multiple queries. In this paper, we develop a framework, where we propose efficient ways of evaluating and maximizing the MI between labels as an objective for batch mode active learning. Our proposed framework efficiently reduces the computational complexity from an order proportional to the batch size, when no approximation is applied, to the linear cost. The performance of this framework is evaluated using data sets from several fields showing that the proposed framework leads to efficient active learning for most of the data sets.

KW - Active learning

KW - Classification

KW - Mutual information

KW - Submodular maximization

UR - http://www.scopus.com/inward/record.url?scp=84960455436&partnerID=8YFLogxK

UR - http://www.scopus.com/inward/citedby.url?scp=84960455436&partnerID=8YFLogxK

U2 - 10.3390/e18020051

DO - 10.3390/e18020051

M3 - Article

AN - SCOPUS:84960455436

SN - 1099-4300

VL - 18

JO - Entropy

JF - Entropy

IS - 2

M1 - 51

ER -

Classification active learning based on mutual information

Abstract

Keywords

ASJC Scopus subject areas

Access to Document

Other files and links

Fingerprint

Cite this