Metric and non-metric proximity transformations at linear costs

Andrej Gisbrecht; Frank-michael Schleif

doi:10.1016/j.neucom.2015.04.017

Metric and non-metric proximity transformations at linear costs

Andrej Gisbrecht, Frank-michael Schleif

Computer Science

Research output: Contribution to journal › Article › peer-review

26 Citations (Scopus)

277 Downloads (Pure)

Abstract

Domain specific (dis-)similarity or proximity measures used e.g. in alignment algorithms of sequence data are popular to analyze complicated data objects and to cover domain specific data properties. Without an underlying vector space these data are given as pairwise (dis-)similarities only. The few available methods for such data focus widely on similarities and do not scale to large datasets. Kernel methods are very effective for metric similarity matrices, also at large scale, but costly transformations are necessary starting with non-metric (dis-) similarities. We propose an integrative combination of Nyström approximation, potential double centering and eigenvalue correction to obtain valid kernel matrices at linear costs in the number of samples. By the proposed approach effective kernel approaches become accessible. Experiments with several larger (dis-)similarity datasets show that the proposed method achieves much better runtime performance than the standard strategy while keeping competitive model accuracy. The main contribution is an efficient and accurate technique, to convert (potentially non-metric) large scale dissimilarity matrices into approximated positive semi-definite kernel matrices at linear costs.

Original language	English
Journal	Neurocomputing
Early online date	21 Apr 2015
DOIs	https://doi.org/10.1016/j.neucom.2015.04.017
Publication status	E-pub ahead of print - 21 Apr 2015

Access to Document

10.1016/j.neucom.2015.04.017

Gisbrecht_Schleif_Metric_non_metric_proximity_Neurocomputing_2015
NOTICE: this is the author’s version of a work that was accepted for publication in Neurocomputing. Changes resulting from the publishing process, such as peer review, editing, corrections, structural formatting, and other quality control mechanisms may not be reflected in this document. Changes may have been made to this work since it was submitted for publication. A definitive version was subsequently published in Neurocomputing, DOI: 10.1016/j.neucom.2015.04.017 Eligibility for repository checked
Accepted author manuscript, 2.57 MBLicence: Other (please specify with Rights Statement)

http://linkinghub.elsevier.com/retrieve/pii/S092523121500452X

Cite this

@article{cbf0bca8b95849809cf3b1f130c11510,

title = "Metric and non-metric proximity transformations at linear costs",

abstract = "Domain specific (dis-)similarity or proximity measures used e.g. in alignment algorithms of sequence data are popular to analyze complicated data objects and to cover domain specific data properties. Without an underlying vector space these data are given as pairwise (dis-)similarities only. The few available methods for such data focus widely on similarities and do not scale to large datasets. Kernel methods are very effective for metric similarity matrices, also at large scale, but costly transformations are necessary starting with non-metric (dis-) similarities. We propose an integrative combination of Nystr{\"o}m approximation, potential double centering and eigenvalue correction to obtain valid kernel matrices at linear costs in the number of samples. By the proposed approach effective kernel approaches become accessible. Experiments with several larger (dis-)similarity datasets show that the proposed method achieves much better runtime performance than the standard strategy while keeping competitive model accuracy. The main contribution is an efficient and accurate technique, to convert (potentially non-metric) large scale dissimilarity matrices into approximated positive semi-definite kernel matrices at linear costs.",

author = "Andrej Gisbrecht and Frank-michael Schleif",

year = "2015",

month = apr,

day = "21",

doi = "10.1016/j.neucom.2015.04.017",

language = "English",

journal = "Neurocomputing",

issn = "0925-2312",

publisher = "Elsevier",

}

TY - JOUR

T1 - Metric and non-metric proximity transformations at linear costs

AU - Gisbrecht, Andrej

AU - Schleif, Frank-michael

PY - 2015/4/21

Y1 - 2015/4/21

N2 - Domain specific (dis-)similarity or proximity measures used e.g. in alignment algorithms of sequence data are popular to analyze complicated data objects and to cover domain specific data properties. Without an underlying vector space these data are given as pairwise (dis-)similarities only. The few available methods for such data focus widely on similarities and do not scale to large datasets. Kernel methods are very effective for metric similarity matrices, also at large scale, but costly transformations are necessary starting with non-metric (dis-) similarities. We propose an integrative combination of Nyström approximation, potential double centering and eigenvalue correction to obtain valid kernel matrices at linear costs in the number of samples. By the proposed approach effective kernel approaches become accessible. Experiments with several larger (dis-)similarity datasets show that the proposed method achieves much better runtime performance than the standard strategy while keeping competitive model accuracy. The main contribution is an efficient and accurate technique, to convert (potentially non-metric) large scale dissimilarity matrices into approximated positive semi-definite kernel matrices at linear costs.

AB - Domain specific (dis-)similarity or proximity measures used e.g. in alignment algorithms of sequence data are popular to analyze complicated data objects and to cover domain specific data properties. Without an underlying vector space these data are given as pairwise (dis-)similarities only. The few available methods for such data focus widely on similarities and do not scale to large datasets. Kernel methods are very effective for metric similarity matrices, also at large scale, but costly transformations are necessary starting with non-metric (dis-) similarities. We propose an integrative combination of Nyström approximation, potential double centering and eigenvalue correction to obtain valid kernel matrices at linear costs in the number of samples. By the proposed approach effective kernel approaches become accessible. Experiments with several larger (dis-)similarity datasets show that the proposed method achieves much better runtime performance than the standard strategy while keeping competitive model accuracy. The main contribution is an efficient and accurate technique, to convert (potentially non-metric) large scale dissimilarity matrices into approximated positive semi-definite kernel matrices at linear costs.

U2 - 10.1016/j.neucom.2015.04.017

DO - 10.1016/j.neucom.2015.04.017

M3 - Article

SN - 0925-2312

JO - Neurocomputing

JF - Neurocomputing

ER -

Metric and non-metric proximity transformations at linear costs

Abstract

Access to Document

Fingerprint

Cite this