Semi-supervised and Incremental VSEARCH for Metagenomic Classification

Emrecan Ozdogan, Adriana Fasino, Rachel Nguyen, Bahrad Sokhansanj, Gail Rosen, Robi Polikar

    Research output: Chapter in Book/Report/Conference proceedingConference contribution

    Abstract

    DNA Sequencing of microbial communities from en-vironmental samples generates large volumes of data, which can be analyzed using various bioinformatics pipelines. Unsupervised clustering algorithms are usually an early and critical step in an analysis pipeline, since much of such data are unlabeled, unstructured, or novel. However, curated reference databases that provide taxonomic label information are also increasing and growing, which can help in the classification of sequences, and not just clustering. In this contribution, we report on our progress in developing a semi-supervised approach for genomic clustering algorithms, such as U/VSEARCH. The primary contribution of this approach is the ability to recognize previously seen or unseen novel sequences using an incremental approach: for sequences whose examples were previously seen by the algorithm, the algorithm can predict a correct label. For previously unseen novel sequences, the algorithm assigns a temporary label and then updates that label with a permanent one if/when such a label is established in a future reference database. The incremental learning aspect of the proposed approach provides the additional benefit and capability to process the data continuously as new datasets become available. This functionality is notable as most sequence data processing platforms are static in nature, designed to run on a single batch of data, whose only other remedy to process additional data is to combine the new and old data and rerun the entire analysis. We report our promising preliminary results on an extended 16S rRNA database.

    Original languageEnglish (US)
    Title of host publicationProceedings of the 2022 IEEE Symposium Series on Computational Intelligence, SSCI 2022
    EditorsHisao Ishibuchi, Chee-Keong Kwoh, Ah-Hwee Tan, Dipti Srinivasan, Chunyan Miao, Anupam Trivedi, Keeley Crockett
    PublisherInstitute of Electrical and Electronics Engineers Inc.
    Pages1119-1126
    Number of pages8
    ISBN (Electronic)9781665487689
    DOIs
    StatePublished - 2022
    Event2022 IEEE Symposium Series on Computational Intelligence, SSCI 2022 - Singapore, Singapore
    Duration: Dec 4 2022Dec 7 2022

    Publication series

    NameProceedings of the 2022 IEEE Symposium Series on Computational Intelligence, SSCI 2022

    Conference

    Conference2022 IEEE Symposium Series on Computational Intelligence, SSCI 2022
    Country/TerritorySingapore
    CitySingapore
    Period12/4/2212/7/22

    All Science Journal Classification (ASJC) codes

    • Artificial Intelligence
    • Computer Science Applications
    • Decision Sciences (miscellaneous)
    • Computational Mathematics
    • Control and Optimization
    • Transportation

    Fingerprint

    Dive into the research topics of 'Semi-supervised and Incremental VSEARCH for Metagenomic Classification'. Together they form a unique fingerprint.

    Cite this