Big data and other challenges in the quest for orthologs

Erik L L Sonnhammer*, Toni Gabaldón, Alan W Sousa da Silva, Maria Martin, Marc Robinson-Rechavi, Brigitte Boeckmann, Paul D Thomas, Christophe Dessimoz, Quest for Orthologs consortium

*Corresponding author for this work

Research output: Contribution to journalArticleAcademicpeer-review

Abstract

UNLABELLED: Given the rapid increase of species with a sequenced genome, the need to identify orthologous genes between them has emerged as a central bioinformatics task. Many different methods exist for orthology detection, which makes it difficult to decide which one to choose for a particular application. Here, we review the latest developments and issues in the orthology field, and summarize the most recent results reported at the third 'Quest for Orthologs' meeting. We focus on community efforts such as the adoption of reference proteomes, standard file formats and benchmarking. Progress in these areas is good, and they are already beneficial to both orthology consumers and providers. However, a major current issue is that the massive increase in complete proteomes poses computational challenges to many of the ortholog database providers, as most orthology inference algorithms scale at least quadratically with the number of proteomes. The Quest for Orthologs consortium is an open community with a number of working groups that join efforts to enhance various aspects of orthology analysis, such as defining standard formats and datasets, documenting community resources and benchmarking.

AVAILABILITY AND IMPLEMENTATION: All such materials are available at http://questfororthologs.org.

Original languageEnglish
Pages (from-to)2993-8
Number of pages6
JournalBioinformatics
Volume30
Issue number21
DOIs
Publication statusPublished - 1 Nov 2014

Keywords

  • Algorithms
  • Genomics
  • Protein Structure, Tertiary
  • Proteome
  • Sequence Analysis, DNA
  • Sequence Analysis, Protein
  • Sequence Homology

Cite this