Identifying Important Pairwise Logratios in Compositional Data with Sparse Principal Component Analysis

Viktorie Nesrstová*, Ines Wilms, Karel Hron, Peter Filzmoser

*Corresponding author for this work

Research output: Contribution to journalArticleAcademicpeer-review

Abstract

Compositional data are characterized by the fact that their elemental information is contained in simple pairwise logratios of the parts that constitute the composition. While pairwise logratios are typically easy to interpret, the number of possible pairs to consider quickly becomes too large even for medium-sized compositions, which may hinder interpretability in further multivariate analysis. Sparse methods can therefore be useful for identifying a few important pairwise logratios (and parts contained in them) from the total candidate set. To this end, we propose a procedure based on the construction of all possible pairwise logratios and employ sparse principal component analysis to identify important pairwise logratios. The performance of the procedure is demonstrated with both simulated and real-world data. In our empirical analysis, we propose three visual tools showing (i) the balance between sparsity and explained variability, (ii) the stability of the pairwise logratios, and (iii) the importance of the original compositional parts to aid practitioners in their model interpretation.
Original languageEnglish
JournalMathematical Geosciences
DOIs
Publication statusE-pub ahead of print - 10 Oct 2024

Keywords

  • compositional data
  • pairwise logratios
  • Sparse PCA
  • geochemical data

Fingerprint

Dive into the research topics of 'Identifying Important Pairwise Logratios in Compositional Data with Sparse Principal Component Analysis'. Together they form a unique fingerprint.

Cite this