Unsupervised Interpretable Basis Extraction for Concept – Based Visual Explanations

Alexandros Doumanoglou; Stylianos Asteriadis; Dimitrios Zarpalas

doi:10.1109/TAI.2023.3338169

Unsupervised Interpretable Basis Extraction for Concept – Based Visual Explanations

Alexandros Doumanoglou, Stylianos Asteriadis, Dimitrios Zarpalas

Research output: Contribution to journal › Article › Academic › peer-review

Abstract

An important line of research attempts to explain CNN image classifier predictions and intermediate layer representations in terms of human understandable concepts. In this work, we expand on previous works in the literature that use annotated concept datasets to extract interpretable feature space directions and propose an unsupervised post-hoc method to extract a disentangling interpretable basis by looking for the rotation of the feature space that explains sparse one-hot thresholded transformed representations of pixel activations. We do experimentation with existing popular CNNs and demonstrate the effectiveness of our method in extracting an interpretable basis across network architectures and training datasets. We make extensions to the existing basis interpretability metrics found in the literature and show that, intermediate layer representations become more interpretable when transformed to the bases extracted with our method. Finally, using the basis interpretability metrics, we compare the bases extracted with our method with the bases derived with a supervised approach and find that, in one aspect, the proposed unsupervised approach has a strength that constitutes a limitation of the supervised one and give potential directions for future research.

Original language	English
Journal	IEEE Transactions on Artificial Intelligence
Issue number	4
DOIs	https://doi.org/10.1109/TAI.2023.3338169
Publication status	E-pub ahead of print - 1 Jan 2023

Keywords

Annotations
Artificial intelligence
Detectors
Explainable Artificial Intelligence (XAI)
Feature extraction
Interpretable Artificial Intelligence (IAI)
Interpretable Basis
Measurement
Semantics
Training
Unsupervised Learning

Access to Document

10.1109/TAI.2023.3338169Licence: CC BY

Cite this

@article{761c66b723b840ccbd9ca53522110280,

title = "Unsupervised Interpretable Basis Extraction for Concept – Based Visual Explanations",

abstract = "An important line of research attempts to explain CNN image classifier predictions and intermediate layer representations in terms of human understandable concepts. In this work, we expand on previous works in the literature that use annotated concept datasets to extract interpretable feature space directions and propose an unsupervised post-hoc method to extract a disentangling interpretable basis by looking for the rotation of the feature space that explains sparse one-hot thresholded transformed representations of pixel activations. We do experimentation with existing popular CNNs and demonstrate the effectiveness of our method in extracting an interpretable basis across network architectures and training datasets. We make extensions to the existing basis interpretability metrics found in the literature and show that, intermediate layer representations become more interpretable when transformed to the bases extracted with our method. Finally, using the basis interpretability metrics, we compare the bases extracted with our method with the bases derived with a supervised approach and find that, in one aspect, the proposed unsupervised approach has a strength that constitutes a limitation of the supervised one and give potential directions for future research.",

keywords = "Annotations, Artificial intelligence, Detectors, Explainable Artificial Intelligence (XAI), Feature extraction, Interpretable Artificial Intelligence (IAI), Interpretable Basis, Measurement, Semantics, Training, Unsupervised Learning",

author = "Alexandros Doumanoglou and Stylianos Asteriadis and Dimitrios Zarpalas",

note = "Publisher Copyright: Authors",

year = "2023",

month = jan,

day = "1",

doi = "10.1109/TAI.2023.3338169",

language = "English",

journal = "IEEE Transactions on Artificial Intelligence",

issn = "2691-4581",

publisher = "IEEE",

number = "4",

}

TY - JOUR

T1 - Unsupervised Interpretable Basis Extraction for Concept – Based Visual Explanations

AU - Doumanoglou, Alexandros

AU - Asteriadis, Stylianos

AU - Zarpalas, Dimitrios

N1 - Publisher Copyright: Authors

PY - 2023/1/1

Y1 - 2023/1/1

N2 - An important line of research attempts to explain CNN image classifier predictions and intermediate layer representations in terms of human understandable concepts. In this work, we expand on previous works in the literature that use annotated concept datasets to extract interpretable feature space directions and propose an unsupervised post-hoc method to extract a disentangling interpretable basis by looking for the rotation of the feature space that explains sparse one-hot thresholded transformed representations of pixel activations. We do experimentation with existing popular CNNs and demonstrate the effectiveness of our method in extracting an interpretable basis across network architectures and training datasets. We make extensions to the existing basis interpretability metrics found in the literature and show that, intermediate layer representations become more interpretable when transformed to the bases extracted with our method. Finally, using the basis interpretability metrics, we compare the bases extracted with our method with the bases derived with a supervised approach and find that, in one aspect, the proposed unsupervised approach has a strength that constitutes a limitation of the supervised one and give potential directions for future research.

AB - An important line of research attempts to explain CNN image classifier predictions and intermediate layer representations in terms of human understandable concepts. In this work, we expand on previous works in the literature that use annotated concept datasets to extract interpretable feature space directions and propose an unsupervised post-hoc method to extract a disentangling interpretable basis by looking for the rotation of the feature space that explains sparse one-hot thresholded transformed representations of pixel activations. We do experimentation with existing popular CNNs and demonstrate the effectiveness of our method in extracting an interpretable basis across network architectures and training datasets. We make extensions to the existing basis interpretability metrics found in the literature and show that, intermediate layer representations become more interpretable when transformed to the bases extracted with our method. Finally, using the basis interpretability metrics, we compare the bases extracted with our method with the bases derived with a supervised approach and find that, in one aspect, the proposed unsupervised approach has a strength that constitutes a limitation of the supervised one and give potential directions for future research.

KW - Annotations

KW - Artificial intelligence

KW - Detectors

KW - Explainable Artificial Intelligence (XAI)

KW - Feature extraction

KW - Interpretable Artificial Intelligence (IAI)

KW - Interpretable Basis

KW - Measurement

KW - Semantics

KW - Training

KW - Unsupervised Learning

U2 - 10.1109/TAI.2023.3338169

DO - 10.1109/TAI.2023.3338169

M3 - Article

SN - 2691-4581

JO - IEEE Transactions on Artificial Intelligence

JF - IEEE Transactions on Artificial Intelligence

IS - 4

ER -