BI-MODAL COMPOSITIONAL NETWORK FOR FEATURE DISENTANGLEMENT

Document Type

Conference Article

Publication Title

Proceedings - International Conference on Image Processing, ICIP

Abstract

The purpose of Compositional Zero-Shot Learning (CZSL) is to recognise unseen (or new) compositions of known objects (e.g. apple, banana) and its states (e.g. peeled, ripe). The CZSL poses a challenging problem as the visual features of objects and states are often difficult to isolate from their composition in images. Also, the features of a state may have wide variation across different compositions. For example, the state sliced presents different visual features in compositions, sliced potato and sliced bread. Hence, for solving CZSL problem, disentanglement of state and object features is the key challenge that we attempt to address in this paper by introducing Bi-modal Compositional Network (BiComNet). The proposed BiComNet consists of two distinct branches for state and object recognition, which helps in better disentanglement of the state and object features. We also propose a novel loss component to better facilitate the disentanglement of the state and object features. Finally, the proposed BiComNet integrates two different modalities of input data, the visual and word features, to better recognise unseen state-object compositions. Extensive experiments are carried out on the benchmark datasets MIT-States & UT-Zappos50K. Our proposed algorithm outperforms competitive methods in most count.

First Page

3051

Last Page

3055

DOI

10.1109/ICIP46576.2022.9897457

Publication Date

1-1-2022

This document is currently not available here.

Share

COinS