(Un)likelihood Training for Interpretable Embedding

Wu, Jiaxin; Ngo, Chong-Wah; Chan, Wing-Kwong; Hou, Zhijian

Computer Science > Computer Vision and Pattern Recognition

arXiv:2207.00282 (cs)

[Submitted on 1 Jul 2022 (v1), last revised 10 Nov 2023 (this version, v3)]

Title:(Un)likelihood Training for Interpretable Embedding

Authors:Jiaxin Wu, Chong-Wah Ngo, Wing-Kwong Chan, Zhijian Hou

View PDF

Abstract:Cross-modal representation learning has become a new normal for bridging the semantic gap between text and visual data. Learning modality agnostic representations in a continuous latent space, however, is often treated as a black-box data-driven training process. It is well-known that the effectiveness of representation learning depends heavily on the quality and scale of training data. For video representation learning, having a complete set of labels that annotate the full spectrum of video content for training is highly difficult if not impossible. These issues, black-box training and dataset bias, make representation learning practically challenging to be deployed for video understanding due to unexplainable and unpredictable results. In this paper, we propose two novel training objectives, likelihood and unlikelihood functions, to unroll semantics behind embeddings while addressing the label sparsity problem in training. The likelihood training aims to interpret semantics of embeddings beyond training labels, while the unlikelihood training leverages prior knowledge for regularization to ensure semantically coherent interpretation. With both training objectives, a new encoder-decoder network, which learns interpretable cross-modal representation, is proposed for ad-hoc video search. Extensive experiments on TRECVid and MSR-VTT datasets show the proposed network outperforms several state-of-the-art retrieval models with a statistically significant performance margin.

Comments:	accepted in ACM Transactions on Information Systems
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR); Multimedia (cs.MM)
Cite as:	arXiv:2207.00282 [cs.CV]
	(or arXiv:2207.00282v3 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2207.00282

Submission history

From: Jiaxin Wu [view email]
[v1] Fri, 1 Jul 2022 09:15:02 UTC (3,006 KB)
[v2] Wed, 17 May 2023 03:07:11 UTC (10,956 KB)
[v3] Fri, 10 Nov 2023 10:18:00 UTC (7,422 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:(Un)likelihood Training for Interpretable Embedding

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:(Un)likelihood Training for Interpretable Embedding

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators