GIMLET: A Unified Graph-Text Model for Instruction-Based Molecule Zero-Shot Learning

Zhao, Haiteng; Liu, Shengchao; Ma, Chang; Xu, Hannan; Fu, Jie; Deng, Zhi-Hong; Kong, Lingpeng; Liu, Qi

Computer Science > Machine Learning

arXiv:2306.13089 (cs)

[Submitted on 28 May 2023 (v1), last revised 22 Oct 2023 (this version, v3)]

Title:GIMLET: A Unified Graph-Text Model for Instruction-Based Molecule Zero-Shot Learning

Authors:Haiteng Zhao, Shengchao Liu, Chang Ma, Hannan Xu, Jie Fu, Zhi-Hong Deng, Lingpeng Kong, Qi Liu

View PDF

Abstract:Molecule property prediction has gained significant attention in recent years. The main bottleneck is the label insufficiency caused by expensive lab experiments. In order to alleviate this issue and to better leverage textual knowledge for tasks, this study investigates the feasibility of employing natural language instructions to accomplish molecule-related tasks in a zero-shot setting. We discover that existing molecule-text models perform poorly in this setting due to inadequate treatment of instructions and limited capacity for graphs. To overcome these issues, we propose GIMLET, which unifies language models for both graph and text data. By adopting generalized position embedding, our model is extended to encode both graph structures and instruction text without additional graph encoding modules. GIMLET also decouples encoding of the graph from tasks instructions in the attention mechanism, enhancing the generalization of graph features across novel tasks. We construct a dataset consisting of more than two thousand molecule tasks with corresponding instructions derived from task descriptions. We pretrain GIMLET on the molecule tasks along with instructions, enabling the model to transfer effectively to a broad range of tasks. Experimental results demonstrate that GIMLET significantly outperforms molecule-text baselines in instruction-based zero-shot learning, even achieving closed results to supervised GNN models on tasks such as toxcast and muv.

Subjects:	Machine Learning (cs.LG); Computation and Language (cs.CL); Biomolecules (q-bio.BM)
Cite as:	arXiv:2306.13089 [cs.LG]
	(or arXiv:2306.13089v3 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2306.13089

Submission history

From: Haiteng Zhao [view email]
[v1] Sun, 28 May 2023 18:27:59 UTC (20,922 KB)
[v2] Fri, 23 Jun 2023 06:26:11 UTC (13,613 KB)
[v3] Sun, 22 Oct 2023 18:13:40 UTC (13,614 KB)

Computer Science > Machine Learning

Title:GIMLET: A Unified Graph-Text Model for Instruction-Based Molecule Zero-Shot Learning

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:GIMLET: A Unified Graph-Text Model for Instruction-Based Molecule Zero-Shot Learning

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators