GenePT: Leveraging LLMs for Gene and Cell Embeddings
This article introduces a model called GenePT, which utilizes large language models (such as GPT-3.5) to generate gene and cell embedding representations for single-cell biology analysis. The authors designed a series of downstream tasks to validate the effectiveness of the GenePT model and compared it with existing single-cell foundation models (such as Geneformer and scGPT). The following is a summary of the downstream tasks, evaluation methods, metrics, and datasets mentioned in the article.
Model Complexity
The specific number of parameters for the scGPT model is not directly mentioned. However, the article mentions some architectural details of the model, such as:
- Pre-trained base model: embedding size of 512, containing 12 stacked transformer blocks, each with 8 attention heads.
- Fully connected layer: hidden layer size of 512.
Downstream Tasks
- Gene Function Category Prediction: Predict which of the 15 most common functional categories a gene belongs to.
- Gene Property Prediction: Including distinguishing dosage-sensitive from insensitive transcription factors, distinguishing different gene methylation states, etc.
- Gene-Gene Interaction Prediction: Predict whether two genes interact.
- Protein-Protein Interaction Prediction: Predict whether there is an interaction between two proteins.
- Cell Type Annotation: Predict cell types based on cell embedding representations.
- Batch Effect Assessment: Evaluate the ability of embedding representations to remove batch effects while retaining biological information.
Evaluation Methods and Metrics
- Accuracy: Number of correct predictions divided by total predictions.
- Precision: Number of correctly predicted positives divided by total predicted positives.
- Recall: Number of correctly predicted positives divided by actual positives.
- F1 Score: Harmonic mean of precision and recall, measuring the balance between accuracy and completeness.
- ROC-AUC: Area under the receiver operating characteristic curve, measuring classifier performance across all possible classification thresholds.
- Adjusted Rand Index (ARI) and Adjusted Mutual Information (AMI): Evaluate consistency between clustering results and true labels.
Author’s Metric Performance
The article mentions that GenePT performs comparably to Geneformer and other models across multiple tasks, and even better on some tasks.
Gene Function Category Prediction
- GenePT: Overall accuracy reaches 96%, with only small misclassifications in class-specific accuracy.
Gene-Gene Interaction Prediction (GGI)
- GenePT: Using the test GGI dataset with shared Gene Ontology (GO) annotations, ROC-AUC is 0.82.
Protein-Protein Interaction Prediction (PPI)
- GenePT: On three different PPI datasets, ROC-AUC values are:
- Literature-derived dataset: ROC-AUC not explicitly given, but mentioned to outperform other models.
- Comprehensive detection dataset: ROC-AUC not explicitly given, but mentioned to outperform other models.
- Tissue-specific protein-protein functional network dataset: ROC-AUC not explicitly given, but mentioned to outperform other models.
Cell Type Annotation
GenePT-w and GenePT-s performance on multiple datasets compared to scGPT and Geneformer:
Aorta Dataset:
- GenePT-w: ARI=0.12, AMI=0.12, ASW=0.01
- GenePT-s: ARI=0.09, AMI=0.12, ASW=-0.04
Artery Dataset:
- GenePT-w: ARI=0.47, AMI=0.64, ASW=0.18
- GenePT-s: ARI=0.36, AMI=0.59, ASW=0.15
Bones Dataset:
- GenePT-w: ARI=0.12, AMI=0.21, ASW=-0.01
- GenePT-s: ARI=0.21, AMI=0.29, ASW=0.02
Myeloid Dataset (cancer types):
- GenePT-w: ARI=0.25, AMI=0.27, ASW=0.02
- GenePT-s: ARI=0.17, AMI=0.17, ASW=0.06
Pancreas Dataset:
- GenePT-w: ARI=0.49, AMI=0.69, ASW=0.15
- GenePT-s: ARI=0.30, AMI=0.50, ASW=0.10
Multiple Sclerosis Dataset (age):
- GenePT-w: ARI=0.07, AMI=0.13, ASW=-0.07
- GenePT-s: ARI=0.06, AMI=0.12, ASW=-0.03
Batch Effect Assessment
- GenePT performs well in removing batch effects while retaining biological information. Compared to scGPT and Geneformer, ARI values decrease on multiple datasets, indicating smaller batch effects.
ARI and AMI are used to measure consistency between clustering results and true labels, while ASW is used to evaluate the cohesion and separation of clustering results. GenePT demonstrates comparable or better performance than existing models on most tasks.
Datasets
- Gene Function Category Prediction: Uses specific gene function category datasets (geneformer).
- Gene Property Prediction: Uses open-source data provided by Theodoris et al.
- Gene-Gene Interaction Prediction: Uses benchmark datasets based on shared Gene Ontology annotations by Du et al.
- Protein-Protein Interaction Prediction: Uses multiple datasets, including HuRI, Lit-BM, and tissue-specific protein-protein functional networks provided by Greene et al.
- Cell Type Annotation: Uses datasets from the circulatory system (Aorta, Artery), bone tissue (Bones, Myeloid), pancreas, and immune cells from healthy individuals and multiple sclerosis patients.
Summary
| Task Type | Metric | GenePT Performance | Dataset Name | Baseline/Comparison Method | Reference Link |
|---|---|---|---|---|---|
| Gene Functionality Class Prediction | Accuracy | 96% | - | Geneformer, etc. | - |
| Gene Property Prediction Tasks | Various | Not specified | - | Gene2vec, scGPT | Theodoris et al. 2023 |
| Gene-Gene Interaction Prediction | ROC-AUC | 0.82 | GEO expression data | Gene2vec/scGPT/Geneformer | Du et al. 2019 |
| Protein-Protein Interaction Prediction | ROC-AUC | Outperforms other models | HuRI, Lit-BM, Tissue-specific PPI networks | Other models | HuRI dataset, Lit-BM dataset, Tissue-specific PPI networks |
| Cell Type Annotation | ARI, AMI, ASW | Varies by dataset | Aorta, Artery, Bones, etc. | scGPT, Geneformer | Chaffin et al. 2022, Li et al. 2020 |
| Batch Effect Assessment | ARI | Reduced, indicating smaller batch effects | Cardiomyocyte dataset, Aorta dataset | Geneformer, scGPT | Chaffin et al. dataset, Li et al. dataset |
Enjoy Reading This Article?
Here are some more articles you might like to read next:
- Google Gemini updates: Flash 1.5, Gemma 2 and Project Astra
- Displaying External Posts on Your al-folio Blog
- GitHub Connection Under Extreme Conditions
- Creating a VM Instance on AutoDL
- Introduction to Deep Learning
- Backpropagation
- Creating a VM Instance on GCP
- scGPT: Model Performance and Downstream Tasks
- LLM and BioSequence: GenePT, scGPT, and Geneformer
- Introduction to Diffusion Models
giscus comments misconfigured
Please follow instructions at http://giscus.app and update your giscus configuration.
Missing required keys:
repo_id, category_id.