DistilBERT fine-tuned on MNLI
This checkpoint was fully fine-tuned (backbone + head) on the full MNLI training set as part of the A3 SVD compression project.
Results
| Metric | Value |
|---|---|
| Accuracy | 0.8211 |
| F1 (macro) | 0.8202 |
Training Details
- Base model:
distilbert-base-uncased - Epochs: 3
- Learning rate: 4e-05
- Batch size: 128
- Train samples: 392,702
- Val samples: 9,815
Usage
from transformers import AutoModelForSequenceClassification, AutoTokenizer
model = AutoModelForSequenceClassification.from_pretrained("TahaaaaM/distilbert-mnli-a3")
tokenizer = AutoTokenizer.from_pretrained("TahaaaaM/distilbert-mnli-a3")
Citation
Part of the A3 SVD transformer compression research framework.
- Downloads last month
- 11