DistilBERT fine-tuned on MNLI

This checkpoint was fully fine-tuned (backbone + head) on the full MNLI training set as part of the A3 SVD compression project.

Results

Metric Value
Accuracy 0.8211
F1 (macro) 0.8202

Training Details

  • Base model: distilbert-base-uncased
  • Epochs: 3
  • Learning rate: 4e-05
  • Batch size: 128
  • Train samples: 392,702
  • Val samples: 9,815

Usage

from transformers import AutoModelForSequenceClassification, AutoTokenizer

model = AutoModelForSequenceClassification.from_pretrained("TahaaaaM/distilbert-mnli-a3")
tokenizer = AutoTokenizer.from_pretrained("TahaaaaM/distilbert-mnli-a3")

Citation

Part of the A3 SVD transformer compression research framework.

Downloads last month
11
Safetensors
Model size
67M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support