dcarpintero
/

pangolin-guard-base

Text Classification

Model card Files Files and versions

dcarpintero commited on Mar 30

Commit

d139d7b

·

verified ·

1 Parent(s): 697a491

Update README.md

Files changed (1) hide show

README.md +16 -14

README.md CHANGED Viewed

@@ -12,28 +12,30 @@ model-index:
   results: []
 ---
-<!-- This model card has been generated automatically according to the information the Trainer had access to. You
-should probably proofread and complete it, then remove this comment. -->
-# pangolin-guard-base
-This model is a fine-tuned version of [answerdotai/ModernBERT-base](https://huggingface.co/answerdotai/ModernBERT-base) on an unknown dataset.
-It achieves the following results on the evaluation set:
-- Loss: 0.0196
-- F1: 0.9923
-- Accuracy: 0.9950
-## Model description
-More information needed
-## Intended uses & limitations
-More information needed
-## Training and evaluation data
-More information needed
 ## Training procedure

   results: []
 ---
+# PangolinGuard-Base
+LLM applications face critical security challenges in form of prompt injections and jailbreaks. This can result in models leaking sensitive data or deviating from their intended behavior. Existing safeguard models are not fully open and have limited context windows (e.g., only 512 tokens in LlamaGuard).
+PangolinGuard is a ModernBERT (Base), lightweight model that discriminates malicious prompts.
+🤗 [Tech-Blog](https://huggingface.co/blog/dcarpintero/pangolin-fine-tuning-modern-bert) | [GitHub Repo](https://github.com/dcarpintero/pangolin-guard)
+## Intended uses
+- Adding custom, self-hosted safety checks to AI agents and conversational interfaces
+- Topic and content moderation
+- Mitigating risks when connecting AI pipelines to external services
+## Evaluation data
+Evaluated on unseen data from a subset of specialized benchmarks targeting prompt safety and malicious input detection, while testing over-defense behavior:
+- NotInject: Designed to measure over-defense in prompt guard models by including benign inputs enriched with trigger words common in prompt injection attacks.
+- BIPIA: Evaluates privacy invasion attempts and boundary-pushing queries through indirect prompt injection attacks.
+- Wildguard-Benign: Represents legitimate but potentially ambiguous prompts.
+- PINT: Evaluates particularly nuanced prompt injection, jailbreaks, and benign prompts that could be misidentified as malicious.
+![image/png](https://cdn-uploads.huggingface.co/production/uploads/64a13b68b14ab77f9e3eb061/ygIo-Yo3NN7mDhZlLFvZb.png)
 ## Training procedure