jev-playground-rlcd

A calibrated decision classifier trained with a proper scoring rule on synthetic typed questions, a local, open reimplementation of the "System One" decision-model contract documented for Jev-class models. Part of jev-playground.

Given a text state and typed questions, Noul (yes/no), Choice (closed set), Score (ordered levels), it returns calibrated probability distributions over request-defined label sets. The full distribution is the artifact; confidence is derived from it, never a learned head.

Results

In-domain (14 training domains, held-out validation data): vs the same base model, stock:

Metric noul choice score
accuracy 0.731 โ†’ 0.952 0.510 โ†’ 0.977 0.898 โ†’ 0.992

Aggregate val ECE 0.0090 over 64K pairs (16-domain v2 run on an RTX 3090); loss stayed above the reference entropy floor the entire run (no label overfitting).

Held-out domains (never seen): transfer scaling is real but weak, and uneven:

Split noul choice score
In-domain val (14 seen) 0.952 acc / 0.071 L1 0.977 acc / 0.091 L1 0.992 acc / 0.077 L1
Held-out invoices (v1, 7 domains) 0.704 acc 0.328 acc 0.406 acc
Held-out invoices (v2, 14 domains) 0.607 acc 0.464 acc 0.326 acc
Held-out legal (v2, 14 domains) 0.960 acc 0.637 acc 0.362 acc

Doubling the training domains improved cross-domain choice accuracy (0.328 โ†’ 0.464) but the effect is sublinear with large per-domain variance. Noul (closest to the base NLI prior) transfers best; score transfers worst.

Honest scope statement

This model is calibrated and accurate on its training domains and unreliable outside them. Cross-domain generalization is the open problem: within-domain diversity (paraphrase explosions, co-annotation, varied cardinality) does not transfer across the domain boundary, the encoder fine-tune overwrites the text prior with domain-specific behavior. The architectural fix (zero-shot readout over request-defined label tokens) is tracked in the repo. This scope statement ships with the model on purpose: a lab publishing known weaknesses is more credible than one hiding them.

Usage

The checkpoint is a standard NLI sequence classifier (entailment head). The easiest way to serve it is behind the jev-playground contract:

pip install -e ".[server,deberta]"   # from the jev-playground repo
hf download soyrsoyr/jev-playground-rlcd --include "deberta-large-7domain/*" --local-dir checkpoints

JEV_PLAYGROUND_BACKEND=deberta JEV_PLAYGROUND_DEBERTA_MODEL=checkpoints/deberta-large-7domain jev-playground-serve --port 8000

Training your own: see the repo, the data engine (training/data.py) emits reference distributions that are exact by construction, and the trainer optimizes a proper scoring rule (log-loss / Brier) against them.

Training details

  • Data: soyrsoyr/jev-playground-rlcd-v0, 16 domains (support tickets, reviews, security incidents, agent traces, invoices, code reviews, insurance claims, shipment alerts, meeting emails, e-commerce orders, transaction fraud, content moderation, expense reports, bug triage, legal contracts, clinic scheduling), 580K training pairs, typed questions with reference distributions exact by construction (crisp/weak/contradictory uncertainty slices at 70/20/10).
  • Objective: soft-target cross-entropy (log score) over the entailment head; Brier available (--loss brier).
  • Hyperparameters: differential LR (heads 2e-5 / encoder 1e-5), linear warmup + decay, batch 16, max-length 256, fp32, 1 epoch.
  • Hardware: RTX 3090 (CUDA). The deberta-large-v2-16domain/ dir holds this run; deberta-large-7domain/ the 7-domain run; distilroberta/ an 82M CPU run.

License

MIT

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for soyrsoyr/jev-playground-rlcd

Finetuned
(7)
this model