YOLOX-Pylon-L

A YOLOX-L detector extended with one additional class, traffic cones (traffic_cone), on top of the 80 COCO classes, for 81 classes total. We train the adaptation so the original COCO capabilities are kept, not traded away. In practice this checkpoint retains 98% of the official YOLOX-L baseline while the added cone class scores higher than 79 of the 80 original classes.

Cones are our public demo class. The same adaptation recipe adds arbitrary custom classes such as defects, parts, or PPE to a proven detector without losing what it already knows.

Part of the YOLOX-Pylon family, S · M · L · XL. This is the strongest checkpoint that still runs comfortably in real time.

Built by Empirisch Tech GmbH (Vienna, Austria) under our Chaperone AI brand. See About Empirisch Tech below.

Results

We evaluate on COCO val2017 plus a held-out traffic-cone split, 81 classes in a single pass, 640×640 input, IoU 0.50:0.95 unless noted.

Metric Value
mAP 50:95 (81 classes) 48.9
mAP 50:95, original 80 COCO classes only 48.5
AP50 / AP75 66.0 / 52.7
AP small / medium / large 31.1 / 53.2 / 63.5
AR@100 60.8
traffic_cone AP / AR 78.6 / 81.4
Inference (forward + NMS, batch 1, A100) 3.71 ms

Two things stand out.

  • Retention held at 98%. The official YOLOX-L val2017 baseline is 49.7 mAP on COCO. After adding the cone class, this checkpoint keeps 48.5 on the same 80 classes, so we traded 1.2 points for an entire new class arriving at top-tier accuracy.
  • The added class is the best class. At 78.6 AP, traffic_cone is second only to bear (79.7) on this checkpoint, ahead of giraffe (76.4) and stop sign (75.8), and far ahead of workhorse classes like person (61.6) and car (58.5).
Full per-class AP (81 classes)
class AP class AP class AP
person 61.618 bicycle 37.547 car 58.494
motorcycle 51.265 airplane 73.783 bus 74.632
train 71.065 truck 53.815 boat 30.934
traffic light 43.080 fire hydrant 75.738 stop sign 75.843
parking meter 51.456 bench 37.430 bird 41.997
cat 74.731 dog 67.190 horse 67.541
sheep 59.677 cow 63.133 elephant 71.051
bear 79.689 zebra 74.188 giraffe 76.406
backpack 20.183 umbrella 47.879 handbag 20.090
tie 43.556 suitcase 48.253 frisbee 69.678
skis 33.713 snowboard 41.683 sports ball 50.590
kite 51.097 baseball bat 39.119 baseball glove 41.565
skateboard 62.194 surfboard 45.487 tennis racket 59.014
bottle 42.903 wine glass 41.440 cup 48.639
fork 49.134 knife 26.992 spoon 26.324
bowl 46.760 banana 28.472 apple 19.796
sandwich 35.887 orange 31.618 broccoli 24.728
carrot 27.921 hot dog 43.810 pizza 59.750
donut 50.739 cake 43.368 chair 38.982
couch 49.349 potted plant 30.830 bed 49.074
dining table 30.917 toilet 64.619 tv 61.952
laptop 68.424 mouse 64.775 remote 43.296
keyboard 56.172 cell phone 44.216 microwave 69.574
oven 43.304 toaster 28.361 sink 43.081
refrigerator 62.442 book 16.598 clock 52.497
vase 41.479 scissors 34.524 teddy bear 52.537
hair drier 10.184 toothbrush 29.280 traffic_cone 78.623

Comparison with the base model

The comparison that matters is against the checkpoint we adapted from, with the same architecture, the same parameter count, the same FLOPs, and one extra class. Baseline figures are the official COCO val2017 numbers from the YOLOX model table.

Model COCO mAP 50:95 Params FLOPs Custom classes License
yolox_pylon_l (this model) 48.5 kept + traffic_cone 78.6 54.2M 155.6G cone added, COCO kept Apache-2.0
YOLOX-L (base) 49.7 54.2M 155.6G COCO only Apache-2.0

Reading that table, the adaptation costs us 1.2 mAP on the original 80 classes and buys an entire new class at 78.6 AP. Nothing else about the model changes. Parameters, FLOPs, and inference cost are the same as stock YOLOX-L, and Apache-2.0 carries over from the base, so the weights can be deployed commercially with no per-deployment license and no obligation to open-source derivative work.

Siblings for scale, same recipe and same eval protocol.

Family member mAP (81 cls) COCO kept Cone AP Inference
yolox_pylon_s 42.0 41.6 74.5 1.7 ms
yolox_pylon_m 47.2 46.8 77.5 2.6 ms
yolox_pylon_l 48.9 48.5 78.6 3.7 ms
yolox_pylon_xl 50.4 50.0 78.8 6.0 ms

We measure inference as forward plus NMS on an A100. Those times are not comparable to the V100 figures published in the official YOLOX table.

Usage

The checkpoint loads with the official YOLOX codebase. The only change from stock YOLOX-L is num_classes = 81, with traffic_cone as class index 80.

import torch
from yolox.exp import get_exp
from yolox.utils import postprocess

# stock yolox-l exp, patched to 81 classes
exp = get_exp(exp_name="yolox-l")
exp.num_classes = 81

model = exp.get_model()
ckpt = torch.load("yolox_pylon_l.pth", map_location="cpu")
model.load_state_dict(ckpt["model"])
model.eval().cuda()

# img is a float32 tensor [1, 3, 640, 640], preprocessed YOLOX-style
with torch.no_grad():
    outputs = model(img)
outputs = postprocess(outputs, num_classes=81, conf_thre=0.25, nms_thre=0.45)

COCO_CLASSES = [...]                        # standard 80-class list
CLASSES = COCO_CLASSES + ["traffic_cone"]   # index 80

Or with the repo's demo tool.

git clone https://github.com/Megvii-BaseDetection/YOLOX && cd YOLOX
python tools/demo.py image \
    -f exps/default/yolox_l.py \
    -c yolox_pylon_l.pth \
    --path your_image.jpg --conf 0.25 --nms 0.45 --tsize 640 --device gpu
# patch exps/default/yolox_l.py with self.num_classes = 81 first

Training

  • Base. YOLOX-L (54.2M params), initialized from COCO-pretrained weights
  • Data. 147k images across 81 classes, COCO train2017 plus roughly 30k traffic-cone images, trained jointly so the original 80 classes stay in the mix during adaptation
  • Eval. COCO val2017 plus a held-out cone split, single 81-class evaluation pass
  • Input. 640×640

Intended use and limitations

We built this for roadside and infrastructure perception where traffic cones matter, such as work zones, lane closures, and autonomous driving research, and as a template for class-extension on YOLOX. The L size is the accuracy pick that still runs comfortably in real time, and it fits server- side inference over multiple camera streams as well as offline analysis and auto-labeling. It also holds the family best on small objects at 31.1 AP, ahead of the larger XL checkpoint.

The model detects boxes for 81 classes. It does not segment, track, or estimate distance. For fixed single-camera deployments, yolox_pylon_m offers most of the accuracy at 2.6 ms. For embedded and edge boards, yolox_pylon_s runs in 1.7 ms. As with any detector, we recommend validating on the target cameras before production use.

About Empirisch Tech

We are Empirisch Tech GmbH, a Vienna-based AI company, and we publish YOLOX-Pylon under our Chaperone AI brand. We run one recipe across three domains. We adapt a proven foundation model to a specific domain, keep what the base already knows, and ship the checkpoint together with the data it was trained on.

  • Language. Thinking-LQ-1.0 (84% MedQA, within 4 points of GPT-4o at ~20GB) and Coder-LQ-1.0
  • Physics. Chaperone-Flow-1.0 (Poseidon-B extended to new CFD regimes, 1.8% wake error) and Palace-LoRA (electromagnetics solver configs)
  • Vision. The YOLOX-Pylon family and a road-scene anomaly segmentation pipeline

These models power our production platforms, including NumericalAI (GPU physics simulation) and Simvera (industrial perception trained in simulation, deployed on real cameras). We self-host everything in our own Vienna datacenter and use no third- party model APIs. We are a member of the NVIDIA Inception and Microsoft for Startups programs, and our open checkpoints have passed 30,000 downloads on Hugging Face.

Custom builds. The cone class took one adaptation run. For other classes, cameras, or datasets, reach out via chaperoneai.com/contact.

License

Apache-2.0, matching the YOLOX base.

Citation

@article{yolox2021,
  title={YOLOX: Exceeding YOLO Series in 2021},
  author={Ge, Zheng and Liu, Songtao and Wang, Feng and Li, Zeming and Sun, Jian},
  journal={arXiv preprint arXiv:2107.08430},
  year={2021}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train empirischtech/yolox-pylon-l

Paper for empirischtech/yolox-pylon-l