YOLOX-Pylon-L
A YOLOX-L detector extended with one additional class, traffic cones (traffic_cone), on top of
the 80 COCO classes, for 81 classes total. We train the adaptation so the original COCO capabilities
are kept, not traded away. In practice this checkpoint retains 98% of the official YOLOX-L
baseline while the added cone class scores higher than 79 of the 80 original classes.
Cones are our public demo class. The same adaptation recipe adds arbitrary custom classes such as defects, parts, or PPE to a proven detector without losing what it already knows.
Part of the YOLOX-Pylon family, S · M · L · XL. This is the strongest checkpoint that still runs comfortably in real time.
Built by Empirisch Tech GmbH (Vienna, Austria) under our Chaperone AI brand. See About Empirisch Tech below.
Results
We evaluate on COCO val2017 plus a held-out traffic-cone split, 81 classes in a single pass,
640×640 input, IoU 0.50:0.95 unless noted.
| Metric | Value |
|---|---|
| mAP 50:95 (81 classes) | 48.9 |
| mAP 50:95, original 80 COCO classes only | 48.5 |
| AP50 / AP75 | 66.0 / 52.7 |
| AP small / medium / large | 31.1 / 53.2 / 63.5 |
| AR@100 | 60.8 |
| traffic_cone AP / AR | 78.6 / 81.4 |
| Inference (forward + NMS, batch 1, A100) | 3.71 ms |
Two things stand out.
- Retention held at 98%. The official YOLOX-L
val2017baseline is 49.7 mAP on COCO. After adding the cone class, this checkpoint keeps 48.5 on the same 80 classes, so we traded 1.2 points for an entire new class arriving at top-tier accuracy. - The added class is the best class. At 78.6 AP,
traffic_coneis second only tobear(79.7) on this checkpoint, ahead ofgiraffe(76.4) andstop sign(75.8), and far ahead of workhorse classes likeperson(61.6) andcar(58.5).
Full per-class AP (81 classes)
| class | AP | class | AP | class | AP |
|---|---|---|---|---|---|
| person | 61.618 | bicycle | 37.547 | car | 58.494 |
| motorcycle | 51.265 | airplane | 73.783 | bus | 74.632 |
| train | 71.065 | truck | 53.815 | boat | 30.934 |
| traffic light | 43.080 | fire hydrant | 75.738 | stop sign | 75.843 |
| parking meter | 51.456 | bench | 37.430 | bird | 41.997 |
| cat | 74.731 | dog | 67.190 | horse | 67.541 |
| sheep | 59.677 | cow | 63.133 | elephant | 71.051 |
| bear | 79.689 | zebra | 74.188 | giraffe | 76.406 |
| backpack | 20.183 | umbrella | 47.879 | handbag | 20.090 |
| tie | 43.556 | suitcase | 48.253 | frisbee | 69.678 |
| skis | 33.713 | snowboard | 41.683 | sports ball | 50.590 |
| kite | 51.097 | baseball bat | 39.119 | baseball glove | 41.565 |
| skateboard | 62.194 | surfboard | 45.487 | tennis racket | 59.014 |
| bottle | 42.903 | wine glass | 41.440 | cup | 48.639 |
| fork | 49.134 | knife | 26.992 | spoon | 26.324 |
| bowl | 46.760 | banana | 28.472 | apple | 19.796 |
| sandwich | 35.887 | orange | 31.618 | broccoli | 24.728 |
| carrot | 27.921 | hot dog | 43.810 | pizza | 59.750 |
| donut | 50.739 | cake | 43.368 | chair | 38.982 |
| couch | 49.349 | potted plant | 30.830 | bed | 49.074 |
| dining table | 30.917 | toilet | 64.619 | tv | 61.952 |
| laptop | 68.424 | mouse | 64.775 | remote | 43.296 |
| keyboard | 56.172 | cell phone | 44.216 | microwave | 69.574 |
| oven | 43.304 | toaster | 28.361 | sink | 43.081 |
| refrigerator | 62.442 | book | 16.598 | clock | 52.497 |
| vase | 41.479 | scissors | 34.524 | teddy bear | 52.537 |
| hair drier | 10.184 | toothbrush | 29.280 | traffic_cone | 78.623 |
Comparison with the base model
The comparison that matters is against the checkpoint we adapted from, with the same architecture,
the same parameter count, the same FLOPs, and one extra class. Baseline figures are the official
COCO val2017 numbers from the YOLOX model table.
| Model | COCO mAP 50:95 | Params | FLOPs | Custom classes | License |
|---|---|---|---|---|---|
| yolox_pylon_l (this model) | 48.5 kept + traffic_cone 78.6 |
54.2M | 155.6G | cone added, COCO kept | Apache-2.0 |
| YOLOX-L (base) | 49.7 | 54.2M | 155.6G | COCO only | Apache-2.0 |
Reading that table, the adaptation costs us 1.2 mAP on the original 80 classes and buys an entire new class at 78.6 AP. Nothing else about the model changes. Parameters, FLOPs, and inference cost are the same as stock YOLOX-L, and Apache-2.0 carries over from the base, so the weights can be deployed commercially with no per-deployment license and no obligation to open-source derivative work.
Siblings for scale, same recipe and same eval protocol.
| Family member | mAP (81 cls) | COCO kept | Cone AP | Inference |
|---|---|---|---|---|
| yolox_pylon_s | 42.0 | 41.6 | 74.5 | 1.7 ms |
| yolox_pylon_m | 47.2 | 46.8 | 77.5 | 2.6 ms |
| yolox_pylon_l | 48.9 | 48.5 | 78.6 | 3.7 ms |
| yolox_pylon_xl | 50.4 | 50.0 | 78.8 | 6.0 ms |
We measure inference as forward plus NMS on an A100. Those times are not comparable to the V100 figures published in the official YOLOX table.
Usage
The checkpoint loads with the official YOLOX
codebase. The only change from stock YOLOX-L is num_classes = 81, with traffic_cone as class
index 80.
import torch
from yolox.exp import get_exp
from yolox.utils import postprocess
# stock yolox-l exp, patched to 81 classes
exp = get_exp(exp_name="yolox-l")
exp.num_classes = 81
model = exp.get_model()
ckpt = torch.load("yolox_pylon_l.pth", map_location="cpu")
model.load_state_dict(ckpt["model"])
model.eval().cuda()
# img is a float32 tensor [1, 3, 640, 640], preprocessed YOLOX-style
with torch.no_grad():
outputs = model(img)
outputs = postprocess(outputs, num_classes=81, conf_thre=0.25, nms_thre=0.45)
COCO_CLASSES = [...] # standard 80-class list
CLASSES = COCO_CLASSES + ["traffic_cone"] # index 80
Or with the repo's demo tool.
git clone https://github.com/Megvii-BaseDetection/YOLOX && cd YOLOX
python tools/demo.py image \
-f exps/default/yolox_l.py \
-c yolox_pylon_l.pth \
--path your_image.jpg --conf 0.25 --nms 0.45 --tsize 640 --device gpu
# patch exps/default/yolox_l.py with self.num_classes = 81 first
Training
- Base. YOLOX-L (54.2M params), initialized from COCO-pretrained weights
- Data. 147k images across 81 classes, COCO
train2017plus roughly 30k traffic-cone images, trained jointly so the original 80 classes stay in the mix during adaptation - Eval. COCO
val2017plus a held-out cone split, single 81-class evaluation pass - Input. 640×640
Intended use and limitations
We built this for roadside and infrastructure perception where traffic cones matter, such as work zones, lane closures, and autonomous driving research, and as a template for class-extension on YOLOX. The L size is the accuracy pick that still runs comfortably in real time, and it fits server- side inference over multiple camera streams as well as offline analysis and auto-labeling. It also holds the family best on small objects at 31.1 AP, ahead of the larger XL checkpoint.
The model detects boxes for 81 classes. It does not segment, track, or estimate distance. For fixed single-camera deployments, yolox_pylon_m offers most of the accuracy at 2.6 ms. For embedded and edge boards, yolox_pylon_s runs in 1.7 ms. As with any detector, we recommend validating on the target cameras before production use.
About Empirisch Tech
We are Empirisch Tech GmbH, a Vienna-based AI company, and we publish YOLOX-Pylon under our Chaperone AI brand. We run one recipe across three domains. We adapt a proven foundation model to a specific domain, keep what the base already knows, and ship the checkpoint together with the data it was trained on.
- Language. Thinking-LQ-1.0 (84% MedQA, within 4 points of GPT-4o at ~20GB) and Coder-LQ-1.0
- Physics. Chaperone-Flow-1.0 (Poseidon-B extended to new CFD regimes, 1.8% wake error) and Palace-LoRA (electromagnetics solver configs)
- Vision. The YOLOX-Pylon family and a road-scene anomaly segmentation pipeline
These models power our production platforms, including NumericalAI (GPU physics simulation) and Simvera (industrial perception trained in simulation, deployed on real cameras). We self-host everything in our own Vienna datacenter and use no third- party model APIs. We are a member of the NVIDIA Inception and Microsoft for Startups programs, and our open checkpoints have passed 30,000 downloads on Hugging Face.
Custom builds. The cone class took one adaptation run. For other classes, cameras, or datasets, reach out via chaperoneai.com/contact.
License
Apache-2.0, matching the YOLOX base.
Citation
@article{yolox2021,
title={YOLOX: Exceeding YOLO Series in 2021},
author={Ge, Zheng and Liu, Songtao and Wang, Feng and Li, Zeming and Sun, Jian},
journal={arXiv preprint arXiv:2107.08430},
year={2021}
}