OpenGate

From perception to post-training, a deployment-audit gate
Given a small deployment vocabulary V, OpenGate audits any detector's episode-level behavior: EpisodicAP (in-vocab accuracy) and OOV-FP (out-of-vocab hallucination) → PASS/FAIL — and reads WorldEngine's own nuPlan eval output through the same gate.

01Real GPU numbers — official 3D stack (nuScenes mini-val)

DETR3D mAP

0.3315
official weights · mmdet3d 1.x · mini-val

DETR3D NDS

0.3956
nuScenes detection score

FCOS3D mAP

0.294
official weights · mini-val

FCOS3D NDS

0.322
nuScenes detection score

Hardware

RTX 5070
CUDA 12.8 · torch 2.11 · mmcv 2.1.0 (Windows build)
Both are real runs of MMDetection3D official weights on the official nuScenes mini-val protocol — the same AlgEngine stack WorldEngine builds on. Recipes: DETR3D_GPU_RESULT · FCOS3D_GPU_RESULT.

02Real driving-scene gate (nuScenes mini, 6 cameras)

EpisodicAP

17.54
full-mini 2,424 views · YOLO-World · 6-camera pool

OOV-FP

0.240
driving scenes are naturally OOV-dense

Gate

PASS
EpisodicAP ≥ 10 · OOV-FP ≤ 0.45

Key insight

Views change conclusions
front-only overestimates; 60-view baseline EpisodicAP 13.47
Full-mini run: 404 keyframes × 6 cams = 2,424 real views, YOLO-World 10-class deploy vocab; 3D GT projected to 2D via the repo's view-index builder. 60-view sampled baseline (EpisodicAP 13.47 / OOV-FP 0.351) documented in DRIVING_SCENE_GATE.

03WorldEngine bridge — post-training eval → the same gate

Open-loop PDMS

88.95 → 59.83
common → rare (official table)

Closed-loop SR

88.89
rare scenarios (official table)

Production

−45.5% collision
200 km zero intervention

Bridge gate

PASS
synthetic eval: AP 85.7 / OOV 0.143; rare 60/0.4
worldengine_reader projects scenario-level PDMS/SR/collision/offroad onto a deployment constraint vocabulary, so EpisodicAP becomes constraint-satisfaction rate and OOV-FP becomes violation rate. See WORLD_ENGINE_BRIDGE.

04Open-vocab detection (COCO, GPU)

RunAdapterVocabEpisodicAPOOV-FPGate
yolo_deployYOLO-World1248.520.051PASS
owl_deployOWL-ViT-v21252.040.103PASS
yolo_wideYOLO-World4056.790.102PASS
n=120 · seed=0 · COCO val2017 subset · same gate, real weights. OWL pays for higher AP with higher OOV-FP. Details in HARD_CURRENCY_RESULTS.

05The paper stack, mounted

StagePapersMounted as
感知Occ3D · DETR3D · FCOS3Dcurated + real GPU numbers
地图HDMapNet · VectorMapNet · NMPnarrative + map-vocab analogy
预测VectorNet · TNT · ViP3Dcurated tables (ViP3D minADE 2.03) + analogy
规划UniAD · VADcurated official tables (avg L2 1.03→0.72 m)
VLMDriveVLMscene-vocab smoke + Dual interface
世界模型·后训练Vista · WorldEnginecurated tables + nuPlan bridge
Machine-readable stack: fixtures/vcad/overview.json.

06Research evidence — closed-loop safety frontier

Safety frontier: perception vocabulary coverage → closed-loop collision / intervention / failure rates with Wilson 95% CI and cliff shading
Safety frontier (nuPlan IDM, real WSL2 sweep, 25 scenarios/tier): the artificial-coverage control cliffs at the first degradation step (0.08 → 0.36). Injecting a real detector's per-class miss profile (miss-only arm, N=75) gives a saturated 0.56 collision plateau — misses are a standing risk, not a coverage-cliff. Wilson 95% CI error bars; amber shading = bootstrap cliff CI. Full analysis in CLOSED_LOOP_SAFETY_FRONTIER.
0.9.1: dual-channel attribution (miss-only / OOV-phantom-only / combined / confusion-only, N=75/tier/arm) shows the miss channel dominates (0.56 at full coverage, flat across tiers); the measured deployment axis (22 synthetic + 10 real YOLO-World closed-loop anchors) locates the gate at n_gt-weighted miss rate ≤ 0.149. H2 grid calibration (conf × profile source, 24 cells) is 24/24 miss-dominant, 48/48 consistent. Full write-up (PDF): OpenGate 0.9.1 Technical Report.

06.5The product, running

OpenGate UI overview — six stages, fourteen-paper stack
Overview: the paper stack by stage, with curated tables per mount.
OpenGate UI live gate — EpisodicAP / OOV-FP → PASS/FAIL
Live gate: EpisodicAP / OOV-FP → PASS/FAIL with threshold configuration.

07Why not just full-class mAP?

Full-class mAP hides what actually breaks in deployment: as the allowed vocabulary shrinks, out-of-vocabulary hallucination dominates. OpenGate audits exactly that.

08Quick start

pip install -e ".[ui]"          # CPU demo + UI
python scripts/init_third_party.py   # optional: clone Zhao/MARS upstream repos
python -m opengate run -c configs/vcad_demo.yaml     # CPU demo gate
python -m opengate run -c configs/vcad_detr3d_official.yaml  # official-format export
python -m opengate run -c configs/vcad_worldengine_posttrain.yaml  # WE eval → gate
python -m opengate ui           # dashboard (overview / mounts / live gate)
Optional GPU driving vocab (YOLO-World + COCO subset) — see docs/vcad/GPU_SETUP.md.

09Research docs

Research statement

H1 cliff · H2 gate

Safety frontier

coverage → risk

Technical report

arXiv style

Reproduce

one-shot

Limitations

scope