One score never tells the story.
Same latent shapes. Four conditions. No shape shared across training and holdout.
A geometry classification release, evaluated across clean, noisy, rotated, and scaled point clouds. Every decision traces back to the data and model that produced it.
Explore two saved evaluations of real model artifacts. Switching updates the evidence below; it does not call a live backend.
Same latent shapes. Four conditions. No shape shared across training and holdout.
Rows: actual class · columns: predicted class
Counts are evaluation samples, not independent shapes.
These drills change real artifacts and re-run the gate. No fabricated model scores.
Candidate features compared with clean training data. Train-only quantile bins; 0.5 pseudo-count smoothing.
| Holdout slice | Mean PSI | Maximum mean shift |
|---|
PSI and these demo thresholds are diagnostic choices for this synthetic fixture. They are not calibrated production or CAD acceptance criteria.
Canonical JSON SHA-256 binds models, dataset, report, and immutable release metadata. Promotion re-evaluates the exact inputs and compares against the active model.
Real softmax models are trained locally on generated cube, sphere, and cylinder surface clouds. Invariant features improve this seeded synthetic task. The dataset does not contain real CAD assemblies, sensor captures, or industrial defects.
Confidence intervals resample latent shape groups. This small, deliberately simple benchmark demonstrates release engineering; it does not establish real-world geometric recognition quality.