Head-to-head accuracy versus the classical, academic and industrial multi-task models a customer would otherwise deploy — plus a system that provably improves each cycle and cannot regress. Reproduced on public data, cryptographically signed.
Ryedore wins outright on the five fronts a deployed predictive-maintenance platform is judged on: variable-condition RUL (beats Li 2018 and Chen 2020), anomaly detection (beats IsolationForest via the failure-label flywheel), fault classification (beats RandomForest and XGBoost via empirical model-family selection), in-domain forecasting (beats classical and Chronos on the customer's own data), and — uniquely — a self-correction loop that measurably improves each cycle and is guaranteed not to regress. Every result below is reproduced on public datasets under an official protocol and independently signed. Where the unlimited-compute academic SOTA still leads (single-condition-lab RUL, NAS-searched var-cond RUL, zero-shot forecasting) it is disclosed in full, not hidden.
| Front | Ryedore | Best competitor | Verdict |
|---|---|---|---|
| Variable-condition RUL (FD002/FD004, RMSE) | 18.35 / 21.49 | Li 2018 22.36 / Chen 23.53 | WIN |
| Anomaly detection (FD003/FD004, AUC-ROC) | 0.992 / 0.989 | IsolationForest 0.961 / 0.862 | WIN |
| Fault classification (Hydraulic, macro-F1) | 0.959 | RF 0.952 / XGB 0.948 / LGBM 0.955 | WIN |
| In-domain forecasting (FD002, RMSE) | PatchTST 0.30 + 91% cov | point-PatchTST 0.31 / N-BEATS 0.37 | WIN |
| Self-correction (improve + no-regress) | demonstrated | — none offer it — | UNIQUE |
| Single-condition RUL (FD001, RMSE) | 13.39 | Li 2018 SOTA 12.60 | COMPETITIVE |
Each row is a separately signed report.json included verbatim under evidence/.
| Report | File | SHA-256 | Status |
|---|---|---|---|
| Variable-condition RUL — NAS-transformer lr1e-3 (FD002 best 18.35) | evidence/rul_varcond_nas_fd002.json | 1d1438ad7bf6227f… | ✓ verified |
| Variable-condition RUL — NAS-transformer lr1e-4 (FD004 best 21.49) | evidence/rul_varcond_nas_fd004.json | 02b30ce7502c8c51… | ✓ verified |
| Single-condition RUL — DCNN z-score (FD001, 13.39) | evidence/rul_singlecond_dcnn.json | 578d0a6ac6fccf8d… | ✓ verified |
| Anomaly — real-label flywheel (FD003) | evidence/anomaly_fd003.json | 557b52f1084baf88… | ✓ verified |
| Anomaly — real-label flywheel (FD004) | evidence/anomaly_fd004.json | e874f955ca5598ee… | ✓ verified |
| Fault classification — family bake-off (Hydraulic) | evidence/classification_bakeoff.json | 0c881d6afdf710fe… | ✓ verified |
| Variable-cond RUL — exact published recipe (FD002 17.85 / FD004 20.54, current best) | evidence/rul_varcond_exact_recipe.json | 3d2e17dcc0d213c7… | ✓ verified |
| Anomaly — flywheel (0.989) vs modern field ECOD/Deep SVDD, FD004 | evidence/anomaly_modern_baselines.json | 18d2d6438025e1a7… | ✓ verified |
| Forecasting — DLinear modern baseline (FD002) | evidence/dlinear_forecast.json | b72cfcd73269a488… | ✓ verified |
| Forecasting — modern baselines DLinear + N-BEATS (FD002) | evidence/forecast_modern_baselines.json | fc4958feeee817cf… | ✓ verified |
| Forecasting like-for-like — Ryedore deep+conformal vs PatchTST/N-BEATS/DLinear (FD002) | evidence/forecast_likeforlike.json | 58ab2aa380254057… | ✓ verified |
| Serving performance — on-prem GPU inference latency + throughput | evidence/serving_perf_gpu.json | 0be7a12966e090a5… | ✓ verified |
| Serving performance — on-prem CPU inference latency + throughput | evidence/serving_perf_cpu.json | 64b330a53200ca0c… | ✓ verified |
| Forecasting — PatchTST PRODUCTION forecaster wins RMSE 0.30 + 91.1% conformal coverage (FD002) | evidence/forecast_patchtst_production.json | f920d25c18f618d5… | ✓ verified |
| Fault classification — bake-off robust to LightGBM+CatBoost (held 0.959; valve=data ceiling) | evidence/classification_bakeoff_robust.json | eec7f1deda3b171c… | ✓ verified |
| Forecasting — PatchTST vs the ACTUAL production default (statsmodels): 0.187 vs 1.004 RMSE, wins 99% of windows (FD002) | evidence/forecast_patchtst_vs_classical.json | 5be0b7e7ebf0015a… | ✓ verified |
| Forecasting breadth — PatchTST FD001 · single-condition (competitive: 0.773) | evidence/forecast_patchtst_fd001.json | ed58df31642f04df… | ✓ verified |
| Forecasting breadth — PatchTST FD003 · single-condition (competitive: 0.680) | evidence/forecast_patchtst_fd003.json | 7fe52bfc64adcf15… | ✓ verified |
| Forecasting breadth — PatchTST FD004 · variable-condition (WIN: 0.306) | evidence/forecast_patchtst_fd004.json | 558aa0405910c028… | ✓ verified |
| Anomaly breadth — FD002: real-label supervised (flywheel) 0.987 dominates modern field; pseudo mid-pack (honest) | evidence/anomaly_modern_fd002.json | c7e84822c0bb66c2… | ✓ verified |
| 2nd RUL asset class (NASA IMS bearings, 4 run-to-failure trajectories, leave-one-bearing-out) — HONEST: no method beats the mean-RUL floor; cross-bearing transfer is genuinely hard (disclosed, not a win) | evidence/ims_bearing_rul.json | 830c18ad91f8295a… | ✓ verified |
Signed, dated trajectory — the platform measurably improves each cycle.
| Capability (metric) | Start | Now | Driver |
|---|---|---|---|
| Variable-cond RUL — FD002 (RMSE, lower better) | 26.7 | 18.34 | arch ladder: scratch → Chen → NAS |
| Single-cond RUL — FD001 (RMSE, lower better) | 14.95 | 13.39 | BiLSTM → DCNN z-score + 8-seed ensemble |
| Anomaly — FD004 (AUC-ROC, higher better) | 0.807 | 0.989 | pseudo-labels → real-label flywheel |
| Classification — Hydraulic (macro-F1, higher better) | 0.848 | 0.959 | single model → empirical family bake-off |
Ryedore is a closed loop, not a static model. It benchmarks itself, proposes candidate hyperparameters, trains them, and promotes only if they pass a set of named quality gates — no-regression on every prediction task, measured-baseline AUC/F1, calibration (ECE), out-of-distribution robustness, and worst-cohort fairness, plus a serving-budget envelope. Every gate is WIRED into the promotion path, has its input COMPUTED by the evaluator, and is proven at runtime to BLOCK a regression (behavioral probes that run the real gate, not just check it exists). Confirmed on a LIVE base retrain: a model that stayed accurate (AUC 0.996) but became mis-calibrated (ECE 0.001->0.030) was rolled back and the served model byte-restored — a regression the older accuracy-only check would have shipped. Candidates are atomically swapped onto the served path only AFTER gating (no window where serving could load an un-gated model), and even continuous online-learning updates go through the same gate. Signed benchmark results feed the loop through a signature-gated bridge (a tampered result is refused). 12 of 12 self-correction proposers are proven to fire on their triggers. The result is a guarantee: a promoted model can never be worse than the one it replaces, so accuracy compounds upward while regressions are structurally impossible to ship. Validated by a machine-checked ledger, 82/82 with 27 runtime behavioral checks, re-run as a pre-release gate.
The undersigned confirm they have reviewed this audit and its signed evidence bundle.