Retire duplicate reproduction paths in favor of the canonical V5 and Legacy runners, while preserving historical tooling in archives and publishing audited summary plots. Co-authored-by: Cursor <cursoragent@cursor.com>
8.4 KiB
V5 Eval Pipeline — Inference & Results Documentation
Scope: Inference infrastructure for V5-trained PPO models on the new CelerisLab solver. Legacy model reproduction is handled by @src/drl_pinball/reproduce/. Companion to
train/TRAIN_PIPELINE.mdfor the training infrastructure.
1. Architecture Overview
scene_manifest.py (Single Source of Truth)
└── TRAIN_SCENES[20 entries] → used by infer_train.py
infer_train.py (GPU 2)
├── Create V5 CFD env
├── Per-seed:
│ ├── VecNormalize.load(seed.pkl)
│ ├── Skeleton PPO + weight inject
│ └── 360-step deterministic roll
├── Output: signals.npz + fields.npz + PNGs
└── Best-seed metrics.json
run_all_with_fields.sh
├── Phase 1: infer_train.py (stagger 120s per config change)
├── Phase 2: collect_baselines.py
├── Phase 3: bridge_to_sr.py
└── Phase 4: data integrity check
2. Scene Manifest (scene_manifest.py)
The single source of truth for all scene configurations used by both pipelines.
2.1 TRAIN_SCENES (V5 PPO, 2000×600, uniform + free-slip)
| scene_id | Type | SI | num_steps | Seeds | Description |
|---|---|---|---|---|---|
kar_re100_sc |
karman | 800 | 360 | 41-45 | Karman Re100, 5-seed study |
kar_re60_tr |
karman | 800 | 360 | 43 | Cross-Re transfer to Re60 |
kar_re200_tr |
karman | 500 | 360 | 43 | Cross-Re transfer to Re200 |
kar_re400_tr |
karman | 400 | 360 | 43 | Cross-Re transfer to Re400 |
kar_d075_sc |
karman | 800 | 360 | 44 | VarDist d=0.75L scratch |
kar_d15_sc |
karman | 800 | 360 | 45 | VarDist d=1.5L scratch |
kar_d2_sc |
karman | 800 | 360 | 45 | VarDist d=2.0L scratch |
ill_075L_sc |
illusion | 400 | 360 | 43 | Illusion 0.75L target |
ill_1L_sc |
illusion | 600 | 360 | 43 | Illusion 1.0L target |
ill_15L_sc |
illusion | 800 | 360 | 43 | Illusion 1.5L target |
ill_2L_sc |
illusion | 800 | 360 | 43 | Illusion 2.0L target (new) — training bug known |
Note on illusion scenes: Training bug — all target diameters were accidentally set to 1L. Transfer results invalid.
2.2 Legacy REPRODUCE_SCENES
Handled by @src/drl_pinball/reproduce/ — not by eval.
3. Train Inference Pipeline (infer_train.py)
3.1 Per-Scene Workflow
For each scene in TRAIN_SCENES:
1. Load calibration.json + target.npy (+ target_harmonics.json for illusion)
2. Create KarmanCloakEnv / IllusionCloakEnv (fresh CFD init)
3. For each seed (seed_label, model_dir):
a. Create skeleton PPO(env=vec_env, Sin, [64,64])
b. Extract policy weights from best_model.zip
c. Load VecNormalize from seed's vec_normalize.pkl (frozen)
d. 360-step deterministic rollout
e. Record: sensors, forces, actions, rewards, per-component r_cd/r_cl/r_sim
4. Pick best seed by tail-180 avg reward
5. Re-create env, load best seed model, run → capture vorticity PNGs
6. Generate target vorticity (dist_cyl only / target cyl only)
7. Generate zero-action baseline vorticity
8. Write: signals.npz, metrics.json, all_seeds.json, vorticity_*.png
3.2 Skeleton Injection Pattern
Due to numpy._core.numeric cloudpickle deserialization issues with SB3 models trained on certain Python versions, PPO.load() may fail. The fallback is skeleton injection:
# Method: skeleton PPO + manual weight injection
skeleton = PPO(
"MlpPolicy",
policy_kwargs={"activation_fn": Sin, "net_arch": [64, 64]},
env=vec_env, device=device,
n_steps=2048, batch_size=64, n_epochs=10,
learning_rate=3e-4, gamma=0.995, verbose=0,
)
# Extract weights from zip
with zipfile.ZipFile("best_model.zip") as zf:
with zf.open("policy.pth") as f:
state_dict = torch.load(io.BytesIO(f.read()), map_location="cpu")
skeleton.policy.load_state_dict(state_dict, strict=False)
# VecNormalize loaded separately
vec_env = VecNormalize.load("vec_normalize.pkl", vec_env)
vec_env.training = False # Frozen statistics
vec_env.norm_reward = False
3.3 Output Files per Scene
eval/output/train/{scene_id}/
├── signals.npz # sensors (360,6), forces (360,6), actions (360,3), rewards (360,)
├── metrics.json # Best seed summary: DTW, reward, action stats
├── all_seeds.json # Per-seed breakdown: reward, r_cd, r_cl, r_sim, sim_raw, dt_sec
├── vorticity_controlled.png # Final frame after full DRL rollout
├── vorticity_target.png # Target state (disturbance only / target cylinder only)
└── vorticity_zero.png # Zero-action baseline (no control)
3.4 Config Switching Delay
When consecutive scenes use different LBM config files (different ν → different kernel compilation), a 120-second delay is inserted to allow the previous GPU context to fully release before the new config triggers PTX recompilation. Scenes sharing the same config file run back-to-back without delay.
4. Baseline Collection (collect_baselines.py)
Collects q_in.npz (background flow) and q_blk.npz (zero-action pinball) for each scene.
Baselines are needed by OID and CCD analysis pipelines.
python collect_baselines.py --device 2 # all scenes
python collect_baselines.py --scene kar_re100_sc # single scene
5. SR Bridge (bridge_to_sr.py)
Converts V5 eval output to the format expected by SR_analysis/:
calibration.json→norm.json(FORCE_SCALE, SENS_SCALE, sens_deviation=0)signals.npz→controlled.npz(sensors, forces, actions, rewards)target.npy→target.npz
6. GPU Scheduling (run_all_with_fields.sh)
Single GPU (2). Config switch waits 120s to avoid PTX compilation conflicts.
GPU2: [kar_re100_sc(5 seeds)] → [vardist_sc+tr] → [ill_*_sc]
--120s--> [kar_re60_sc] --120s--> [kar_re200_sc] --120s--> [kar_re400_sc]
--120s--> [cross-re transfers]
Run Commands
# Full pipeline (infer + baselines + bridge):
bash run_all_with_fields.sh
# Single scene:
bash run_all_with_fields.sh --scene kar_re100_sc --gpu 2
python infer_train.py --scene kar_re100_sc --device-id 2
python viz_flow.py
6. Eval Results
6.1 Train Pipeline — Completed Cases
| Scene | Best Seed | DTW sim_raw | Reward | r_cd | r_cl | r_sim |
|---|---|---|---|---|---|---|
| kar_re100_sc | 45 | 0.926 | 0.941 | 0.984 | 0.990 | 0.872 |
| kar_d075_sc | 44 | 0.911 | 0.949 | 0.966 | 0.974 | 0.919 |
| kar_re60_tr | 43 | 0.187 | 0.312 | 0.659 | 0.334 | 0.091 |
| kar_re200_tr | 43 | 0.506 | 0.367 | 0.669 | 0.212 | 0.278 |
| kar_re400_tr | 43 | 0.428 | 0.399 | 0.737 | 0.400 | 0.176 |
6.2 Train Pipeline — Re100 5-Seed Breakdown
| Seed | DTW sim_raw | r_cd | r_cl | r_sim |
|---|---|---|---|---|
| 41 | 0.634 | 0.307 | 0.241 | 0.388 |
| 42 | 0.893 | 0.959 | 0.762 | 0.788 |
| 43 | 0.850 | 0.908 | 0.469 | 0.629 |
| 44 | 0.523 | 0.282 | 0.177 | 0.322 |
| 45 | 0.926 | 0.984 | 0.990 | 0.872 |
Mean DTW across seeds: 0.765 ± 0.172. Best-worst spread: 0.403.
6.3 DTW Formula Note
Important: V5 uses a different DTW formula than Legacy pipelines. See recompute_unified_dtw.py for fair comparison using the same formula.
7. Visualization & Reporting Tools
7.1 viz_signals.py
Per-scene diagnostic plots (4×2 subplot grid): sensors, forces, actions, reward, FFT, phase portrait.
7.2 viz_flow.py
Cross-scene vorticity comparison panels with unified [-0.003, 0.003] colormap.
7.3 generate_report.py
Master comparison report with DTW bar chart.
8. Known Limitations
- Illusion training bug: All target diameters were accidentally set to 1L. Transfer results invalid.
- VecNormalize compat:
PPO.load()may fail with cloudpickle on Python 3.10+. Skeleton injection fallback handles this. - GPU cleanup: Between scenes,
pkill -f pycuda; sleep 180may be needed.
9. Companion Documents
| File | Content |
|---|---|
eval/README.md |
Quick start + self-review checklist |
eval/scene_manifest.py |
TRAIN_SCENES single source of truth |
eval/infer_train.py |
V5 train model inference |
eval/recompute_unified_dtw.py |
Fair DTW comparison (unified formula) |
train/TRAIN_PIPELINE.md |
Training pipeline documentation |
reproduce/REPRODUCE_KNOWLEDGE.md |
Reproduce module (separate) |
../SR_analysis/README.md |
SR analysis pipeline overview |
Document last updated: 2026-07-13.