refactor(eval): consolidate drl-pinball reproduction

Retire duplicate reproduction paths in favor of the canonical V5 and Legacy runners, while preserving historical tooling in archives and publishing audited summary plots.

Co-authored-by: Cursor <cursoragent@cursor.com>
This commit is contained in:
Frank14f
2026-08-08 16:19:11 +08:00
co-authored by Cursor
parent 61e82ec90a
commit 326d642751
79 changed files with 3363 additions and 3996 deletions
@@ -0,0 +1,234 @@
# V5 Eval Pipeline — Inference & Results Documentation
> **Scope**: Inference infrastructure for V5-trained PPO models on the new CelerisLab solver.
> Legacy model reproduction is handled by @src/drl_pinball/reproduce/.
> Companion to `train/TRAIN_PIPELINE.md` for the training infrastructure.
---
## 1. Architecture Overview
```
scene_manifest.py (Single Source of Truth)
└── TRAIN_SCENES[20 entries] → used by infer_train.py
infer_train.py (GPU 2)
├── Create V5 CFD env
├── Per-seed:
│ ├── VecNormalize.load(seed.pkl)
│ ├── Skeleton PPO + weight inject
│ └── 360-step deterministic roll
├── Output: signals.npz + fields.npz + PNGs
└── Best-seed metrics.json
run_all_with_fields.sh
├── Phase 1: infer_train.py (stagger 120s per config change)
├── Phase 2: collect_baselines.py
├── Phase 3: bridge_to_sr.py
└── Phase 4: data integrity check
```
---
## 2. Scene Manifest (`scene_manifest.py`)
The single source of truth for all scene configurations used by both pipelines.
### 2.1 TRAIN_SCENES (V5 PPO, 2000×600, uniform + free-slip)
| scene_id | Type | SI | num_steps | Seeds | Description |
|----------|------|:--:|:---------:|:-----:|-------------|
| `kar_re100_sc` | karman | 800 | 360 | 41-45 | Karman Re100, 5-seed study |
| `kar_re60_tr` | karman | 800 | 360 | 43 | Cross-Re transfer to Re60 |
| `kar_re200_tr` | karman | 500 | 360 | 43 | Cross-Re transfer to Re200 |
| `kar_re400_tr` | karman | 400 | 360 | 43 | Cross-Re transfer to Re400 |
| `kar_d075_sc` | karman | 800 | 360 | 44 | VarDist d=0.75L scratch |
| `kar_d15_sc` | karman | 800 | 360 | 45 | VarDist d=1.5L scratch |
| `kar_d2_sc` | karman | 800 | 360 | 45 | VarDist d=2.0L scratch |
| `ill_075L_sc` | illusion | 400 | 360 | 43 | Illusion 0.75L target |
| `ill_1L_sc` | illusion | 600 | 360 | 43 | Illusion 1.0L target |
| `ill_15L_sc` | illusion | 800 | 360 | 43 | Illusion 1.5L target |
| `ill_2L_sc` | illusion | 800 | 360 | 43 | Illusion 2.0L target (new) — training bug known |
> **Note on illusion scenes**: Training bug — all target diameters were accidentally set to 1L. Transfer results invalid.
### 2.2 Legacy REPRODUCE_SCENES
Handled by @src/drl_pinball/reproduce/ — not by eval.
---
## 3. Train Inference Pipeline (`infer_train.py`)
### 3.1 Per-Scene Workflow
```
For each scene in TRAIN_SCENES:
1. Load calibration.json + target.npy (+ target_harmonics.json for illusion)
2. Create KarmanCloakEnv / IllusionCloakEnv (fresh CFD init)
3. For each seed (seed_label, model_dir):
a. Create skeleton PPO(env=vec_env, Sin, [64,64])
b. Extract policy weights from best_model.zip
c. Load VecNormalize from seed's vec_normalize.pkl (frozen)
d. 360-step deterministic rollout
e. Record: sensors, forces, actions, rewards, per-component r_cd/r_cl/r_sim
4. Pick best seed by tail-180 avg reward
5. Re-create env, load best seed model, run → capture vorticity PNGs
6. Generate target vorticity (dist_cyl only / target cyl only)
7. Generate zero-action baseline vorticity
8. Write: signals.npz, metrics.json, all_seeds.json, vorticity_*.png
```
### 3.2 Skeleton Injection Pattern
Due to `numpy._core.numeric` cloudpickle deserialization issues with SB3 models trained on certain Python versions, `PPO.load()` may fail. The fallback is **skeleton injection**:
```python
# Method: skeleton PPO + manual weight injection
skeleton = PPO(
"MlpPolicy",
policy_kwargs={"activation_fn": Sin, "net_arch": [64, 64]},
env=vec_env, device=device,
n_steps=2048, batch_size=64, n_epochs=10,
learning_rate=3e-4, gamma=0.995, verbose=0,
)
# Extract weights from zip
with zipfile.ZipFile("best_model.zip") as zf:
with zf.open("policy.pth") as f:
state_dict = torch.load(io.BytesIO(f.read()), map_location="cpu")
skeleton.policy.load_state_dict(state_dict, strict=False)
# VecNormalize loaded separately
vec_env = VecNormalize.load("vec_normalize.pkl", vec_env)
vec_env.training = False # Frozen statistics
vec_env.norm_reward = False
```
### 3.3 Output Files per Scene
```
eval/output/train/{scene_id}/
├── signals.npz # sensors (360,6), forces (360,6), actions (360,3), rewards (360,)
├── metrics.json # Best seed summary: DTW, reward, action stats
├── all_seeds.json # Per-seed breakdown: reward, r_cd, r_cl, r_sim, sim_raw, dt_sec
├── vorticity_controlled.png # Final frame after full DRL rollout
├── vorticity_target.png # Target state (disturbance only / target cylinder only)
└── vorticity_zero.png # Zero-action baseline (no control)
```
### 3.4 Config Switching Delay
When consecutive scenes use different LBM config files (different ν → different kernel compilation), a 120-second delay is inserted to allow the previous GPU context to fully release before the new config triggers PTX recompilation. Scenes sharing the same config file run back-to-back without delay.
---
## 4. Baseline Collection (`collect_baselines.py`)
Collects `q_in.npz` (background flow) and `q_blk.npz` (zero-action pinball) for each scene.
Baselines are needed by OID and CCD analysis pipelines.
```bash
python collect_baselines.py --device 2 # all scenes
python collect_baselines.py --scene kar_re100_sc # single scene
```
## 5. SR Bridge (`bridge_to_sr.py`)
Converts V5 eval output to the format expected by `SR_analysis/`:
- `calibration.json``norm.json` (FORCE_SCALE, SENS_SCALE, sens_deviation=0)
- `signals.npz``controlled.npz` (sensors, forces, actions, rewards)
- `target.npy``target.npz`
## 6. GPU Scheduling (`run_all_with_fields.sh`)
Single GPU (2). Config switch waits 120s to avoid PTX compilation conflicts.
```
GPU2: [kar_re100_sc(5 seeds)] → [vardist_sc+tr] → [ill_*_sc]
--120s--> [kar_re60_sc] --120s--> [kar_re200_sc] --120s--> [kar_re400_sc]
--120s--> [cross-re transfers]
```
### Run Commands
```bash
# Full pipeline (infer + baselines + bridge):
bash run_all_with_fields.sh
# Single scene:
bash run_all_with_fields.sh --scene kar_re100_sc --gpu 2
python infer_train.py --scene kar_re100_sc --device-id 2
python viz_flow.py
```
---
## 6. Eval Results
### 6.1 Train Pipeline — Completed Cases
| Scene | Best Seed | DTW sim_raw | Reward | r_cd | r_cl | r_sim |
|-------|:---------:|:-----------:|:------:|:----:|:----:|:-----:|
| kar_re100_sc | 45 | **0.926** | 0.941 | 0.984 | 0.990 | 0.872 |
| kar_d075_sc | 44 | **0.911** | 0.949 | 0.966 | 0.974 | 0.919 |
| kar_re60_tr | 43 | 0.187 | 0.312 | 0.659 | 0.334 | 0.091 |
| kar_re200_tr | 43 | 0.506 | 0.367 | 0.669 | 0.212 | 0.278 |
| kar_re400_tr | 43 | 0.428 | 0.399 | 0.737 | 0.400 | 0.176 |
### 6.2 Train Pipeline — Re100 5-Seed Breakdown
| Seed | DTW sim_raw | r_cd | r_cl | r_sim |
|:----:|:-----------:|:----:|:----:|:-----:|
| 41 | 0.634 | 0.307 | 0.241 | 0.388 |
| 42 | 0.893 | 0.959 | 0.762 | 0.788 |
| 43 | 0.850 | 0.908 | 0.469 | 0.629 |
| 44 | 0.523 | 0.282 | 0.177 | 0.322 |
| 45 | **0.926** | 0.984 | 0.990 | 0.872 |
Mean DTW across seeds: 0.765 ± 0.172. Best-worst spread: 0.403.
### 6.3 DTW Formula Note
**Important**: V5 uses a different DTW formula than Legacy pipelines. See `recompute_unified_dtw.py` for fair comparison using the same formula.
---
## 7. Visualization & Reporting Tools
### 7.1 `viz_signals.py`
Per-scene diagnostic plots (4×2 subplot grid): sensors, forces, actions, reward, FFT, phase portrait.
### 7.2 `viz_flow.py`
Cross-scene vorticity comparison panels with unified [-0.003, 0.003] colormap.
### 7.3 `generate_report.py`
Master comparison report with DTW bar chart.
---
## 8. Known Limitations
- **Illusion training bug**: All target diameters were accidentally set to 1L. Transfer results invalid.
- **VecNormalize compat**: `PPO.load()` may fail with cloudpickle on Python 3.10+. Skeleton injection fallback handles this.
- **GPU cleanup**: Between scenes, `pkill -f pycuda; sleep 180` may be needed.
---
## 9. Companion Documents
| File | Content |
|------|---------|
| `eval/README.md` | Quick start + self-review checklist |
| `eval/scene_manifest.py` | TRAIN_SCENES single source of truth |
| `eval/infer_train.py` | V5 train model inference |
| `eval/recompute_unified_dtw.py` | Fair DTW comparison (unified formula) |
| `train/TRAIN_PIPELINE.md` | Training pipeline documentation |
| `reproduce/REPRODUCE_KNOWLEDGE.md` | Reproduce module (separate) |
| `../SR_analysis/README.md` | SR analysis pipeline overview |
---
*Document last updated: 2026-07-13.*