Deep RL
Training & sessions
Everything runs through python main.py. A training session resolves the settings, writes a snapshot, builds vectorised envs, trains with Stable-Baselines3, and records the result.
Commands
python main.py list # every command and tool
python main.py train # config/network.yaml + orbit.yaml + train.yaml
python main.py train --network config/network_quick.yaml # smoke run (20k steps, 4 envs)
python main.py train --algo sac --arch 3x32 --timesteps 500000 --n-envs 16 --seed 7
python main.py train --variant v11 --physics high --gsi cll --controller quaternion
python main.py train --name storm --set orbit.training.ap=[100,200] --set network.ppo.learning_rate=1e-4
python main.py train --set 'train.reward_w={dE_weight: 3.0}'
python main.py train --backend v12 --arch 4x18 --timesteps 600000 --set train.v12.dual_eval=true
python main.py train --variant v10r6 # six rods, 6 on/off gates
python main.py train --from-snapshot outputs/snapshots/<id>.yaml [--timesteps N --seed S]
python main.py sweep --config config/sweep.yaml [--force]
python main.py snapshots
python main.py quantize --in outputs/models/<id>/models/ppo_<id>.zip --out int8.zip
python -m arlamx_v2.sc6mq --config config/sc_v2_6mq.yaml --steps 500000 # SolarCat six-rod trainer
Precedence: CLI flags and --set section.key=value (value parsed as YAML; sections network, orbit, train) override the YAML files. --network/--orbit/--train FILE swaps in a whole file.
Pipeline
- Resolve (
session.load_settings): network, orbit and train YAMLs + overrides + the plant configs (physics, reward, gains, power, estimator, sensors) fromconfig/plant/. - Snapshot: write
outputs/snapshots/<MM-DD-HH-MM>_<ALGO>_<LxW>[_name].yamlbefore training. It holds status, command, replay command, versions, every resolved setting and, after the run, the result. Crashed runs are recorded too. - Pin (
config.pin): plant configs are frozen in each worker (ArlamxV2Env(frozen=…)), so editing YAMLs mid-run has no effect. - Envs:
SubprocVecEnv(n_envs > 1) ofMonitor(ArlamxV2Env(seed=seed+rank, variant, **env_kw)), wrapped inVecNormalize(reward-only by default). - Model: SB3
MlpPolicywithnet_arch= [width] × layers (PPO: separate π and V of the same shape). - Outputs:
outputs/models/<id>/models/<algo>_<id>.zip,vecnormalize.pkl,logs/(TensorBoard),metrics.json. The v12 backend addsbest/,ckpts/,eval_curve.csv,BEST.md.
Backends
| Backend | What it does |
|---|---|
train | Plain SB3 (PPO / SAC / TD3) on the env reward for the chosen variant. |
v12 | PPO only. Env v10 + V12Wrapper (slew term, MPC-scored comparator, w_env, wrong-way clip, climate mix, low-SoC terms), with a checkpoint and a 10-orbit eval against the MPC references every 50k steps. Knobs are under train.v12. |
Default hyperparameters (config/network.yaml)
| Block | Values |
|---|---|
| session | algo ppo, 4 × 16, 300 000 steps, 32 envs, seed 42 |
| ppo | lr 3e-4, n_steps 128, batch 64, 10 epochs, γ 0.99, ent 0.003, clip 0.2, λ 0.95 |
| sac / td3 | lr 3e-4, buffer 100 000, batch 256, γ 0.99, τ 0.005 |
| vecnormalize | norm_obs false, norm_reward true, clip_reward 10, γ 0.99 |
The networks are deliberately tiny (4×16 … 4×20) because the target is an onboard STM32U575 (see inference_budget.py). Training is FP32. INT8 is a separate step: quantize.py gives dynamic INT8 of the actor Linear layers, and int8.py gives an integer-only actor with a tanh LUT.
Sweeps
# config/sweep.yaml
name: example
base: {network.timesteps: 100000, network.n_envs: 8, train.variant: v8a}
grid: {network.algo: [ppo, sac], train.controller: [mrp, quaternion]} # 4 sessions
# arms: [{train.physics: high}, {network.width: 32}] # or an explicit list
Each arm is a normal session with its own snapshot. The ledger is outputs/results/sweeps/<name>.csv, and arms already marked done are skipped unless you pass --force.
Replaying a snapshot
--from-snapshot replays exactly, with the plant values pinned. Network or orbit changes keep the snapshot's plant. Changing train.variant, backend, physics, gsi or reward_w re-reads the plant from config/plant/.
Physics fidelity for training
| Preset | SH degree | Lunisolar | GSI | Field | MSIS interp | Use |
|---|---|---|---|---|---|---|
fast | 2 | off | Sentman | dipole | off | smoke tests, quick sweeps |
standard | 4 | off | Sentman α_E 0.93 | WMM | on | default training |
high | 8 | on | Walker-CLL α_N 0.93 | WMM | on | fidelity checks, final evals |
standard with domain randomisation (weather envelope, damage, SRP flashes, sensors), then evaluate on high. Once materials exist (see the plan), randomise accommodation and optical coefficients per episode within their uncertainty bands rather than training on one point value.Other tools
| Group | Tools |
|---|---|
sim | decay_run, aoa_decay, lift_drag_study, srp_assess, srp_f107_sweep, optics_run, advisor_run, validate_attitudes, plot_simplify, geometry |
eval | mc_decay, mc_v12, mc_four, damage_eval, analyze_ood, eval_v4, check_mpc_ref, inference_budget |
plot | plot_v8, plot_v12, plot_v14fix, plot_production, plot_gradient70, showcase_plots, slides, ipc_animation, report_v8, build_mpc_report |
legacy | bench_v8, campaign, tune_mpc, gradient70 (write into outputs/.old) |
These tools run immediately with no flags: inference_budget, plot_v14fix, plot_production, plot_gradient70, showcase_plots, slides, ipc_animation, report_v8, build_mpc_report, gradient70.