Deep RL
Gym env, variants, rewards
ArlamxV2Env (python/arlamx_v2/env.py) is a Gymnasium env around the C++ plant. One step is one 300 s advisor decision.
Constructor
ArlamxV2Env(geom_path=None, ggm_path=None, seed=None, variant="v3", gate_floor=None, kp=None, kd=None,
obs_extra="observer", reward_w=None, controller=None, physics="standard", gsi=None,
orbit=None, frozen=None, vehicle="vehicle", sensors_cfg="sensors_solarcat",
mtq_duty=1.0, cmd_frame="inertial", advisor_s=None, inner_dt=None)
| Argument | Meaning |
|---|---|
variant | v3, v4a, v4b, v5a, v5b, v6, v7, v8a, v8b, v9a–d, v10, v11, v10r6 |
physics | fast | standard | high | path to a physics YAML | dict |
gsi, controller | override aero.gsi and attitude.law |
orbit | envelope dict: altitude_km, inc_deg, ecc, f107, ap, raan_deg, argp_deg, nu_deg, omega_dps, mass_kg, soc, episode_orbits |
frozen | pinned plant configs (from a snapshot) |
obs_extra | none | ctrl | observer | observer_fore. Zero-masks the v7 block for ablations without changing the shape. |
cmd_frame, mtq_duty, advisor_s, inner_dt | plant command frame, coil duty, decision period, substep |
Reset
- Orbit: altitude U(300, 500) km, inclination U(20, 40)° (v10+), eccentricity U(0, 0.01), RAAN/argp/ν uniform. Converted with
coe_to_rv. - Attitude: random MRP |σ| ≤ 0.3. Tumble ω U(−5, 5)°/s per axis. Forced brownout gives 3–5°/s.
- Space weather per the variant envelope (see Atmosphere). Mass: the vehicle card (v8+) or U(0.5, 0.75) kg. SoC 0.5.
- Episode length: 40 / 120 (v10) / 160 (v11) orbits of a 400 km period. Terminated below 250 km geodetic.
- Deterministic evaluation:
reset(seed=…, options={altitude_km, inc_deg, ecc, f107, ap, omega_dps, soc, mass_kg, damage_kind, damage_step_frac, force_brownout, …}).
Observation
| Block | Size | Content |
|---|---|---|
| base | 26 | r/7000 km (3), v/8 km/s (3), ω (3), B̂_B (3), |B|/50 µT, v̂_B (3), sin ν, cos ν, mass, log₁₀ρ, SoC, P_gen_norm, ground-station direction + visible (4) |
| v7 IPC (+9) | 35 | τ_ctrl/cap (3), disturbance estimate/cap (3), mean gate, Sun foresight (2) |
| v10r6 (+6, appended last) | 89 | current on/off state of the 6 rods (0/1) |
| v9 temporal | up to 83 | future (FP32 forecast): Δalt, Δsma, SoC_proj, offset per offset; history: alt, SoC, log₁₀ρ, |ω| per offset. v9d / v10+: 6 + 6 offsets. |
The policy sees plant truth for r, v, ω, B. Sensor noise only affects the torque KF. The same holds for the MPC baselines, so advisor-vs-advisor comparisons are fair, but the input side is not flight-representative.
Action
| Variant | Dim | Content |
|---|---|---|
| v3–v5 | 4 | target quaternion (normalised, slerp-clipped) |
| v6 | 5 | + torque scale 0.15–1 (150 s step) |
| v7 | 7 | + per-axis authority gates in [gate_floor, 1] |
| v8b, v9, v10, v11 | 7 | + delegation split s ∈ [0, 1]³, gate = 1 − s(1 − floor). In v8a the split is ignored (control arm). |
| v10r6 | 10 | quaternion + 6 rod on/off gates (> 0 = on) |
Reward families
| Family | File | Main terms |
|---|---|---|
| SC_v3–v7 | reward.py | exponential-longevity-weighted ΔE vs baseline, SoC-modulated power, ground-station alignment, smoothness, ω, momentum, altitude cliff; v4 power band; v5 lift/SRP work, torque thrift; v7 composite |
| SC_v8+ | reward_v8.py, config/plant/reward_v*.yaml | dE_vs_baseline (weight 2.6, the decisive knob), power band 0.4–0.6, tiered GS pointing, feasibility, stability margin (320/250 km), dE stability, delegation power/accuracy/boost, v9 trend adaptation, v10 power drop, v11 adapt_consistency / deleg_hold |
| v12–v14 | train_v12.py::V12Wrapper | slew penalty, env-torque alignment, clip penalty, MPC-scored comparator, wrong-way clip, climate mix, low-SoC and generation terms |
Full equations are in docs/modules/14_reward_v8.md and docs/handoff/06_reward_and_delegation.md. Note the documented bug: before 2026-09-13, deleg_power was identically zero in every run.
Power model
Housekeeping 7 mW, +205 mW sunlit avionics (and GNSS 205 mW), +400 mW during a pointed lit pass. Generation 0.78 W × 0.85 × |ẑ·ŝ| (two-sided cells). 0.53 Wh supercap. Actuation is I²R-quadratic in the applied dipole (v8+). At zero SoC the env enters brownout: detumble, then a Sun-search recovery loop.
Baselines
advisors/attitudes.py: min drag, fixed AoA, axis pointing.advisors/heuristic.py: power-aware heuristic bank.advisors/mpc.py,mpc_v3.py: sampling MPC over the C++ plant (flow-frame, actuator-limited in v3).duo.py: arbiter between two advisors by FP32 forward propagation.