Documents

Guide

Control loop and deep RL

The learned network does not command coil current directly. It chooses a target attitude, on a slow clock. A classical law and the magnetorquers track that target on a fast clock. The same split is how you design the onboard loop, whether the advisor is a script or a network.

Two clocks

  1. Advisor step, 300 s (150 s on variant v6). The advisor emits a quaternion. Later variants also emit a torque-authority split, or on/off gates for six torque rods. The quaternion is normalised. The change from the previous command is limited to about 40°.
  2. Plant step, 2 s (integrator.rk4_step_s in the physics YAML). Classical RK4 integrates position, velocity, modified Rodrigues parameters, and body rate. Aero, radiation pressure, gravity-gradient, and control torques are held across the substep.

Three plant modes:

ModeWhat it does
pointClosed loop. The attitude law tracks the advisor’s quaternion. This is the training mode.
detumbleB-dot. Used when the battery model browns out.
prescribedThe attitude is set to the target and the rate is zero. Decay studies use this so the orbit change is the aerodynamics.

Inner law

Pick the law with --controller mrp or --controller quaternion, or with attitude.law in the physics file. Both are PD laws with a per-axis torque clip and a body-rate limit. Gains are in config/plant/gains_mrp.yaml.

The torque command then goes through the magnetic field.

  • Three coils. The dipole is the component of the demanded torque that is perpendicular to B, clipped to dipole_max in config/plant/power_mtq.yaml. The torque the spacecraft actually gets is m × B. There is no torque along B.
  • Six torque rods on variant v10r6. The SolarCat cards config/plant/vehicle_v3*.yaml supply the rod axes, the dipoles, and the plate file. The demanded torque is shared across rod axes by least squares, then scaled so each rod stays inside its dipole. The advisor’s extra six numbers gate rods on or off.
  • attitude.ideal_torque: true applies the PD torque as a body couple and skips the magnetic constraint. The earlier campaign gains were tuned that way. The file default is false, which is the magnetic plant. Read that switch before you treat a trained policy as a flight law. The measurement is written up in docs/modules/15_control_magnetorquer.md.

cmd_frame is inertial by default (hold the quaternion in inertial axes). flow holds it in the airflow frame and refreshes the target every substep. An optional onboard two-body state can feed that flow frame when you are studying navigation without GNSS.

What you change while designing the loop

QuestionWhere
MRP or quaternion--controller or attitude.law
Magnetic plant or ideal coupleattitude.ideal_torque
Gains, rate limit, gate floorconfig/plant/gains_mrp.yaml
Dipole and powerconfig/plant/power_mtq.yaml
Mass, inertia, rods, plate fileconfig/plant/vehicle.yaml and vehicle_v3*.yaml (variant v10r6)
Advisor periodArlamxV2Env(advisor_s=...). The command-line trainer uses 300 s, or 150 s for v6
Inner stepintegrator.rk4_step_s, or ArlamxV2Env(inner_dt=...)
Disturbances the loop must reject--physics and --gsi

Scripted advisors live in python/arlamx_v2/advisors/: minimum drag, a fixed angle of attack, a power-aware heuristic, and a sampling MPC. python main.py sim advisor_run flies the closed-loop campaign. Those are the baselines a trained network is compared with. duo.py picks between two advisors by propagating both forward.

Training the advisor

ArlamxV2Env in python/arlamx_v2/env.py is the Gymnasium environment. One step is one advisor decision. Under it, the plant runs the 2 s loop, MSIS updates the density, a power model updates the state of charge, and a torque Kalman filter watches the gyros. The policy’s observation of position, velocity, and rate is the plant truth. Sensor noise enters the torque filter. That is the same observation the MPC baseline sees.

python main.py train --network config/network_quick.yaml
python main.py train --algo ppo --arch 4x16 --timesteps 300000 --n-envs 32
python main.py train --variant v10 --physics standard --controller mrp
python main.py train --backend v12 --arch 4x18 --timesteps 600000
VariantAction
v3–v54 numbers: the target quaternion
v6Quaternion plus a torque scale. Advisor step is 150 s
v7–v11Quaternion plus three authority numbers. From v8 those are a delegation split per axis
v10r6Quaternion plus six rod gates. Plates and inertia come from the vehicle card

The default reward from v8 onward is dominated by orbital-energy change against a drag baseline, with terms for the power band, ground-station pointing, and altitude. Weights are config/plant/reward_*.yaml. Override one of them with --set 'train.reward_w={dE_weight: 3.0}'.

The network is an MLP, width repeated layers times. The default in config/network.yaml is PPO, 4×16, 300 000 steps, 32 environments. SAC and TD3 are the other algorithms. VecNormalize scales the reward and leaves the observation raw. The small width is deliberate: inference_budget sizes the actor for an STM32U575-class microcontroller.

Training is FP32. After it finishes:

python main.py quantize --in outputs/models/<id>/models/ppo_<id>.zip --out int8.zip

Each session freezes the plant it actually used inside outputs/snapshots/. Replay with --from-snapshot. A sweep in config/sweep.yaml launches several of those sessions and writes outputs/results/sweeps/<name>.csv.

A different spacecraft python main.py train loads the hex sail, except variant v10r6, which loads the geom: path in the vehicle YAML (a file under data/). To train on plates you simplified yourself, build the environment with ArlamxV2Env(geom_path="craft_1pct.geom", ...) and pass that environment to the same Stable-Baselines3 call python/arlamx_v2/train.py uses. Mass, inertia, and coil limits still have to match the vehicle, or the reward is scoring a different spacecraft than the one you will fly.

A design pass that stays small

  1. Simplify the exterior mesh (STL page) and decide mass, inertia, and dipole.
  2. Run python main.py decay on the closest shipped geometry, or the plate snippet on the flow page, so you know the aerodynamic torque the loop has to beat.
  3. Pick ideal_torque on purpose. Leave it false when the question is magnetic authority.
  4. Fly advisor_run (heuristic and MPC) on that physics file. That is the bar.
  5. Train a 4×16 PPO for a short budget (network_quick.yaml), read the snapshot, then lengthen the run only if the short run is healthy.
  6. Quantize the actor you intend to time on the microcontroller, and keep the FP32 zip beside it.