Documents

Reference

Architecture & the step loop

How one RL decision turns into 300 s of integrated orbit and attitude motion, and which module owns each part.

Repository layout

main.py              front door (train, sweep, decay, snapshots, quantize, sim/eval/plot/legacy, run, test, verify)
cpp/include/arlamx/  headers: types, constants, api, aero/, srp/, orbit/, attitude/, control/, mag/, onboard/
cpp/src/             implementations + bindings.cpp (pybind11 module arlamx_cpp)
python/arlamx_v2/    env.py (Gym), physics.py (presets), atmosphere.py (MSIS), geometry.py, train*.py,
                     session.py (snapshots), sweep.py, reward*.py, sensors.py, estimators.py, power.py,
                     advisors/ (heuristic, sampling MPC), propagator.py + nav.py (onboard), quantize.py, int8.py
config/              network.yaml, orbit.yaml, train.yaml, sweep.yaml, sc_v*_6mq*.yaml, plant/*.yaml
data/                GGM03S, WMM COF, .geom plate models, STL meshes
tests/               per-module tests + basilisk_ref + integration + validation
docs/                module notes (docs/modules), handoff notes, this site (docs/html)
outputs/             snapshots/, models/, results/, .old/ (pre-v2.7 campaigns)

Who does what

ConcernOwnerLanguage
Orbit + attitude integration, all forces and torquesSimulator::step (cpp/src/api.cpp)C++
Atmosphere density, temperature, mean mass, speciesatmosphere.query_msis → Simulator.set_atmosphere[_end]Python (pymsis)
Space-weather draws and jumps, SRP flashes, damage eventsArlamxV2Env.reset / _apply_jumps / stepPython
Physics preset → SimParamsphysics.apply_to_paramsPython
Observations, reward, power, sensors, torque KFenv.py, reward*.py, sensors.py, estimators.py, power.pyPython
Onboard forecast (FP32 two-body + J2 + exp. drag + held SRP)propagator.py ↔ onboard/propagator_f32.cppboth (mirror)
PoliciesSB3 (PPO/SAC/TD3), advisors/ (heuristic, sampling MPC), duo.pyPython

One env step

Python · ArlamxV2Env.step(action) decode action → q_cmd (4),gates / split / rod on-off weather jumps (F10.7, Ap),damage strike, SRP flash scale MSIS at geodetic point, t0(+ MSIS at predicted t0+300 sif atmosphere.interp) set_atmosphere(ρ,T,m̄,χ)set_atmosphere_end(...) power / battery, sensors,torque KF, ground stations,forecast, observation, reward→ (obs, r, term, trunc, info) C++ · Simulator::step(q_cmd) ×1 per env step slerp-clip q_cmd to max_slew (40–45°) · Sun/Moon for third body once per step substep loop: nsub = advisor_step_s / dt_s = 300 / 2 = 150 target σ, ω (inertial or flow frame, optional onboard-nav) · lerp atmosphere aero: Sentman | Walker-CLLF_B, τ_B per panel sum+ min-drag baseline C_D SRP: optical | cannonballEarth IR + albedoeclipse (cylindrical) gravity-gradient torque3μ/r³ û × Iû control: PD (MRP|quat) → coils /rods through B (WMM|dipole) | B-dot energy book-keeping: dE_actual, dE_baseline, dE_drag, dE_lift, drag_N(F·v dt / m, inertial frame) RK4 over dt_s in ⌈dt_s / rk4_step_s⌉ steps: ṙ = v, v̇ = a_grav(SH + 3rd body) + C_BNᵀF_B/m,σ̇ = ¼B(σ)ω, ω̇ = I⁻¹(τ − ω×Iω) (F_B, τ_B held over the substep) step(q) StepOut
One env step is one advisor step: 300 s (150 s for v6), 150 substeps of 2 s, RK4 inside each substep. External forces and torques are evaluated once per substep and held (zero-order hold) across its RK4 stages. Gravity is re-evaluated at every RK4 stage.

Timing constants

QuantityDefaultWhere
advisor step (env step)300 s (v6: 150 s; override advisor_s)SimParams.advisor_step_s, env ctor
substep (force/torque evaluation)2 s (override inner_dt)SimParams.dt_s
RK4 step2 s (≤ dt_s)integrator.rk4_step_s
slew clip per step45° (v3), 40° (v4+)SimParams.max_slew_rad
epochJD 2461060.5 = 2026-01-15 00:00 UTCenv.EPOCH_JD, SimParams.epoch_jd
hard stopgeodetic altitude < 80 km (plant); env terminal altitudes are per variantapi.cpp
step cost≈1.23 ms (v10 standard physics)CHANGELOG_v2.7

Frames

FrameDefinition in code
N (inertial)Earth-centred, z = rotation axis. Earth-fixed is reached by a z-rotation of IAU-1982 GMST (dcm_EN), with no precession, nutation or polar motion. The analytic Sun and Moon are given in the mean equator of date and used as N. SPICE returns J2000 (see limits).
E (Earth-fixed)r_E = dcm_EN(gmst) r_N, used for gravity SH, WMM and MSIS longitude
B (body)MRP σ_BN, C_BN = mrp_to_dcm(σ). Panels, rods and the inertia are in B.
F (flow)e₁ = v_rel, e₃ = orbit normal (or ⟂B for flowB, ⟂Sun for flowS), e₂ = e₃ × e₁. Used when cmd_frame ≠ inertial.
L (LVLH)dcm_LN: z = −r̂, y = −ĥ, x = y × z

StepOut (what Simulator.step returns)

KeyMeaning
r, v, sigma, omega, tstate at the end of the step (N frame, MRP shadow set, body rates)
altitude_km, sma_mBowring geodetic height; vis-viva semi-major axis
eclipse, sun_B, B_Bshadow flag at the end, Sun and field in body axes
Cd, Cllast-substep coefficients on A_ref (half the total panel area if one_sided_ref)
drag_Nmean along-track drag magnitude (N)
dE_actual, dE_baseline, dE_drag, dE_liftspecific orbital-energy change (J/kg) from aero + SRP + ERP, from the min-drag counterfactual, and the drag and lift parts
tracking_err_radmean attitude error to the target
tau_ctrl_mean, tau_aero_mean, tau_srp_mean, tau_erp_mean, tau_gg_mean, tau_env_meanmean torques (N·m, body)
tau_cmd_mean, tau_demand_mean, tau_demand_absmean, tau_shortfall_frac, m_mean, gyro_meancontroller demand before and after clips, dipole used, magnetic shortfall, gyroscopic term
a_srp_sun, n_sunlitmean along-track SRP acceleration while lit; count of lit substeps
rod_duty2_meanper-rod mean of (d/d_max)² × duty (I²R power proxy)
n_substepssubsteps completed (fewer if the 80 km stop fired)