Documents

Deep RL

Gym env, variants, rewards

ArlamxV2Env (python/arlamx_v2/env.py) is a Gymnasium env around the C++ plant. One step is one 300 s advisor decision.

Constructor

ArlamxV2Env(geom_path=None, ggm_path=None, seed=None, variant="v3", gate_floor=None, kp=None, kd=None,
            obs_extra="observer", reward_w=None, controller=None, physics="standard", gsi=None,
            orbit=None, frozen=None, vehicle="vehicle", sensors_cfg="sensors_solarcat",
            mtq_duty=1.0, cmd_frame="inertial", advisor_s=None, inner_dt=None)
ArgumentMeaning
variantv3, v4a, v4b, v5a, v5b, v6, v7, v8a, v8b, v9a–d, v10, v11, v10r6
physicsfast | standard | high | path to a physics YAML | dict
gsi, controlleroverride aero.gsi and attitude.law
orbitenvelope dict: altitude_km, inc_deg, ecc, f107, ap, raan_deg, argp_deg, nu_deg, omega_dps, mass_kg, soc, episode_orbits
frozenpinned plant configs (from a snapshot)
obs_extranone | ctrl | observer | observer_fore. Zero-masks the v7 block for ablations without changing the shape.
cmd_frame, mtq_duty, advisor_s, inner_dtplant command frame, coil duty, decision period, substep

Reset

  • Orbit: altitude U(300, 500) km, inclination U(20, 40)° (v10+), eccentricity U(0, 0.01), RAAN/argp/ν uniform. Converted with coe_to_rv.
  • Attitude: random MRP |σ| ≤ 0.3. Tumble ω U(−5, 5)°/s per axis. Forced brownout gives 3–5°/s.
  • Space weather per the variant envelope (see Atmosphere). Mass: the vehicle card (v8+) or U(0.5, 0.75) kg. SoC 0.5.
  • Episode length: 40 / 120 (v10) / 160 (v11) orbits of a 400 km period. Terminated below 250 km geodetic.
  • Deterministic evaluation: reset(seed=…, options={altitude_km, inc_deg, ecc, f107, ap, omega_dps, soc, mass_kg, damage_kind, damage_step_frac, force_brownout, …}).

Observation

BlockSizeContent
base26r/7000 km (3), v/8 km/s (3), ω (3), B̂_B (3), |B|/50 µT, v̂_B (3), sin ν, cos ν, mass, log₁₀ρ, SoC, P_gen_norm, ground-station direction + visible (4)
v7 IPC (+9)35τ_ctrl/cap (3), disturbance estimate/cap (3), mean gate, Sun foresight (2)
v10r6 (+6, appended last)89current on/off state of the 6 rods (0/1)
v9 temporalup to 83future (FP32 forecast): Δalt, Δsma, SoC_proj, offset per offset; history: alt, SoC, log₁₀ρ, |ω| per offset. v9d / v10+: 6 + 6 offsets.

The policy sees plant truth for r, v, ω, B. Sensor noise only affects the torque KF. The same holds for the MPC baselines, so advisor-vs-advisor comparisons are fair, but the input side is not flight-representative.

Action

VariantDimContent
v3–v54target quaternion (normalised, slerp-clipped)
v65+ torque scale 0.15–1 (150 s step)
v77+ per-axis authority gates in [gate_floor, 1]
v8b, v9, v10, v117+ delegation split s ∈ [0, 1]³, gate = 1 − s(1 − floor). In v8a the split is ignored (control arm).
v10r610quaternion + 6 rod on/off gates (> 0 = on)

Reward families

FamilyFileMain terms
SC_v3–v7reward.pyexponential-longevity-weighted ΔE vs baseline, SoC-modulated power, ground-station alignment, smoothness, ω, momentum, altitude cliff; v4 power band; v5 lift/SRP work, torque thrift; v7 composite
SC_v8+reward_v8.py, config/plant/reward_v*.yamldE_vs_baseline (weight 2.6, the decisive knob), power band 0.4–0.6, tiered GS pointing, feasibility, stability margin (320/250 km), dE stability, delegation power/accuracy/boost, v9 trend adaptation, v10 power drop, v11 adapt_consistency / deleg_hold
v12–v14train_v12.py::V12Wrapperslew penalty, env-torque alignment, clip penalty, MPC-scored comparator, wrong-way clip, climate mix, low-SoC and generation terms

Full equations are in docs/modules/14_reward_v8.md and docs/handoff/06_reward_and_delegation.md. Note the documented bug: before 2026-09-13, deleg_power was identically zero in every run.

Power model

Housekeeping 7 mW, +205 mW sunlit avionics (and GNSS 205 mW), +400 mW during a pointed lit pass. Generation 0.78 W × 0.85 × |ẑ·ŝ| (two-sided cells). 0.53 Wh supercap. Actuation is I²R-quadratic in the applied dipole (v8+). At zero SoC the env enters brownout: detumble, then a Sun-search recovery loop.

Baselines

  • advisors/attitudes.py: min drag, fixed AoA, axis pointing.
  • advisors/heuristic.py: power-aware heuristic bank.
  • advisors/mpc.py, mpc_v3.py: sampling MPC over the C++ plant (flow-frame, actuator-limited in v3).
  • duo.py: arbiter between two advisors by FP32 forward propagation.