Learn / Practice
Embodied AI VI: what do recovery gating and replanning each solve?
EMBODIED-AI · Engineering cases
A 30-rollout ablation reduces first target jumps from 50–97 mm to about 0.25 mm, without improving placement outcomes in this fixed matrix. Separate recovery evidence, command continuity and task success.
View count unavailable
On this page 9
The previous experiment added fresh grasp confirmation and remaining-path replanning together. A comparison of policy bundles cannot tell us which change produced which effect.
This experiment separates the two switches. Across four fixed pairs, replanning reduced the first post-resume target jump from 50.285–97.055 mm to 0.245–0.263 mm. Yet all four active-recovery combinations placed the object in all five non-drop conditions. No improvement in placement success was observed in this matrix.
These findings concern different questions: whether to resume, what target to command after resuming, and whether the task eventually succeeds.
Pinned code · 30 raw rollouts · Independent metrics
1. Separate evidence from the next command
After communication returns, an old packet may be mistaken for evidence that the cube is still held. Alternatively, the grasp may be valid, but resuming the old absolute-time trajectory skips the portion that was not executed during the hold. A gate addresses the first problem; path handling addresses the second.
| Gate | wallclock |
replan |
|---|---|---|
receipt |
Resume the original schedule on any valid delivery | Regenerate the remaining path on any valid delivery |
revalidate |
Resume the original schedule after fresh evidence persists | Regenerate the path after fresh evidence persists |
A fifth variant, latched/wallclock, never resumes. A second path variant would add nothing to that reference. Here wallclock means the original trajectory's absolute simulation time, not the computer's physical clock.
2. What the 30 rollouts control
Five combinations × six conditions produce 30 deterministic cells, each run once without a random seed. A separate set of 18 E5 regression rollouts checks compatibility with the earlier experiment; those are not additional E6 samples.
| Item | Frozen setting |
|---|---|
| Scene | Same Cartesian gripper, cube, tray and initial state |
| Physics/control interval | 2 ms |
| Observation interval | 20 ms |
| Fault onset | 4.8 s, during transfer |
| Delivery gap | [4.8, 5.04), 240 ms |
| Stale threshold | Capture age reaches 60 ms |
| Recovery confirmation | At least 100 ms of valid captures, gaps ≤20 ms |
| Maximum hold | 600 ms; deadline wins over same-tick recovery |
| Regenerated transfer | 1.5 s, then lower/release/retreat/settle |
| Episode horizon | 10.8 s |
Conditions are clean, constant 40 ms delay, every third packet delayed 80 ms to cause reordering, a delivery gap, old-good-packet replay after the gap, and a gap with forced gripper opening. Opening changes actuator commands; contact and gravity move the cube. Its position is never teleported.
The supplied simulations ran on Linux with Python 3.12.14, MuJoCo 3.3.7 and NumPy 2.2.6. Source, logs and tests were independently checked on Windows during integration. That audit is not another physical simulation batch.
3. A target jump is not a gripper teleport
The resume decision occurs after the physics step at tick r. That row still uses the held target; tick r+1 applies the first command from the resumed path.
J = || target[r+1] - target[r] ||₂
E = || target[r+1] - measured_hand[r] ||₂J measures the change in the commanded target. E measures its distance from the actual gripper position at the decision. Both use XYZ only, excluding finger opening. Actual movement depends on actuator tracking and physics; a target jump does not imply instantaneous movement of the same distance. Logged hand speed is a 2 ms finite difference, not a complete velocity, acceleration or impact safety assessment.
The relevant implementation is small:
if path_mode == "wallclock":
return path, None
previous = np.array([*hand, float(grip)])
rest = [("transfer", 1.5, [.24, .12, .18, .033])] + base.schedule(0)[5:]
return segments(rest, tick * DT, previous), previousReplanning starts at the measured position and resets the remaining schedule. It is neither obstacle avoidance nor inverse kinematics. It changes timing as well as geometry.
4. Same gate, different recovery path
These comparisons hold the condition and gate fixed and examine the first resume only. The audit verifies matching state/control histories through the decision.
| Condition | Gate | Original schedule J/mm | Replanned J/mm |
|---|---|---|---|
| Gap | Receipt | 50.285 | 0.263 |
| Gap | Revalidate | 71.433 | 0.245 |
| Stale replay | Receipt | 50.285 | 0.263 |
| Stale replay | Revalidate | 97.055 | 0.245 |
The replanned/original ratios are 0.25%–0.52%, below the preregistered one-half threshold in all four pairs. These are descriptive deterministic comparisons, without a confidence interval or a claim about other robots.

The panels use different horizontal ranges to keep small values readable. Compare the printed values, not bar lengths across panels. Both the figure and table come from metrics.json.
Why does revalidation produce a larger jump when paired with the old schedule? It waits longer while the original trajectory clock keeps advancing. Under stale replay, revalidation resumes at 5.302 s and targets a later point on the old path. Better evidence does not automatically produce continuous commands.

This is a log plot, not camera footage. The upper panels show world x; the lower panels show latest capture age. target and measured refer to commanded and actual positions.
5. A fresh good packet is not a full confirmation window
The legacy unsupported_resumes field counts a stale latest packet (age ≥60 ms) or an invalid grasp predicate. It does not include every requirement of the 100 ms window.
At the end of a gap, receipt-based recovery sees a fresh good packet, making that count zero. But one packet spans zero milliseconds. missing_full_window_resumes records the missing confirmation separately.
| Condition/gate | Resumes per path | Stale/bad resumes | Missing full window |
|---|---|---|---|
| Gap / receipt | 1 | 0 | 1 |
| Replay / receipt | 9 | 8 | 9 |
| Gap / revalidate | 1 | 0 | 0 |
| Replay / revalidate | 1 | 0 | 0 |
Revalidation requires captures after the current hold, strictly increasing capture order, fresh arrival, gaps ≤20 ms and a span ≥100 ms. The latest packet must still be fresh and support the grasp at resume. The grasp predicate requires bilateral contact, cube height >0.12 m and gripper-center distance <0.05 m. Zero in one failure counter does not establish every other prerequisite.
6. Replanning does not recover a dropped cube
| Gap-and-drop combination | Resumes | Stale/bad resumes | Placement |
|---|---|---|---|
| Receipt + original schedule | 6 | 6 | Not completed |
| Receipt + replan | 72 | 72 | Not completed |
| Revalidate + original schedule | 0 | 0 | Not completed; abort at deadline |
| Revalidate + replan | 0 | 0 | Not completed; abort at deadline |
It is tempting to interpret 72 versus 6 as a pure penalty from replanning. However, the inherited gate monitors only transfer. The original schedule reaches lower sooner, after which transfer monitoring no longer triggers holds. Every replan resets the transfer segment and keeps the system exposed to that loop longer.
The totals therefore combine monitoring scope, phase duration and closed-loop feedback. They reveal a real limitation—monitoring one phase is not whole-task protection—but cannot isolate a geometric effect. Both revalidation variants abort after the 600 ms hold deadline without resuming. A protocol-compliant abort is neither placement success nor regrasping.
7. Full matrix and timing cost
| Condition | Latched | Receipt/original | Receipt/replan | Revalidate/original | Revalidate/replan |
|---|---|---|---|---|---|
| Clean | Placed | Placed | Placed | Placed | Placed |
| 40 ms delay | Placed | Placed | Placed | Placed | Placed |
| Reordering | Placed | Placed | Placed | Placed | Placed |
| Gap | Aborted | Placed | Placed | Placed | Placed |
| Stale replay | Aborted | Placed | Placed | Placed | Placed |
| Gap and opening | Aborted | Not placed | Not placed | Aborted | Aborted |
Placement checks the final 0.5 s against the original tray-position and height tolerances, speed <0.02 m/s, no finger contact and a previous lift. Reaching the neighborhood of the target is insufficient.
All four active combinations placed the cube in all five non-drop conditions: this matrix shows no placement-rate improvement. Replanning also delayed some completions. For stale replay with revalidation, the endpoint of the first sustained placement window moved from 7.322 s to 8.622 s. This estimate has 20 ms sampling resolution.
8. Audit and reproduction
The independent audit does not import Gate. It recalculates packet, state and command evidence: 30 cells, 180 raw episode files, 12 source/protocol fingerprints, 24 latched-reference prefixes and 12 same-gate path prefixes. Maximum prefix difference is zero. It checks 102 resume records; all four actual revalidation resumes have valid confirmation windows.
The 18 E5 regression episodes match the old discrete outcomes. Maximum qpos/qvel/ctrl difference is approximately 1.09×10⁻¹⁴, below the 1×10⁻¹⁰ threshold. Artifact validity is separate from hypothesis support: valid negative results must be retained.
From the repository root, audit saved evidence first. The portable wrapper restores only line endings that match known hashes, in a temporary copy:
.venv/Scripts/python -m unittest discover -s experiments/vl01_recovery_ablation -p 'test_*.py'
.venv/Scripts/python experiments/vl01_recovery_ablation/portable-audit.py evidence/recovery-ablation-20260930/e6-finalAll 30 local tests passed. To simulate again, choose an output directory that does not exist:
.venv/Scripts/python experiments/vl01_recovery_ablation/run.py --out runs/e6-my-run
.venv/Scripts/python experiments/vl01_recovery_ablation/audit.py runs/e6-my-run
.venv/Scripts/python experiments/vl01_recovery_ablation/render.py runs/e6-my-runOn Linux use .venv/bin/python. control.csv records every 2 ms; trajectory.csv and states.jsonl every 20 ms. Inspect decisions in events.json and new path starts in replans.json. The archive commit pins the delivered artifacts; manifests retain the original baseline commit and dirty status rather than rewriting their provenance.
9. What should change next?
Define monitoring and recovery rules for lowering and release, then test a resume limit or cooldown and different fault phases under a frozen protocol. These are open questions, not implemented protections.
The experiment still uses one scene, privileged simulator observations, one clock and fixed faults. It does not validate vision, ROS2, clock skew, regrasping, learned policies or physical robots. M4/M5 remain incomplete. Gating determines whether continuation is supported; replanning improves command continuity in these records. Neither alone establishes safe task completion.
A note to your future selfBefore you move on, keep one thought of your own.
Your private reading note stays in this browser; it is never uploaded or published. Clearing browser data removes it, so export a copy to keep. This article’s translations share the same note.
KEEP EXPLORING
Related content
- Engineering cases
Embodied practice 5: communication is back—may transfer resume?
18 MuJoCo rollouts reveal unsupported resumptions hidden by eventual success; validate fresh evidence and replan the remaining path.
- Engineering cases
Embodied practice IV: messages arrive, but observations expire
45 MuJoCo rollouts separate sampling intervals, bad-sample counts and capture age, showing how old messages can hide grip loss and why duration thresholds still have limits.
- LAB / 012
E6 / Recovery gate and path ablation
Separate revalidation from replanning: which differences concern evidence and which concern the resumed target?
- PROJECT / 003
Hohoo Embodied Experiments
Seven MuJoCo teaching experiments, from pickup bias to stage contracts, with raw trajectories, decisions and independent audits.
Which way next?
- Embodied practice IV: messages arrive, but observations expire →
45 MuJoCo rollouts separate sampling intervals, bad-sample counts and capture age, showing how old messages can hide grip loss and why duration thresholds still have limits.
- Embodied practice 5: communication is back—may transfer resume? →
18 MuJoCo rollouts reveal unsupported resumptions hidden by eventual success; validate fresh evidence and replan the remaining path.
Based on article relationships and published same-topic content, not random recommendations.