Back to Insights
Lars Eriksson

Failure Recovery for Industrial Robots: Detecting and Responding to Unexpected States

A dropped part, a slipped grasp, a component that rotated mid-transfer. Recovery strategies matter as much as initial task execution. Here is how VLA models handle failure states.

Industrial robot arm paused mid-operation with warning indicator, factory safety lighting in background

The Gap Between Task Success and Production Uptime

Most robot programming evaluations measure task success rate under nominal conditions. The arm picks the part, places it correctly, the counter ticks up. But production uptime is determined by a different question: what happens when the task does not go nominally?

Parts drop. Grasps slip. Components shift in the gripper mid-transfer. A preceding station delivers a part with different orientation than expected. The arm bumps a fixture and the target position is now 8 mm off the planned path. Each of these events is a failure state: the robot's world model no longer matches reality, and continuing to execute the original motion plan will either damage the part, miss the target, or trigger an e-stop.

Rule-based systems handle this with exception handlers: if sensor X reads out of range, halt and alert. The coverage is whatever the programmer thought to enumerate. The states that were not enumerated are not handled. When I was commissioning automation systems earlier in my career, "not enumerated" was a polite way of saying "whoever's on shift at 2am handles it manually."

The question EmbodyX addresses is whether a VLA model can do better: not just execute tasks but detect when task state has diverged from expectation and respond productively.

Failure State Detection: The Taxonomy We Work With

For practical recovery planning, we categorize failure states into four types, based on how they are detected and what recovery actions are available.

Grasp failure: The gripper closed but the part was not acquired, or was acquired with insufficient grip force to maintain control. Detection sources: finger-position feedback from the gripper controller (if available), wrist-mounted force-torque sensor reading that does not match expected grasp contact signature, camera verification showing target object still in its pre-grasp location after gripper close. The most reliable detection combines at least two of these signals.

In-hand slip: The part was grasped but shifted orientation or position during transport. This is harder to detect than grasp failure because the object is in the gripper, not visible to the workspace cameras in most configurations. We use wrist-mounted cameras where the workspace configuration permits, but in tightly constrained cells this is not always feasible. Force-torque signatures during transport can also indicate slip events, particularly for parts with defined contact geometry.

Placement failure: The arm executed the place motion but the part did not land in the target state. Detection: camera verification of the target location after nominal place. If the scene graph does not register the part in the expected position and orientation within a time window after the place motion completed, a placement failure is flagged.

Environment state change: Something else in the workspace changed in a way that invalidates the current plan. A preceding process deposited a part in the target location. A fixture plate shifted. Another arm occupied the approach path. These are detected by comparing the current scene graph against the expected state the planning layer had at task initiation.

What the Recovery Layer Actually Does

Recovery is not a single action. It is a re-planning problem. When a failure state is detected, the VLA model's reasoning layer re-evaluates the current scene graph against the original task specification and generates a recovery action sequence rather than a fresh task execution sequence. The distinction matters: recovery actions are constrained by the current physical state, not the ideal starting state.

For grasp failures, the most common recovery is a revised grasp attempt with adjusted approach parameters. The model receives the updated scene (including the still-present target object) and the information that the previous grasp failed with the specific parameters it used. It generates a new grasp pose, typically with a modified approach angle or grip width. In our early-access pilot evaluations, grasp-retry success on first recovery attempt was substantially higher when the retry used an adjusted approach compared to retrying with identical parameters.

For in-hand slip, recovery depends on whether the slip was detected before or after the place motion began. Pre-place detection: the model can interrupt the transfer, return to a stable position, open the gripper to release and re-grasp. Post-place detection: the part was placed in an incorrect state, which now becomes an environment state change problem for the next action in the sequence.

For environment state changes, the recovery scope can be larger. If a fixture is occupied by a part from a previous failed cycle, the recovery sequence may need to clear the fixture before the current cycle can proceed. The model checks whether a clearing action is within its task scope, and if not, raises a human verification request rather than attempting an out-of-scope action.

The Human-in-the-Loop Threshold

Not all failures should be recovered autonomously. There are two categories of failure where the right response is stopping and requesting operator input.

First, repeated failure on the same action. If a grasp attempt fails twice with different approach parameters, something is wrong that the perception and planning data is not capturing. Continuing to retry is not productive and risks part damage. The system flags the task and waits. This is configurable: some tasks warrant more retries before escalation, others (high-value parts, precision assembly) should escalate after a single failure.

Second, failure states that require out-of-scope actions. The arm's task scope is defined when the task is specified. If recovery from an unexpected state would require the arm to interact with equipment outside its defined workspace, the system should not make that decision autonomously. Safety clearance is the operator's call, not the model's.

We surface this through the SDK as two distinct events: RecoveryAttempt (the system is handling it automatically) and HumanVerificationRequired (the system needs operator input before proceeding). Both events carry the current scene state and the reason for the event, so the operator has the context they need to make a decision without walking to the cell.

Force-Torque Integration and Why It Is Not Optional

Camera-based failure detection has a coverage gap: anything that happens while the part is in the gripper and not visible. Force-torque sensing at the wrist fills this gap. We strongly recommend wrist-mounted force-torque sensors for any deployment where in-hand slip is a meaningful failure mode: small parts, polished surfaces, parts with low friction coefficients relative to the gripper material.

Force-torque data adds a detection channel that is not dependent on line-of-sight. It also improves the quality of grasp success confirmation. A camera can see that the part appears to be in the gripper; a force-torque sensor can confirm that the contact forces match the expected signature for a stable grasp. These two signals together are significantly more reliable than either alone.

The cost of wrist-mounted force-torque sensors varies by arm type. For UR arms, the integrated force-torque option is straightforward. For FANUC and KUKA, external sensors are typically required and add integration complexity. It is worth the complexity for assembly tasks with tight tolerances and high-value parts. For bulk bin picking where part damage is less critical and throughput is the primary metric, camera-only detection may be an acceptable tradeoff.

Calibrating Recovery Confidence Against Production Risk

One thing worth being direct about: autonomous failure recovery is not always safer than immediate halt-and-alert. Recovery actions involve motion in an unexpected state. A recovery grasp attempt in a failure state can move a part further out of position if the recovery parameters are also wrong. For tasks involving fragile parts, hazardous materials, or tight spatial clearances, the risk-adjusted decision may favor halt-and-alert over recovery attempt.

EmbodyX supports configuring recovery behavior per task, including disabling autonomous recovery entirely for designated high-risk tasks. The goal is to give facilities the tools to match recovery autonomy to their specific risk tolerance, not to assert that autonomous recovery is always the right answer.

See EmbodyX on your arms

Schedule a pilot evaluation with your existing FANUC, KUKA, UR, or ABB arms. No new hardware required.