The Closed-World Assumption
Rule-based robot programming is built on a closed-world assumption: everything the arm needs to know about its environment is enumerated before deployment. The programmer specifies the target objects, their positions, their orientations, the acceptable tolerance ranges for each parameter, and the exception handlers for each out-of-range condition. Within that enumerated space, the arm executes reliably. Outside it, the arm halts.
Factory floors are not closed worlds. They are open worlds with a continuous stream of perturbations: supplier changes, operator behavior, equipment wear, seasonal temperature effects on part dimensions and conveyor speeds, shifts in lighting as building fixtures age. Every one of these perturbations is either within the enumerated space (handled correctly), within the tolerance bands (handled, sometimes incorrectly), or outside the enumerated space entirely (halts).
The programmers who built those rule trees knew all of this. They did their best to enumerate a large space. The problem is not that they enumerated too few cases out of laziness. The problem is that enumerating the full variance of a real production environment is not a tractable engineering task. The space is too large, and it changes over time.
Where the Failures Actually Come From
I have spent years working through integration projects in factory operations. The failure patterns are remarkably consistent across different industries and controller types. They cluster into four categories.
Part variation accumulation. Suppliers change part geometry slightly between production runs, staying within their own spec tolerances but outside the robot programmer's assumed range. An M8 bolt from Supplier A has a head diameter of 12.8 mm. The replacement order from Supplier B delivers heads at 13.2 mm. Both are within M8 spec. The robot's grasp width parameter is tuned for 12.8 mm with a 0.5 mm tolerance. At 13.2 mm, the gripper closes on the bolt's rim rather than the body. This is not an edge case. It is the baseline behavior of a supply chain over time.
Fixture drift. Mechanical fixtures shift over thousands of production cycles. Bolts loosen. Wear patterns develop on locating pins. A fixture that placed a part at position (245.3, 118.7, 50.0) mm in September is now placing it at (246.1, 119.0, 49.8) mm in February. The robot's approach vector was programmed to 0.5 mm tolerance. The cumulative drift is now outside that tolerance. The programmer sees a reliable station that has "suddenly" developed intermittent failures with no apparent cause.
Environmental sensor interference. Machine vision systems are sensitive to lighting. A facility installs new LED fixtures in one bay. The spectral distribution shifts. A vision system that was calibrated under the previous metal halide lighting now produces systematically biased intensity readings. The threshold that correctly distinguished part-present from part-absent under the old lighting now generates false positives at a 4% rate under the new lighting. Four percent sounds small until you are running 1,500 cycles per shift.
Unenumerated task states. An operator places an empty container in the staging area rather than a full one, because a delivery was late. The robot's rule tree has no state for "empty container." It attempts to pick from an empty bin and either faults on a grasp failure or, worse, picks the container itself. The programmer never enumerated "empty container" as a handled state because it was not supposed to happen during normal operations.
Why Widening the Tolerances Does Not Solve the Problem
The first-order response to rule-based failures is usually to widen the tolerance bands. If the system is failing at 0.5 mm tolerance, make it 1.5 mm. This solves the immediate failure but creates a new problem: the task now runs with wider tolerances, which degrades precision on the cases where precision matters. For assembly operations with tight fitting clearances, wider tolerances on part pickup translate directly to higher insertion failure rates downstream.
The second-order response is to add more exception handlers. If you can characterize the failure modes you are seeing, you can add handling for them. This works until the next failure mode appears, and the rule tree grows with each new exception. Mature rule trees in production facilities are typically 30 to 50% exception-handling code by line count. They are difficult to maintain, difficult to debug when interactions occur between handlers, and brittle in ways that are hard to predict.
Neither response addresses the root cause: the rule tree's closed-world assumption does not match the reality of the environment it is running in.
What a Different Class of Solution Looks Like
Vision-language-action models take a different approach. Instead of enumerating the state space and writing rules for each state, a VLA model learns a mapping from raw sensor observations to task-appropriate actions from demonstration data. The model's implicit representation of the task is learned, not enumerated. This means it generalizes to states it was not explicitly shown, as long as those states are within the distribution of variation the model encountered during training and fine-tuning.
This is not the same as saying VLA models never fail. They fail differently. A rule-based system fails sharply at the boundary of its enumerated space. A VLA model's failure rate increases gradually as the task state moves further from the training distribution. That graduated failure mode is, in many contexts, more manageable than a sharp cliff: you see task completion rate declining in the telemetry before you see outright failures, which gives you time to retrain rather than reacting to a production halt.
EmbodyX is a VLA platform specifically built for industrial manipulation. The base model was trained on diverse industrial manipulation data. When you run EmbodyX on your facility, fine-tuning on your specific parts and task conditions brings the model's behavior closer to your operational envelope. The result is a system that handles part variation, lighting variation, and fixture drift without requiring a programmer to enumerate every case.
Where Rule-Based Programming Still Makes Sense
This is worth being direct about. Rule-based robot programming is not inherently bad engineering. It is the right tool for specific contexts: tasks with truly constrained, stable environments; high-volume operations with narrow part variation tightly controlled by supplier agreements and incoming inspection; tasks where the cost of a failure is so high that the narrow, well-characterized behavior of a rule-based system is preferable to a probabilistic model with a tighter but non-zero tail. Automotive welding on a fixed-geometry jig is not a problem that needs VLA.
The failure mode we are describing applies specifically to tasks with meaningful real-world variation: bin picking with supplier-sourced parts, mixed-SKU sortation, flexible assembly with varying product lines. These are the tasks where the closed-world assumption creates an increasing maintenance burden over time, and where the investment in a perception-first approach pays back in reduced downtime and reduced reprogramming cycles.
The Practical Migration Path
You do not need to replace your entire robot programming stack to start moving away from brittle rule trees. The most productive approach is to identify the top two or three tasks in your facility that generate the most e-stop events and manual interventions, and pilot a perception-first approach on those specific tasks while leaving the rest of your automation running as-is. If those tasks run more reliably with fewer interventions, you have a concrete data set for the decision about whether to expand.
That is the exact structure of our pilot program: one facility, one to three tasks, 30 days of real production data. By the end you know whether the task completion rate on your specific parts in your specific environment is better than what you were getting from the rule tree. The comparison is against your actual current performance, not a lab benchmark.