The Catalog Dependency Problem
Traditional industrial vision systems work off a model library. You train or scan each object you intend to handle, add it to a catalog, and the vision controller matches incoming frames against that catalog to locate and pose-estimate the target. This works well when your part set is stable, your supplier is consistent, and your operations team has time to maintain the library when anything changes.
Most manufacturing floors do not have all three. Part revisions slip in mid-production run. Supplier substitutions arrive with slightly different geometry than the original drawing. Consumer-electronics assembly operations introduce new components every few months. Each one requires adding to the catalog, which means photography sessions, CAD-to-scan alignment, and revalidation of the vision system on the new object. In practice, there is always a backlog of uncatalogued parts waiting on a robotics engineer's desk.
The downstream cost is not just engineering time. When an uncatalogued object appears on the belt or in the bin, the vision system either fails to detect it and the arm skips the task, or it generates a false-positive match against the closest catalog entry and attempts a grasp with the wrong parameters. Either outcome stops production and requires manual intervention.
What Foundation Model Perception Changes
Foundation models trained on large, diverse image datasets have learned feature representations that generalize across object categories. They can characterize the shape, surface texture, and geometric properties of an object they have never seen in isolation, because those properties map onto shared feature space learned from millions of training examples. This is qualitatively different from classical template matching or even purpose-trained CNN classifiers.
For novel object recognition in manipulation, this means the perception module can describe an object it has never seen before in terms that are useful for planning a grasp. It cannot always name the object, but it can estimate the object's graspable geometry: convex hull, principal axes, likely center of mass, stable surface normals for gripper placement. For a pick-and-place task, that is often sufficient.
The model can also accept a natural language description as the grounding signal. Instead of a catalog match, the operator types "pick the cylindrical black plastic cap" and the system identifies the closest matching object in the scene using the text description as a semantic anchor. No scan, no CAD model, no training samples of that specific part.
Where This Works Well and Where It Does Not
Novel object recognition works best on objects that are distinct enough in the scene to be localized without ambiguity. A single unfamiliar component on a mostly clear workbench is the ideal case. The foundation model can describe its geometry, estimate grasp poses, and execute. We have tested this with components ranging from irregular cast metal housings to unfamiliar consumer product packaging, with consistent results in our internal benchmarks during early-access pilot evaluations.
It works less reliably in three categories of situations. First, dense cluttered bins where multiple similar-looking novel objects are present simultaneously. The scene graph reasoning layer handles some of this by maintaining object identities across frames, but when 20 identical-looking novel fasteners are tumbled together in a bin, the perception module's grasp quality estimates degrade. Second, objects with highly specular surfaces, including polished metal and glossy plastic, where depth estimation from structured light is unreliable and the RGB channel alone is insufficient for stable pose estimation. Third, very small objects: below approximately 15 mm in the longest dimension, our current pipeline loses reliable pose estimates at typical working distances.
These are not novel findings. They are constraints that apply broadly to vision-based manipulation systems. The difference is that with a catalog-based system, you know in advance which objects are supported. With open-vocabulary recognition, the failure boundary is fuzzier. An operator may try to run a task on an object that happens to fall near a failure mode without expecting it to fail. We try to address this with confidence-score outputs that trigger a human verification request when the grasp confidence falls below a task-configurable threshold.
The Grounding Interface
The interface between the operator and the perception module matters as much as the perception model itself. There are three grounding modes we expose through the EmbodyX SDK.
Text grounding: The operator provides a natural language description of the target object. The system returns the top-k candidates from the current scene with confidence scores. This is the most flexible mode but requires the description to be specific enough to discriminate among scene objects.
Visual exemplar grounding: The operator provides a reference image of the target object, either a photograph or a screenshot. The system uses this as a visual anchor rather than a text description. This is useful when the object is easy to photograph but hard to describe in precise words, for example an irregular casting with no standard geometric name.
Bounding box grounding: For cases where an operator is watching a live camera feed and wants to direct the arm to a specific visible object, they can draw a bounding box around it. The perception module treats the boxed region as the grasp target without requiring any classification or matching. This is the most deterministic grounding mode and is useful for debugging and initial deployment testing.
In practice, most early-access pilot users defaulted to text grounding for routine operations and bounding box grounding for initial system setup and edge cases.
Integration with the Task Planner
Novel object recognition is most useful when it is integrated tightly with the task planning layer, not just the perception layer. The reason is that recognizing an object and planning a task-appropriate grasp are different problems. The same object may need to be picked by its base flange for stable transport or by its lateral surface for insertion into a fixture. Without task context, the grasp planner can only optimize for generic stable grasps.
EmbodyX passes the natural language task description into both the perception and planning layers. The task description is the grounding signal for object recognition and simultaneously the conditioning input for the motion planner. "Pick the valve housing and orient it for insertion in slot B" carries information that affects which surface to grasp, the approach angle, and the insertion alignment target. Keeping these two grounding signals synchronized is one of the architectural decisions that distinguishes a VLA system from a perception system bolted onto a traditional motion planner.
A Note on Catalog Maintenance
Open-vocabulary recognition does not eliminate catalog maintenance entirely. For high-volume repetitive tasks where the same part is handled thousands of times per shift, a few demonstration runs on the specific part still improve grasp reliability compared to zero-shot generalization. Fine-tuning the perception model on task-specific data is worthwhile when you have that data. The value of open-vocabulary capability is not that fine-tuning is no longer useful. It is that fine-tuning is no longer required before you can run the arm at all. The system is operational on novel parts from the first attempt, and it gets better as you provide feedback and demonstrations.