What a Scene Graph Is (and What It Is Not)
A scene graph is a directed graph where nodes represent objects and edges represent relationships between them. In a factory workspace, a scene graph might encode: "valve_housing_A is above conveyor_section_3, which is adjacent to the insertion_fixture, which is currently occupied." The relationships are typed: spatial (above, left_of, behind), functional (supports, occludes, contains), and state-based (grasped, at_rest, in_motion).
This is not a semantic segmentation map. Segmentation gives you object boundaries and class labels per pixel. A scene graph gives you the relational structure of the scene. The difference matters when the task you are executing depends not on which objects are present but on how they are arranged relative to each other.
We are also not talking about 3D scene reconstruction. A scene graph does not need to be a complete geometric model of the workspace. It needs to carry the relationships that are task-relevant, and nothing more. Keeping the graph compact is a real engineering goal, not an afterthought. Overloading the graph with every detectable relationship in the scene increases inference time and tends to introduce noise that complicates planning.
Why Pixels Alone Create Fragile Grasp Planning
Raw pixel streams from RGB-D cameras give the arm a lot of information, but they do not tell it which information is task-relevant. A planning system working directly from pixel tensors has to learn, implicitly, that the position of object B matters only when object A is nearby and oriented a certain way. That implicit learning is possible, but it is data-hungry and the learned dependencies are difficult to inspect when something goes wrong.
The more common failure mode in practice looks like this: an arm trained to pick small cylindrical components from a fixed-position bin handles the task cleanly under normal conditions. A supplier then changes the bin's flange geometry slightly. The bin now sits at a different tilt angle. The approach vector that was learned from thousands of demonstrations puts the gripper 3 cm off the component centerline. The arm does not know the bin's orientation changed because it was never given a representation that separated "bin position and orientation" from "component position and orientation." Everything was encoded jointly in the pixel stream.
Scene graphs impose explicit separation. The nodes are objects. Their positions and orientations are attributes of those nodes. The edges encode how they relate. A change to the bin's flange geometry does not corrupt the arm's understanding of the component's position within the bin. Those are different nodes with a containment edge between them.
How EmbodyX Builds the Graph
EmbodyX's perception module runs on every frame from the wrist-mounted and overhead RGB-D cameras. The first stage is object detection and class assignment using a fine-tuned foundation model. The second stage is pose estimation for each detected object, producing a 6-DoF pose in the robot's coordinate frame. The third stage is relationship inference: for each pair of nearby objects, the model outputs the relevant edge types and their confidence scores.
The graph updates at approximately 8 to 12 Hz in our current deployments. This is sufficient for the task planning loop but is not frame-rate video. For tasks that require tracking moving objects on a running conveyor, we add a velocity attribute to moving object nodes. The planner then accounts for where the object will be at the start of the approach trajectory, not just where it is when planning begins.
One point worth emphasizing: the scene graph is an intermediate representation, not the final output. The VLA model's reasoning module consumes the graph along with the natural language task specification and produces the motion plan. The graph is not handed to the arm controller directly. The controller receives Cartesian motion commands, not graph queries.
Which Relationship Types Matter on the Factory Floor
Based on our work across three pilot facilities, the relationship types that appear most often in task-conditional planning are the following.
Containment: "part is in bin," "assembly is on fixture plate," "fastener is at insertion point." These anchor the approach vector and determine which side of the object the arm should approach from.
Adjacency with orientation: Not just "next to" but "to the left of, rotated 90 degrees." The orientation component often determines which grasp pose is feasible. A bracket that is adjacent to a wall at 0 degrees rotation is accessible from the top. The same bracket at 45 degrees rotation may only be accessible from a diagonal approach that requires collision clearance checks.
Occlusion: "part B is partially hidden by part A." This drives reachability decisions and sometimes requires clearing an obstacle before attempting the primary grasp. The perception module assigns occlusion edges based on visible fraction of the target object's bounding volume. When the visible fraction drops below a task-specific threshold, the planner considers a clearing action.
Kinematic constraint: "bolt is constrained to slot on bracket." This is less common but critical for assembly tasks where a part can only be removed along a specific axis. Representing this explicitly saves the planner from learning it implicitly through failed approach attempts.
We keep the edge vocabulary to 12 core relationship types in the current model version. A graph with 40 relationship types is technically richer but practically harder for the reasoning module to use consistently. The tradeoff between representation completeness and planning tractability is real, and we land on the side of tractability.
Where Scene Graphs Fall Short
We are not claiming scene graphs resolve all perceptual problems in manipulation. There are two categories of tasks where our current representation struggles.
First, highly deformable materials. Cables, fabric, flexible tubing: these do not map cleanly to rigid object nodes with stable pose estimates. The graph representation assumes objects hold their shape well enough to have a meaningful 6-DoF pose. We are working on deformable object extensions, but this is current research, not current product capability.
Second, dense clutter above a certain threshold. A bin containing more than approximately 40 to 50 randomly overlapping small components stresses the relationship inference stage because occlusion relationships become layered and the per-object pose estimates degrade in reliability. The perception module's edge confidence scores drop, and the planner must make decisions with partially reliable structural information. In these scenarios we fall back to simpler representations and rely on uncertainty-aware motion planning rather than rich scene structure.
A Concrete Task: Shaft Collar Insertion with Keyway Alignment
One of our pilot facilities runs a subassembly task where an arm picks shaft collars from a gravity-fed chute and inserts them onto motor shafts staged in a rotary indexer. The collars arrive with random rotational orientation because the chute feed does not control rotation about the shaft axis. The insertion axis is fixed by the indexer, but the collar's keyway must align to the shaft's keyway before insertion or the part jams.
Without explicit relational representation, a perception system has to learn the relationship between collar orientation, keyway position, and insertion outcome implicitly from demonstrations. That requires a large number of demonstrations covering the full rotation distribution, and the learned policy is hard to inspect when it fails.
With the scene graph, the planner receives an explicit edge: "collar_keyway is at X degrees relative to shaft_keyway" on every planning cycle. The task instruction ("insert collar with keyways aligned") grounds directly onto that edge. The planner adjusts the wrist rotation in the approach trajectory accordingly. This does not require the model to have learned the keyway alignment requirement from data; it is encoded as a relationship the planner can read directly.
In our internal benchmarks during the pilot evaluation period, explicit relational encoding improved task completion rate on this specific subtask compared to a baseline approach without scene graph output. We are withholding specific numbers until we have data from a wider set of facilities to contextualize them.
The Representation Question Is Not Settled
Some VLA researchers prefer end-to-end models that learn relational structure implicitly from large demonstration datasets, without explicit scene graph construction as an intermediate step. That approach has real strengths: it avoids the engineering overhead of maintaining a graph schema, and it can capture relationships that are difficult to name explicitly. We built EmbodyX around explicit scene graphs because we believe inspectability matters more to industrial operators than it does to research benchmarks. When a task fails, an operator can look at the scene graph and see: "the part was 12 mm off the expected insertion position." That is a diagnosis, not just an error code. Whether implicit or explicit relational representation is the right long-term architecture is an open question worth following if you work in this space.