The EmbodyX Platform

From pixels to motion in a single stack

EmbodyX runs a compact vision-language-action model on-arm. The robot sees the workspace, reasons about the task, and generates motion commands directly. No waypoint scripting, no separate perception node, no re-deployment when parts change.

Abstract visualization of vision-to-action neural processing with geometric depth
38 ms Inference latency
94% Novel object grasp success, internal benchmark
12+ OEM arm SDKs supported
<4 hr Typical onboarding
Architecture

Three layers, one inference pass

Every EmbodyX deployment runs the same pipeline. Perception ingests the scene, the VLA model reasons over it, and the action decoder outputs joint-space commands at control frequency.

Perception Layer RGB-D Camera Point Cloud Scene embedding encoder VLA Model Core Language instruction tokenizer Transformer backbone 7B param, quantized INT8 Task-conditioned token routing Action Decoder Joint-space trajectory generation @ 25 Hz OEM SDK adapter FANUC / KUKA / UR / ABB Safety envelope monitor
Capabilities

What the platform handles

Four core capabilities that traditional scripted automation cannot deliver reliably in production.

Novel object handling

The model grasps objects it has never been explicitly trained on. Foundation-model perception generalizes from visual features, not a fixed SKU catalog.

Mixed bin-picking, unstructured infeeds

Adaptive re-grasping

When a grasp attempt fails or a part shifts, the model detects the new state and replans automatically. No operator intervention, no line stop.

Recovery on slip, drop, or tilt events

Language task conditioning

Operators send natural-language instructions that the model parses into task context. Changing what the arm does requires a text update, not a re-teach.

Task handoff between shifts, variant selection

Real-time inference

The quantized model runs on-arm at 38 ms average latency. No cloud round-trip, no network dependency during production. The arm acts at its own control rate.

25 Hz joint command generation
Integration

Connects to arms you already run

EmbodyX ships with adapter modules for the major industrial arm platforms. The SDK translates VLA action output into each OEM's native motion protocol. Typical integration time is under four hours per arm type.

  • FANUC R-30iB and R-30iB Plus controllers
  • KUKA KR C4 and KR C5 via KUKA.Connect
  • Universal Robots UR3e, UR5e, UR10e via URCap
  • ABB OmniCore via RobotWare SDK
  • Yaskawa Motoman DX200 / YRC1000
  • Custom URDF models via open adapter interface
See Integration Quickstart

Supported OEM platforms

FANUC
KUKA
Universal Robots
ABB
Yaskawa
Custom URDF
Python
from embodyx import EmbodyXClient

client = EmbodyXClient(
  arm="ur5e",
  task="pick and place"
)

# Task updates via natural language
client.set_task("sort metal cylinders by diameter")
client.run()
Performance

Benchmark results against rule-based baselines

Based on internal benchmarks across 3 pilot facilities: mixed-SKU industrial grasping, 500 trials per system, 6-week evaluation period.

Metric Rule-based baseline EmbodyX VLA
Grasp success rate (novel objects) 41% 94%
Recovery after grasp failure Manual reset required Automatic, 97% success
Re-teach time on part change 4-8 hours per variant Text update, <5 min
Inference latency (p95) N/A 52 ms
Unplanned downtime (30-day pilot) Baseline -62% vs baseline

Run the platform on your arm

The Evaluation tier is free for 30 days on one robot. Bring your own arm, your own parts, your own task. No integration contract needed to start.