What Fine-Tuning Actually Means for a VLA Model
Fine-tuning in the context of vision-language-action models is not the same as retraining a classifier from scratch. The base model has already learned a broad representation of physical manipulation from diverse training data. Fine-tuning updates the model's weights on your specific task data, preserving the broad capability while specializing the model's behavior on your parts, your lighting conditions, your grasp geometries, and your task sequences.
For EmbodyX specifically, fine-tuning operates on the perception and action generation layers together. You are not separately tuning the vision encoder and the motion planner. The VLA architecture means that task-specific data improves both how the model interprets your scene and what motions it generates in response to it. This is one of the places where VLA architecture gives you something that decoupled perception-plus-planner systems cannot match as efficiently.
The practical result is that a fine-tuned model on your specific task handles edge cases that the base model struggles with: unusual part orientations, your facility's specific lighting spectrum, gripper-part contact geometry that your tooling requires. On our internal benchmarks across three pilot facility evaluations, fine-tuned models showed meaningfully better task completion on the specific trained tasks compared to the base model, with the gap being most pronounced on tasks involving parts with low visual distinctiveness (e.g., many similar small components in a bin).
Demonstration Data: What You Collect and How Much
The primary input to fine-tuning is demonstration data: the arm performing the target task correctly, with the scene cameras recording each attempt. The data consists of synchronized streams: RGB-D frames from the scene cameras, joint state and Cartesian pose logs from the arm controller, gripper state, and the natural language task instruction associated with each demonstration run.
The minimum viable demonstration count for a single task is approximately 50 to 80 complete demonstrations. Below 50, the model does not have enough variance coverage to generalize to the full distribution of part positions and orientations it will encounter in production. Above 200 demonstrations, you typically see diminishing returns unless the task involves unusual complexity or very high part variation. For tasks with two or more distinct subtasks in sequence, scale accordingly: a task with three chained subtasks benefits from 150 to 250 demonstrations that cover the full sequence, not just each subtask in isolation.
The distribution of demonstrations matters as much as the count. If all your demonstrations have the target part in a similar position range, the fine-tuned model will overfit to that range. Deliberately vary the part position, orientation, and position within the bin across your demonstration runs. Cover the full workspace volume that the arm will encounter in production. If your facility runs multiple part variants, collect demonstrations on all variants rather than one representative part.
Collecting Good Demonstration Data in Practice
The SDK includes a demonstration recording mode that handles data capture automatically. You set up the task specification, put the arm in demonstration mode, and execute the task manually (via kinesthetic teaching or teleoperation). The SDK records all sensor streams, timestamps them against the controller clock, and stores them in a format ready for fine-tuning pipeline ingestion.
A few things we have learned from watching people collect data in pilot facilities:
Consistency in task instruction language matters. If you describe the task as "pick the bolt and insert it" in some demonstrations and "grasp the M6 bolt and place it into slot A" in others, the model sees these as partially different tasks. Settle on the canonical instruction text before you start collecting and use it exactly every time. You can add variation later in a second data collection pass if you want the model to handle instruction paraphrases.
Failed attempts are useful. The demonstration recording mode has an option to flag a demonstration as a failure and record it anyway. Failure data with labeled failure reasons (slip, miss, incorrect orientation) is valuable for training the recovery detection layer. Do not discard failed attempts.
Lighting variation is easy to neglect. Factory lighting changes with shift changes, weather, and equipment status. Collect demonstrations across at least two different lighting conditions: typical production lighting and the lower-light condition that occurs at shift change or when overhead fixtures are partially off. A model trained only under ideal lighting may not transfer to the 6am shift without this.
The Fine-Tuning Process: What Happens on Our Side
Once you upload your demonstration dataset through the EmbodyX platform, the fine-tuning pipeline runs on our infrastructure. For customers on the Team tier, two tasks are included per billing period. The process takes between 4 and 12 hours depending on dataset size and current compute queue depth, and we send a notification when the fine-tuned model version is available for deployment.
The fine-tuned model is versioned separately from the base model and can be rolled back to the base model or a previous fine-tuned version without re-uploading data. We maintain each facility's fine-tuned model versions for 6 months. This matters for regression scenarios: if a part revision changes the task and the fine-tuned model now underperforms the base model on the new part, you can revert to the base while collecting new demonstrations on the revised part.
We run automated validation on your fine-tuned model before making it available for deployment. The validation checks that the fine-tuned model has not catastrophically forgotten base capabilities (a known risk with aggressive fine-tuning), and that its task-specific performance on a held-out slice of your demonstration data meets a minimum threshold. If validation fails, we flag it and contact you before deploying. This has happened in a small number of cases where the demonstration dataset had systematic labeling issues; we have not yet seen a case where the fine-tuning itself produced a clearly worse model on a clean dataset.
Fine-Tuning for Multiple Tasks
If your facility runs multiple distinct tasks on the same arm, you have two options for fine-tuning: task-specific models (a separate fine-tuned model per task), or a multi-task fine-tuned model that handles all tasks from a single model version.
Task-specific models have higher precision on each task but require managing multiple model versions and switching between them at runtime. The SDK supports runtime model selection per task specification, so this is manageable, but it adds operational overhead. Multi-task fine-tuning requires combined demonstration data from all tasks in a single fine-tuning run and is typically appropriate when the tasks are related enough to benefit from shared feature representations (e.g., several variants of a bin picking task on related part families) or when model switching overhead is not acceptable in your task scheduling.
For most facilities in our pilot set, we have recommended starting with task-specific models for the two to three highest-volume tasks, validating those, and then evaluating whether multi-task fine-tuning makes sense once the single-task models are working well. Multi-task fine-tuning is harder to debug when something goes wrong; narrowing the failure scope is easier with task-specific models during initial deployment.
When Fine-Tuning Is Not the Right Answer
Fine-tuning requires a meaningful upfront investment in data collection and process time. For tasks that run only a few hundred cycles per month or tasks where the part variation is very low, the base model may perform adequately and the fine-tuning investment is not justified. We try to help customers assess this during the evaluation phase rather than defaulting to fine-tuning as the solution to every performance gap. Sometimes the right answer is better lighting, a different camera mounting position, or a task instruction that is more precisely worded. These are cheaper and faster than a fine-tuning run.