Reward AI Releases OM-1: A Robot Policy Trained on Human Demonstrations Only, With No Teleoperation or On-Robot Data
Reward AI, a robotics startup whose staff’s prior work consists of DexCap, HumanPlus, and ALOHA, has launched OM-1, quick for Omnibody Model 1. OM-1 is a general-purpose manipulation coverage that learns from people carrying a sensorized glove, then runs on industrial arms and humanoids at human velocity. The key findings that stands out: no teleoperation knowledge and no on-robot knowledge go into coaching. The system follows one precept, ‘One Model, One Data Interface, Any Body,’.
Is it deployable? No, OM-1 is Reward AI’s in-house coverage. No weights, code, dataset, or API have been launched, so builders can’t run it on their very own {hardware} but.
Why Skip Robot Data?
Most robotic basis insurance policies prepare on teleoperated or self-collected robotic knowledge, which binds the dataset to 1 embodiment. Reward AI argues that human-level manipulation won’t come from extra of this knowledge or extra compute, citing Anderson’s “More Is Different.” Instead, seize, studying, and management are designed as one pipeline, so demonstrations recorded at the moment can prepare robotic our bodies that don’t exist but.
Omnibody Hand: A 7-DoF Wearable
The stack begins with Omnibody Hand, a wearable that extends the staff’s earlier DexCap work on transportable movement seize. Rather than copying the human hand joint by joint, it’s a seven-degree-of-freedom design constructed across the features that matter: selecting contact factors, reorienting objects in-hand, and transferring between precision and energy grasps. It captures thumb-index pinching, thumb and index flexion, and paired movement of the center, ring, and little fingers on the MCP joints.
Ergonomics is handled as a data-quality subject: a tool that slips or constrains the wearer produces a compensated grasp. A distal flexion mechanism absorbs variations in finger size, so no per-user adjustment is required.
One Data Interface: Capturing Contact at Human Speed
One Data Interface turns wearer movement into coaching knowledge with no staged setup and no supervisor. The design goal is conveyor-belt sorting, the place an individual spots, grasps, and tosses an object in a fraction of a second. To cowl the entire interplay, the glove combines high-frequency tactile sensing, proximity sensing for the pre-contact strategy, and global-shutter in-hand cameras that maintain context via fast movement.
Hand pose monitoring is the place Reward AI studies its first quantitative outcome. Visual-inertial monitoring is the frequent default, however its accuracy at quick reversals is capped by the visible replace price. Reward AI augments it with electromagnetic sensing plus disturbance compensation. Moving each trackers between two mechanical stops at eight speeds from 3 to 67 cm/s, averaged over ten runs every, electromagnetic monitoring rose from about 0.4 mm to 9.5 mm of imply overshoot error, whereas visual-inertial rose from about 2.1 mm to 24.9 mm: a 60% discount on the highest velocity, with a narrower run-to-run unfold. Force is recorded alongside the identical trajectory, so demonstrations carry effort in addition to path.
OM-1: One Policy, Single-Stage Training
OM-1 learns to generate robotic actions straight from human movement somewhat than routing conduct via an intermediate robotic. Because each demonstration arrives in the identical format, there is no such thing as a break up between pre-training and post-training: the primary demonstration ever recorded and the latest one prepare a single coverage in a single stage.
Inputs are the glove’s multimodal streams: photographs, tactile indicators, inter-finger proximity, and hand pose trajectories. Each modality is processed at its sensor’s native sampling price somewhat than downsampled to a typical frequency, so high-frequency tactile and movement cues survive alongside lower-frequency imaginative and prescient. Outputs carry movement course, velocity, drive, and the timing of occasions reminiscent of grasp initiation. Reward AI says it constructed a novel structure for environment friendly inference, although architectural particulars and parameter counts are usually not disclosed.
Control Any Body: An RL Layer on Its Own Clock
Below the coverage sits a high-frequency management layer skilled with reinforcement studying in simulation to deal with velocity- and acceleration-dependent dynamics, exterior disturbances, and system delays. Where a classical controller pushed off its reference by an sudden load by no means recovers, this layer holds the reference and settles again, which is what lets a robotic open a completely closed fridge door or raise containers of unknown weight.
The management layer runs on its personal clock, persevering with whereas the coverage computes the following actions, so inference latency by no means stalls movement. Because successive predictions could not be a part of easily, it optimizes the transition between them on-line. The identical motion area covers manipulation and navigation for cellular robots.
Results
Reward AI studies that OM-1 picks up a brand-new process, together with difficult dynamics and lengthy horizons, from lower than half-hour of human knowledge, and attributes this to the built-in stack somewhat than the coverage alone. Its about page states that every one printed clips run at 1x velocity and that the mannequin spans arms, legged humanoids, and wheeled cellular manipulators. No success charges, public-baseline comparisons, or paper have been launched, so these claims are demonstration-backed somewhat than benchmark-backed.
Key Takeaways
- OM-1 trains solely on human demonstrations from a wearable 7-DoF glove; no teleoperation or robotic knowledge.
- Electromagnetic monitoring lower imply overshoot error 60% at 67 cm/s versus visual-inertial (9.5 mm vs 24.9 mm).
- Each sensor stream retains its native price; the RL management layer runs on its personal clock.
- A new process is discovered from below half-hour of information, per Reward AI.
Check out the Technical details here. All credit score goes to the researcher of this venture. Also, be at liberty to comply with us on Twitter and don’t neglect to affix our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Need to associate with us for selling your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar and many others.? Connect with us
The put up Reward AI Releases OM-1: A Robot Policy Trained on Human Demonstrations Only, With No Teleoperation or On-Robot Data appeared first on MarkTechPost.
