Axis Robotics Open-Sources One of the Largest Franka Arm Simulation Datasets for Physical AI

Axis Robotics has launched Axis Sim Dataset V1, one of the largest open-source simulation datasets for Franka arm manipulation, with the full dataset, coaching code, and benchmarks publicly out there. V1 is constructed from greater than 50,000 human-teleoperated simulation trajectories throughout 207 manipulation duties and 60,000+ scene variants on a simulated Franka Research 3 arm.

This dataset drew over 160,000 downloads, making it the most downloaded open-source simulation Franka manipulation dataset on Hugging Face. In benchmarks, continuous pretraining on V1 lifted π0.5 and beat a volume-matched RoboCasa baseline, with each end result open and verifiable.

Axis Robotics is constructing the final compounding information engine for Physical AI, a vertically built-in system spanning large-scale simulation, selfish real-world seize, humanoid loco-manipulation, and human-gated DAgger post-training. The firm raised $12 million in seed funding led by Hack VC, with participation from Nomad Capital, Pi Network Ventures, 10K Ventures, and angel buyers.

A Bet Against “Clean Data Only”

A standard assumption in robotics is that demonstrations have to be near-optimal to start with — filter right down to skilled trajectories, standardize the setup, and discard something noisy earlier than it’s secure to mimic. Axis’s thesis runs the different method: information high quality lives at the distribution stage, not the single trajectory. When a big and various sufficient crowd produces noisy, suboptimal trajectories and their errors are uncorrelated, the noise averages out and a working coverage survives throughout coaching.

Axis Sim Dataset V1 places that thesis to a public take a look at. Its trajectories span pick-and-place, stacking, pouring, articulated-object manipulation, and power use, all collected via Axis’s browser-based teleoperation platform, Axis Hub, by a distributed crowd quite than a single skilled crew. The dataset was constructed with researchers from UC Berkeley, Johns Hopkins, the University of Michigan, and different establishments.

Results That Scale

On LIBERO-Plus, continuous pretraining on V1 lifts π0.5 from 83.9% to 88.8% success and outperforms a volume-matched RoboCasa365 baseline by 37.3%. Performance improves persistently as pretraining information scales from 25% to 100% of the dataset, with no saturation in sight, proof that the beneficial properties come from variety and protection quite than a one-off bump. The largest enhancements seem underneath digicam, sensor-noise, and structure perturbations, the precise axes Axis randomizes throughout era.

The crew says V2 is already underway, scaling to 1.2 million trajectories throughout 1,200 duties, with cross-embodiment generalization and outcomes throughout a number of VLA fashions exhibiting that suboptimal simulation information trains strong insurance policies.

The Engine Behind the Dataset

The dataset is one output of a bigger, actively compounding information engine. Where a conventional information vendor collects to a hard and fast spec and stops, Axis makes use of mannequin efficiency and failure instances to find out what must be collected subsequent, so each coaching spherical informs the subsequent. That engine runs on a hybrid technique throughout 4 information traces, and all 4 now run at scale:

  • Simulation: over 200,000 distributed contributors on Axis Hub, a top-3 dApp on Base, producing 4.7M+ trajectories throughout 13 embodiments.
  • Egocentric: a managed community of 1,000+ full-time, QC-trained collectors capturing first-person exercise in actual properties and companies throughout 14 industries: 200,000+ hours already banked and rising by 4,000+ hours every single day, with Vicon-verified hand pose.
  • Loco-manipulation: 500+ hours combining mobility and dexterity on actual humanoids (Unitree G1, Booster T2) via hardware-agnostic teleoperation.
  • Human-gated DAgger post-training: 500+ hours of human-in-the-loop correction focused at deployment edge instances.

Every job and trajectory is recorded on-chain on Base for provenance, and contributors are rewarded for verified work high quality.

From Open Data to Commercial Deployment

Beyond open-sourcing simulation information, Axis works straight with robotic embodiment firms to construct custom-made, embodiment-specific information pipelines and mannequin priors.

As Booster Robotics’ first sim-data associate, Axis rebuilt Booster’s actual workspace as a task-aligned digital twin, had distributed contributors accumulate 42,000+ simulation episodes on it, and distilled them right into a Booster-specific mannequin prior. With simply 30 real-robot demos, that prior reached 87.5% success versus 37.5% for an out-of-the-box π0.5, matching π0.5 utilizing half the real-world demonstrations.

Other companions span embodiment firms (Feagine Robotics), mannequin firms (Manycore Tech, Dexmal) and industrial automation (Lotus Cars, Geely Auto). Axis additionally provides on-chain robotics networks: BitRobot on Solana and OpenRoboto on Bittensor.

Redefining Physical AI’s Data Foundation

“The future of Physical AI isn’t a static dataset you obtain as soon as,” mentioned Chris Feng, founder of Axis Robotics. “It’s an engine that retains producing the information the mannequin wants subsequent. Scale will get you broad protection. Diversity retains the noise unbiased. The closed loop turns each failure into progress. That’s what compounds.”

Axis was based by researchers from UC Berkeley, CMU, Georgia Tech, and SJTU, alongside serial founders who’ve scaled shopper platforms to over 30 million customers. Its analysis is suggested by Jiachen Li, Assistant Professor at Georgia Tech.

 

Paper Link: https://arxiv.org/abs/2607.21588

Project Page: https://axisaiorg.github.io/AXIS-V1/

Dataset Link: https://huggingface.co/datasets/axisrobotics/Franka-Dataset

Github Codebase: https://github.com/AxisAIOrg/Axis-V1-Training

Similar Posts