Trossen integrates stereo vision into Physical AI platforms

Trossen integrates stereo vision into Physical AI platforms

Trossen integrates Stereolabs stereo vision into new Physical AI platforms. Workbench and Rivet combine synchronised GMSL2 cameras, Jetson AGX Orin compute, and bimanual manipulation hardware for robot-learning data collection.


IN Brief:

  • Workbench and Rivet integrate one ZED X Mini scene camera and two ZED X Nano wrist cameras.
  • GMSL2 links synchronise the three-camera system with onboard Jetson AGX Orin compute for RGB and depth data collection.
  • The platforms target repeatable datasets for imitation learning, reinforcement learning, and sim-to-real robotics development.

Trossen Robotics has integrated Stereolabs ZED X Mini and ZED X Nano stereo cameras into its Workbench and Rivet Physical AI platforms, creating a factory-calibrated three-camera vision system for robot-learning data collection. Workbench is a stationary bimanual platform, while Rivet adds the same development approach to a mobile manipulator.

Both systems use dual WidowX Pro six-degree-of-freedom arms, with configurations covering 700mm to 1,000mm reach, payload capacities from 4kg to 6kg, and specified repeatability of 1mm. Onboard processing is provided by an NVIDIA Jetson AGX Orin with 64GB of memory, allowing data collection and inference to run on the platform.

The camera architecture combines a centre-mounted ZED X Mini with two wrist-mounted ZED X Nano units. The scene camera provides a wider stereo-depth view of the workspace, while the wrist cameras record close-range information around the grippers and manipulated objects.

ZED X Nano uses two 2.3-megapixel 1,920 × 1,200 global-shutter sensors operating at up to 60 frames per second. Stereolabs specifies depth sensing from around 3cm, placing the cameras within the working distance needed for close manipulation tasks rather than only wider environmental perception.

The three cameras connect through GMSL2 using locking, EMI-resistant cabling and deterministic synchronisation. Trossen and Stereolabs also describe a zero-copy pipeline on the Jetson platform, allowing recording, encoding, and inference to operate concurrently without repeatedly moving image data through host memory.

Those characteristics address a practical weakness in robot-learning rigs assembled from separate cameras and development hardware. If scene and wrist views are not aligned in time, a recorded action can correspond to slightly different moments across the dataset. Calibration drift, motion blur, dropped frames, and inconsistent drivers can add further variation before any model training begins.

Factory integration does not remove every source of dataset error, but it moves camera selection, mounting, calibration, cabling, and synchronisation into the platform itself. That gives research teams a more repeatable baseline when collecting demonstrations across multiple sessions or comparing policies on different units.

Workbench and Rivet support ROS 2 and NVIDIA Isaac Sim and Isaac Lab, while the broader Trossen hardware suite includes Glide passive leader arms and the Cockpit operator station for teleoperated data collection. In that configuration, demonstrations can be captured as synchronised robot motion, RGB imagery, and depth information for subsequent imitation-learning or reinforcement-learning workflows.

The choice of GMSL2 reflects a wider convergence between automotive camera technology and robotics. Long cable runs, moving joints, electrical noise, and the need for stable timing are normal constraints in vehicles and become similarly relevant when cameras are mounted on mobile robots or at the end of articulated manipulators.

USB remains useful for laboratory vision systems, but its cabling and synchronisation limitations become more visible as cameras are distributed across larger moving machines. GMSL2 offers a more physically robust connection for those applications, although system designers still have to account for power distribution, camera calibration, software timing, thermal behaviour, and failure handling across the complete robot.

Better sensor integration also changes where engineering effort is spent. A calibrated camera stack will not make a manipulation policy reliable by itself, but it can reduce the amount of time spent assembling and debugging the recording platform before useful training data is collected.

Trossen is showing the suite during its Physical AI Residency in San Francisco and at the Actuate conference. The longer-term test will be whether a standardised sensing and compute architecture can make datasets more reproducible across multiple robots and research teams — an issue that becomes harder to ignore as robot-learning programmes expand beyond single laboratory prototypes.


Stories for you


  • Aetina launches palm-sized four-camera edge AI systems

    Aetina launches palm-sized four-camera edge AI systems

    Aetina has launched palm-sized edge AI systems for mobile deployments. The Jetson Orin platforms combine four GMSL2 camera links, up to 100 TOPS, and rugged vehicle-oriented power and environmental features.


  • Avicena ships 1Tbps microLED optical evaluation kits

    Avicena ships 1Tbps microLED optical evaluation kits

    Avicena is now shipping 1Tbps LightBundle kits for customer evaluation. The microLED platform combines 335 optical channels at up to 3Gbps each, giving AI hardware developers a laboratory route to test short-reach optical links.