Skip to main content

Image

Take physical AI from pilot to production

Build physical AI that performs under real-world pressure.

Turn real-world complexity into confident physical AI

While physical AI surpasses expectations in the lab, it fails in the real world when environments change beyond their training distribution. To enhance model production performance, we offer multimodal data systems designed to capture the full spectrum of real-world conditions, including those that are rare and unpredictable.

Access in-depth sensor expertise

Capture and calibrate across diverse edge conditions using LiDAR, radar, mapping, dashcam, and 360° imagery.

Build with multi-sensor fusion

Ground your models in accurate spatial and temporal representations of the physical world using 3D point clouds.

Train models on real-world dynamics

Go beyond a single frame and capture physics, motion, and scenario diversity to understand movement in real-world conditions.

Simplify multi-sensor labelling

Streamline annotation across sensor types and build production-grade data pipelines without additional operational burden.

Get consistent and accurate data with multi-sensor labelling

Combine your 2D image data and 3D point cloud data into a single view for effortless labelling using Segments.ai by Uber.

Grid of twelve screens displaying green battery icons and contour lines on black backgrounds, with a white cursor in one panel.

Image labelling

Create pixel-perfect annotations with ML-assisted tools.

Lidar point cloud visualisation with detected vehicles highlighted in green and orange bounding boxes, interface elements visible.

3D point cloud annotation

Accelerate 3D point cloud labelling with ML-assisted features.

Urban street with vehicles highlighted by green detection boxes on both sides of the carriageway.

Multi-sensor data fusion

Overlay 2D images with 3D point clouds to speed up labelling.

Lidar point cloud view of a suburban junction with cars, houses, and trees detected by autonomous vehicle sensors.

Start your 14-day free trial

See how we enable physical AI across industries

Frequently asked questions

What is physical AI?

Physical AI collects real-world data using sensors, such as cameras, light detection and ranging (LiDAR), and radar, to enable AI models deployed for autonomous systems and robots to understand physical space so they can operate in real-world conditions.

What is sensor fusion?

Sensor fusion is the process of combining data from multiple sensors to improve the accuracy and reliability of the information. By combining information from different types of sensors (such as LiDAR, radar, and cameras), a system can create a more complete picture of the environment. Sensor fusion can also reduce the impact of errors or failures in any individual sensor. This technology is used in various applications, including autonomous vehicles (AV) and robotics.

One of the most common applications of sensor fusion is the combination of various sensors mounted on a robotaxi, including 3D point clouds from LiDARs and 2D images from front, side, and rear cameras.

What are the different types of 3D sensors?

Sensors are devices that detect and respond to physical stimuli, such as light, heat, and sound. Different sensors use different technologies to perform their functions. For example, LiDAR and radar sensors use lasers and radio waves to sense their environments, while ultrasonic sensors use sound waves.

What types of 3D sensors are used in automotive, autonomous vehicles (AV), and robotics?

LiDAR (Light Detection and Ranging)

High accuracy, long range, and rapid data acquisition. LiDAR is ideal for mapping and obstacle avoidance in AVs and robots.

RADAR (Radio Detection and Ranging)

Radar sensors can detect objects in a range of weather conditions and at long distances, making them ideal for applications such as collision avoidance and autonomous driving.

SONAR (Sound Navigation and Ranging)

Sonar sensors emit sound waves and measure the time it takes for the waves to bounce back after hitting an object. In robotics and automotive, they can be used to determine the distance to the object.

Structured light

Structured light sensors use a 3D scanner to measure the 3D dimensions of an object. High resolution and accuracy make them suitable for 3D scanning and mapping applications.

Time-of-Flight (ToF) sensor

ToF sensors are used for measuring distance with depth sensing technology. They are fast and reliable, so they are a good choice for gesture recognition, object tracking, and robot navigation applications.

Stereo vision sensor

Stereo vision sensors use two (or more) sensors to simulate human binocular vision. High-precision depth sensing, good spatial resolution, and low cost make these suitable for obstacle avoidance, 3D mapping, and robot navigation applications.

Ultrasonic sensor

Ultrasonic sensors are inexpensive, easy to use, and can detect a wide range of materials, making them suitable for applications such as parking assistance and object detection.

What are best practices for labelling multi-sensor data?

Step 1: Overlay the data

  • The first step is to calibrate and align all the sensor views, then overlay all your data into a single view. This fused view gives your labeler more context, allowing it to determine what groups of point clouds represent.

Step 2: Label the 3D point clouds

  • We recommend labelling in 3D before projecting to 2D for several reasons, even though it may seem counterintuitive. Labelling in 3D first is often considerably more efficient, even if your main interest is in obtaining 2D labels.
  • Imagine you are driving past a stationary object such as a road sign. Equipped with multiple cameras, this road sign remains visible in three of them as you pass, spanning approximately 100 frames.
  • Annotating this scenario using 2D bounding boxes would require you to label a total of 300 instances. However, in the 3D space, annotating this static, non-moving traffic sign would involve a single cuboid annotation.
  • While labelling a 3D cuboid does take three times as long as annotating a 2D bounding box, it is still 100 times more efficient than labelling directly in the images.

Step 3: Project to 2D images

  • Once you’ve labelled your 3D point clouds, you can calibrate and align them to your 2D image data. Segments.ai will automatically copy over object IDs to the 2D data, saving you hours of drawing bounding boxes. You only need to make minor adjustments during your quality control check.
What’s early fusion vs. late fusion?

The shift from a late fusion approach to an early fusion approach is becoming increasingly common among AV and robotics companies in their machine learning (ML) models.

Late fusion involves the use of separate ML models for each sensor used, which generate individual outputs. These outputs are then merged or fused to create a coherent 3D representation of the scene. It involves running multiple models independently and combining their results afterwards.

Early fusion takes a more modern approach. Instead of using separate models for each sensor, all sensor data is fed into a single ML model. This unified model is designed to make predictions directly within the 3D space.

For early fusion to be effective, voxel grids prove to be an advantageous representation of the scene. Voxel grids exhibit a regular structure, making them well-suited to this approach. They can be conceptualised as tensors, allowing for end-to-end prediction. The entire process, from input to output, can be predicted using a single model.

Let’s build better
AI together

Tell us about your project. We’ll show you
the data that gets you there.

Let’s build better
AI together

Tell us about your project. We’ll show you
the data that gets you there.