Skip to main content

Image

Take physical AI from pilot to production

Build physical AI that performs under real-world pressure.

Turn real-world complexity into confident physical AI

While physical AI exceeds expectations in the lab, it breaks down in the real world when environments shift outside of their training distribution. To improve model production performance, we provide multimodal data systems built to capture the full range of real-world conditions, including the rare and unpredictable.

Access deep sensor expertise

Capture and calibrate across diverse edge conditions using LiDAR, radar, mapping, dashcam, and 360° imagery.

Build with multi-sensor fusion

Ground your models in accurate spatial and temporal representations of the physical world using 3D point clouds.

Train models on real dynamics

Go beyond a single frame and capture physics, motion, and scenario diversity to understand movement in real-world conditions.

Simplify multi-sensor labeling

Streamline annotation across sensor types and build production-grade data pipelines without added operational burden.

Get consistent and accurate data with multi-sensor labeling

Combine your 2D image data and 3D point cloud data into a single view for effortless labeling using Segments.ai by Uber.

Image labeling

Create pixel-perfect annotations with ML-assisted tools.

3D point cloud annotation

Accelerate 3D point cloud labeling with ML-assisted features.

Multi-sensor data fusion

Overlay 2D images with 3D point clouds to speed up labeling.

Start your 14-day free trial

See how we power physical AI across industries

Frequently asked questions

What is physical AI?

Physical AI gathers real-world data using sensors, such as camera, light detection and ranging (LiDAR), and radar, to enable AI models deployed for autonomous systems and robots to understand physical space so they can perform in real-world conditions.

What is sensor fusion?

Sensor fusion is the process of combining data from multiple sensors to improve the accuracy and reliability of the information. By combining information from different types of sensors (such as LiDAR, radar, and cameras), a system can create a more complete picture of the environment. Sensor fusion can also reduce the impact of errors or failures in any individual sensor. This technology is used in various applications, including autonomous vehicles (AV) and robotics.

One of the most common applications of sensor fusion is the combination of various sensors mounted on a robotaxi including 3D point clouds from LiDARs and 2D images from front, side and rear cameras.

What are the different types of 3D sensors?

Sensors are devices that detect and respond to physical stimuli, such as light, heat, and sound. Different sensors use different technologies to perform their functions. For example, LiDAR and radar sensors use lasers and radio waves to sense their environments, while ultrasonic sensors use sound waves.

What types of 3D sensors are used in automotive, autonomous vehicles (AV), and robotics?

LiDAR (Light Detection and Ranging)

High accuracy, long-range, and fast data acquisition. LiDAR is ideal for mapping and obstacle avoidance in AV and robots.

RADAR (Radio Detection and Ranging)

Radar sensors can detect objects through various weather conditions and at long distances, making them ideal for applications such as collision avoidance and autonomous driving.

SONAR (Sound Navigation and Ranging)

Sonar sensors emit sound waves and measure the time it takes for the waves to bounce back after hitting an object. In robotics and automotive they can be used to determine the distance to the object.

Structured light

Structured light sensors use a 3D scanner to measure the 3D dimensions of an object. High resolution and accuracy make them suitable for 3D scanning and mapping applications.

Time-of-Flight (ToF) sensor

ToF sensors are used for measuring distance with depth sensing technology. They are fast and reliable, so they are a good choice for gesture recognition, object tracking, and robot navigation applications.

Stereo vision sensor

Stereo vision sensors use two (or more) sensors to simulate human binocular vision. High precision depth sensing, good spatial resolution, and low cost make these suitable for obstacle avoidance, 3D mapping, and robot navigation applications.

Ultrasonic sensor

Ultrasonic sensors are low-cost, easy to use, and can detect a wide range of materials, making them suitable for applications such as parking assistance and object detection.

What are best practices for labeling multi-sensor data?

Step 1: Overlay the data

  • The first step is to calibrate and align all the sensor views, then overlay all of your data into a single view. This fused view gives your labeler more context, allowing it to tell what groups of point clouds represent.

Step 2: Label the 3D point clouds

  • We advise that you label in 3D before projecting to 2D for a few reasons, even though it may seem counterintuitive. Labeling in 3D first often proves to be significantly more efficient, even if your primary interest is in obtaining 2D labels.
  • Imagine you are driving past a stationary object like a traffic sign. Equipped with multiple cameras, this traffic sign remains visible in three of them as you pass by, spanning approximately 100 frames.
  • Annotating this scenario using 2D bounding boxes would require you to label a total of 300 instances. However, in the 3D space, annotating this static, non-moving traffic sign would involve a single cuboid annotation.
  • While labeling a 3D cuboid does take three times as long as annotating a 2D bounding box, it is still 100 times more efficient than labeling directly in the images.

Step 3: Project to 2D images

  • Once you’ve labeled your 3D point clouds, you can calibrate and align them to your 2D image data. Segments.ai will automatically copy over object IDs to the 2D data, saving you hours of drawing bounding boxes. You only need to make minor adjustments during your quality control check.
What’s early fusion vs. late fusion?

The transition from a late fusion approach to an early fusion approach is becoming increasingly prevalent among AV and robotics companies in their machine learning (ML) models.

Late fusion involves the utilization of separate ML models for each sensor employed, which generates individual outputs. These outputs are then merged or fused to create a coherent 3D representation of the scene. It involves running multiple models independently and combining their results afterward.

Early fusion takes a more contemporary approach. Instead of using separate models for each sensor, all sensor data is fed into a single ML model. This unified model is designed to directly make predictions within the 3D space.

For early fusion to be effective, voxel grids prove to be an advantageous representation of the scene. Voxel grids exhibit a regular structure, making them well-suited for this approach. They can be conceptualized as tensors, allowing for end-to-end prediction. The entire process, from input to output, can be predicted using a single model.

Let’s build better
AI together

Tell us about your project. We'll show you
the data that gets you there.

Let’s build better
AI together

Tell us about your project. We'll show you
the data that gets you there.