What Is Synthetic Data Generation?
Synthetic data generation is the process of creating data virtually. Synthetic data can be based on data collected from real-world events that has been altered to create variations, or it can be created wholly from digital representations of real-world vehicles, people, sensors and environments. Either way, synthetic data generation realistically reproduces conditions and behaviors to help train, test and validate intelligent systems.
Synthetic data can be multimodal, spanning text, audio, images or video, depending on the system being developed and tested. Generation may occur offline, where datasets can be created in advance, or in real time, where virtual environments can produce new data as systems interact with them.
Why Synthetic Data Generation Is Important
All sorts of devices and systems rely on data to make decisions. Vehicles, industrial robots, drones, telecom substations — all of them capture information as they operate, analyze the input and make choices. A self-driving vehicle relies on sensors https://www.aptiv.com/en/solutions/intelligent-perception analyze images and sensor readings to ensure that pickle jars are properly filled and sealed , and the sensors in a cell phone tower monitor traffic to optimize network performance.
Those systems need a suitable amount of data from a wide range of scenarios before they can make credible decisions, and they need a dependable data source. Manufacturers can acquire plenty of data from millions of sensors and other sources. But collecting real-world data is expensive, time-consuming and, in some cases, unsafe. As intelligent systems become more complex and incorporate more features — and as regulatory requirements continue to evolve — organizations must test a growing number of scenarios and edge cases. Traditional testing methods struggle to perform at that scale.
The most important situations can be the hardest to capture. A manufacturer may need examples of equipment failures that occur only a few times each year, or a telecommunications provider may want to understand how a network will respond during a large-scale outage, before such an event has actually happened.
Synthetic Data Generation’s Place in Testing and Validation
Manufacturers and OEMs must anticipate potential risks before products reach the market. They develop requirements for near-collisions, extreme weather conditions, unusual driver behavior and other edge cases and then task suppliers with proving that their systems can respond appropriately.
Instead of collecting large amounts of real-world data, engineers can create virtual roads, factories, robots, vehicles, telecommunication networks, people and environmental conditions in a simulation. Engineers can then expose systems to those scenarios as if they were real.
Synthetic data is not intended to fully replace real-world data. Engineers still rely on real-world testing to confirm that systems perform as expected outside of a simulation and to identify potential biases that may exist in virtual datasets. Real-world testing and validation remains the gold standard, while synthetic data generation helps speed up development and test scenarios that might otherwise be difficult to capture.
Reasons to Adopt Synthetic Data Generation
There are several key advantages to generating synthetic data.
Synthetic data generation reduces bottlenecks. Traditional data-gathering is slow. Each step — collecting data, labeling data, training models, and validating performance across internal, customer and regulatory requirements — can create a bottleneck. Engineers may spend weeks gathering data on snowstorms, dense fog or other conditions that may restrict a camera view, or they may wait months to capture a rare equipment failure in an industrial environment. Creating controlled testing environments to reproduce those rare situations requires significant investments of time and resources. Real-world testing is limited because it can capture only what occurs during a specific test.
This challenge becomes even greater in emerging industries, where organizations may not yet have enough real-world data to support testing, development and validation efforts.
Synthetic data generates new scenarios, environments and vehicle setups whenever needed. By reducing reliance on real-world testing, vehicle availability and physical test conditions, synthetic data can scale development in domains where real-world datasets are scarce, speed up project timelines and expand what engineers can test.
Synthetic data generation drastically reduces costs. Physical, real-life testing requires vehicles, equipment, facilities, employees and test participants, all of which are significant resources. By testing and validating results in virtual environments with synthetic data, organizations can reduce the time, money and resources required for testing while still maintaining high levels of test coverage.
Synthetic data generation is flexible. A developer can test a vehicle in daylight, rain, fog or snow without leaving the lab, or a warehouse operator can evaluate how an autonomous mobile robot responds to various obstacles without disrupting daily warehouse operations. Synthetic data generates thousands of variations for each scenario, adjusting factors such as lighting, weather, traffic conditions and road layout.
Synthetic data generation can increase safety and regulation resilience. Distracted driving, severe weather, near-collisions and other rare occurrences may be dangerous to test repeatedly with physical vehicles and human test participants. Synthetic data generation can be used to expand datasets of various driver behaviors, helping engineers without putting actual humans at risk.
Other industries face similar safety concerns. Utility providers, for example, cannot trigger real outages to test response procedures, and manufacturers incur heavy costs when they intentionally damage expensive equipment just to understand how automated systems react.
Virtual environments make it possible to safely test these scenarios repeatedly — without putting people, equipment or infrastructure at risk — before moving to real-world validation.
Synthetic Data Generation: An Example
Consider the challenge of creating and testing driver-monitoring systems.
Detecting drowsiness is difficult because engineers need realistic examples of tired drivers, but collecting the data can require test participants to spend long hours in a vehicle or simulator while engineers wait for the right moments where exhaustion becomes visible. This data can be expensive and potentially unsafe to collect.
Synthetic data generation allows engineers to take a small set of real-world recordings and create thousands of variations, mapping behaviors and facial expressions onto virtual individuals with different appearances and characteristics. This expands coverage, reduces testing costs and helps engineers build more diverse datasets without placing participants in uncomfortable or unsafe situations.
Roads, vehicles, warehouses, pedestrians and human behavior can be simulated inside a computer-generated environment — similar to a video game designed to look and behave like the real world.
Synthetic data is valuable only if it accurately reflects the real world. Aptiv’s approach is built on decades of experience developing and validating systems. In applications such as interior sensing and driver monitoring, Aptiv has accumulated millions of hours of real-world recordings and decades of testing expertise, allowing synthetic environments to be compared against real-world recordings to help ensure accuracy and reliability.
As automated systems become more advanced, the demand for testing, validation and high-quality data will continue to grow — and synthetic data helps meet that demand. While real-world testing will remain the last step of system validation, synthetic data is a valuable tool for helping intelligent systems learn, adapt and perform to meet user and manufacturer expectations.