Creating data for artificial intelligence can mean preparing examples of situations a model needs to recognize. Parallel Domain says its new Data Lab API gives customers a way to build those examples inside virtual worlds, using prompts and Python code to shape what the system generates.
The startup is targeting autonomy, drone and robotics companies that need large datasets to train machine-learning models. Its approach could make it easier to explore unusual scenarios and adjust training data as engineers refine their models.
Engineers can define scenarios through code
Data Lab builds on Parallel Domain’s 3D simulation technology. Engineers can use prompts to add objects and circumstances that were not already available in the company’s asset library, then generate datasets that bring more of the variability of the real world into the simulation.
The source gives examples that illustrate the range of scenarios: a highway with a cab flipped over across two lanes, or a person wearing an inflatable dinosaur outfit. Such cases can help teams create training examples for events that are difficult to plan around using only familiar, standard scenes.
Getting started is designed to be self-serve: customers install the API from GitHub and write Python code to generate datasets. That puts iteration in the hands of machine-learning engineers, who can translate an idea for a scenario into an API call and see what data it produces.
From weeks of work to near real time
Parallel Domain has customers among major automakers developing advanced driver assistance systems, as well as autonomous driving companies. In the past, creating datasets to meet a customer’s specific requirements could take the startup weeks or months. CEO Kevin McNamara said customers can now form datasets in “near real time.”
That change could help teams explore more possibilities while developing a model. Instead of waiting for a new dataset to be made around a requested condition, engineers can use the API to express a new idea in code and generate corresponding examples. The practical pace still depends on how quickly they decide what to test and turn that idea into an API call.
The company says the system supports a broad range of prompts. That flexibility is intended to let customers tailor virtual scenarios to their needs, including uncommon objects or combinations that may matter to a particular model.
Testing synthetic data against real examples
Parallel Domain also points to a test involving strollers. McNamara said the company compared certain autonomous vehicle models using synthetic datasets of strollers with models using real-world stroller datasets, and found that the models performed better when trained on synthetic data.
The result suggests synthetic examples may be useful for training, but the source does not give details about the test setup or its results beyond that comparison. It presents the finding as an early indication of how simulated data might contribute to autonomous driving development.
Data Lab draws on generative AI components, though Parallel Domain is not using OpenAI APIs such as ChatGPT. McNamara said the company uses open-source foundation models, including Stable Diffusion, which it can fine-tune to support image and content generation from text prompts. The team has also developed custom technology to label objects as it generates them.
A platform with ambitions beyond driving
Data Lab builds on Reactor, Parallel Domain’s synthetic data generation engine, which launched in May for internal use and beta testing with trusted customers. Offering Reactor through an API may also change how the startup charges for its platform.
The company’s current commercial approach has customers buy allotments of data and use credits during the year. McNamara said Data Lab could support a software-as-a-service model, with subscriptions and charges based on usage. That would connect payment more directly to customers’ access to and use of the platform.
Parallel Domain sees possible applications beyond autonomous vehicles. It has named agriculture, retail and manufacturing as areas where computer vision technology could help improve efficiency. The broader ambition is to offer a place to start whenever a business needs to train AI to interpret the world through a sensor.
For now, the API’s core promise is more direct control over synthetic dataset creation. If customers can describe new scenarios and generate examples on demand, they may be able to investigate model needs more quickly and adapt their training data without waiting for a custom dataset to be prepared.