Fake it until you make it: Can synthetic data help train your AI model?

Yes and no. It's complicated.

The saying "data is the new oil," was reportedly coined by British mathematician and marketing whiz Clive Humby in 2006. Humby's remark rings true more now than ever with the rise of deep learning.

Data is the fuel powering modern AI models; without enough of it the performance of these systems will sputter and fail. And like oil, the resource is scarce and controlled by big businesses. What do you do if you're a small computer vision company? You can turn to fake data to train your models, and if you're lucky it might just work.

The market for synthetic data generation grew to over $110 million in 2021 and is expected to increase to $1.15 billion by the end of 2027, according to a report published by research firm Cognilytica.

Numerous startups have built tools to spin up synthetic images to help companies train their machine learning algorithms.

There are many benefits to using computer-generated data, Gil Elbaz, co-founder and CTO of Datagen, explained to The Register.

The startup, founded in 2018 and based in Israel, has built a software platform that allows customers to easily create mock images at the click of a button. Synthetic data provides a way to scale up datasets and automatically annotate each picture with the necessary metadata without much human labor.

Issues of privacy and bias can be avoided too. "Privacy for human faces is very, very hard, and it's not ideal to even hold [that kind of data] in your servers," Elbaz says.

"With our data, there's no [personally identifiable information]. This is not a real person. This is completely synthetic, so there's no privacy issues. And bias-wise we can generate whatever distribution of ethnicities, ages, genders you want in your data, so we are not biased in any way," he says as he shows us a three-dimensional fake face.

Datagen works with companies to train computer vision models for different tasks. Simulated data is used by the automotive industry to develop AI software that automatically detects driver behavior, such as when they're distracted or falling asleep at the wheel.

Fake data has also been used by surveillance camera companies to flag whenever packages have been delivered outside people's homes. AI applications in augmented and virtual reality also benefit from ingesting copious amounts of synthetic data.

Rendering fake data is a complicated process. Datagen uses multiple methods to create computer-made images, from physics-based ray tracing algorithms to generative adversarial networks (GANs). Making the data is the easy part. Getting a model trained on false images to work in the real world is the challenge. Ideally, companies should have some real data to hand and can't just rely on fake data.

"What we see working really well is to train the network on a large amount of synthetic data, and then fine tune it on the small amount of real data. This last step is optional.

"It's really not a must, but it does improve the performance to do a small fine tune on the real world. What this means in practice is that you need much less real world data. So you don't need as much, you can use like 1/20th or 1/50th of the amount of real data and use mostly synthetic data for your training," Elbaz says.

Models trained on fake images have to be robust enough to work in real-life settings. Synthetic data has been successful in training self-driving cars to recognise ​​things like cars, road signs, and pedestrians in its environment and simulate driving the same roads in different weather conditions. It has proven useful in robotics too in limited scenarios, like getting mechanical grippers to rotate or pick up objects.

Simulation to reality

Developers relying on synthetic data have to test and tweak their models rigorously to make sure they'll work.

"If you test your models in a good way, the idea is that your test should validate that the performance will be of high quality or have a quality that you expect. If your testing is not as good or if you don't have enough test data, then you can find a gap in the performance," says Elbaz.

"We can do testing to see where the neural network is weak, pretty much by trying to ask it, for example, what do you think about this guy? And if I make him darker, or if I make him further away, or if I change him to look more angry? What do you think about that? And I can ask the network all of these different things and see where it's weaker, and really map out the weaknesses of the network itself," Elbaz says.

But in some cases the real world is too difficult to model, and synthesizing data samples won't be worthwhile. "There is a very high effort that's required in order to build [for niche things]. Say, if you're trying to understand where a dog's nose is in an image, we don't do synthetic data for dog noses. Trying to pick something like that out on your own is extremely hard." 

These gaps open up opportunities for startups that use synthetic data in a different way. Synthetaic, based in Wisconsin and founded in 2019 by Corey Jaskolski, doesn't sell computer-generated images to customers. Instead, it uses generative models like GANs or transformers to help image detection algorithms automatically label objects.

"We're still building AI that is capable of generating synthetic data. However, the novel piece that we're doing is we're not using it to generate synthetic data to then use to train an AI. We're using this generative capability to create, effectively, a way to look at real world data that allows us to do things like this auto-labeling," Jaskolski tells The Register.

"What's going on behind the scenes here is using a transformer technology that is usually used to generate imagery, but because it's so powerful and good at generating imagery, it's actually also so powerful and good at describing real world imagery in a way that lets you click on a single image and [detect others like it]."

Synthetaic showed El Reg a demo, where its Rapid Automatic Image Categorization (RAIC) technology could zero in on specific frames in a video feed. Jaskolski fed the system a photograph of a cheetah, and RAIC was able to find instances where a cheetah popped up in the video.

Real is always better

Real data is still more important for Synthetaic, despite the company's somewhat confusing name. "There are lots of examples in defense and in other industry applications where just adding 3D data or synthetic data doesn't fix the problem. I think because every situation is different, and AI always has trouble transferring from domains. It might not transfer well to the real world."

Generating synthetic data is a great way to create a larger and more diverse dataset, but it's only effective for training machine learning algorithms that perform jobs that aren't too simple and aren't too complex either. Easy computer vision tasks doesn't always require fake data, and AI. Difficult tasks require a high level of detail in simulated images and expert knowledge is needed to assess its quality. 

"I think that medical data is a really good example of a use case that we don't want to work on," says Elbaz.

"In order to model medical diseases, you need real doctors to help you.

"There's a lot of specialized knowledge that you would need in order to create this medical synthetic data. Even though medical data is extremely valuable. It's something that I think requires a separate company. It's just too hard. Anything that requires very, very, specialized knowledge is hard," he concluded. ®

Similar topics

Broader topics

Other stories you might like

  • Lonestar plans to put datacenters in the Moon's lava tubes
    How? Founder tells The Register 'Robots… lots of robots'

    Imagine a future where racks of computer servers hum quietly in darkness below the surface of the Moon.

    Here is where some of the most important data is stored, to be left untouched for as long as can be. The idea sounds like something from science-fiction, but one startup that recently emerged from stealth is trying to turn it into a reality. Lonestar Data Holdings has a unique mission unlike any other cloud provider: to build datacenters on the Moon backing up the world's data.

    "It's inconceivable to me that we are keeping our most precious assets, our knowledge and our data, on Earth, where we're setting off bombs and burning things," Christopher Stott, founder and CEO of Lonestar, told The Register. "We need to put our assets in place off our planet, where we can keep it safe."

    Continue reading
  • Conti: Russian-backed rulers of Costa Rican hacktocracy?
    Also, Chinese IT admin jailed for deleting database, and the NSA promises no more backdoors

    In brief The notorious Russian-aligned Conti ransomware gang has upped the ante in its attack against Costa Rica, threatening to overthrow the government if it doesn't pay a $20 million ransom. 

    Costa Rican president Rodrigo Chaves said that the country is effectively at war with the gang, who in April infiltrated the government's computer systems, gaining a foothold in 27 agencies at various government levels. The US State Department has offered a $15 million reward leading to the capture of Conti's leaders, who it said have made more than $150 million from 1,000+ victims.

    Conti claimed this week that it has insiders in the Costa Rican government, the AP reported, warning that "We are determined to overthrow the government by means of a cyber attack, we have already shown you all the strength and power, you have introduced an emergency." 

    Continue reading
  • China-linked Twisted Panda caught spying on Russian defense R&D
    Because Beijing isn't above covert ops to accomplish its five-year goals

    Chinese cyberspies targeted two Russian defense institutes and possibly another research facility in Belarus, according to Check Point Research.

    The new campaign, dubbed Twisted Panda, is part of a larger, state-sponsored espionage operation that has been ongoing for several months, if not nearly a year, according to the security shop.

    In a technical analysis, the researchers detail the various malicious stages and payloads of the campaign that used sanctions-related phishing emails to attack Russian entities, which are part of the state-owned defense conglomerate Rostec Corporation.

    Continue reading
  • FTC signals crackdown on ed-tech harvesting kid's data
    Trade watchdog, and President, reminds that COPPA can ban ya

    The US Federal Trade Commission on Thursday said it intends to take action against educational technology companies that unlawfully collect data from children using online educational services.

    In a policy statement, the agency said, "Children should not have to needlessly hand over their data and forfeit their privacy in order to do their schoolwork or participate in remote learning, especially given the wide and increasing adoption of ed tech tools."

    The agency says it will scrutinize educational service providers to ensure that they are meeting their legal obligations under COPPA, the Children's Online Privacy Protection Act.

    Continue reading
  • Mysterious firm seeks to buy majority stake in Arm China
    Chinese joint venture's ousted CEO tries to hang on - who will get control?

    The saga surrounding Arm's joint venture in China just took another intriguing turn: a mysterious firm named Lotcap Group claims it has signed a letter of intent to buy a 51 percent stake in Arm China from existing investors in the country.

    In a Chinese-language press release posted Wednesday, Lotcap said it has formed a subsidiary, Lotcap Fund, to buy a majority stake in the joint venture. However, reporting by one newspaper suggested that the investment firm still needs the approval of one significant investor to gain 51 percent control of Arm China.

    The development comes a couple of weeks after Arm China said that its former CEO, Allen Wu, was refusing once again to step down from his position, despite the company's board voting in late April to replace Wu with two co-chief executives. SoftBank Group, which owns 49 percent of the Chinese venture, has been trying to unentangle Arm China from Wu as the Japanese tech investment giant plans for an initial public offering of the British parent company.

    Continue reading
  • SmartNICs power the cloud, are enterprise datacenters next?
    High pricing, lack of software make smartNICs a tough sell, despite offload potential

    SmartNICs have the potential to accelerate enterprise workloads, but don't expect to see them bring hyperscale-class efficiency to most datacenters anytime soon, ZK Research's Zeus Kerravala told The Register.

    SmartNICs are widely deployed in cloud and hyperscale datacenters as a means to offload input/output (I/O) intensive network, security, and storage operations from the CPU, freeing it up to run revenue generating tenant workloads. Some more advanced chips even offload the hypervisor to further separate the infrastructure management layer from the rest of the server.

    Despite relative success in the cloud and a flurry of innovation from the still-limited vendor SmartNIC ecosystem, including Mellanox (Nvidia), Intel, Marvell, and Xilinx (AMD), Kerravala argues that the use cases for enterprise datacenters are unlikely to resemble those of the major hyperscalers, at least in the near term.

    Continue reading

Biting the hand that feeds IT © 1998–2022