Home
Entrar
Cadastrar
The AI Economy Is DEAD. 6 Billion Images Now POISONED. - Video học tiếng Anh
Prática de escuta
Prática de escuta
/
Video
/
The Infographics Show
/
The AI Economy Is DEAD. 6 Billion Images Now POISONED.
The AI Economy Is DEAD. 6 Billion Images Now POISONED.
Selecione o modo de aprendizagem:
Ver legendas
Escolher palavra
Reescrever palavra
Highlight:
3000 Oxford Words
4000 IELTS Words
5000 Oxford Words
3000 Common Words
1000 TOEIC Words
5000 TOEFL Words
Legendas (213)
0:00
The AI takeover is inevitable… or at least that’s what you’ve been told.
0:04
Once you upload a photo or voice note, it’s gone forever, scraped into AI systems
0:08
feeding an insatiable hunger for data. But what if that’s not the full story?
0:13
Right now, a digital insurgency is taking place. Traps are being laid inside the data these models
0:17
depend on. Researchers have figured out how to“lobotomize” AI models using just a handful
0:23
of modified JPEGs. And it works.
0:26
So what happens when the scrapers don’t just learn from the internet…
0:29
but start breaking because of it? Let’s say someone types a simple request into
0:33
an AI model. They ask for a photorealistic dog, running through a park. It’s the kind of thing
0:38
a competent model nails every single time. But this time, the results are different.
0:42
The legs don’t connect properly. Extra joints appear where they shouldn’t exist. The face subtly
0:47
drifts into something that’s from the Uncanny Valley. The fur loses structure and definition.
0:52
It looks like something that’s seen a dog a thousand times,
0:56
but doesn’t actually understand what it is. That’s the work of a poisoned AI model.
1:00
It’s not a bug. The model is still doing what it was trained to do.
1:04
The problem is what it was trained on A poisoned AI model is what happens when
1:08
someone messes with the dataset it learns from. They slip in manipulated or misleading examples
1:13
during training so the model starts picking up the wrong patterns. No one has to hack anything or
1:18
break into a system. Artists and creators simply post their work online like they always have.
1:23
Then the scrapers arrive. These are automated bots that crawl the internet,
1:27
scooping up massive amounts of images, text, and audio from websites. They don’t understand
1:32
what they’re gathering, they just vacuum it up to build training datasets for AI models.
1:37
That’s where the problem starts. Mixed in with all that normal content,
1:40
the poisoned examples get collected too. Nothing looks wrong at the time. The trap only reveals
1:45
itself later, during the next training run. A team at the University of Chicago put
1:49
the theory to the test. They fed Stable Diffusion about 50 altered pictures of dogs.
1:54
Then they put the AI model to work. Almost immediately, the images were warped. Every
1:59
single one had something wrong with it. The team kept going. Once they got to
2:03
around 300 poisoned images the model started producing images of a cat.
2:07
A model trains on billions of images, but for any single concept, say, a dog, it only really relies
2:13
on a few thousand of them. Corrupt a small slice and it all starts to fall apart. You don’t need
2:18
millions of bad files. You need less than 1% of the right ones, aimed at the right concept.
2:23
These models aren’t isolated. They’re powering Midjourney, DALL·E, and the image tools built
2:28
into your phone. So if one system becomes corrupted, the effects don’t stay contained.
2:33
They spread. It’s the opening shot of a new kind of fight.
2:36
And the people doing it aren’t rival labs or bad actors. They’re illustrators, photographers,
2:41
creatives and ordinary users with a free app… and they all have a reason to be furious.
2:46
But before we go any further, imagine this.
2:48
You’re online every single day. You check your email, open a few apps, look things up,
2:52
stream videos, maybe even do all of that while traveling. And to you, it feels totally normal.
2:57
But behind the scenes, your internet provider, advertisers, network admins, and sometimes
3:02
even governments can build a surprisingly detailed picture of what you’re doing online.
3:06
Now, to be clear, no VPN can protect you from everything. You still have to be smart online.
3:11
Don’t click suspicious links, don’t hand over personal information to sketchy emails,
3:15
and definitely don’t trust the so-called Prince of Nigeria.
3:18
But a good VPN is an important layer of privacy, because privacy shouldn’t be something you only
3:22
think about after something goes wrong. It should be the default.
3:25
And that’s exactly what Proton VPN is built for. Proton VPN helps keep your browsing private
3:30
wherever you are, whether you’re at home, traveling, or just trying to stop your
3:34
online activity from being tracked. Their no-logs policy has been verified by independent auditors,
3:39
and Proton is backed by a foundation dedicated to privacy, transparency, and user rights.
3:44
Proton VPN is also fully open source, meaning its code is publicly available for inspection. So
3:49
instead of just asking you to trust them, Proton gives people a way to verify what they’re doing.
3:54
And privacy doesn’t have to mean slow. Proton VPN offers high-speed connections
3:58
with VPN Accelerator, plus NetShield, which blocks malicious ads and trackers
4:02
before they ruin your browsing experience. It also gives you more freedom online. You can
4:06
connect through servers around the world, access content more securely while traveling, and use it
4:11
for streaming, gaming, and P2P file sharing, all while keeping your online activity more private.
4:16
So whether you’re trying to stop advertisers from following you, keep your browsing private
4:20
from your ISP, or just take back control of your online life, Proton VPN makes privacy simple.
4:26
Click the link in the description or go to protonvpn.com/theinfographicsshow
4:30
to get up to 70% discount when you sign up to Proton VPN 2-year plan.
4:35
Nobody poisons their own work for fun. They do it because they felt like their
4:39
work was taken from them. That feeling doesn’t come
4:42
from nowhere. It all started with a theft. For years, the firms building these models
4:46
treated the open web as a warehouse. They scraped billions of images, voices,
4:51
and paragraphs from it. The original creators were never asked, and they were never paid.
4:55
It was a heist, plain and simple. Much of that haul went into one open dataset
5:00
called LAION-5B, which held a massive dataset, made of up to 6 billion images and accompanying
5:06
text. Stable Diffusion learned from it, and so did most of the art tools that followed. It was
5:11
a compressed version of the visible internet, folded into something machines learn from.
5:16
Then the artists paying attention. They started recognizing their own work in the
5:20
output. Styles and techniques that took decades to develop were being reproduced in seconds. By 2023,
5:26
a generator could mimic the brushstrokes and talent of an artist in a heartbeat. A career’s
5:31
worth of experience was reduced to a prompt and undercut by a tool that cost nothing to use.
5:36
Conceptual artist Karla Ortiz, along with other artists, sued Stability AI in 2023.
5:42
Getty Images filed its own case around the same time. The lawsuits dragged out, and all the while,
5:47
the scraping continued and the models improved. The companies behind the models freely admitted
5:52
to their practices. OpenAI told the British parliament that building today's top models
5:56
without copyrighted work would be impossible. Sure, there were opt-out forms but it put the
6:01
onus on the creator. Instead of asking permission upfront, artists had to hunt down models, file
6:06
requests one by one, and hope they were honored. The balance of power remained with the developers.
6:12
So the artists adapted. Ordinary uploads turned
6:15
into weapons, one file at a time. In 2023, the University of Chicago
6:19
released a program called Glaze and it was aimed directly at the AI models.
6:23
Type an artist’s name into an AI model and it’ll generate a painting based on that
6:28
particular style and the data it’s learned on. Glaze changes that. It works by exploiting how
6:33
a computer sees an image. To an AI, a painting isn’t really a painting at all. It’s a series
6:38
of numbers that make up the style and structure. Glaze distorts the numbers slightly. It’s subtle,
6:43
something that isn’t visible to the human eye, but the machine learns from the cloaked information.
6:48
It works like an optical illusion. Two viewers can look at the same image and
6:52
see completely different things. In this case, the viewers are you and the machine.
6:56
You see your painting exactly as it was intended. The model sees something else.
7:01
An oil painting might seem like it’s charcoal. A watercolor might be seen as a completely
7:05
different medium. The visual style the scraper came to learn from has effectively been moved.
7:10
Word spread fast, and Glaze has now passed 6 million downloads. But there are limits.
7:15
Glaze could hide a style, but it couldn’t stop someone from taking the image itself. And every
7:20
time people found a new way to alter the image, scrapers came back with better ways to recover it.
7:24
It was like playing defense forever. The creators need something
7:27
that could end the fight. They needed a poison pill.
7:30
It came in January 2024, and was Nightshade. Another product of the University of Chicago,
7:37
it had one goal. To infect. An image run through Nightshade
7:40
will look completely normal, to both human eyes and a machine. But it’s like a trojan horse. If
7:45
a machine is trained on that one image, and the model starts learning nonsense.
7:50
The results are almost comical. Hats turn into cakes, handbags into toasters,
7:54
and a car learns to draw itself as a cow. For artists and creators, that felt like payback.
7:59
For years, AI companies had scraped artwork to train their models. Now the very thing being taken
8:04
could be turned against the people taking it. The demand was instant.
8:08
Nightshade hit 250,000 downloads in just 5 days. University servers buckled under the traffic.
8:14
The team behind it had to scramble to post backup download links as thousands more rushed to get it.
8:19
Most people assume this is a form of hacking. It’s not. No firewall gets breached, and the
8:24
tool never touches a single company server. Researchers call it adversarial machine learning,
8:29
but a simpler way to look at it is that Nightshade is a magic trick for machines.
8:33
The model isn't attacked from the outside. It's fooled into teaching itself the wrong thing, and
8:38
then it trusts that mistake as if it were true. It doesn't take much to start the process.
8:43
Fewer than 100 carefully crafted images can be enough to poison a single concept inside
8:48
a top AI model. And if you pair Nightshade with Glaze, the same image can hide an artist's style
8:54
and poison the training data at the same time. But those results came from models researchers
8:59
could test directly. The bigger AI systems are harder to study,
9:02
and nobody outside those companies knows exactly how vulnerable they are. But that's not really the
9:08
point. Nightshade took the idea from a theory on paper to something that actually worked.
9:12
The damage doesn't stay contained either. Poison the concept of a dog, and related ideas like
9:17
huskies, puppies, and wolves can start drifting with it. In some tests, researchers fed hundreds
9:22
of corrupted images in a single model until it could barely generate recognizable images at all.
9:27
And that's where this stopped being purely defensive. Every poisoned image uploaded to
9:31
the internet becomes a potential sleeper cell. It can remain unnoticed inside a
9:35
dataset for months or years, waiting for the next training run. The person who uploaded it
9:40
might never know if it was collected, or which model eventually learned from it.
9:45
Images were only the start. The same trick works on anything a machine learns from,
9:49
from your writing to your face. And your voice is the next target.
9:53
A service like ElevenLabs needs only a few clean seconds of your voice to build a convincing copy.
9:58
A podcast clip can be enough. So can a YouTube video, a few seconds of footage on social media or
10:03
even an old voicemail buried in someone's phone. It’s becoming a favorite tool of scammers.
10:08
And it isn’t a hypothetical risk. In 2024, criminals used cloned voices
10:12
and faces to steal about $25 million in a single faked video call. Banks and law
10:18
enforcement agencies now warn people about calls from cloned relatives asking for money.
10:22
A tool called SafeSpeech was built by security researchers for exactly that reason.
10:27
Your voice has its own signature, a kind of audio fingerprint that AI systems use
10:31
to recognize and copy you. SafeSpeech subtly smudges that fingerprint. To another person,
10:36
you sound exactly the same. But to a voice-cloning model, you become much harder to copy.
10:41
To the machine, it works a bit like radio interference. You hear the voice clearly,
10:45
but the AI doesn't. The details it needs most get scrambled. Train a voice clone on that recording,
10:51
and the result comes out wrong. They’re close enough to sound human, but missing the subtle
10:55
traits that make you… sound like you. The timing is what makes it clever.
10:59
This isn't a filter slapped onto a fake after it's generated. It’s in the original recording
11:04
itself. The AI learns from the protected version, which means the cloning process
11:09
breaks before it ever succeeds. There is no clean copy for the model to learn from.
11:13
And that's what makes this different from protecting artwork. A stolen portfolio is
11:17
one kind of loss. Your voice is another. It's the sound your family and friends recognize instantly,
11:22
the thing a scammer wants most when they're trying to impersonate you.
11:26
It turns the tables. The target gets to set the trap.
11:29
For the first time, artists, writers, and everyday people had a way to fight back against systems
11:34
trained on their work and their identities. It sounded like the next step in digital security.
11:39
Then the AI labs responded. They weren't about to lose
11:42
access to the data that fueled their models. If artists could poison the training set,
11:46
the labs would try to remove the poison. And that kicked off something bigger than a security tool.
11:51
An arms race. They couldn’t just accept corrupted datasets,
11:54
and they couldn’t afford to throw most of them away either. So they tried a straightforward fix.
11:59
They began checking each scraped image against its caption, then dropped anything that didn’t match.
12:05
In theory, poisoned data should stand out. In reality, it only catches part of it.
12:10
Tests on these systems show it flags maybe 40 to 60% of poisoned images.
12:14
The rest slips through. And there’s a second problem. It also deletes plenty of clean data,
12:20
leaving the model with either contaminated date or not enough to learn from.
12:24
So they tried something else. They started scrubbing every image through a
12:28
cleaning model before training, trying to wash out anything suspicious. It sounds clever, but it’s
12:33
slow and expensive. And it still has limitations. Because the poisoning adapts. Artists design it
12:39
to survive that cleaning step, so it reappears on the other side like a stain bleeding back through.
12:44
Not long after Glaze launched, one research group said it had already broken through the cloaking,
12:48
so the Chicago team pushed out a tougher version. Whenever the program was breached,
12:52
they developed a new patch. It was a war of attrition.
12:55
Even if a voice has protection built in, there are ways to strip parts of
12:59
it out. Push it through enough processing and some of that protection starts to fade.
13:03
In some cases, it gets noticeably weaker. Not gone, but damaged enough to matter.
13:08
The poisoners are winning, for now, and the reason is lopsided.
13:11
You have small teams constantly trying new tricks in the open. On the other,
13:15
a handful of AI labs trying to patch holes as they appear. So every time a lab rolls out a fix, it
13:21
gets tested immediately. And almost immediately, a new version shows up, adjusted just enough to
13:26
get around it, usually within weeks. And this all comes with a price tag.
13:30
It doesn’t land on the labs first. It lands on the people training the models with the data.
13:35
The whole AI business model was built on the promise that data would be free,
13:39
endless, and clean. The poison breaks the idea of clean. Once that goes,
13:43
the other two stop feeling true as well. Because now nothing can be trusted at face value.
13:48
A poisoned file looks identical to a real one. There’s no label or warning. So every dataset
13:54
turns into something that has to be checked manually or with tools that still miss a lot.
13:58
Training a top model already costs hundreds of millions of dollars, and every extra cleanup
14:03
step adds to that bill. A single poisoned batch can set a project back by weeks.
14:08
That’s the point. The goal was never
14:10
to wreck one model for fun. It was to make stolen data more expensive than paid data.
14:15
There’s already a legal, cleaner method being used. Adobe trained its Firefly tool on Midjourney
14:21
Images. Shutterstock cut similar deals. If you want clean, quality data, you need to pay for it.
14:27
That is a serious problem for the industry. Investors poured
14:30
tens of billions into AI on a single bet, that the data would stay cheap forever.
14:35
That bet isn’t looking like a sure thing anymore. For a long time, the belief was that if you
14:40
created something, you own it. The scrapers tore that deal up without asking anyone.
14:44
The AI companies acted as if human work was free, endless, and owned by no one.
14:49
That was the original mistake. Paintings were scraped. Photos were downloaded. Voices, faces,
14:53
videos were collected in huge amounts. If it was online, it was fair game.
14:58
Consent seemed optional. In every other industry, the opposite
15:01
is standard. A photographer licenses an image before a brand uses it. Musicians clear samples
15:06
before a track goes live. Even film studios pay for every frame of stock footage they use.
15:12
The work of millions of people got treated as if it belonged to no one.
15:16
This battle wasn’t created by the artists. It was the AI labs and their trainers.The moment work
15:21
was taken without asking, the balance changed. Everything after that was just a reaction. You
15:26
can’t build an empire on the idea that people do not count, then act shocked when those same people
15:32
fight back and question how the system works. Because the problem isn’t sitting in a contract
15:36
or a policy page. It’s built into the way the models learn in the first place. Every safeguard
15:41
before this tried to politely control behavior through copyright notices or opt-out forms.
15:47
But none of it really stopped the scraping This is different.
15:50
For the first time, a single person can influence what the system learns next.
15:55
And the system doesn’t get to ignore it. With people turning the tables on AI scrapers,
15:59
digital theft is getting harder, especially when it comes to scams.
16:03
But this isn’t the first time people have pushed back. Watch “Times Scammers Messed
16:07
With the Wrong People” to find out what happens when karma is dished out. Or click on this video.