Logo
Home
language
Kebijakan Privasi·Ketentuan Layanan

Latihan mendengarkan

Latihan mendengarkan/Video/The Infographics Show/The AI Economy Is DEAD. 6 Billion Images Now POISONED.

The AI Economy Is DEAD. 6 Billion Images Now POISONED.

Pilih mode belajar:

Highlight:

3000 Oxford Words4000 IELTS Words5000 Oxford Words3000 Common Words1000 TOEIC Words5000 TOEFL Words

Subtitle (213)

0:00The AI takeover is inevitable… or at  least that’s what you’ve been told. 
0:04Once you upload a photo or voice note,  it’s gone forever, scraped into AI systems  
0:08feeding an insatiable hunger for data. But what if that’s not the full story? 
0:13Right now, a digital insurgency is taking place.  Traps are being laid inside the data these models  
0:17depend on. Researchers have figured out how  to“lobotomize” AI models using just a handful  
0:23of modified JPEGs. And it works. 
0:26So what happens when the scrapers  don’t just learn from the internet…  
0:29but start breaking because of it? Let’s say someone types a simple request into  
0:33an AI model. They ask for a photorealistic dog,  running through a park. It’s the kind of thing  
0:38a competent model nails every single time. But this time, the results are different. 
0:42The legs don’t connect properly. Extra joints  appear where they shouldn’t exist. The face subtly  
0:47drifts into something that’s from the Uncanny  Valley. The fur loses structure and definition. 
0:52It looks like something that’s  seen a dog a thousand times,  
0:56but doesn’t actually understand what it is. That’s the work of a poisoned AI model. 
1:00It’s not a bug. The model is still  doing what it was trained to do. 
1:04The problem is what it was trained on A poisoned AI model is what happens when  
1:08someone messes with the dataset it learns from.  They slip in manipulated or misleading examples  
1:13during training so the model starts picking up  the wrong patterns. No one has to hack anything or  
1:18break into a system. Artists and creators simply  post their work online like they always have. 
1:23Then the scrapers arrive. These are automated bots that crawl the internet,  
1:27scooping up massive amounts of images, text,  and audio from websites. They don’t understand  
1:32what they’re gathering, they just vacuum it  up to build training datasets for AI models. 
1:37That’s where the problem starts. Mixed in with all that normal content,  
1:40the poisoned examples get collected too. Nothing  looks wrong at the time. The trap only reveals  
1:45itself later, during the next training run. A team at the University of Chicago put  
1:49the theory to the test. They fed Stable  Diffusion about 50 altered pictures of dogs.  
1:54Then they put the AI model to work. Almost  immediately, the images were warped. Every  
1:59single one had something wrong with it. The team kept going. Once they got to  
2:03around 300 poisoned images the model  started producing images of a cat. 
2:07A model trains on billions of images, but for any  single concept, say, a dog, it only really relies  
2:13on a few thousand of them. Corrupt a small slice  and it all starts to fall apart. You don’t need  
2:18millions of bad files. You need less than 1%  of the right ones, aimed at the right concept. 
2:23These models aren’t isolated. They’re powering  Midjourney, DALL·E, and the image tools built  
2:28into your phone. So if one system becomes  corrupted, the effects don’t stay contained. 
2:33They spread. It’s the opening shot of a new kind of fight.  
2:36And the people doing it aren’t rival labs or  bad actors. They’re illustrators, photographers,  
2:41creatives and ordinary users with a free app…  and they all have a reason to be furious.
2:46But before we go any further, imagine this. 
2:48You’re online every single day. You check  your email, open a few apps, look things up,  
2:52stream videos, maybe even do all of that while  traveling. And to you, it feels totally normal. 
2:57But behind the scenes, your internet provider,  advertisers, network admins, and sometimes  
3:02even governments can build a surprisingly  detailed picture of what you’re doing online. 
3:06Now, to be clear, no VPN can protect you from  everything. You still have to be smart online.  
3:11Don’t click suspicious links, don’t hand  over personal information to sketchy emails,  
3:15and definitely don’t trust the  so-called Prince of Nigeria. 
3:18But a good VPN is an important layer of privacy,  because privacy shouldn’t be something you only  
3:22think about after something goes  wrong. It should be the default. 
3:25And that’s exactly what Proton VPN is built for. Proton VPN helps keep your browsing private  
3:30wherever you are, whether you’re at home,  traveling, or just trying to stop your  
3:34online activity from being tracked. Their no-logs  policy has been verified by independent auditors,  
3:39and Proton is backed by a foundation dedicated  to privacy, transparency, and user rights.
3:44Proton VPN is also fully open source, meaning  its code is publicly available for inspection. So  
3:49instead of just asking you to trust them, Proton  gives people a way to verify what they’re doing. 
3:54And privacy doesn’t have to mean slow.  Proton VPN offers high-speed connections  
3:58with VPN Accelerator, plus NetShield,  which blocks malicious ads and trackers  
4:02before they ruin your browsing experience. It also gives you more freedom online. You can  
4:06connect through servers around the world, access  content more securely while traveling, and use it  
4:11for streaming, gaming, and P2P file sharing, all  while keeping your online activity more private. 
4:16So whether you’re trying to stop advertisers  from following you, keep your browsing private  
4:20from your ISP, or just take back control of your  online life, Proton VPN makes privacy simple. 
4:26Click the link in the description or  go to protonvpn.com/theinfographicsshow  
4:30to get up to 70% discount when you  sign up to Proton VPN 2-year plan. 
4:35Nobody poisons their own work for fun.  They do it because they felt like their  
4:39work was taken from them. That feeling doesn’t come  
4:42from nowhere. It all started with a theft. For years, the firms building these models  
4:46treated the open web as a warehouse.  They scraped billions of images, voices,  
4:51and paragraphs from it. The original creators  were never asked, and they were never paid. 
4:55It was a heist, plain and simple. Much of that haul went into one open dataset  
5:00called LAION-5B, which held a massive dataset,  made of up to 6 billion images and accompanying  
5:06text. Stable Diffusion learned from it, and so  did most of the art tools that followed. It was  
5:11a compressed version of the visible internet,  folded into something machines learn from. 
5:16Then the artists paying attention. They started recognizing their own work in the  
5:20output. Styles and techniques that took decades to  develop were being reproduced in seconds. By 2023,  
5:26a generator could mimic the brushstrokes and  talent of an artist in a heartbeat. A career’s  
5:31worth of experience was reduced to a prompt and  undercut by a tool that cost nothing to use. 
5:36Conceptual artist Karla Ortiz, along with  other artists, sued Stability AI in 2023.  
5:42Getty Images filed its own case around the same  time. The lawsuits dragged out, and all the while,  
5:47the scraping continued and the models improved. The companies behind the models freely admitted  
5:52to their practices. OpenAI told the British  parliament that building today's top models  
5:56without copyrighted work would be impossible.  Sure, there were opt-out forms but it put the  
6:01onus on the creator. Instead of asking permission  upfront, artists had to hunt down models, file  
6:06requests one by one, and hope they were honored.  The balance of power remained with the developers. 
6:12So the artists adapted. Ordinary uploads turned  
6:15into weapons, one file at a time. In 2023, the University of Chicago  
6:19released a program called Glaze and it  was aimed directly at the AI models. 
6:23Type an artist’s name into an AI model and  it’ll generate a painting based on that  
6:28particular style and the data it’s learned on.  Glaze changes that. It works by exploiting how  
6:33a computer sees an image. To an AI, a painting  isn’t really a painting at all. It’s a series  
6:38of numbers that make up the style and structure.  Glaze distorts the numbers slightly. It’s subtle,  
6:43something that isn’t visible to the human eye, but  the machine learns from the cloaked information. 
6:48It works like an optical illusion. Two  viewers can look at the same image and  
6:52see completely different things. In this  case, the viewers are you and the machine.  
6:56You see your painting exactly as it was intended. The model sees something else. 
7:01An oil painting might seem like it’s charcoal.  A watercolor might be seen as a completely  
7:05different medium. The visual style the scraper  came to learn from has effectively been moved. 
7:10Word spread fast, and Glaze has now passed  6 million downloads. But there are limits.  
7:15Glaze could hide a style, but it couldn’t stop  someone from taking the image itself. And every  
7:20time people found a new way to alter the image,  scrapers came back with better ways to recover it.  
7:24It was like playing defense forever. The creators need something  
7:27that could end the fight. They needed a poison pill. 
7:30It came in January 2024, and was Nightshade. Another product of the University of Chicago,  
7:37it had one goal. To infect. An image run through Nightshade  
7:40will look completely normal, to both human eyes  and a machine. But it’s like a trojan horse. If  
7:45a machine is trained on that one image,  and the model starts learning nonsense. 
7:50The results are almost comical. Hats  turn into cakes, handbags into toasters,  
7:54and a car learns to draw itself as a cow. For artists and creators, that felt like payback.  
7:59For years, AI companies had scraped artwork to  train their models. Now the very thing being taken  
8:04could be turned against the people taking it. The demand was instant. 
8:08Nightshade hit 250,000 downloads in just 5 days.  University servers buckled under the traffic.  
8:14The team behind it had to scramble to post backup  download links as thousands more rushed to get it. 
8:19Most people assume this is a form of hacking.  It’s not. No firewall gets breached, and the  
8:24tool never touches a single company server.  Researchers call it adversarial machine learning,  
8:29but a simpler way to look at it is that  Nightshade is a magic trick for machines. 
8:33The model isn't attacked from the outside. It's  fooled into teaching itself the wrong thing, and  
8:38then it trusts that mistake as if it were true. It doesn't take much to start the process. 
8:43Fewer than 100 carefully crafted images can  be enough to poison a single concept inside  
8:48a top AI model. And if you pair Nightshade with  Glaze, the same image can hide an artist's style  
8:54and poison the training data at the same time. But those results came from models researchers  
8:59could test directly. The bigger  AI systems are harder to study,  
9:02and nobody outside those companies knows exactly  how vulnerable they are. But that's not really the  
9:08point. Nightshade took the idea from a theory  on paper to something that actually worked. 
9:12The damage doesn't stay contained either. Poison  the concept of a dog, and related ideas like  
9:17huskies, puppies, and wolves can start drifting  with it. In some tests, researchers fed hundreds  
9:22of corrupted images in a single model until it  could barely generate recognizable images at all. 
9:27And that's where this stopped being purely  defensive. Every poisoned image uploaded to  
9:31the internet becomes a potential sleeper  cell. It can remain unnoticed inside a  
9:35dataset for months or years, waiting for the  next training run. The person who uploaded it  
9:40might never know if it was collected, or  which model eventually learned from it. 
9:45Images were only the start. The same trick  works on anything a machine learns from,  
9:49from your writing to your face. And your voice is the next target. 
9:53A service like ElevenLabs needs only a few clean  seconds of your voice to build a convincing copy.  
9:58A podcast clip can be enough. So can a YouTube  video, a few seconds of footage on social media or  
10:03even an old voicemail buried in someone's phone. It’s becoming a favorite tool of scammers. 
10:08And it isn’t a hypothetical risk. In 2024, criminals used cloned voices  
10:12and faces to steal about $25 million in  a single faked video call. Banks and law  
10:18enforcement agencies now warn people about  calls from cloned relatives asking for money. 
10:22A tool called SafeSpeech was built by  security researchers for exactly that reason.  
10:27Your voice has its own signature, a kind  of audio fingerprint that AI systems use  
10:31to recognize and copy you. SafeSpeech subtly  smudges that fingerprint. To another person,  
10:36you sound exactly the same. But to a voice-cloning  model, you become much harder to copy. 
10:41To the machine, it works a bit like radio  interference. You hear the voice clearly,  
10:45but the AI doesn't. The details it needs most get  scrambled. Train a voice clone on that recording,  
10:51and the result comes out wrong. They’re close  enough to sound human, but missing the subtle  
10:55traits that make you… sound like you. The timing is what makes it clever. 
10:59This isn't a filter slapped onto a fake after  it's generated. It’s in the original recording  
11:04itself. The AI learns from the protected  version, which means the cloning process  
11:09breaks before it ever succeeds. There is  no clean copy for the model to learn from. 
11:13And that's what makes this different from  protecting artwork. A stolen portfolio is  
11:17one kind of loss. Your voice is another. It's the  sound your family and friends recognize instantly,  
11:22the thing a scammer wants most when  they're trying to impersonate you. 
11:26It turns the tables. The target gets to set the trap. 
11:29For the first time, artists, writers, and everyday  people had a way to fight back against systems  
11:34trained on their work and their identities. It  sounded like the next step in digital security. 
11:39Then the AI labs responded. They weren't about to lose  
11:42access to the data that fueled their models.  If artists could poison the training set,  
11:46the labs would try to remove the poison. And that  kicked off something bigger than a security tool. 
11:51An arms race. They couldn’t just accept corrupted datasets,  
11:54and they couldn’t afford to throw most of them  away either. So they tried a straightforward fix.  
11:59They began checking each scraped image against its  caption, then dropped anything that didn’t match. 
12:05In theory, poisoned data should stand out. In reality, it only catches part of it. 
12:10Tests on these systems show it flags  maybe 40 to 60% of poisoned images.  
12:14The rest slips through. And there’s a second  problem. It also deletes plenty of clean data,  
12:20leaving the model with either contaminated  date or not enough to learn from. 
12:24So they tried something else. They started scrubbing every image through a  
12:28cleaning model before training, trying to wash out  anything suspicious. It sounds clever, but it’s  
12:33slow and expensive. And it still has limitations.  Because the poisoning adapts. Artists design it  
12:39to survive that cleaning step, so it reappears on  the other side like a stain bleeding back through. 
12:44Not long after Glaze launched, one research group  said it had already broken through the cloaking,  
12:48so the Chicago team pushed out a tougher  version. Whenever the program was breached,  
12:52they developed a new patch. It was a war of attrition. 
12:55Even if a voice has protection built  in, there are ways to strip parts of  
12:59it out. Push it through enough processing  and some of that protection starts to fade.  
13:03In some cases, it gets noticeably weaker.  Not gone, but damaged enough to matter. 
13:08The poisoners are winning, for  now, and the reason is lopsided. 
13:11You have small teams constantly trying  new tricks in the open. On the other,  
13:15a handful of AI labs trying to patch holes as they  appear. So every time a lab rolls out a fix, it  
13:21gets tested immediately. And almost immediately,  a new version shows up, adjusted just enough to  
13:26get around it, usually within weeks. And this all comes with a price tag. 
13:30It doesn’t land on the labs first. It lands on  the people training the models with the data. 
13:35The whole AI business model was built  on the promise that data would be free,  
13:39endless, and clean. The poison breaks  the idea of clean. Once that goes,  
13:43the other two stop feeling true as well. Because now nothing can be trusted at face value. 
13:48A poisoned file looks identical to a real one.  There’s no label or warning. So every dataset  
13:54turns into something that has to be checked  manually or with tools that still miss a lot.  
13:58Training a top model already costs hundreds of  millions of dollars, and every extra cleanup  
14:03step adds to that bill. A single poisoned  batch can set a project back by weeks. 
14:08That’s the point. The goal was never  
14:10to wreck one model for fun. It was to make  stolen data more expensive than paid data. 
14:15There’s already a legal, cleaner method being  used. Adobe trained its Firefly tool on Midjourney  
14:21Images. Shutterstock cut similar deals. If you  want clean, quality data, you need to pay for it. 
14:27That is a serious problem for  the industry. Investors poured  
14:30tens of billions into AI on a single bet,  that the data would stay cheap forever. 
14:35That bet isn’t looking like a sure thing anymore. For a long time, the belief was that if you  
14:40created something, you own it. The scrapers  tore that deal up without asking anyone. 
14:44The AI companies acted as if human work  was free, endless, and owned by no one.  
14:49That was the original mistake. Paintings were  scraped. Photos were downloaded. Voices, faces,  
14:53videos were collected in huge amounts. If it was online, it was fair game. 
14:58Consent seemed optional. In every other industry, the opposite  
15:01is standard. A photographer licenses an image  before a brand uses it. Musicians clear samples  
15:06before a track goes live. Even film studios  pay for every frame of stock footage they use. 
15:12The work of millions of people got  treated as if it belonged to no one. 
15:16This battle wasn’t created by the artists. It was  the AI labs and their trainers.The moment work  
15:21was taken without asking, the balance changed.  Everything after that was just a reaction. You  
15:26can’t build an empire on the idea that people do  not count, then act shocked when those same people  
15:32fight back and question how the system works. Because the problem isn’t sitting in a contract  
15:36or a policy page. It’s built into the way the  models learn in the first place. Every safeguard  
15:41before this tried to politely control behavior  through copyright notices or opt-out forms.  
15:47But none of it really stopped the scraping This is different. 
15:50For the first time, a single person can  influence what the system learns next. 
15:55And the system doesn’t get to ignore it. With people turning the tables on AI scrapers,  
15:59digital theft is getting harder,  especially when it comes to scams.  
16:03But this isn’t the first time people have  pushed back. Watch “Times Scammers Messed  
16:07With the Wrong People” to find out what happens  when karma is dished out. Or click on this video.