AI chatbot Grok can’t stop talking about 'white genocide', admits it's by design

geneva_convenience@lemmy.ml · 18 days ago

AI chatbot Grok can’t stop talking about 'white genocide', admits it's by design

MartianSands@sh.itjust.works · 18 days ago

Training LLMs on text which has been generated by an LLM is actually pretty problematic. The model can easily collapse, becoming completely useless. That’s why they always try and source really clean training data, which is becoming increasingly difficult

brucethemoose@lemmy.world · edit-2 12 days ago

Removed by mod

WhatsTheHoldup@lemmy.ml · 18 days ago

You’re not training an LLM on text generated by an LLM. You’re training it on 98% real data, and intentionally biasing it by sprinkling in the fake data intermittently.

brucethemoose@lemmy.world · edit-2 12 days ago

Removed by mod

WhatsTheHoldup@lemmy.ml · 18 days ago

Oh, yeah then I agree with above commenter. This would collapse the model.

brucethemoose@lemmy.world · edit-2 12 days ago

Removed by mod

WhatsTheHoldup@lemmy.ml · 18 days ago

Open LLMs are finetuned on partially or fully synthetic data all the time

That’s what I was suggesting.

You explained to me you weren’t talking about “finetuning”, but training on completely synthetic data.

(Fine-tuning happens after the LLM has already been trained)

brucethemoose@lemmy.world · edit-2 12 days ago

Removed by mod

WhatsTheHoldup@lemmy.ml · 18 days ago

The point I was trying to raise that wasn’t semantics was that if the majority of the full training data were synthetic, it could lead to model collapse.

But luckily (or not?) a small amount of finetuning can be very effective in correcting the range of responses.

queermunist she/her@lemmy.ml · 18 days ago

Where do you get the real data, though? They just scrap data from websites, but now that chatbots have proliferated this will only introduce contaminated data. Keeping it clean would require hiring people to scrub contamination from the data sets.

WhatsTheHoldup@lemmy.ml · edit-2 18 days ago

Where do you get the real data, though? They just scrap data from websites

Great question… Do they “just” scrape data from websites?

https://www.theatlantic.com/technology/archive/2025/03/libgen-meta-openai/682093/

Keeping it clean would require hiring people to scrub contamination from the data sets.

That’s exactly right.

https://time.com/6247678/openai-chatgpt-kenya-workers/

queermunist she/her@lemmy.ml · 18 days ago

Big problem with the 3rd world cubical farms - how do you evaluate their performance? You’d have to hire even more people to double-check their work, otherwise people will do the smart thing and cut corners to make their job easier.

Using books is definitely a way to keep out contamination, though.

50MYT@lemmy.world · 17 days ago

It’s also fantastic that there are ai honey pot mazes that exist to suck up the AI crawler with data links and bogus data to absolutely screw with their databases

And there are many of them up and working now.