AI Lie: Machines Don’t Learn Like Humans (And Don’t Have the Right To)

RickRussell_CA@beehaw.org · 3 years ago

AI Lie: Machines Don’t Learn Like Humans (And Don’t Have the Right To)

RickRussell_CA@beehaw.org · 3 years ago

Two things:

Many of these LLMs – perhaps all of them – have been trained on datasets that include books that were absolutely NOT released into the public domain.
Ethically, we would ask any author who parrots the work of others to provide citations to original references. That rarely happens with AI language models, and if they do provide citations, they often do it wrong.

lily33@lemm.ee · 3 years ago

I’m sick and tired of this “parrots the works of others” narrative. Here’s a challenge for you: go to https://huggingface.co/chat/, input some prompt (for example, “Write a three paragraphs scene about Jason and Carol playing hide and seek with some other kids. Jason gets injured, and Carol has to help him.”). And when you get the response, try to find the author that it “parroted”. You won’t be able to - because it wouldn’t just reproduce someone else’s already made scene. It’ll mesh maaany things from all over the training data in such a way that none of them will be even remotely recognizable.

state_electrician@discuss.tchncs.de · 3 years ago

Well, I think that these models learn in a way similar to humans as in it’s basically impossible to tell where parts of the model came from. And as such the copyright claims are ridiculous. We need less copyright, not more. But, on the other hand, LLMs are not humans, they are tools created by and owned by corporations and I hate to see them profiting off of other people’s work without proper compensation.

I am fine with public domain models being trained on anything and being used for noncommercial purposes without being taken down by copyright claims.

RickRussell_CA@beehaw.org · 3 years ago

it’s basically impossible to tell where parts of the model came from

AIs are deterministic.

Train the AI on data without the copyrighted work.
Train the same AI on data with the copyrighted work.
Ask the two instances the same question.
The difference is the contribution of the copyrighted work.

There may be larger questions of precisely how an AI produces one answer when trained with a copyrighted work, and another answer when not trained with the copyrighted work. But we know why the answers are different, and we can show precisely what contribution the copyrighted work makes to the response to any prompt, just by running the AI twice.

RickRussell_CA@beehaw.org · 3 years ago

And yet, we know that the work is mechanically derivative.

keegomatic@kbin.social · 3 years ago

So is your comment. And mine. What do you think our brains do? Magic?

edit: This may sound inflammatory but I mean no offense

conciselyverbose@kbin.social · 3 years ago

So is literally every human work in the last 1000 years in every context.

Nothing is “original”. It’s all derivative. Feeding copyrighted work into an algorithm does not in any way violate any copyright law, and anyone telling you otherwise is a liar and a piece of shit. There is no valid interpretation anywhere close.

RandoCalrandian@kbin.social · 3 years ago

Is there a meaningful difference between reproducing the work and giving a summary? Because I’ll absolutely be using AI to filter all the editorial garbage out of news, setup and trained myself to surface what is meaningful to me stripped of all advertising, sponsorships, and detectable bias

RickRussell_CA@beehaw.org · 3 years ago

When you figure out how to train an AI without bias, let us know.

RandoCalrandian@kbin.social · 3 years ago

You’re confusing ai with chatgpt, but to answer your question: if it’s my own bias, why would I care that it’s in my personal ai? That’s kind of the point: using my personal lens (bias) to determine what info I would be interested in being alerted of

RickRussell_CA@beehaw.org · 3 years ago

The bias is in the AI design and the training dataset.

Ilandar@aussie.zone · 3 years ago

You’re confusing ai with chatgpt

???

RaleighEnt@kbin.social · 3 years ago

oooh I dunno man having an AI feed you shit based on what fits your personal biases is basically what social media already does and I do not think that’s something we need more of.