nexa
By thread
nexa@server-nexa.polito.it
By month
Messages by month
- ----- 2026 -----
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2025 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2024 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2023 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2022 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2021 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2020 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2019 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2018 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2017 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2016 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2015 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2014 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2013 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2012 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2011 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2010 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2009 -----
- December
- November
- October
- September
- August
- July
- June
- May
September 2023
- 49 participants
- 220 messages
R: R: ‘Biggest act of copyright theft in history’: thousands of Australian books allegedly used to train AI model | Australia news | The Guardian
by Lorenzo Albertini
§§ 54-64 della citazione in giudizio (facilmente reperibile , ad es. qui <https://www.google.com/url?sa=t&rct=j&q=&esrc=s&source=web&cd=&ved=2ahUKEwi…> ):
<<54. Recent generative AI systems designed to recognize input text and generate
output text are built on “large language models” or “LLMs.”
55. LLMs use predictive algorithms that are designed to detect statistical patterns in
the text datasets on which they are “trained” and, on the basis of these patterns, generate
responses to user prompts. “Training” an LLM refers to the process by which the parameters that
define an LLM’s behavior are adjusted through the LLM’s ingestion and analysis of large
“training” datasets.
56. Once “trained,” the LLM analyzes the relationships among words in an input
prompt and generates a response that is an approximation of similar relationships among words
in the LLM’s “training” data. In this way, LLMs can be capable of generating sentences,
paragraphs, and even complete texts, from cover letters to novels.
57. “Training” an LLM requires supplying the LLM with large amounts of text for
the LLM to ingest—the more text, the better. That is, in part, the large in large language model.
58. As the U.S. Patent and Trademark Office has observed, LLM “training” “almost
by definition involve[s] the reproduction of entire works or substantial portions thereof.”4
59. “Training” in this context is therefore a technical-sounding euphemism for
“copying and ingesting.”
60. The quality of the LLM (that is, its capacity to generate human-seeming responses
to prompts) is dependent on the quality of the datasets used to “train” the LLM.
61. Professionally authored, edited, and published books—such as those authored by
Plaintiffs here—are an especially important source of LLM “training” data.
62. As one group of AI researchers (not affiliated with Defendants) has observed,
“[b]ooks are a rich source of both fine-grained information, how a character, an object or a scene
looks like, as well as high-level semantics, what someone is thinking, feeling and how these
states evolve through a story.”5
63. In other words, books are the high-quality materials Defendants want, need, and
have therefore outright pilfered to develop generative AI products that produce high-quality
results: text that appears to have been written by a human writer.
64. This use is highly commercial.>>.
_______________
Le informazioni contenute nella presente comunicazione e nei documenti ad essa allegati potrebbero essere tutelate dal segreto professionale e sono comunque confidenziali e ad uso esclusivo del destinatario sopra indicato. Qualora la presente comunicazione non fosse destinata a Voi, Vi preghiamo di tener presente che la divulgazione, distribuzione o riproduzione di qualunque informazione contenuta nella presente comunicazione o nei documenti ad essa allegati sono vietate. Se avete ricevuto la presente comunicazione per errore, Vi preghiamo di volerci avvertire immediatamente e di distruggere quanto ricevuto senza leggerlo. Grazie per la collaborazione.
The information contained in this email and any documents attached to it may be legally privileged and confidential. The information is intended only for the use of the individual or entity named above. If you are not the intended recipient, you are hereby notified that any use, dissemination, distribution or reproduction of any information contained in or attached to this email is prohibited. If you have received this email in error, please immediately notify us by reply email or by telephone, and destroy the original transmission and its attachments without reading them. Thank you.
>>-----Messaggio originale-----
>>Da: nexa <nexa-bounces(a)server-nexa.polito.it> Per conto di Rossana
>>Morriello
>>Inviato: venerdì 29 settembre 2023 16:08
>>A: Nexa <nexa(a)server-nexa.polito.it>
>>Oggetto: [nexa] R: ‘Biggest act of copyright theft in history’: thousands of
>>Australian books allegedly used to train AI model | Australia news | The
>>Guardian
>>
>>Non sono una giurista ma credo che questa rassegna possa essere utile alla
>>discussione
>>
>>https://www.thefashionlaw.com/from-chatgpt-to-deepfake-creating-apps-a- <https://www.thefashionlaw.com/from-chatgpt-to-deepfake-creating-apps-a-runn…>
>>running-list-of-key-ai-lawsuits/
>>
>>
>>Saluti
>>Rossana Morriello
>>
>>
>>
>>
>>-----Messaggio originale-----
>>Da: nexa <nexa-bounces(a)server-nexa.polito.it <mailto:nexa-bounces@server-nexa.polito.it> > Per conto di Stefano
>>Quintarelli
>>Inviato: venerdì 29 settembre 2023 15:21
>>Cc: Nexa <nexa(a)server-nexa.polito.it <mailto:nexa@server-nexa.polito.it> >
>>Oggetto: Re: [nexa] ‘Biggest act of copyright theft in history’: thousands of
>>Australian books allegedly used to train AI model | Australia news | The
>>Guardian
>>
>>Ho una domanda per i giuristi (anzi, piu' di una)
>>
>>per allenare un modello, ho bisogno di un file con la versione digitale di un
>>testo.
>>(cosnsidero ovviamente testi non PD, CC0, ecc.)
>>
>>la versione digitale di un testo la posso ottenere da un ebook (gia' digitale),
>>togliendo il probabile DRM.
>>ma un ebook non e' unbene ma e' un servizio soggetto a licenza d'uso, quindi
>>se non e'
>>prevista nella licenza d'uso la facolta' di estrarre il testo digitale per allenarci un
>>modello, mi sembra che ci sia gia' una violazione della licenza, per cui, credo,
>>non possa essere usato come base di un allenamento, tanto piu' se il fine di
>>tale allenamento e'
>>commerciale (se vendo un servizio basato su quel modello).
>>
>>se e' cosi', per allenare il mio modello devo allora prednere il testo digitale
>>facendo scan/ocr di un testo cartaceo.
>>ma cio' e' possibile, se non erro, solo per uso personale e non commerciale.
>>
>>se questo e' corretto, non mi pare ci sia un modo per prendere un testo digitale
>>senza infrangere una licenza d'uso/copyright
>>
>>dove e' la fallacia del ragionamento ?
>>
>>grazie, s.
>>
>>On 29/09/23 15:00, Stefano Borroni Barale wrote:
>>> Buongiorno lista,
>>>
>>>> L'idea che istruire un modello su dei testi coperti da copyright sia
>>>> una violazione del suddetto copyright è altamente opinabile
>>>
>>> Fin qui, ho l'impressione che tutti i legali in lista concorderanno.
>>>
>>>> ragionamento è in realtà abbastanza semplice: se istruirsi su un
>>>> testo ne violasse il copyright, saremmo tutti dei criminali.
>>>
>>> Ma siccome noi siamo umani e quello che produciamo non è - salvo i discorsi
>>dei politici(*) - ontologicamente identico alla produzione di esseri tecnici non
>>viventi, logica vuole che quanto si applica a noi non possa applicarsi a un LLM,
>>tanto quanto la legge sul copyright non si applica pedissequamente all'utilizzo
>>di testi umani per creare modelli linguistici.
>>>
>>> Questo è il motivo per il quale tutti i tentativi di "proteggere via copyright" il
>>prodotto di software generativi sono falliti miseramente, e con motivazioni
>>scritte in sentenze; che per il diritto credo abbiano un peso assai maggiore del
>>sito di CC.
>>>
>>> La mia impressione è che la questione terrà impegnati legali, informatici,
>>filosofi e società ancora moooooolto a lungo.
>>> SBB
>>>
>>> (*) Come sanno bene i bambini degli anni '80 che hanno giocato con
>>> questo spassoso giocattolo:
>>> https://www.enricodalbosco.it/giochi/tubolario/
>>>
>>>
>>> Di quei testi
>>>> non c'è fisicamente traccia all'interno dei modelli, non viene
>>>> copiato niente. I modelli sono un'opera trasformativa di quei testi,
>>>> non derivativa.
>>>>
>>>> Lo argomenta molto bene Creative Commons:
>>>> https://creativecommons.org/2023/02/17/fair-use-training-generative-a
>>>> i/
>>>>
>>>> Detto questo, cito le parole di un altro autore, Jeff Jarvis:
>>>>
>>https://www.facebook.com/jeff.jarvis/posts/pfbid0LMFeqdTYoxnGHQAZwp5 <https://www.facebook.com/jeff.jarvis/posts/pfbid0LMFeqdTYoxnGHQAZwp5H>
>>H
>>>> MmeeVqgMSjL2dkcwMcBojkb2cinBpgYTHyc7Fhq1B9NPl
>>>>
>>>> «I, for one, am not complaining about my books being in in large
>>>> language model training sets. I write to enter ideas into public
>>>> discourse. I prefer informed over ignorant AI. I believe it is fair
>>>> use for anyone to read & use books for transformative work. In fact,
>>>> I'd probably feel snubbed if my books were not there. I'm happy when
>>>> they are in libraries. I'm fine that they're here.»
>>>>
>>>> Fabio
>>>>
>>>> Il giorno ven 29 set 2023 alle ore 07:52 Alberto Cammozzo via nexa
>>>> nexa(a)server-nexa.polito.it <mailto:nexa@server-nexa.polito.it> ha scritto:
>>>>
>>>>> https://www.theguardian.com/australia-news/2023/sep/28/australian- <https://www.theguardian.com/australia-news/2023/sep/28/australian-bo>
>>bo
>>>>> oks-training-ai-books3-stolen-pirated
>>>>>
>>>>> Thousands of books from some of Australia’s most celebrated authors
>>have potentially been caught up in what Booker prize-winning novelist Richard
>>Flanagan has called “the biggest act of copyright theft in history”.
>>>>>
>>>>> The works have allegedly been pirated by the US-based Books3 dataset
>>and used to train generative AI for corporations such as Meta and Bloomberg.
>>>>>
>>>>> Flanagan, who found 10 of his works, including the multi-international
>>award-winning 2013 novel The Narrow Road to the Deep North, on the
>>Books3 dataset, told Guardian Australia he was deeply shocked by the
>>discovery made several days ago.
>>>>>
>>>>> “I felt as if my soul had been strip mined and I was powerless to stop it,”
>>he said in a statement.
>>>>>
>>>>> “This is the biggest act of copyright theft in history.”
>>>>>
>>>>> AI could ‘turbo-charge fraud’ and be monopolised by tech companies,
>>>>> Andrew Leigh warns
>>>>>
>>>>> The Australian Publishers Association confirmed to Guardian Australia on
>>Wednesday that as many as 18,000 fiction and nonfiction titles with
>>Australian ISBNs (unique international standard book numbers) appeared to
>>be affected by the copyright infringement, although it is not yet clear what
>>proportion of these are Australian editions of internationally authored books.
>>>>>
>>>>> “We’re still working through [the data] to work out the impact in terms of
>>Australian authors,” APA spokesperson Stuart Glover said.
>>>>>
>>>>> “This is a massive legal and ethical challenge for the publishing industry
>>and for authors globally.”
>>>>>
>>>>> A search tool published on Monday by US media platform The Atlantic and
>>uploaded by the US Authors Guild on Wednesday revealed the works of Peter
>>Carey, Helen Garner, Kate Grenville, Anna Funder, Christos Tsiolkas and
>>Thomas Keneally, as well as Flanagan and dozens of other high-profile
>>Australian authors, were included in the pirated dataset containing more than
>>180,000 titles.
>>>>>
>>>>> On Thursday, the Australian Society of Authors issued a statement saying
>>it was “horrified” to learn that the works of Australian writers were being used
>>to train artificial intelligence without permission from the authors.
>>>>>
>>>>> ASA chief executive, Olivia Lanchester, described the Books3 dataset as
>>piracy on an industrial scale.
>>>>>
>>>>> “Authors appropriately feel outraged,” Lanchester said. “The fact is this
>>technology relies upon books, journals, essays written by authors, yet
>>permission was not sought nor compensation granted.”
>>>>>
>>>>> Lanchester said the Australian literary industry, while not objecting per se
>>to emerging technologies such as AI, was deeply concerned about the lack of
>>transparency evident in the development and monetisation of AI by global
>>tech companies.
>>>>>
>>>>> “Turning a blind eye to the legitimate rights of copyright owners threatens
>>to diminish already precarious creative careers,” she said.
>>>>>
>>>>> “The enrichment of a few powerful companies is at the cost of thousands
>>of individual creators. This is not how a fair market functions.”
>>>>>
>>>>> Josephine Johnston, chief executive of Australia’s Copyright Agency,
>>described the Books3 development as “a free kick to big tech” at the expense
>>of Australia’s creative and cultural life.
>>>>>
>>>>> “We’re going to need greater transparency – how these tools have been
>>developed, trained, how they operate – before people can truly understand
>>what their legal rights might be,” she said.
>>>>>
>>>>> “We seem to be in this terrible position now where content owners –
>>remembering that the vast majority of them will be individual authors – may
>>actually have to take out court cases to enforce their rights.”
>>>>>
>>>>> Australian copyright law protects creators of original content from data
>>scraping.
>>>>>
>>>>> Litigation in the US against ChatGPT creator OpenAI over use of allegedly
>>pirated book datasets, Books1 and Books2 (which do not appear to be
>>affiliated with Books3) has already commenced.
>>>>>
>>>>> In July, North American horror/fantasy writers Mona Awad (author of
>>Bunny) and Paul Tremblay (author of The Cabin at the End of the World) filed a
>>lawsuit in a San Francisco federal court, alleging ChatGPT unlawfully digested
>>their books as part of its AI training data.
>>>>>
>>>>> On 28 August, OpenAI filed a motion to dismiss the lawsuit, arguing that
>>the authors “misconceive the scope of copyright, failing to take into account
>>the limitations and exceptions (including fair use) that properly leave room for
>>innovations like the large language models now at the forefront of artificial
>>intelligence”.
>>>>>
>>>>> On 19 September the Writers Guild and 17 of its members, including
>>bestselling novelists John Grisham, George RR Martin and Jodi Picoult, filed a
>>complaint in a New York district court against OpenAI, seeking redress for
>>“flagrant and harmful infringements” of guild members’ registered copyrights.
>>>>>
>>>>> In a statement on its website, the guild says while it is aware that
>>companies such as Meta and Bloomberg have used the Books3 dataset to
>>train their LLMs, it is not yet clear whether OpenAI is using Books3 to train its
>>ChatGPT models GPT 3.5 or GPT 4.
>>>>>
>>>>> Democracies face ‘truth decay’ as AI blurs fact and fiction, warns
>>>>> head of Australia’s military
>>>>>
>>>>> Guardian Australia has sought comment from OpenAI, which has yet to
>>officially respond to the guild’s complaint, and Meta.
>>>>>
>>>>> On 4 September, US technology magazine Wired reported that a Danish
>>anti-piracy group called Rights Alliance had been told by Bloomberg that the
>>company did not plan to train future versions of its BloombergGPT using
>>Books3.
>>>>>
>>>>> Bloomberg declined to respond to the Guardian’s queries.
>>>>>
>>>>> The APA said the global nature of the issue would present significant
>>challenges in enforcement and prosecution, and has joined the authors’
>>society in calling for AI technologies to be regulated.
>>>>>
>>>>> Consultation closed last month for a Department of Industry, Science and
>>Resources discussion paper on supporting responsible AI.
>>>>>
>>>>> A parliamentary inquiry is under way examining the use of generative
>>artificial intelligence in the Australian education system.
>>>>>
>>>>> Flanagan said it was up to the Australian government to act to protect
>>Australia’s writers.
>>>>>
>>>>> “It has power and we do not,” he said.
>>>>>
>>>>> “If it cares for our culture it must now stand up and fight for it.”
>>>>>
>>>>> _______________________________________________
>>>>> nexa mailing list
>>>>> nexa(a)server-nexa.polito.it <mailto:nexa@server-nexa.polito.it>
>>>>> https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
>>>>
>>>> _______________________________________________
>>>> nexa mailing list
>>>> nexa(a)server-nexa.polito.it <mailto:nexa@server-nexa.polito.it>
>>>> https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
>>> _______________________________________________
>>> nexa mailing list
>>> nexa(a)server-nexa.polito.it <mailto:nexa@server-nexa.polito.it>
>>> https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
>>_______________________________________________
>>nexa mailing list
>>nexa(a)server-nexa.polito.it <mailto:nexa@server-nexa.polito.it>
>>https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
>>_______________________________________________
>>nexa mailing list
>>nexa(a)server-nexa.polito.it <mailto:nexa@server-nexa.polito.it>
>>https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
Sept. 29, 2023
R: ‘Biggest act of copyright theft in history’: thousands of Australian books allegedly used to train AI model | Australia news | The Guardian
by Rossana Morriello
Non sono una giurista ma credo che questa rassegna possa essere utile alla discussione
https://www.thefashionlaw.com/from-chatgpt-to-deepfake-creating-apps-a-runn…
Saluti
Rossana Morriello
-----Messaggio originale-----
Da: nexa <nexa-bounces(a)server-nexa.polito.it> Per conto di Stefano Quintarelli
Inviato: venerdì 29 settembre 2023 15:21
Cc: Nexa <nexa(a)server-nexa.polito.it>
Oggetto: Re: [nexa] ‘Biggest act of copyright theft in history’: thousands of Australian books allegedly used to train AI model | Australia news | The Guardian
Ho una domanda per i giuristi (anzi, piu' di una)
per allenare un modello, ho bisogno di un file con la versione digitale di un testo.
(cosnsidero ovviamente testi non PD, CC0, ecc.)
la versione digitale di un testo la posso ottenere da un ebook (gia' digitale), togliendo il probabile DRM.
ma un ebook non e' unbene ma e' un servizio soggetto a licenza d'uso, quindi se non e'
prevista nella licenza d'uso la facolta' di estrarre il testo digitale per allenarci un modello, mi sembra che ci sia gia' una violazione della licenza, per cui, credo, non possa essere usato come base di un allenamento, tanto piu' se il fine di tale allenamento e'
commerciale (se vendo un servizio basato su quel modello).
se e' cosi', per allenare il mio modello devo allora prednere il testo digitale facendo scan/ocr di un testo cartaceo.
ma cio' e' possibile, se non erro, solo per uso personale e non commerciale.
se questo e' corretto, non mi pare ci sia un modo per prendere un testo digitale senza infrangere una licenza d'uso/copyright
dove e' la fallacia del ragionamento ?
grazie, s.
On 29/09/23 15:00, Stefano Borroni Barale wrote:
> Buongiorno lista,
>
>> L'idea che istruire un modello su dei testi coperti da copyright sia
>> una violazione del suddetto copyright è altamente opinabile
>
> Fin qui, ho l'impressione che tutti i legali in lista concorderanno.
>
>> ragionamento è in realtà abbastanza semplice: se istruirsi su un
>> testo ne violasse il copyright, saremmo tutti dei criminali.
>
> Ma siccome noi siamo umani e quello che produciamo non è - salvo i discorsi dei politici(*) - ontologicamente identico alla produzione di esseri tecnici non viventi, logica vuole che quanto si applica a noi non possa applicarsi a un LLM, tanto quanto la legge sul copyright non si applica pedissequamente all'utilizzo di testi umani per creare modelli linguistici.
>
> Questo è il motivo per il quale tutti i tentativi di "proteggere via copyright" il prodotto di software generativi sono falliti miseramente, e con motivazioni scritte in sentenze; che per il diritto credo abbiano un peso assai maggiore del sito di CC.
>
> La mia impressione è che la questione terrà impegnati legali, informatici, filosofi e società ancora moooooolto a lungo.
> SBB
>
> (*) Come sanno bene i bambini degli anni '80 che hanno giocato con
> questo spassoso giocattolo:
> https://www.enricodalbosco.it/giochi/tubolario/
>
>
> Di quei testi
>> non c'è fisicamente traccia all'interno dei modelli, non viene
>> copiato niente. I modelli sono un'opera trasformativa di quei testi,
>> non derivativa.
>>
>> Lo argomenta molto bene Creative Commons:
>> https://creativecommons.org/2023/02/17/fair-use-training-generative-a
>> i/
>>
>> Detto questo, cito le parole di un altro autore, Jeff Jarvis:
>> https://www.facebook.com/jeff.jarvis/posts/pfbid0LMFeqdTYoxnGHQAZwp5H
>> MmeeVqgMSjL2dkcwMcBojkb2cinBpgYTHyc7Fhq1B9NPl
>>
>> «I, for one, am not complaining about my books being in in large
>> language model training sets. I write to enter ideas into public
>> discourse. I prefer informed over ignorant AI. I believe it is fair
>> use for anyone to read & use books for transformative work. In fact,
>> I'd probably feel snubbed if my books were not there. I'm happy when
>> they are in libraries. I'm fine that they're here.»
>>
>> Fabio
>>
>> Il giorno ven 29 set 2023 alle ore 07:52 Alberto Cammozzo via nexa
>> nexa(a)server-nexa.polito.it ha scritto:
>>
>>> https://www.theguardian.com/australia-news/2023/sep/28/australian-bo
>>> oks-training-ai-books3-stolen-pirated
>>>
>>> Thousands of books from some of Australia’s most celebrated authors have potentially been caught up in what Booker prize-winning novelist Richard Flanagan has called “the biggest act of copyright theft in history”.
>>>
>>> The works have allegedly been pirated by the US-based Books3 dataset and used to train generative AI for corporations such as Meta and Bloomberg.
>>>
>>> Flanagan, who found 10 of his works, including the multi-international award-winning 2013 novel The Narrow Road to the Deep North, on the Books3 dataset, told Guardian Australia he was deeply shocked by the discovery made several days ago.
>>>
>>> “I felt as if my soul had been strip mined and I was powerless to stop it,” he said in a statement.
>>>
>>> “This is the biggest act of copyright theft in history.”
>>>
>>> AI could ‘turbo-charge fraud’ and be monopolised by tech companies,
>>> Andrew Leigh warns
>>>
>>> The Australian Publishers Association confirmed to Guardian Australia on Wednesday that as many as 18,000 fiction and nonfiction titles with Australian ISBNs (unique international standard book numbers) appeared to be affected by the copyright infringement, although it is not yet clear what proportion of these are Australian editions of internationally authored books.
>>>
>>> “We’re still working through [the data] to work out the impact in terms of Australian authors,” APA spokesperson Stuart Glover said.
>>>
>>> “This is a massive legal and ethical challenge for the publishing industry and for authors globally.”
>>>
>>> A search tool published on Monday by US media platform The Atlantic and uploaded by the US Authors Guild on Wednesday revealed the works of Peter Carey, Helen Garner, Kate Grenville, Anna Funder, Christos Tsiolkas and Thomas Keneally, as well as Flanagan and dozens of other high-profile Australian authors, were included in the pirated dataset containing more than 180,000 titles.
>>>
>>> On Thursday, the Australian Society of Authors issued a statement saying it was “horrified” to learn that the works of Australian writers were being used to train artificial intelligence without permission from the authors.
>>>
>>> ASA chief executive, Olivia Lanchester, described the Books3 dataset as piracy on an industrial scale.
>>>
>>> “Authors appropriately feel outraged,” Lanchester said. “The fact is this technology relies upon books, journals, essays written by authors, yet permission was not sought nor compensation granted.”
>>>
>>> Lanchester said the Australian literary industry, while not objecting per se to emerging technologies such as AI, was deeply concerned about the lack of transparency evident in the development and monetisation of AI by global tech companies.
>>>
>>> “Turning a blind eye to the legitimate rights of copyright owners threatens to diminish already precarious creative careers,” she said.
>>>
>>> “The enrichment of a few powerful companies is at the cost of thousands of individual creators. This is not how a fair market functions.”
>>>
>>> Josephine Johnston, chief executive of Australia’s Copyright Agency, described the Books3 development as “a free kick to big tech” at the expense of Australia’s creative and cultural life.
>>>
>>> “We’re going to need greater transparency – how these tools have been developed, trained, how they operate – before people can truly understand what their legal rights might be,” she said.
>>>
>>> “We seem to be in this terrible position now where content owners – remembering that the vast majority of them will be individual authors – may actually have to take out court cases to enforce their rights.”
>>>
>>> Australian copyright law protects creators of original content from data scraping.
>>>
>>> Litigation in the US against ChatGPT creator OpenAI over use of allegedly pirated book datasets, Books1 and Books2 (which do not appear to be affiliated with Books3) has already commenced.
>>>
>>> In July, North American horror/fantasy writers Mona Awad (author of Bunny) and Paul Tremblay (author of The Cabin at the End of the World) filed a lawsuit in a San Francisco federal court, alleging ChatGPT unlawfully digested their books as part of its AI training data.
>>>
>>> On 28 August, OpenAI filed a motion to dismiss the lawsuit, arguing that the authors “misconceive the scope of copyright, failing to take into account the limitations and exceptions (including fair use) that properly leave room for innovations like the large language models now at the forefront of artificial intelligence”.
>>>
>>> On 19 September the Writers Guild and 17 of its members, including bestselling novelists John Grisham, George RR Martin and Jodi Picoult, filed a complaint in a New York district court against OpenAI, seeking redress for “flagrant and harmful infringements” of guild members’ registered copyrights.
>>>
>>> In a statement on its website, the guild says while it is aware that companies such as Meta and Bloomberg have used the Books3 dataset to train their LLMs, it is not yet clear whether OpenAI is using Books3 to train its ChatGPT models GPT 3.5 or GPT 4.
>>>
>>> Democracies face ‘truth decay’ as AI blurs fact and fiction, warns
>>> head of Australia’s military
>>>
>>> Guardian Australia has sought comment from OpenAI, which has yet to officially respond to the guild’s complaint, and Meta.
>>>
>>> On 4 September, US technology magazine Wired reported that a Danish anti-piracy group called Rights Alliance had been told by Bloomberg that the company did not plan to train future versions of its BloombergGPT using Books3.
>>>
>>> Bloomberg declined to respond to the Guardian’s queries.
>>>
>>> The APA said the global nature of the issue would present significant challenges in enforcement and prosecution, and has joined the authors’ society in calling for AI technologies to be regulated.
>>>
>>> Consultation closed last month for a Department of Industry, Science and Resources discussion paper on supporting responsible AI.
>>>
>>> A parliamentary inquiry is under way examining the use of generative artificial intelligence in the Australian education system.
>>>
>>> Flanagan said it was up to the Australian government to act to protect Australia’s writers.
>>>
>>> “It has power and we do not,” he said.
>>>
>>> “If it cares for our culture it must now stand up and fight for it.”
>>>
>>> _______________________________________________
>>> nexa mailing list
>>> nexa(a)server-nexa.polito.it
>>> https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
>>
>> _______________________________________________
>> nexa mailing list
>> nexa(a)server-nexa.polito.it
>> https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
> _______________________________________________
> nexa mailing list
> nexa(a)server-nexa.polito.it
> https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
_______________________________________________
nexa mailing list
nexa(a)server-nexa.polito.it
https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
Sept. 29, 2023
Re: [nexa] ‘Biggest act of copyright theft in history’: thousands of Australian books allegedly used to train AI model | Australia news | The Guardian
by Stefano Quintarelli
Ho una domanda per i giuristi (anzi, piu' di una)
per allenare un modello, ho bisogno di un file con la versione digitale di un testo.
(cosnsidero ovviamente testi non PD, CC0, ecc.)
la versione digitale di un testo la posso ottenere da un ebook (gia' digitale), togliendo
il probabile DRM.
ma un ebook non e' unbene ma e' un servizio soggetto a licenza d'uso, quindi se non e'
prevista nella licenza d'uso la facolta' di estrarre il testo digitale per allenarci un
modello, mi sembra che ci sia gia' una violazione della licenza, per cui, credo, non possa
essere usato come base di un allenamento, tanto piu' se il fine di tale allenamento e'
commerciale (se vendo un servizio basato su quel modello).
se e' cosi', per allenare il mio modello devo allora prednere il testo digitale facendo
scan/ocr di un testo cartaceo.
ma cio' e' possibile, se non erro, solo per uso personale e non commerciale.
se questo e' corretto, non mi pare ci sia un modo per prendere un testo digitale senza
infrangere una licenza d'uso/copyright
dove e' la fallacia del ragionamento ?
grazie, s.
On 29/09/23 15:00, Stefano Borroni Barale wrote:
> Buongiorno lista,
>
>> L'idea che istruire un modello su dei testi coperti da copyright sia una violazione del suddetto copyright è altamente opinabile
>
> Fin qui, ho l'impressione che tutti i legali in lista concorderanno.
>
>> ragionamento è in realtà abbastanza semplice: se istruirsi su un
>> testo ne violasse il copyright, saremmo tutti dei criminali.
>
> Ma siccome noi siamo umani e quello che produciamo non è - salvo i discorsi dei politici(*) - ontologicamente identico alla produzione di esseri tecnici non viventi, logica vuole che quanto si applica a noi non possa applicarsi a un LLM, tanto quanto la legge sul copyright non si applica pedissequamente all'utilizzo di testi umani per creare modelli linguistici.
>
> Questo è il motivo per il quale tutti i tentativi di "proteggere via copyright" il prodotto di software generativi sono falliti miseramente, e con motivazioni scritte in sentenze; che per il diritto credo abbiano un peso assai maggiore del sito di CC.
>
> La mia impressione è che la questione terrà impegnati legali, informatici, filosofi e società ancora moooooolto a lungo.
> SBB
>
> (*) Come sanno bene i bambini degli anni '80 che hanno giocato con questo spassoso giocattolo: https://www.enricodalbosco.it/giochi/tubolario/
>
>
> Di quei testi
>> non c'è fisicamente traccia all'interno dei modelli, non viene copiato
>> niente. I modelli sono un'opera trasformativa di quei testi, non
>> derivativa.
>>
>> Lo argomenta molto bene Creative Commons:
>> https://creativecommons.org/2023/02/17/fair-use-training-generative-ai/
>>
>> Detto questo, cito le parole di un altro autore, Jeff Jarvis:
>> https://www.facebook.com/jeff.jarvis/posts/pfbid0LMFeqdTYoxnGHQAZwp5HMmeeVq…
>>
>> «I, for one, am not complaining about my books being in in large
>> language model training sets. I write to enter ideas into public
>> discourse. I prefer informed over ignorant AI. I believe it is fair
>> use for anyone to read & use books for transformative work. In fact,
>> I'd probably feel snubbed if my books were not there. I'm happy when
>> they are in libraries. I'm fine that they're here.»
>>
>> Fabio
>>
>> Il giorno ven 29 set 2023 alle ore 07:52 Alberto Cammozzo via nexa
>> nexa(a)server-nexa.polito.it ha scritto:
>>
>>> https://www.theguardian.com/australia-news/2023/sep/28/australian-books-tra…
>>>
>>> Thousands of books from some of Australia’s most celebrated authors have potentially been caught up in what Booker prize-winning novelist Richard Flanagan has called “the biggest act of copyright theft in history”.
>>>
>>> The works have allegedly been pirated by the US-based Books3 dataset and used to train generative AI for corporations such as Meta and Bloomberg.
>>>
>>> Flanagan, who found 10 of his works, including the multi-international award-winning 2013 novel The Narrow Road to the Deep North, on the Books3 dataset, told Guardian Australia he was deeply shocked by the discovery made several days ago.
>>>
>>> “I felt as if my soul had been strip mined and I was powerless to stop it,” he said in a statement.
>>>
>>> “This is the biggest act of copyright theft in history.”
>>>
>>> AI could ‘turbo-charge fraud’ and be monopolised by tech companies, Andrew Leigh warns
>>>
>>> The Australian Publishers Association confirmed to Guardian Australia on Wednesday that as many as 18,000 fiction and nonfiction titles with Australian ISBNs (unique international standard book numbers) appeared to be affected by the copyright infringement, although it is not yet clear what proportion of these are Australian editions of internationally authored books.
>>>
>>> “We’re still working through [the data] to work out the impact in terms of Australian authors,” APA spokesperson Stuart Glover said.
>>>
>>> “This is a massive legal and ethical challenge for the publishing industry and for authors globally.”
>>>
>>> A search tool published on Monday by US media platform The Atlantic and uploaded by the US Authors Guild on Wednesday revealed the works of Peter Carey, Helen Garner, Kate Grenville, Anna Funder, Christos Tsiolkas and Thomas Keneally, as well as Flanagan and dozens of other high-profile Australian authors, were included in the pirated dataset containing more than 180,000 titles.
>>>
>>> On Thursday, the Australian Society of Authors issued a statement saying it was “horrified” to learn that the works of Australian writers were being used to train artificial intelligence without permission from the authors.
>>>
>>> ASA chief executive, Olivia Lanchester, described the Books3 dataset as piracy on an industrial scale.
>>>
>>> “Authors appropriately feel outraged,” Lanchester said. “The fact is this technology relies upon books, journals, essays written by authors, yet permission was not sought nor compensation granted.”
>>>
>>> Lanchester said the Australian literary industry, while not objecting per se to emerging technologies such as AI, was deeply concerned about the lack of transparency evident in the development and monetisation of AI by global tech companies.
>>>
>>> “Turning a blind eye to the legitimate rights of copyright owners threatens to diminish already precarious creative careers,” she said.
>>>
>>> “The enrichment of a few powerful companies is at the cost of thousands of individual creators. This is not how a fair market functions.”
>>>
>>> Josephine Johnston, chief executive of Australia’s Copyright Agency, described the Books3 development as “a free kick to big tech” at the expense of Australia’s creative and cultural life.
>>>
>>> “We’re going to need greater transparency – how these tools have been developed, trained, how they operate – before people can truly understand what their legal rights might be,” she said.
>>>
>>> “We seem to be in this terrible position now where content owners – remembering that the vast majority of them will be individual authors – may actually have to take out court cases to enforce their rights.”
>>>
>>> Australian copyright law protects creators of original content from data scraping.
>>>
>>> Litigation in the US against ChatGPT creator OpenAI over use of allegedly pirated book datasets, Books1 and Books2 (which do not appear to be affiliated with Books3) has already commenced.
>>>
>>> In July, North American horror/fantasy writers Mona Awad (author of Bunny) and Paul Tremblay (author of The Cabin at the End of the World) filed a lawsuit in a San Francisco federal court, alleging ChatGPT unlawfully digested their books as part of its AI training data.
>>>
>>> On 28 August, OpenAI filed a motion to dismiss the lawsuit, arguing that the authors “misconceive the scope of copyright, failing to take into account the limitations and exceptions (including fair use) that properly leave room for innovations like the large language models now at the forefront of artificial intelligence”.
>>>
>>> On 19 September the Writers Guild and 17 of its members, including bestselling novelists John Grisham, George RR Martin and Jodi Picoult, filed a complaint in a New York district court against OpenAI, seeking redress for “flagrant and harmful infringements” of guild members’ registered copyrights.
>>>
>>> In a statement on its website, the guild says while it is aware that companies such as Meta and Bloomberg have used the Books3 dataset to train their LLMs, it is not yet clear whether OpenAI is using Books3 to train its ChatGPT models GPT 3.5 or GPT 4.
>>>
>>> Democracies face ‘truth decay’ as AI blurs fact and fiction, warns head of Australia’s military
>>>
>>> Guardian Australia has sought comment from OpenAI, which has yet to officially respond to the guild’s complaint, and Meta.
>>>
>>> On 4 September, US technology magazine Wired reported that a Danish anti-piracy group called Rights Alliance had been told by Bloomberg that the company did not plan to train future versions of its BloombergGPT using Books3.
>>>
>>> Bloomberg declined to respond to the Guardian’s queries.
>>>
>>> The APA said the global nature of the issue would present significant challenges in enforcement and prosecution, and has joined the authors’ society in calling for AI technologies to be regulated.
>>>
>>> Consultation closed last month for a Department of Industry, Science and Resources discussion paper on supporting responsible AI.
>>>
>>> A parliamentary inquiry is under way examining the use of generative artificial intelligence in the Australian education system.
>>>
>>> Flanagan said it was up to the Australian government to act to protect Australia’s writers.
>>>
>>> “It has power and we do not,” he said.
>>>
>>> “If it cares for our culture it must now stand up and fight for it.”
>>>
>>> _______________________________________________
>>> nexa mailing list
>>> nexa(a)server-nexa.polito.it
>>> https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
>>
>> _______________________________________________
>> nexa mailing list
>> nexa(a)server-nexa.polito.it
>> https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
> _______________________________________________
> nexa mailing list
> nexa(a)server-nexa.polito.it
> https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
Sept. 29, 2023
Re: [nexa] ‘Biggest act of copyright theft in history’: thousands of Australian books allegedly used to train AI model | Australia news | The Guardian
by Stefano Borroni Barale
Buongiorno lista,
> L'idea che istruire un modello su dei testi coperti da copyright sia una violazione del suddetto copyright è altamente opinabile
Fin qui, ho l'impressione che tutti i legali in lista concorderanno.
> ragionamento è in realtà abbastanza semplice: se istruirsi su un
> testo ne violasse il copyright, saremmo tutti dei criminali.
Ma siccome noi siamo umani e quello che produciamo non è - salvo i discorsi dei politici(*) - ontologicamente identico alla produzione di esseri tecnici non viventi, logica vuole che quanto si applica a noi non possa applicarsi a un LLM, tanto quanto la legge sul copyright non si applica pedissequamente all'utilizzo di testi umani per creare modelli linguistici.
Questo è il motivo per il quale tutti i tentativi di "proteggere via copyright" il prodotto di software generativi sono falliti miseramente, e con motivazioni scritte in sentenze; che per il diritto credo abbiano un peso assai maggiore del sito di CC.
La mia impressione è che la questione terrà impegnati legali, informatici, filosofi e società ancora moooooolto a lungo.
SBB
(*) Come sanno bene i bambini degli anni '80 che hanno giocato con questo spassoso giocattolo: https://www.enricodalbosco.it/giochi/tubolario/
Di quei testi
> non c'è fisicamente traccia all'interno dei modelli, non viene copiato
> niente. I modelli sono un'opera trasformativa di quei testi, non
> derivativa.
>
> Lo argomenta molto bene Creative Commons:
> https://creativecommons.org/2023/02/17/fair-use-training-generative-ai/
>
> Detto questo, cito le parole di un altro autore, Jeff Jarvis:
> https://www.facebook.com/jeff.jarvis/posts/pfbid0LMFeqdTYoxnGHQAZwp5HMmeeVq…
>
> «I, for one, am not complaining about my books being in in large
> language model training sets. I write to enter ideas into public
> discourse. I prefer informed over ignorant AI. I believe it is fair
> use for anyone to read & use books for transformative work. In fact,
> I'd probably feel snubbed if my books were not there. I'm happy when
> they are in libraries. I'm fine that they're here.»
>
> Fabio
>
> Il giorno ven 29 set 2023 alle ore 07:52 Alberto Cammozzo via nexa
> nexa(a)server-nexa.polito.it ha scritto:
>
> > https://www.theguardian.com/australia-news/2023/sep/28/australian-books-tra…
> >
> > Thousands of books from some of Australia’s most celebrated authors have potentially been caught up in what Booker prize-winning novelist Richard Flanagan has called “the biggest act of copyright theft in history”.
> >
> > The works have allegedly been pirated by the US-based Books3 dataset and used to train generative AI for corporations such as Meta and Bloomberg.
> >
> > Flanagan, who found 10 of his works, including the multi-international award-winning 2013 novel The Narrow Road to the Deep North, on the Books3 dataset, told Guardian Australia he was deeply shocked by the discovery made several days ago.
> >
> > “I felt as if my soul had been strip mined and I was powerless to stop it,” he said in a statement.
> >
> > “This is the biggest act of copyright theft in history.”
> >
> > AI could ‘turbo-charge fraud’ and be monopolised by tech companies, Andrew Leigh warns
> >
> > The Australian Publishers Association confirmed to Guardian Australia on Wednesday that as many as 18,000 fiction and nonfiction titles with Australian ISBNs (unique international standard book numbers) appeared to be affected by the copyright infringement, although it is not yet clear what proportion of these are Australian editions of internationally authored books.
> >
> > “We’re still working through [the data] to work out the impact in terms of Australian authors,” APA spokesperson Stuart Glover said.
> >
> > “This is a massive legal and ethical challenge for the publishing industry and for authors globally.”
> >
> > A search tool published on Monday by US media platform The Atlantic and uploaded by the US Authors Guild on Wednesday revealed the works of Peter Carey, Helen Garner, Kate Grenville, Anna Funder, Christos Tsiolkas and Thomas Keneally, as well as Flanagan and dozens of other high-profile Australian authors, were included in the pirated dataset containing more than 180,000 titles.
> >
> > On Thursday, the Australian Society of Authors issued a statement saying it was “horrified” to learn that the works of Australian writers were being used to train artificial intelligence without permission from the authors.
> >
> > ASA chief executive, Olivia Lanchester, described the Books3 dataset as piracy on an industrial scale.
> >
> > “Authors appropriately feel outraged,” Lanchester said. “The fact is this technology relies upon books, journals, essays written by authors, yet permission was not sought nor compensation granted.”
> >
> > Lanchester said the Australian literary industry, while not objecting per se to emerging technologies such as AI, was deeply concerned about the lack of transparency evident in the development and monetisation of AI by global tech companies.
> >
> > “Turning a blind eye to the legitimate rights of copyright owners threatens to diminish already precarious creative careers,” she said.
> >
> > “The enrichment of a few powerful companies is at the cost of thousands of individual creators. This is not how a fair market functions.”
> >
> > Josephine Johnston, chief executive of Australia’s Copyright Agency, described the Books3 development as “a free kick to big tech” at the expense of Australia’s creative and cultural life.
> >
> > “We’re going to need greater transparency – how these tools have been developed, trained, how they operate – before people can truly understand what their legal rights might be,” she said.
> >
> > “We seem to be in this terrible position now where content owners – remembering that the vast majority of them will be individual authors – may actually have to take out court cases to enforce their rights.”
> >
> > Australian copyright law protects creators of original content from data scraping.
> >
> > Litigation in the US against ChatGPT creator OpenAI over use of allegedly pirated book datasets, Books1 and Books2 (which do not appear to be affiliated with Books3) has already commenced.
> >
> > In July, North American horror/fantasy writers Mona Awad (author of Bunny) and Paul Tremblay (author of The Cabin at the End of the World) filed a lawsuit in a San Francisco federal court, alleging ChatGPT unlawfully digested their books as part of its AI training data.
> >
> > On 28 August, OpenAI filed a motion to dismiss the lawsuit, arguing that the authors “misconceive the scope of copyright, failing to take into account the limitations and exceptions (including fair use) that properly leave room for innovations like the large language models now at the forefront of artificial intelligence”.
> >
> > On 19 September the Writers Guild and 17 of its members, including bestselling novelists John Grisham, George RR Martin and Jodi Picoult, filed a complaint in a New York district court against OpenAI, seeking redress for “flagrant and harmful infringements” of guild members’ registered copyrights.
> >
> > In a statement on its website, the guild says while it is aware that companies such as Meta and Bloomberg have used the Books3 dataset to train their LLMs, it is not yet clear whether OpenAI is using Books3 to train its ChatGPT models GPT 3.5 or GPT 4.
> >
> > Democracies face ‘truth decay’ as AI blurs fact and fiction, warns head of Australia’s military
> >
> > Guardian Australia has sought comment from OpenAI, which has yet to officially respond to the guild’s complaint, and Meta.
> >
> > On 4 September, US technology magazine Wired reported that a Danish anti-piracy group called Rights Alliance had been told by Bloomberg that the company did not plan to train future versions of its BloombergGPT using Books3.
> >
> > Bloomberg declined to respond to the Guardian’s queries.
> >
> > The APA said the global nature of the issue would present significant challenges in enforcement and prosecution, and has joined the authors’ society in calling for AI technologies to be regulated.
> >
> > Consultation closed last month for a Department of Industry, Science and Resources discussion paper on supporting responsible AI.
> >
> > A parliamentary inquiry is under way examining the use of generative artificial intelligence in the Australian education system.
> >
> > Flanagan said it was up to the Australian government to act to protect Australia’s writers.
> >
> > “It has power and we do not,” he said.
> >
> > “If it cares for our culture it must now stand up and fight for it.”
> >
> > _______________________________________________
> > nexa mailing list
> > nexa(a)server-nexa.polito.it
> > https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
>
> _______________________________________________
> nexa mailing list
> nexa(a)server-nexa.polito.it
> https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
Sept. 29, 2023
Ending Human-Dependent Peer Review
by maurizio lana
«The scholarly kitchen» pubblica oggi questo articolo:
Ending Human-Dependent Peer Review
by Haseeb Irfanullah
Human-dependent peer review is inequitable, suffers from injustice, and
is potentially unsustainable.
Here’s why we should replace it (eventually) with AI-based peer review.
https://scholarlykitchen.sspnet.org/2023/09/29/ending-human-dependent-peer-…
colpisce nel sottotitolo quel "human-dependent peer review is
inequitable, suffers from injustice", laddove iniquità e ingiustizia
sono proprio due arcinoti e discussi difetti dei sistemi di valutazione
basati su IA che in sostanza nessuno sa come evitare
potremmo pensare che il posizionamento geo-politico-culturale
dell'autore non ne faccia un portavoce di posizioni fortemente condivise
nel contesto biblioteconomico internazionale, ma rimane il fatto che
«The scholarly kitchen» è una sede importante per gli orientamenti su
cui si incontrano aziendalismo e biblioteconomia - mi esprimo in chiaro:
di fronte al peso insostenibile delle review, Elsevier, Springer, et
al., stanno probabilmente pensando a questo: peer review fatte da
sistemi di IA. ma chi può affermare che un sistema di IA sia un peer di
un/a studiosa/o? (questione già posta da H. Schönenberger
nell'introduzione del libro seminale Beta Writer. Lithium-Ion Batteries.
A Machine-Generated Summary of Current Research. New York, NY: Springer,
2019. https://doi.org/10.1007/978-3-030-16800-1)
Maurizio
------------------------------------------------------------------------
A questo punto devo fare una confessione:
come il mio amico Erri De Luca, sono un europeista estremista.
questo significa che, per me, l’Europa unita è l’unica utopia politica
ragionevole che noi europei abbiamo coniato
xavier cercas, salone del libro 2018
------------------------------------------------------------------------
Maurizio Lana - 347 7370925
Sept. 29, 2023
Re: [nexa] ‘Biggest act of copyright theft in history’: thousands of Australian books allegedly used to train AI model | Australia news | The Guardian
by Angelo Raffaele Meo
caro Giacomo e cari tutti,
mi congratulo con Giacomo per la chiarezza e profondità della sua riflessione. A Giacomo, e a lui soltanto perchè non è bello farsi belli con meriti altrui, invio il mio curriculum vitae secondo chatgpt, che mi attribuisce dieci anni di meno, un luogo di nascita dove non sono mai stato e moltissimi meriti scientifici che sono di almeno altri tre autori.
Raf Meo
________________________________
From: nexa <nexa-bounces(a)server-nexa.polito.it> on behalf of Giacomo Tesio <giacomo(a)tesio.it>
Sent: Friday, September 29, 2023 11:53 AM
To: nexa(a)server-nexa.polito.it <nexa(a)server-nexa.polito.it>; Fabio Alemagna <falemagn(a)gmail.com>; Alberto Cammozzo <ac+nexa(a)zeromx.net>
Cc: Nexa <nexa(a)server-nexa.polito.it>
Subject: Re: [nexa] ‘Biggest act of copyright theft in history’: thousands of Australian books allegedly used to train AI model | Australia news | The Guardian
Ciao Fabio,
Il 29 Settembre 2023 08:24:57 UTC, Fabio Alemagna <falemagn(a)gmail.com> ha scritto:
> L'idea che istruire un modello...
Purtroppo l'idea di "istruire" una macchina è di per sé un'allucinazione.
Le macchine si costruiscono e (se sono programmabili) di programmano.
Non c'è nessuna mente che possa imparare lì dentro, perché le macchine non pensano.
Il fatto che possano essere programmate statisticamente per ingannare chi
non ne comprende il funzionamento ci dice che il loro studio andrebbe riservato
a chi lo comprende appieno (tanto da poterle ricostruire da zero) e la loro applicazione
a persone inconsapevoli o fragili semplicemente vietato.
Ciò che chiami "modello" non è stato istruito ma programmato statisticamente
usando determinati testi "sorgente".
Il "modello" rappresenta una codifica parziale (o se peferisci, una compressione con
perdita di informazione) con interferenze (le varie sorgenti casuali utilizzate durante la
programmazione statistica o durante l'esecuzione del programma
e poi scartate per poter fingere che l'output non sia deterministico).
Dunque il modello CONTIENE, seppur in forma difficile da estrarre e non
necessariamente corrispondente all'intento comunicativo dei rispettivi autori,
ampie parti dei testi originali.
Un esempio particolarmente lampante di questo meccanismo fu evidenziato con
Microsoft Copilot (aka CopyALot) che distribuì codice sotto GPL in violazione della stessa,
copiando alla lettera il sorgente ma (guarda caso) attribuendogli una licenza permissiva
ed un autore inesistente.
Quel codice, distribuito attraverso l'editor per programmatori di Microsoft chiamato
Visual Studio Code è stato riconosciuto perché particolarmente famoso, ma è inevitabile
che analogje violazioni avvengano continuamente senza che nessuno se ne accorga.
Violazioni particolarmente gravi perché il codice GPL viene poi incluso in prodotti proprietari.
> cito le parole di un altro autore, Jeff Jarvis:
> https://www.facebook.com/jeff.jarvis/posts/pfbid0LMFeqdTYoxnGHQAZwp5HMmeeVq…
Il fatto che Facebook propini e diffonda le parole di autori felici che le proprie
opere vengano sfruttate in questo modo e ricostruite secondo gli interessi
propagandistici di questa o quella società statunitense, non significa molto.
Piuttosto, evidenzia la scarsa consapevolezza del mezzo facebook (intermediario
interessato e notoriamente senza scrupoli) di chi se le beve e le diffonde.
Personalmente sarei felicissimo di scoprire che fare uno zip di windows o office
è sufficiente a far decadere i diritti di Microsoft a su di esso.
E scommetto che lo sarebbero anche molti suoi dipendenti, che potrebbero distribuire
zip dei sorgenti su GitHub (magari sotto GPL, tanto poi CopyALot li suggerirà
ai concorrenti di Microsoft stessa con una licenza permissiva e attribuzione ad mentula).
L'importante è che l'abolizione dei cosiddetti "diritti di proprietà intellettuale" valga per
chiunque passi un contenuto soggetto agli stessi attraverso un programma software.
Se però questa abolizione non vale per i singoli esseri umani non deve valere
neanche per le aziende.
Perché nota bene: qui non siamo di fronte ad una primordiale intelligenza aliena cui
potremmo anche decidere generosamente di fornire accesso alla nostra cultura.
Qui siamo di fronte ad aziende che approfittano della straordinaria ignoranza informatica
cui è costretta la stragrande maggioranza della popolazione per comportarsi da legibus soluti,
violando per gli altri le stesse leggi che pretendono siano rispettate per sé.
Mi spiace che tu ti sia bevuto la favoletta della "intelligenza artificiale".
Non che sia colpa tua: la propaganda è potente e personalizzata.
(soprattutto se usi GMail! ;-)
Ma dentro un LLM non opera alcuna intelligenza, solo rappresentazioni vettoriali di
testi attraversate lungo tracciati statisticamente probabili selezionati in modo (pseudo)
casuale entro un errore accettabile... e tipicamente post-processati
per scartare gli output politicamente (NON eticamente!) problematici e problematizzanti.
Niente di più.
Giacomo
_______________________________________________
nexa mailing list
nexa(a)server-nexa.polito.it
https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
Sept. 29, 2023
Re: [nexa] ‘Biggest act of copyright theft in history’: thousands of Australian books allegedly used to train AI model | Australia news | The Guardian
by Giacomo Tesio
Ciao Fabio,
Il 29 Settembre 2023 08:24:57 UTC, Fabio Alemagna <falemagn(a)gmail.com> ha scritto:
> L'idea che istruire un modello...
Purtroppo l'idea di "istruire" una macchina è di per sé un'allucinazione.
Le macchine si costruiscono e (se sono programmabili) di programmano.
Non c'è nessuna mente che possa imparare lì dentro, perché le macchine non pensano.
Il fatto che possano essere programmate statisticamente per ingannare chi
non ne comprende il funzionamento ci dice che il loro studio andrebbe riservato
a chi lo comprende appieno (tanto da poterle ricostruire da zero) e la loro applicazione
a persone inconsapevoli o fragili semplicemente vietato.
Ciò che chiami "modello" non è stato istruito ma programmato statisticamente
usando determinati testi "sorgente".
Il "modello" rappresenta una codifica parziale (o se peferisci, una compressione con
perdita di informazione) con interferenze (le varie sorgenti casuali utilizzate durante la
programmazione statistica o durante l'esecuzione del programma
e poi scartate per poter fingere che l'output non sia deterministico).
Dunque il modello CONTIENE, seppur in forma difficile da estrarre e non
necessariamente corrispondente all'intento comunicativo dei rispettivi autori,
ampie parti dei testi originali.
Un esempio particolarmente lampante di questo meccanismo fu evidenziato con
Microsoft Copilot (aka CopyALot) che distribuì codice sotto GPL in violazione della stessa,
copiando alla lettera il sorgente ma (guarda caso) attribuendogli una licenza permissiva
ed un autore inesistente.
Quel codice, distribuito attraverso l'editor per programmatori di Microsoft chiamato
Visual Studio Code è stato riconosciuto perché particolarmente famoso, ma è inevitabile
che analogje violazioni avvengano continuamente senza che nessuno se ne accorga.
Violazioni particolarmente gravi perché il codice GPL viene poi incluso in prodotti proprietari.
> cito le parole di un altro autore, Jeff Jarvis:
> https://www.facebook.com/jeff.jarvis/posts/pfbid0LMFeqdTYoxnGHQAZwp5HMmeeVq…
Il fatto che Facebook propini e diffonda le parole di autori felici che le proprie
opere vengano sfruttate in questo modo e ricostruite secondo gli interessi
propagandistici di questa o quella società statunitense, non significa molto.
Piuttosto, evidenzia la scarsa consapevolezza del mezzo facebook (intermediario
interessato e notoriamente senza scrupoli) di chi se le beve e le diffonde.
Personalmente sarei felicissimo di scoprire che fare uno zip di windows o office
è sufficiente a far decadere i diritti di Microsoft a su di esso.
E scommetto che lo sarebbero anche molti suoi dipendenti, che potrebbero distribuire
zip dei sorgenti su GitHub (magari sotto GPL, tanto poi CopyALot li suggerirà
ai concorrenti di Microsoft stessa con una licenza permissiva e attribuzione ad mentula).
L'importante è che l'abolizione dei cosiddetti "diritti di proprietà intellettuale" valga per
chiunque passi un contenuto soggetto agli stessi attraverso un programma software.
Se però questa abolizione non vale per i singoli esseri umani non deve valere
neanche per le aziende.
Perché nota bene: qui non siamo di fronte ad una primordiale intelligenza aliena cui
potremmo anche decidere generosamente di fornire accesso alla nostra cultura.
Qui siamo di fronte ad aziende che approfittano della straordinaria ignoranza informatica
cui è costretta la stragrande maggioranza della popolazione per comportarsi da legibus soluti,
violando per gli altri le stesse leggi che pretendono siano rispettate per sé.
Mi spiace che tu ti sia bevuto la favoletta della "intelligenza artificiale".
Non che sia colpa tua: la propaganda è potente e personalizzata.
(soprattutto se usi GMail! ;-)
Ma dentro un LLM non opera alcuna intelligenza, solo rappresentazioni vettoriali di
testi attraversate lungo tracciati statisticamente probabili selezionati in modo (pseudo)
casuale entro un errore accettabile... e tipicamente post-processati
per scartare gli output politicamente (NON eticamente!) problematici e problematizzanti.
Niente di più.
Giacomo
Sept. 29, 2023
Re: [nexa] ‘Biggest act of copyright theft in history’: thousands of Australian books allegedly used to train AI model | Australia news | The Guardian
by Alberto Cammozzo
Caro Fabio,
convengo con te che il (c) abbia dei limiti in tale contesto, e
soprattutto non credo che il /fair use/ sia l'unico criterio utile.
Le corti in varie parti del mondo saranno sempre più investite in
merito, vedremo che orientamento prenderanno e se la legge sul (c) sarà
l'unico strumento azionato. Gli LLM sfidano un apparato giuridico che
non è pensato per la produzione industriale di testi, fenomeno che
finora non esisteva.
In merito a quanto dici sull'equivalenza per macchine ed umani di
'/istruirsi/' sullo stesso testo, per parte mia credo che
l'/apprendimento/ umano e il /training/ del modello LLM abbiano una
enorme differenza.
L'umano è in grado di esprimersi col linguaggio anche senza i testi in
questione, e di estrarre le 'idee' in esso contenute prescindendo dalla
formulazione esatta, mentre la macchina produce linguaggio
statisticamente correlato con la semantica associata a quelle idee solo
seguendo la formulazione linguistica dei testi pertinenti, e solo con
tali testi. Non potrei dire la stessa cosa pensando a uno studente che
/si istruisce/ dai libri.
Per la macchina l'idea (che per noi è significante, contenuto) non
esiste, ma solo il linguaggio (il significante): anche se questa non
produrrà frasi che copiano letteralmente l'input di training, lo
specifico input è essenziale alla produzione di testi con la semantica
dell'input in questione.
Vedrei poi altri aspetti che le corti potrebbero tenere in
considerazione, che emergeranno forse maggiormente in futuro, ma che
meriterebbero approfondimento ora.
- l'appropriazione del lavoro linguistico non riconosciuto dell'autore,
anche come collettività autorale. Questo mi pare si veda già con la
produzione automatica di codice informatico attingendo ai repository;
- la responsabilità sulle conseguenze del contenuto dei testi e
eventuali danni derivanti dalla scarsa qualità dello stesso (già vediamo
ricette velenose, fake news, suicidi, induzione a comportamenti
pericolosi, bug nel codice ...);
- l'inquinamento ambientale. Non solo quello energetico (per la
produzione e l'aggiornamento degli LLM), ma inquinamento informativo
dell'ecosistema linguistico (o in generale simbolico). Ammettendo il
linguaggio come bene sistemico e patrimonio comune, l'immissione massiva
di testi (o immagini) generati artificialmente interferisce con
l'ecosistema e la sua naturale evoluzione.
Questo ultimo aspetto è quello meno immediatamente visibile ma sarà il
più intenso, e investirà per primi i motori di ricerca e gli altri
attori dell'ecosistema digitale, che vedranno diluirsi il rapporto
segnale/rumore all'aumentare dei testi generati artificialmente, e di
conseguenza il valore del loro servizio. Dovranno decidere da che parte
stare... Anche la produzione di software a codice aperto risentirà dello
stesso problema.
In generale le collettività che producono testi, codice e immagini e che
li riversano nei commons ne subiscono un danno dal momento che la
produzione industriale artificiale sommergerà il loro ecosistema con
prodotti di qualità dubbia e che sottraggono loro lavoro e riconoscimento.
Stiamo (anche) a vedere...
Alberto
On 29/09/23 10:24, Fabio Alemagna wrote:
> L'idea che istruire un modello su dei testi coperti da copyright sia
> una violazione del suddetto copyright è altamente opinabile, e il
> ragionamento è in realtà abbastanza semplice: se istruirsi su un testo
> ne violasse il copyright, saremmo tutti dei criminali. Di quei testi
> non c'è fisicamente traccia all'interno dei modelli, non viene copiato
> niente. I modelli sono un'opera trasformativa di quei testi, non
> derivativa.
>
> Lo argomenta molto bene Creative Commons:
> https://creativecommons.org/2023/02/17/fair-use-training-generative-ai/
>
> Detto questo, cito le parole di un altro autore, Jeff Jarvis:
> https://www.facebook.com/jeff.jarvis/posts/pfbid0LMFeqdTYoxnGHQAZwp5HMmeeVq…
>
> «I, for one, am not complaining about my books being in in large
> language model training sets. I write to enter ideas into public
> discourse. I prefer informed over ignorant AI. I believe it is fair
> use for anyone to read & use books for transformative work. In fact,
> I'd probably feel snubbed if my books were not there. I'm happy when
> they are in libraries. I'm fine that they're here.»
>
> Fabio
>
> Il giorno ven 29 set 2023 alle ore 07:52 Alberto Cammozzo via nexa
> <nexa(a)server-nexa.polito.it> ha scritto:
>> <https://www.theguardian.com/australia-news/2023/sep/28/australian-books-tra…>
>>
>>
>> Thousands of books from some of Australia’s most celebrated authors have potentially been caught up in what Booker prize-winning novelist Richard Flanagan has called “the biggest act of copyright theft in history”.
>>
>> The works have allegedly been pirated by the US-based Books3 dataset and used to train generative AI for corporations such as Meta and Bloomberg.
>>
>> Flanagan, who found 10 of his works, including the multi-international award-winning 2013 novel The Narrow Road to the Deep North, on the Books3 dataset, told Guardian Australia he was deeply shocked by the discovery made several days ago.
>>
>> “I felt as if my soul had been strip mined and I was powerless to stop it,” he said in a statement.
>>
>> “This is the biggest act of copyright theft in history.”
>>
>> AI could ‘turbo-charge fraud’ and be monopolised by tech companies, Andrew Leigh warns
>>
>>
>> The Australian Publishers Association confirmed to Guardian Australia on Wednesday that as many as 18,000 fiction and nonfiction titles with Australian ISBNs (unique international standard book numbers) appeared to be affected by the copyright infringement, although it is not yet clear what proportion of these are Australian editions of internationally authored books.
>>
>> “We’re still working through [the data] to work out the impact in terms of Australian authors,” APA spokesperson Stuart Glover said.
>>
>> “This is a massive legal and ethical challenge for the publishing industry and for authors globally.”
>>
>> A search tool published on Monday by US media platform The Atlantic and uploaded by the US Authors Guild on Wednesday revealed the works of Peter Carey, Helen Garner, Kate Grenville, Anna Funder, Christos Tsiolkas and Thomas Keneally, as well as Flanagan and dozens of other high-profile Australian authors, were included in the pirated dataset containing more than 180,000 titles.
>>
>> On Thursday, the Australian Society of Authors issued a statement saying it was “horrified” to learn that the works of Australian writers were being used to train artificial intelligence without permission from the authors.
>>
>> ASA chief executive, Olivia Lanchester, described the Books3 dataset as piracy on an industrial scale.
>>
>> “Authors appropriately feel outraged,” Lanchester said. “The fact is this technology relies upon books, journals, essays written by authors, yet permission was not sought nor compensation granted.”
>>
>> Lanchester said the Australian literary industry, while not objecting per se to emerging technologies such as AI, was deeply concerned about the lack of transparency evident in the development and monetisation of AI by global tech companies.
>>
>> “Turning a blind eye to the legitimate rights of copyright owners threatens to diminish already precarious creative careers,” she said.
>>
>> “The enrichment of a few powerful companies is at the cost of thousands of individual creators. This is not how a fair market functions.”
>>
>> Josephine Johnston, chief executive of Australia’s Copyright Agency, described the Books3 development as “a free kick to big tech” at the expense of Australia’s creative and cultural life.
>>
>> “We’re going to need greater transparency – how these tools have been developed, trained, how they operate – before people can truly understand what their legal rights might be,” she said.
>>
>> “We seem to be in this terrible position now where content owners – remembering that the vast majority of them will be individual authors – may actually have to take out court cases to enforce their rights.”
>>
>> Australian copyright law protects creators of original content from data scraping.
>>
>> Litigation in the US against ChatGPT creator OpenAI over use of allegedly pirated book datasets, Books1 and Books2 (which do not appear to be affiliated with Books3) has already commenced.
>>
>>
>> In July, North American horror/fantasy writers Mona Awad (author of Bunny) and Paul Tremblay (author of The Cabin at the End of the World) filed a lawsuit in a San Francisco federal court, alleging ChatGPT unlawfully digested their books as part of its AI training data.
>>
>> On 28 August, OpenAI filed a motion to dismiss the lawsuit, arguing that the authors “misconceive the scope of copyright, failing to take into account the limitations and exceptions (including fair use) that properly leave room for innovations like the large language models now at the forefront of artificial intelligence”.
>>
>> On 19 September the Writers Guild and 17 of its members, including bestselling novelists John Grisham, George RR Martin and Jodi Picoult, filed a complaint in a New York district court against OpenAI, seeking redress for “flagrant and harmful infringements” of guild members’ registered copyrights.
>>
>> In a statement on its website, the guild says while it is aware that companies such as Meta and Bloomberg have used the Books3 dataset to train their LLMs, it is not yet clear whether OpenAI is using Books3 to train its ChatGPT models GPT 3.5 or GPT 4.
>>
>> Democracies face ‘truth decay’ as AI blurs fact and fiction, warns head of Australia’s military
>>
>>
>> Guardian Australia has sought comment from OpenAI, which has yet to officially respond to the guild’s complaint, and Meta.
>>
>> On 4 September, US technology magazine Wired reported that a Danish anti-piracy group called Rights Alliance had been told by Bloomberg that the company did not plan to train future versions of its BloombergGPT using Books3.
>>
>> Bloomberg declined to respond to the Guardian’s queries.
>>
>> The APA said the global nature of the issue would present significant challenges in enforcement and prosecution, and has joined the authors’ society in calling for AI technologies to be regulated.
>>
>> Consultation closed last month for a Department of Industry, Science and Resources discussion paper on supporting responsible AI.
>>
>> A parliamentary inquiry is under way examining the use of generative artificial intelligence in the Australian education system.
>>
>> Flanagan said it was up to the Australian government to act to protect Australia’s writers.
>>
>> “It has power and we do not,” he said.
>>
>> “If it cares for our culture it must now stand up and fight for it.”
>>
>> _______________________________________________
>> nexa mailing list
>> nexa(a)server-nexa.polito.it
>> https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
Sept. 29, 2023
Re: [nexa] ‘Biggest act of copyright theft in history’: thousands of Australian books allegedly used to train AI model | Australia news | The Guardian
by M. Fioretti
On Fri, Sep 29, 2023 10:24:57 AM +0200, Fabio Alemagna wrote:
> L'idea che istruire un modello su dei testi coperti da copyright sia
> una violazione del suddetto copyright è altamente opinabile, e il
> ragionamento è in realtà abbastanza semplice: se istruirsi su un
> testo ne violasse il copyright, saremmo tutti dei criminali.
esatto, grazie. Sono sorpreso che questo argomento non venga fatto
piu' spesso, mi pare cruciale. Inoltre, a sostegno di quanto dice Jarvis:
> I prefer informed over ignorant AI.
ora non trovo il link, ma giorni fa leggevo qualcun altro osservare
esplicitamente che se si lasciano addestrare i LLM solo con spazzatura
razzista, fascista, sessista, omofoba o altro, poi non ci si puo'
lamentare se creano testi razzisti eccetera.
Che poi, allargando un attimo il discorso, e' lo stesso problema
gigante che abbiamo gia' da vent'anni, il blocco da copyright per gli
LLM lo aumenterebbe soltanto:
https://stop.zona-m.net/2021/04/free-news-make-extremists/
Soluzioni non ne ho ma fa ridere, tristemente, vedere giornalisti e
intellettuali progressisti lamentarsi delle masse becere che credono
alla spazzatura online, spiegandogli pazientemente quanto son
becere... o combattere la disinformazione su vaccini, cambiamenti
climatici, Ucraina eccetera... ma solo in articoli dietro paywall
Infine, senza smentire quanto sopra, giusto per completezza:
> Di quei testi non c'è fisicamente traccia all'interno dei modelli,
> non viene copiato niente. I modelli sono un'opera trasformativa di
> quei testi, non derivativa.
volendo essere pignoli, questo non e' sempre **completamente vero**,
vedi il caso del "knotting",
https://www.wired.com/story/fanfiction-omegaverse-sex-trope-artificial-inte…
Marco
Sept. 29, 2023
Re: [nexa] ‘Biggest act of copyright theft in history’: thousands of Australian books allegedly used to train AI model | Australia news | The Guardian
by Fabio Alemagna
L'idea che istruire un modello su dei testi coperti da copyright sia
una violazione del suddetto copyright è altamente opinabile, e il
ragionamento è in realtà abbastanza semplice: se istruirsi su un testo
ne violasse il copyright, saremmo tutti dei criminali. Di quei testi
non c'è fisicamente traccia all'interno dei modelli, non viene copiato
niente. I modelli sono un'opera trasformativa di quei testi, non
derivativa.
Lo argomenta molto bene Creative Commons:
https://creativecommons.org/2023/02/17/fair-use-training-generative-ai/
Detto questo, cito le parole di un altro autore, Jeff Jarvis:
https://www.facebook.com/jeff.jarvis/posts/pfbid0LMFeqdTYoxnGHQAZwp5HMmeeVq…
«I, for one, am not complaining about my books being in in large
language model training sets. I write to enter ideas into public
discourse. I prefer informed over ignorant AI. I believe it is fair
use for anyone to read & use books for transformative work. In fact,
I'd probably feel snubbed if my books were not there. I'm happy when
they are in libraries. I'm fine that they're here.»
Fabio
Il giorno ven 29 set 2023 alle ore 07:52 Alberto Cammozzo via nexa
<nexa(a)server-nexa.polito.it> ha scritto:
>
> <https://www.theguardian.com/australia-news/2023/sep/28/australian-books-tra…>
>
>
> Thousands of books from some of Australia’s most celebrated authors have potentially been caught up in what Booker prize-winning novelist Richard Flanagan has called “the biggest act of copyright theft in history”.
>
> The works have allegedly been pirated by the US-based Books3 dataset and used to train generative AI for corporations such as Meta and Bloomberg.
>
> Flanagan, who found 10 of his works, including the multi-international award-winning 2013 novel The Narrow Road to the Deep North, on the Books3 dataset, told Guardian Australia he was deeply shocked by the discovery made several days ago.
>
> “I felt as if my soul had been strip mined and I was powerless to stop it,” he said in a statement.
>
> “This is the biggest act of copyright theft in history.”
>
> AI could ‘turbo-charge fraud’ and be monopolised by tech companies, Andrew Leigh warns
>
>
> The Australian Publishers Association confirmed to Guardian Australia on Wednesday that as many as 18,000 fiction and nonfiction titles with Australian ISBNs (unique international standard book numbers) appeared to be affected by the copyright infringement, although it is not yet clear what proportion of these are Australian editions of internationally authored books.
>
> “We’re still working through [the data] to work out the impact in terms of Australian authors,” APA spokesperson Stuart Glover said.
>
> “This is a massive legal and ethical challenge for the publishing industry and for authors globally.”
>
> A search tool published on Monday by US media platform The Atlantic and uploaded by the US Authors Guild on Wednesday revealed the works of Peter Carey, Helen Garner, Kate Grenville, Anna Funder, Christos Tsiolkas and Thomas Keneally, as well as Flanagan and dozens of other high-profile Australian authors, were included in the pirated dataset containing more than 180,000 titles.
>
> On Thursday, the Australian Society of Authors issued a statement saying it was “horrified” to learn that the works of Australian writers were being used to train artificial intelligence without permission from the authors.
>
> ASA chief executive, Olivia Lanchester, described the Books3 dataset as piracy on an industrial scale.
>
> “Authors appropriately feel outraged,” Lanchester said. “The fact is this technology relies upon books, journals, essays written by authors, yet permission was not sought nor compensation granted.”
>
> Lanchester said the Australian literary industry, while not objecting per se to emerging technologies such as AI, was deeply concerned about the lack of transparency evident in the development and monetisation of AI by global tech companies.
>
> “Turning a blind eye to the legitimate rights of copyright owners threatens to diminish already precarious creative careers,” she said.
>
> “The enrichment of a few powerful companies is at the cost of thousands of individual creators. This is not how a fair market functions.”
>
> Josephine Johnston, chief executive of Australia’s Copyright Agency, described the Books3 development as “a free kick to big tech” at the expense of Australia’s creative and cultural life.
>
> “We’re going to need greater transparency – how these tools have been developed, trained, how they operate – before people can truly understand what their legal rights might be,” she said.
>
> “We seem to be in this terrible position now where content owners – remembering that the vast majority of them will be individual authors – may actually have to take out court cases to enforce their rights.”
>
> Australian copyright law protects creators of original content from data scraping.
>
> Litigation in the US against ChatGPT creator OpenAI over use of allegedly pirated book datasets, Books1 and Books2 (which do not appear to be affiliated with Books3) has already commenced.
>
>
> In July, North American horror/fantasy writers Mona Awad (author of Bunny) and Paul Tremblay (author of The Cabin at the End of the World) filed a lawsuit in a San Francisco federal court, alleging ChatGPT unlawfully digested their books as part of its AI training data.
>
> On 28 August, OpenAI filed a motion to dismiss the lawsuit, arguing that the authors “misconceive the scope of copyright, failing to take into account the limitations and exceptions (including fair use) that properly leave room for innovations like the large language models now at the forefront of artificial intelligence”.
>
> On 19 September the Writers Guild and 17 of its members, including bestselling novelists John Grisham, George RR Martin and Jodi Picoult, filed a complaint in a New York district court against OpenAI, seeking redress for “flagrant and harmful infringements” of guild members’ registered copyrights.
>
> In a statement on its website, the guild says while it is aware that companies such as Meta and Bloomberg have used the Books3 dataset to train their LLMs, it is not yet clear whether OpenAI is using Books3 to train its ChatGPT models GPT 3.5 or GPT 4.
>
> Democracies face ‘truth decay’ as AI blurs fact and fiction, warns head of Australia’s military
>
>
> Guardian Australia has sought comment from OpenAI, which has yet to officially respond to the guild’s complaint, and Meta.
>
> On 4 September, US technology magazine Wired reported that a Danish anti-piracy group called Rights Alliance had been told by Bloomberg that the company did not plan to train future versions of its BloombergGPT using Books3.
>
> Bloomberg declined to respond to the Guardian’s queries.
>
> The APA said the global nature of the issue would present significant challenges in enforcement and prosecution, and has joined the authors’ society in calling for AI technologies to be regulated.
>
> Consultation closed last month for a Department of Industry, Science and Resources discussion paper on supporting responsible AI.
>
> A parliamentary inquiry is under way examining the use of generative artificial intelligence in the Australian education system.
>
> Flanagan said it was up to the Australian government to act to protect Australia’s writers.
>
> “It has power and we do not,” he said.
>
> “If it cares for our culture it must now stand up and fight for it.”
>
> _______________________________________________
> nexa mailing list
> nexa(a)server-nexa.polito.it
> https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
Sept. 29, 2023