nexa
By thread
nexa@server-nexa.polito.it
By month
Messages by month
- ----- 2026 -----
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2025 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2024 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2023 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2022 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2021 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2020 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2019 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2018 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2017 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2016 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2015 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2014 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2013 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2012 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2011 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2010 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2009 -----
- December
- November
- October
- September
- August
- July
- June
- May
September 2023
- 49 participants
- 220 messages
Europol Sought Unlimited Data Access in Online Child Sexual Abuse Regulation
by J.C. DE MARTIN
*Europol Sought Unlimited Data Access in Online Child Sexual Abuse
Regulation
*
/According to minutes released under FOI, the European police agency
pushed for unfiltered access to data that would be obtained under a
proposed new scanning system for detecting child sexual abuse images on
messaging apps, with a view, experts say, to training AI algorithms.
/
Apostolis Fotiadis, Luděk Stavinoha and Giacomo Zandonini
Athens, Norwich, Rome
BIRN
September 29, 202314:56
The European police agency, Europol, has requested unfiltered access to
data that would be harvested under a controversial EU proposal to scan
online content for child sexual abuse images and for the AI technology
behind it to be applied to other crimes too, according to minutes of a
high-level meeting in mid-2022.
The meeting, involving Europol Executive Director Catherine de Bolle and
the European Commission’s Director-General for Migration and Home
Affairs, Monique Pariat, took place in July last year, weeks after the
Commission unveiled a proposed regulation that would require digital
chat providers to scan client content for child sexual abuse material,
or CSAM.
The regulation, put forward by European Commissioner for Home Affairs
Ylva Johansson, would also create a new EU agency - the EU centre to
prevent and counter child sexual abuse. It has stirred heated debate,
with critics warning it risks opening the door to mass surveillance of
EU citizens.
[...]
continua qui:
https://balkaninsight.com/2023/09/29/europol-sought-unlimited-data-access-i…
Sept. 30, 2023
Re: [nexa] ‘Biggest act of copyright theft in history’: thousands of Australian books allegedly used to train AI model | Australia news | The Guardian
by GC F
Non capisco il senso di questa risposta quando la questione è tecnica e
riguarda l'applicazione delle norme sul diritto d'autore, la dicotomia
idea/espressione e l'applicazione di potenziali eccezioni e limitazioni o
usi privilegiati in base alla giurisdizione di riferimento.
In principio, condivido la posizione di Fabio, poichè la teoria generale
del diritto d'autore vorrebbe che si proteggessero espressioni e non dati o
informazioni estratte per fini ulteriori e trasformativi. Mi rendo però poi
anche conto delle complessità nell'applicare quel principio generale in
diritto europeo. Ho pochi dubbi invece che la dottrina del "fair use"
dovrebbe giustificare gli usi di contenuti protetti in processi di
machine learning. I casi ora pendenti mi smentiranno probabilmente ma mi
sembra che la posizione statunitense sia chiara dai tempi di Baker v Selden
(1879). Sono però anche d'accordo che un qualche soluzione, forse endogena
al diritto d'autore, dovrebbe essere proposta per evitare esternalità
negative rilevanti, anche se forse solo nel breve-medio periodo, sul
mercato della creatività.
Ne parlo in maniera esaustiva qui (anche per rispondere alle molteplici
domande che questo thread contiene):
'Generative AI in Court
<https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4558865>' in Nikos
Koutras and Niloufer Selvadurai (eds), Recreating Creativity, Reinventing
Inventiveness - International Perspectives on AI and IP Governance
(Routledge, Forthcoming)
la proposta menzionata sopra invece è qui:
'Should We Ban Generative AI, Incentivise it or Make it a Medium for
Inclusive Creativity?
<https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4527461>' in Enrico
Bonadio and Caterina Sganga (eds), A Research Agenda for EU Copyright Law
(Edward Elgar, Forthcoming)
Giancarlo
PS Il dibattito circa l'antropomorfizzazione linguistica di quel che fa la
macchina è ormai vecchio e stantio. Ci siano accordati nel dire che la
macchina "genera" e viene "istruita" tramite processi di machine learning,
anche perché ai fini del diritto d'autore chi istruisce potenzialmente
usando contenuti protetti in violazione di privativa altrui è un agente
umano, come poi definire l'effetto di tale processo di "istruzione" sulla
macchina mi pare davvero irrilevante--e probabilmente pretestuoso--nel
contesto di cui qui si discute.
On Fri, Sep 29, 2023 at 10:53 AM Giacomo Tesio <giacomo(a)tesio.it> wrote:
> Ciao Fabio,
>
> Il 29 Settembre 2023 08:24:57 UTC, Fabio Alemagna <falemagn(a)gmail.com> ha
> scritto:
> > L'idea che istruire un modello...
>
> Purtroppo l'idea di "istruire" una macchina è di per sé un'allucinazione.
>
> Le macchine si costruiscono e (se sono programmabili) di programmano.
>
> Non c'è nessuna mente che possa imparare lì dentro, perché le macchine non
> pensano.
>
> Il fatto che possano essere programmate statisticamente per ingannare chi
> non ne comprende il funzionamento ci dice che il loro studio andrebbe
> riservato
> a chi lo comprende appieno (tanto da poterle ricostruire da zero) e la
> loro applicazione
> a persone inconsapevoli o fragili semplicemente vietato.
>
>
> Ciò che chiami "modello" non è stato istruito ma programmato
> statisticamente
> usando determinati testi "sorgente".
>
> Il "modello" rappresenta una codifica parziale (o se peferisci, una
> compressione con
> perdita di informazione) con interferenze (le varie sorgenti casuali
> utilizzate durante la
> programmazione statistica o durante l'esecuzione del programma
> e poi scartate per poter fingere che l'output non sia deterministico).
>
> Dunque il modello CONTIENE, seppur in forma difficile da estrarre e non
> necessariamente corrispondente all'intento comunicativo dei rispettivi
> autori,
> ampie parti dei testi originali.
>
> Un esempio particolarmente lampante di questo meccanismo fu evidenziato
> con
> Microsoft Copilot (aka CopyALot) che distribuì codice sotto GPL in
> violazione della stessa,
> copiando alla lettera il sorgente ma (guarda caso) attribuendogli una
> licenza permissiva
> ed un autore inesistente.
>
> Quel codice, distribuito attraverso l'editor per programmatori di
> Microsoft chiamato
> Visual Studio Code è stato riconosciuto perché particolarmente famoso, ma
> è inevitabile
> che analogje violazioni avvengano continuamente senza che nessuno se ne
> accorga.
> Violazioni particolarmente gravi perché il codice GPL viene poi incluso in
> prodotti proprietari.
>
> > cito le parole di un altro autore, Jeff Jarvis:
> >
> https://www.facebook.com/jeff.jarvis/posts/pfbid0LMFeqdTYoxnGHQAZwp5HMmeeVq…
>
> Il fatto che Facebook propini e diffonda le parole di autori felici che le
> proprie
> opere vengano sfruttate in questo modo e ricostruite secondo gli interessi
> propagandistici di questa o quella società statunitense, non significa
> molto.
>
> Piuttosto, evidenzia la scarsa consapevolezza del mezzo facebook
> (intermediario
> interessato e notoriamente senza scrupoli) di chi se le beve e le diffonde.
>
>
> Personalmente sarei felicissimo di scoprire che fare uno zip di windows o
> office
> è sufficiente a far decadere i diritti di Microsoft a su di esso.
>
> E scommetto che lo sarebbero anche molti suoi dipendenti, che potrebbero
> distribuire
> zip dei sorgenti su GitHub (magari sotto GPL, tanto poi CopyALot li
> suggerirà
> ai concorrenti di Microsoft stessa con una licenza permissiva e
> attribuzione ad mentula).
>
> L'importante è che l'abolizione dei cosiddetti "diritti di proprietà
> intellettuale" valga per
> chiunque passi un contenuto soggetto agli stessi attraverso un programma
> software.
>
> Se però questa abolizione non vale per i singoli esseri umani non deve
> valere
> neanche per le aziende.
>
> Perché nota bene: qui non siamo di fronte ad una primordiale intelligenza
> aliena cui
> potremmo anche decidere generosamente di fornire accesso alla nostra
> cultura.
>
> Qui siamo di fronte ad aziende che approfittano della straordinaria
> ignoranza informatica
> cui è costretta la stragrande maggioranza della popolazione per
> comportarsi da legibus soluti,
> violando per gli altri le stesse leggi che pretendono siano rispettate per
> sé.
>
>
> Mi spiace che tu ti sia bevuto la favoletta della "intelligenza
> artificiale".
> Non che sia colpa tua: la propaganda è potente e personalizzata.
> (soprattutto se usi GMail! ;-)
>
> Ma dentro un LLM non opera alcuna intelligenza, solo rappresentazioni
> vettoriali di
> testi attraversate lungo tracciati statisticamente probabili selezionati
> in modo (pseudo)
> casuale entro un errore accettabile... e tipicamente post-processati
> per scartare gli output politicamente (NON eticamente!) problematici e
> problematizzanti.
>
>
> Niente di più.
>
>
> Giacomo
> _______________________________________________
> nexa mailing list
> nexa(a)server-nexa.polito.it
> https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
>
Sept. 30, 2023
R: R: R: R: ‘Biggest act of copyright theft in history’: thousands of Australian books allegedly used to train AI model | Australia news | The Guardian
by Lorenzo Albertini
Circa il contratto su ebook, valgono le clausole ivi inserite .
Circa l'ocr da cartaceo, si torna a dover verificare se ciò comporti o meno <riproduzione>. In caso positivo, vedrei poche eccezioni/limitazioni/controdiritti azionabili da OpenAI o simili allenatori di AI ...
>>-----Messaggio originale-----
>>Da: nexa <nexa-bounces(a)server-nexa.polito.it> Per conto di Stefano Zacchiroli
>>Inviato: venerdì 29 settembre 2023 19:48
>>A: 'Nexa' <nexa(a)server-nexa.polito.it>
>>Oggetto: Re: [nexa] R: R: R: ‘Biggest act of copyright theft in history’:
>>thousands of Australian books allegedly used to train AI model | Australia
>>news | The Guardian
>>
>>On Fri, Sep 29, 2023 at 05:33:04PM +0200, Lorenzo Albertini wrote:
>>> Non so in usa (fair use, che ha ambito applicativo alquanto vasto?)
>>
>>La posizione (per ora non verificata in tribunale) di Microsoft/GitHub per
>>Copilot è esattamente che il training di modelli ML secondo il copyright
>>americano costituisca fair use. Nel caso dei libri questo non evacua però
>>(credo, IANAL, etc.) la domanda di Stefano Quintarelli su se un contratto d'uso
>>di un ebook impedisca comunque il training. (Nel caso di Microsoft/GitHub la
>>domanda invece non si poneva, perché avevano già tutto il codice che
>>"gentilmente" milioni di sviluppatori gli hanno chiesto di ospitare...)
>>
>>--
>>Stefano Zacchiroli . zack(a)upsilon.cc . https://upsilon.cc/zack _. ^ ._
>>Full professor of Computer Science o o o \/|V|\/
>>Télécom Paris, Polytechnic Institute of Paris o o o </> <\>
>>Co-founder & CTO Software Heritage o o o o /\|^|/\
>>https://twitter.com/zacchiro . https://mastodon.xyz/@zacchiro '" V "'
>>_______________________________________________
>>nexa mailing list
>>nexa(a)server-nexa.polito.it
>>https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
Sept. 29, 2023
Re: [nexa] ‘Biggest act of copyright theft in history’: thousands of Australian books allegedly used to train AI model | Australia news | The Guardian
by Maurizio Borghi
È una riproduzione. In UE è permessa solo per scopo di text and data mining
per ricerca senza finalità commerciali, oppure commerciali se l’autore non
ha riservato il diritto. In USA si entra nel limbo del fair use, su cui
nessuna corte di è ancora pronunciata.
M.
On Fri, 29 Sep 2023 at 20:03, Stefano Quintarelli <stefano(a)quintarelli.it>
wrote:
> grazie
>
> e per quanto riguarda l'uso di testi generati facendo scansione ed OCR ?
>
> grazie!, s.
>
> On 29/09/23 19:58, Maurizio Borghi wrote:
> > Caro Stefano, il tuo ragionamento è corretto. Aggiungo che la rimozione
> del DRM è già
> > perseguibile come violazione del diritto d’autore, non solo nei paesi
> UE, ma anche in USA.
> > Se l’opera non è protetta da DRM il discorso è un po’ più complicato: in
> UE, il contratto
> > di licenza non può escludere certi usi consentiti, in particolare il
> text and data mining
> > per scopi non commerciali. Può però escludere lo stesso utilizzo se per
> scopi commerciali
> > e se l’uso è espressamente riservato. In USA non ci sono regole precise,
> ma la libertà
> > contrattuale tende di solito a prevalere sulla disponibilità di
> eccezioni (fair use). Non
> > è un caso che nella class action contro GitHub / Copilot i claim si
> basino interamente su
> > violazione dei contratti di licenza (open source) e sulla rimozione dei
> DRM, anziché sulla
> > violazione del copyright nel software utilizzato per addestrare
> l’algoritmo.
> > Un caro saluto
> > Maurizio
> >
> >
> >
> > On Fri, 29 Sep 2023 at 15:21, Stefano Quintarelli <
> stefano(a)quintarelli.it
> > <mailto:stefano@quintarelli.it>> wrote:
> >
> > Ho una domanda per i giuristi (anzi, piu' di una)
> >
> > per allenare un modello, ho bisogno di un file con la versione
> digitale di un testo.
> > (cosnsidero ovviamente testi non PD, CC0, ecc.)
> >
> > la versione digitale di un testo la posso ottenere da un ebook (gia'
> digitale), togliendo
> > il probabile DRM.
> > ma un ebook non e' unbene ma e' un servizio soggetto a licenza
> d'uso, quindi se non e'
> > prevista nella licenza d'uso la facolta' di estrarre il testo
> digitale per allenarci un
> > modello, mi sembra che ci sia gia' una violazione della licenza, per
> cui, credo, non
> > possa
> > essere usato come base di un allenamento, tanto piu' se il fine di
> tale allenamento e'
> > commerciale (se vendo un servizio basato su quel modello).
> >
> > se e' cosi', per allenare il mio modello devo allora prednere il
> testo digitale facendo
> > scan/ocr di un testo cartaceo.
> > ma cio' e' possibile, se non erro, solo per uso personale e non
> commerciale.
> >
> > se questo e' corretto, non mi pare ci sia un modo per prendere un
> testo digitale senza
> > infrangere una licenza d'uso/copyright
> >
> > dove e' la fallacia del ragionamento ?
> >
> > grazie, s.
> >
> > On 29/09/23 15:00, Stefano Borroni Barale wrote:
> > > Buongiorno lista,
> > >
> > >> L'idea che istruire un modello su dei testi coperti da copyright
> sia una
> > violazione del suddetto copyright è altamente opinabile
> > >
> > > Fin qui, ho l'impressione che tutti i legali in lista
> concorderanno.
> > >
> > >> ragionamento è in realtà abbastanza semplice: se istruirsi su un
> > >> testo ne violasse il copyright, saremmo tutti dei criminali.
> > >
> > > Ma siccome noi siamo umani e quello che produciamo non è - salvo
> i discorsi dei
> > politici(*) - ontologicamente identico alla produzione di esseri
> tecnici non viventi,
> > logica vuole che quanto si applica a noi non possa applicarsi a un
> LLM, tanto quanto
> > la legge sul copyright non si applica pedissequamente all'utilizzo
> di testi umani per
> > creare modelli linguistici.
> > >
> > > Questo è il motivo per il quale tutti i tentativi di "proteggere
> via copyright" il
> > prodotto di software generativi sono falliti miseramente, e con
> motivazioni scritte in
> > sentenze; che per il diritto credo abbiano un peso assai maggiore
> del sito di CC.
> > >
> > > La mia impressione è che la questione terrà impegnati legali,
> informatici, filosofi
> > e società ancora moooooolto a lungo.
> > > SBB
> > >
> > > (*) Come sanno bene i bambini degli anni '80 che hanno giocato
> con questo spassoso
> > giocattolo: https://www.enricodalbosco.it/giochi/tubolario/
> > <https://www.enricodalbosco.it/giochi/tubolario/>
> > >
> > >
> > > Di quei testi
> > >> non c'è fisicamente traccia all'interno dei modelli, non viene
> copiato
> > >> niente. I modelli sono un'opera trasformativa di quei testi, non
> > >> derivativa.
> > >>
> > >> Lo argomenta molto bene Creative Commons:
> > >>
> https://creativecommons.org/2023/02/17/fair-use-training-generative-ai/
> > <
> https://creativecommons.org/2023/02/17/fair-use-training-generative-ai/>
> > >>
> > >> Detto questo, cito le parole di un altro autore, Jeff Jarvis:
> > >>
> >
> https://www.facebook.com/jeff.jarvis/posts/pfbid0LMFeqdTYoxnGHQAZwp5HMmeeVq…
> <
> https://www.facebook.com/jeff.jarvis/posts/pfbid0LMFeqdTYoxnGHQAZwp5HMmeeVq…
> >
> > >>
> > >> «I, for one, am not complaining about my books being in in large
> > >> language model training sets. I write to enter ideas into public
> > >> discourse. I prefer informed over ignorant AI. I believe it is
> fair
> > >> use for anyone to read & use books for transformative work. In
> fact,
> > >> I'd probably feel snubbed if my books were not there. I'm happy
> when
> > >> they are in libraries. I'm fine that they're here.»
> > >>
> > >> Fabio
> > >>
> > >> Il giorno ven 29 set 2023 alle ore 07:52 Alberto Cammozzo via
> nexa
> > >> nexa(a)server-nexa.polito.it <mailto:nexa@server-nexa.polito.it>
> ha scritto:
> > >>
> > >>>
> >
> https://www.theguardian.com/australia-news/2023/sep/28/australian-books-tra…
> <
> https://www.theguardian.com/australia-news/2023/sep/28/australian-books-tra…
> >
> > >>>
> > >>> Thousands of books from some of Australia’s most celebrated
> authors have
> > potentially been caught up in what Booker prize-winning novelist
> Richard Flanagan has
> > called “the biggest act of copyright theft in history”.
> > >>>
> > >>> The works have allegedly been pirated by the US-based Books3
> dataset and used to
> > train generative AI for corporations such as Meta and Bloomberg.
> > >>>
> > >>> Flanagan, who found 10 of his works, including the
> multi-international
> > award-winning 2013 novel The Narrow Road to the Deep North, on the
> Books3 dataset,
> > told Guardian Australia he was deeply shocked by the discovery made
> several days ago.
> > >>>
> > >>> “I felt as if my soul had been strip mined and I was powerless
> to stop it,” he
> > said in a statement.
> > >>>
> > >>> “This is the biggest act of copyright theft in history.”
> > >>>
> > >>> AI could ‘turbo-charge fraud’ and be monopolised by tech
> companies, Andrew Leigh
> > warns
> > >>>
> > >>> The Australian Publishers Association confirmed to Guardian
> Australia on
> > Wednesday that as many as 18,000 fiction and nonfiction titles with
> Australian ISBNs
> > (unique international standard book numbers) appeared to be affected
> by the copyright
> > infringement, although it is not yet clear what proportion of these
> are Australian
> > editions of internationally authored books.
> > >>>
> > >>> “We’re still working through [the data] to work out the impact
> in terms of
> > Australian authors,” APA spokesperson Stuart Glover said.
> > >>>
> > >>> “This is a massive legal and ethical challenge for the
> publishing industry and
> > for authors globally.”
> > >>>
> > >>> A search tool published on Monday by US media platform The
> Atlantic and uploaded
> > by the US Authors Guild on Wednesday revealed the works of Peter
> Carey, Helen Garner,
> > Kate Grenville, Anna Funder, Christos Tsiolkas and Thomas Keneally,
> as well as
> > Flanagan and dozens of other high-profile Australian authors, were
> included in the
> > pirated dataset containing more than 180,000 titles.
> > >>>
> > >>> On Thursday, the Australian Society of Authors issued a
> statement saying it was
> > “horrified” to learn that the works of Australian writers were being
> used to train
> > artificial intelligence without permission from the authors.
> > >>>
> > >>> ASA chief executive, Olivia Lanchester, described the Books3
> dataset as piracy on
> > an industrial scale.
> > >>>
> > >>> “Authors appropriately feel outraged,” Lanchester said. “The
> fact is this
> > technology relies upon books, journals, essays written by authors,
> yet permission was
> > not sought nor compensation granted.”
> > >>>
> > >>> Lanchester said the Australian literary industry, while not
> objecting per se to
> > emerging technologies such as AI, was deeply concerned about the
> lack of transparency
> > evident in the development and monetisation of AI by global tech
> companies.
> > >>>
> > >>> “Turning a blind eye to the legitimate rights of copyright
> owners threatens to
> > diminish already precarious creative careers,” she said.
> > >>>
> > >>> “The enrichment of a few powerful companies is at the cost of
> thousands of
> > individual creators. This is not how a fair market functions.”
> > >>>
> > >>> Josephine Johnston, chief executive of Australia’s Copyright
> Agency, described
> > the Books3 development as “a free kick to big tech” at the expense
> of Australia’s
> > creative and cultural life.
> > >>>
> > >>> “We’re going to need greater transparency – how these tools
> have been developed,
> > trained, how they operate – before people can truly understand what
> their legal rights
> > might be,” she said.
> > >>>
> > >>> “We seem to be in this terrible position now where content
> owners – remembering
> > that the vast majority of them will be individual authors – may
> actually have to take
> > out court cases to enforce their rights.”
> > >>>
> > >>> Australian copyright law protects creators of original content
> from data scraping.
> > >>>
> > >>> Litigation in the US against ChatGPT creator OpenAI over use of
> allegedly pirated
> > book datasets, Books1 and Books2 (which do not appear to be
> affiliated with Books3)
> > has already commenced.
> > >>>
> > >>> In July, North American horror/fantasy writers Mona Awad
> (author of Bunny) and
> > Paul Tremblay (author of The Cabin at the End of the World) filed a
> lawsuit in a San
> > Francisco federal court, alleging ChatGPT unlawfully digested their
> books as part of
> > its AI training data.
> > >>>
> > >>> On 28 August, OpenAI filed a motion to dismiss the lawsuit,
> arguing that the
> > authors “misconceive the scope of copyright, failing to take into
> account the
> > limitations and exceptions (including fair use) that properly leave
> room for
> > innovations like the large language models now at the forefront of
> artificial
> > intelligence”.
> > >>>
> > >>> On 19 September the Writers Guild and 17 of its members,
> including bestselling
> > novelists John Grisham, George RR Martin and Jodi Picoult, filed a
> complaint in a New
> > York district court against OpenAI, seeking redress for “flagrant
> and harmful
> > infringements” of guild members’ registered copyrights.
> > >>>
> > >>> In a statement on its website, the guild says while it is aware
> that companies
> > such as Meta and Bloomberg have used the Books3 dataset to train
> their LLMs, it is not
> > yet clear whether OpenAI is using Books3 to train its ChatGPT models
> GPT 3.5 or GPT 4.
> > >>>
> > >>> Democracies face ‘truth decay’ as AI blurs fact and fiction,
> warns head of
> > Australia’s military
> > >>>
> > >>> Guardian Australia has sought comment from OpenAI, which has
> yet to officially
> > respond to the guild’s complaint, and Meta.
> > >>>
> > >>> On 4 September, US technology magazine Wired reported that a
> Danish anti-piracy
> > group called Rights Alliance had been told by Bloomberg that the
> company did not plan
> > to train future versions of its BloombergGPT using Books3.
> > >>>
> > >>> Bloomberg declined to respond to the Guardian’s queries.
> > >>>
> > >>> The APA said the global nature of the issue would present
> significant challenges
> > in enforcement and prosecution, and has joined the authors’ society
> in calling for AI
> > technologies to be regulated.
> > >>>
> > >>> Consultation closed last month for a Department of Industry,
> Science and
> > Resources discussion paper on supporting responsible AI.
> > >>>
> > >>> A parliamentary inquiry is under way examining the use of
> generative artificial
> > intelligence in the Australian education system.
> > >>>
> > >>> Flanagan said it was up to the Australian government to act to
> protect
> > Australia’s writers.
> > >>>
> > >>> “It has power and we do not,” he said.
> > >>>
> > >>> “If it cares for our culture it must now stand up and fight for
> it.”
> > >>>
> > >>> _______________________________________________
> > >>> nexa mailing list
> > >>> nexa(a)server-nexa.polito.it <mailto:nexa@server-nexa.polito.it>
> > >>> https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
> > <https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa>
> > >>
> > >> _______________________________________________
> > >> nexa mailing list
> > >> nexa(a)server-nexa.polito.it <mailto:nexa@server-nexa.polito.it>
> > >> https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
> > <https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa>
> > > _______________________________________________
> > > nexa mailing list
> > > nexa(a)server-nexa.polito.it <mailto:nexa@server-nexa.polito.it>
> > > https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
> > <https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa>
> > _______________________________________________
> > nexa mailing list
> > nexa(a)server-nexa.polito.it <mailto:nexa@server-nexa.polito.it>
> > https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
> > <https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa>
> >
>
Sept. 29, 2023
Re: [nexa] ‘Biggest act of copyright theft in history’: thousands of Australian books allegedly used to train AI model | Australia news | The Guardian
by Stefano Quintarelli
grazie
e per quanto riguarda l'uso di testi generati facendo scansione ed OCR ?
grazie!, s.
On 29/09/23 19:58, Maurizio Borghi wrote:
> Caro Stefano, il tuo ragionamento è corretto. Aggiungo che la rimozione del DRM è già
> perseguibile come violazione del diritto d’autore, non solo nei paesi UE, ma anche in USA.
> Se l’opera non è protetta da DRM il discorso è un po’ più complicato: in UE, il contratto
> di licenza non può escludere certi usi consentiti, in particolare il text and data mining
> per scopi non commerciali. Può però escludere lo stesso utilizzo se per scopi commerciali
> e se l’uso è espressamente riservato. In USA non ci sono regole precise, ma la libertà
> contrattuale tende di solito a prevalere sulla disponibilità di eccezioni (fair use). Non
> è un caso che nella class action contro GitHub / Copilot i claim si basino interamente su
> violazione dei contratti di licenza (open source) e sulla rimozione dei DRM, anziché sulla
> violazione del copyright nel software utilizzato per addestrare l’algoritmo.
> Un caro saluto
> Maurizio
>
>
>
> On Fri, 29 Sep 2023 at 15:21, Stefano Quintarelli <stefano(a)quintarelli.it
> <mailto:stefano@quintarelli.it>> wrote:
>
> Ho una domanda per i giuristi (anzi, piu' di una)
>
> per allenare un modello, ho bisogno di un file con la versione digitale di un testo.
> (cosnsidero ovviamente testi non PD, CC0, ecc.)
>
> la versione digitale di un testo la posso ottenere da un ebook (gia' digitale), togliendo
> il probabile DRM.
> ma un ebook non e' unbene ma e' un servizio soggetto a licenza d'uso, quindi se non e'
> prevista nella licenza d'uso la facolta' di estrarre il testo digitale per allenarci un
> modello, mi sembra che ci sia gia' una violazione della licenza, per cui, credo, non
> possa
> essere usato come base di un allenamento, tanto piu' se il fine di tale allenamento e'
> commerciale (se vendo un servizio basato su quel modello).
>
> se e' cosi', per allenare il mio modello devo allora prednere il testo digitale facendo
> scan/ocr di un testo cartaceo.
> ma cio' e' possibile, se non erro, solo per uso personale e non commerciale.
>
> se questo e' corretto, non mi pare ci sia un modo per prendere un testo digitale senza
> infrangere una licenza d'uso/copyright
>
> dove e' la fallacia del ragionamento ?
>
> grazie, s.
>
> On 29/09/23 15:00, Stefano Borroni Barale wrote:
> > Buongiorno lista,
> >
> >> L'idea che istruire un modello su dei testi coperti da copyright sia una
> violazione del suddetto copyright è altamente opinabile
> >
> > Fin qui, ho l'impressione che tutti i legali in lista concorderanno.
> >
> >> ragionamento è in realtà abbastanza semplice: se istruirsi su un
> >> testo ne violasse il copyright, saremmo tutti dei criminali.
> >
> > Ma siccome noi siamo umani e quello che produciamo non è - salvo i discorsi dei
> politici(*) - ontologicamente identico alla produzione di esseri tecnici non viventi,
> logica vuole che quanto si applica a noi non possa applicarsi a un LLM, tanto quanto
> la legge sul copyright non si applica pedissequamente all'utilizzo di testi umani per
> creare modelli linguistici.
> >
> > Questo è il motivo per il quale tutti i tentativi di "proteggere via copyright" il
> prodotto di software generativi sono falliti miseramente, e con motivazioni scritte in
> sentenze; che per il diritto credo abbiano un peso assai maggiore del sito di CC.
> >
> > La mia impressione è che la questione terrà impegnati legali, informatici, filosofi
> e società ancora moooooolto a lungo.
> > SBB
> >
> > (*) Come sanno bene i bambini degli anni '80 che hanno giocato con questo spassoso
> giocattolo: https://www.enricodalbosco.it/giochi/tubolario/
> <https://www.enricodalbosco.it/giochi/tubolario/>
> >
> >
> > Di quei testi
> >> non c'è fisicamente traccia all'interno dei modelli, non viene copiato
> >> niente. I modelli sono un'opera trasformativa di quei testi, non
> >> derivativa.
> >>
> >> Lo argomenta molto bene Creative Commons:
> >> https://creativecommons.org/2023/02/17/fair-use-training-generative-ai/
> <https://creativecommons.org/2023/02/17/fair-use-training-generative-ai/>
> >>
> >> Detto questo, cito le parole di un altro autore, Jeff Jarvis:
> >>
> https://www.facebook.com/jeff.jarvis/posts/pfbid0LMFeqdTYoxnGHQAZwp5HMmeeVq… <https://www.facebook.com/jeff.jarvis/posts/pfbid0LMFeqdTYoxnGHQAZwp5HMmeeVq…>
> >>
> >> «I, for one, am not complaining about my books being in in large
> >> language model training sets. I write to enter ideas into public
> >> discourse. I prefer informed over ignorant AI. I believe it is fair
> >> use for anyone to read & use books for transformative work. In fact,
> >> I'd probably feel snubbed if my books were not there. I'm happy when
> >> they are in libraries. I'm fine that they're here.»
> >>
> >> Fabio
> >>
> >> Il giorno ven 29 set 2023 alle ore 07:52 Alberto Cammozzo via nexa
> >> nexa(a)server-nexa.polito.it <mailto:nexa@server-nexa.polito.it> ha scritto:
> >>
> >>>
> https://www.theguardian.com/australia-news/2023/sep/28/australian-books-tra… <https://www.theguardian.com/australia-news/2023/sep/28/australian-books-tra…>
> >>>
> >>> Thousands of books from some of Australia’s most celebrated authors have
> potentially been caught up in what Booker prize-winning novelist Richard Flanagan has
> called “the biggest act of copyright theft in history”.
> >>>
> >>> The works have allegedly been pirated by the US-based Books3 dataset and used to
> train generative AI for corporations such as Meta and Bloomberg.
> >>>
> >>> Flanagan, who found 10 of his works, including the multi-international
> award-winning 2013 novel The Narrow Road to the Deep North, on the Books3 dataset,
> told Guardian Australia he was deeply shocked by the discovery made several days ago.
> >>>
> >>> “I felt as if my soul had been strip mined and I was powerless to stop it,” he
> said in a statement.
> >>>
> >>> “This is the biggest act of copyright theft in history.”
> >>>
> >>> AI could ‘turbo-charge fraud’ and be monopolised by tech companies, Andrew Leigh
> warns
> >>>
> >>> The Australian Publishers Association confirmed to Guardian Australia on
> Wednesday that as many as 18,000 fiction and nonfiction titles with Australian ISBNs
> (unique international standard book numbers) appeared to be affected by the copyright
> infringement, although it is not yet clear what proportion of these are Australian
> editions of internationally authored books.
> >>>
> >>> “We’re still working through [the data] to work out the impact in terms of
> Australian authors,” APA spokesperson Stuart Glover said.
> >>>
> >>> “This is a massive legal and ethical challenge for the publishing industry and
> for authors globally.”
> >>>
> >>> A search tool published on Monday by US media platform The Atlantic and uploaded
> by the US Authors Guild on Wednesday revealed the works of Peter Carey, Helen Garner,
> Kate Grenville, Anna Funder, Christos Tsiolkas and Thomas Keneally, as well as
> Flanagan and dozens of other high-profile Australian authors, were included in the
> pirated dataset containing more than 180,000 titles.
> >>>
> >>> On Thursday, the Australian Society of Authors issued a statement saying it was
> “horrified” to learn that the works of Australian writers were being used to train
> artificial intelligence without permission from the authors.
> >>>
> >>> ASA chief executive, Olivia Lanchester, described the Books3 dataset as piracy on
> an industrial scale.
> >>>
> >>> “Authors appropriately feel outraged,” Lanchester said. “The fact is this
> technology relies upon books, journals, essays written by authors, yet permission was
> not sought nor compensation granted.”
> >>>
> >>> Lanchester said the Australian literary industry, while not objecting per se to
> emerging technologies such as AI, was deeply concerned about the lack of transparency
> evident in the development and monetisation of AI by global tech companies.
> >>>
> >>> “Turning a blind eye to the legitimate rights of copyright owners threatens to
> diminish already precarious creative careers,” she said.
> >>>
> >>> “The enrichment of a few powerful companies is at the cost of thousands of
> individual creators. This is not how a fair market functions.”
> >>>
> >>> Josephine Johnston, chief executive of Australia’s Copyright Agency, described
> the Books3 development as “a free kick to big tech” at the expense of Australia’s
> creative and cultural life.
> >>>
> >>> “We’re going to need greater transparency – how these tools have been developed,
> trained, how they operate – before people can truly understand what their legal rights
> might be,” she said.
> >>>
> >>> “We seem to be in this terrible position now where content owners – remembering
> that the vast majority of them will be individual authors – may actually have to take
> out court cases to enforce their rights.”
> >>>
> >>> Australian copyright law protects creators of original content from data scraping.
> >>>
> >>> Litigation in the US against ChatGPT creator OpenAI over use of allegedly pirated
> book datasets, Books1 and Books2 (which do not appear to be affiliated with Books3)
> has already commenced.
> >>>
> >>> In July, North American horror/fantasy writers Mona Awad (author of Bunny) and
> Paul Tremblay (author of The Cabin at the End of the World) filed a lawsuit in a San
> Francisco federal court, alleging ChatGPT unlawfully digested their books as part of
> its AI training data.
> >>>
> >>> On 28 August, OpenAI filed a motion to dismiss the lawsuit, arguing that the
> authors “misconceive the scope of copyright, failing to take into account the
> limitations and exceptions (including fair use) that properly leave room for
> innovations like the large language models now at the forefront of artificial
> intelligence”.
> >>>
> >>> On 19 September the Writers Guild and 17 of its members, including bestselling
> novelists John Grisham, George RR Martin and Jodi Picoult, filed a complaint in a New
> York district court against OpenAI, seeking redress for “flagrant and harmful
> infringements” of guild members’ registered copyrights.
> >>>
> >>> In a statement on its website, the guild says while it is aware that companies
> such as Meta and Bloomberg have used the Books3 dataset to train their LLMs, it is not
> yet clear whether OpenAI is using Books3 to train its ChatGPT models GPT 3.5 or GPT 4.
> >>>
> >>> Democracies face ‘truth decay’ as AI blurs fact and fiction, warns head of
> Australia’s military
> >>>
> >>> Guardian Australia has sought comment from OpenAI, which has yet to officially
> respond to the guild’s complaint, and Meta.
> >>>
> >>> On 4 September, US technology magazine Wired reported that a Danish anti-piracy
> group called Rights Alliance had been told by Bloomberg that the company did not plan
> to train future versions of its BloombergGPT using Books3.
> >>>
> >>> Bloomberg declined to respond to the Guardian’s queries.
> >>>
> >>> The APA said the global nature of the issue would present significant challenges
> in enforcement and prosecution, and has joined the authors’ society in calling for AI
> technologies to be regulated.
> >>>
> >>> Consultation closed last month for a Department of Industry, Science and
> Resources discussion paper on supporting responsible AI.
> >>>
> >>> A parliamentary inquiry is under way examining the use of generative artificial
> intelligence in the Australian education system.
> >>>
> >>> Flanagan said it was up to the Australian government to act to protect
> Australia’s writers.
> >>>
> >>> “It has power and we do not,” he said.
> >>>
> >>> “If it cares for our culture it must now stand up and fight for it.”
> >>>
> >>> _______________________________________________
> >>> nexa mailing list
> >>> nexa(a)server-nexa.polito.it <mailto:nexa@server-nexa.polito.it>
> >>> https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
> <https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa>
> >>
> >> _______________________________________________
> >> nexa mailing list
> >> nexa(a)server-nexa.polito.it <mailto:nexa@server-nexa.polito.it>
> >> https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
> <https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa>
> > _______________________________________________
> > nexa mailing list
> > nexa(a)server-nexa.polito.it <mailto:nexa@server-nexa.polito.it>
> > https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
> <https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa>
> _______________________________________________
> nexa mailing list
> nexa(a)server-nexa.polito.it <mailto:nexa@server-nexa.polito.it>
> https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
> <https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa>
>
Sept. 29, 2023
Re: [nexa] ‘Biggest act of copyright theft in history’: thousands of Australian books allegedly used to train AI model | Australia news | The Guardian
by Maurizio Borghi
Caro Stefano, il tuo ragionamento è corretto. Aggiungo che la rimozione del
DRM è già perseguibile come violazione del diritto d’autore, non solo nei
paesi UE, ma anche in USA.
Se l’opera non è protetta da DRM il discorso è un po’ più complicato: in
UE, il contratto di licenza non può escludere certi usi consentiti, in
particolare il text and data mining per scopi non commerciali. Può però
escludere lo stesso utilizzo se per scopi commerciali e se l’uso è
espressamente riservato. In USA non ci sono regole precise, ma la libertà
contrattuale tende di solito a prevalere sulla disponibilità di eccezioni
(fair use). Non è un caso che nella class action contro GitHub / Copilot i
claim si basino interamente su violazione dei contratti di licenza (open
source) e sulla rimozione dei DRM, anziché sulla violazione del copyright
nel software utilizzato per addestrare l’algoritmo.
Un caro saluto
Maurizio
On Fri, 29 Sep 2023 at 15:21, Stefano Quintarelli <stefano(a)quintarelli.it>
wrote:
> Ho una domanda per i giuristi (anzi, piu' di una)
>
> per allenare un modello, ho bisogno di un file con la versione digitale di
> un testo.
> (cosnsidero ovviamente testi non PD, CC0, ecc.)
>
> la versione digitale di un testo la posso ottenere da un ebook (gia'
> digitale), togliendo
> il probabile DRM.
> ma un ebook non e' unbene ma e' un servizio soggetto a licenza d'uso,
> quindi se non e'
> prevista nella licenza d'uso la facolta' di estrarre il testo digitale per
> allenarci un
> modello, mi sembra che ci sia gia' una violazione della licenza, per cui,
> credo, non possa
> essere usato come base di un allenamento, tanto piu' se il fine di tale
> allenamento e'
> commerciale (se vendo un servizio basato su quel modello).
>
> se e' cosi', per allenare il mio modello devo allora prednere il testo
> digitale facendo
> scan/ocr di un testo cartaceo.
> ma cio' e' possibile, se non erro, solo per uso personale e non
> commerciale.
>
> se questo e' corretto, non mi pare ci sia un modo per prendere un testo
> digitale senza
> infrangere una licenza d'uso/copyright
>
> dove e' la fallacia del ragionamento ?
>
> grazie, s.
>
> On 29/09/23 15:00, Stefano Borroni Barale wrote:
> > Buongiorno lista,
> >
> >> L'idea che istruire un modello su dei testi coperti da copyright sia
> una violazione del suddetto copyright è altamente opinabile
> >
> > Fin qui, ho l'impressione che tutti i legali in lista concorderanno.
> >
> >> ragionamento è in realtà abbastanza semplice: se istruirsi su un
> >> testo ne violasse il copyright, saremmo tutti dei criminali.
> >
> > Ma siccome noi siamo umani e quello che produciamo non è - salvo i
> discorsi dei politici(*) - ontologicamente identico alla produzione di
> esseri tecnici non viventi, logica vuole che quanto si applica a noi non
> possa applicarsi a un LLM, tanto quanto la legge sul copyright non si
> applica pedissequamente all'utilizzo di testi umani per creare modelli
> linguistici.
> >
> > Questo è il motivo per il quale tutti i tentativi di "proteggere via
> copyright" il prodotto di software generativi sono falliti miseramente, e
> con motivazioni scritte in sentenze; che per il diritto credo abbiano un
> peso assai maggiore del sito di CC.
> >
> > La mia impressione è che la questione terrà impegnati legali,
> informatici, filosofi e società ancora moooooolto a lungo.
> > SBB
> >
> > (*) Come sanno bene i bambini degli anni '80 che hanno giocato con
> questo spassoso giocattolo:
> https://www.enricodalbosco.it/giochi/tubolario/
> >
> >
> > Di quei testi
> >> non c'è fisicamente traccia all'interno dei modelli, non viene copiato
> >> niente. I modelli sono un'opera trasformativa di quei testi, non
> >> derivativa.
> >>
> >> Lo argomenta molto bene Creative Commons:
> >> https://creativecommons.org/2023/02/17/fair-use-training-generative-ai/
> >>
> >> Detto questo, cito le parole di un altro autore, Jeff Jarvis:
> >>
> https://www.facebook.com/jeff.jarvis/posts/pfbid0LMFeqdTYoxnGHQAZwp5HMmeeVq…
> >>
> >> «I, for one, am not complaining about my books being in in large
> >> language model training sets. I write to enter ideas into public
> >> discourse. I prefer informed over ignorant AI. I believe it is fair
> >> use for anyone to read & use books for transformative work. In fact,
> >> I'd probably feel snubbed if my books were not there. I'm happy when
> >> they are in libraries. I'm fine that they're here.»
> >>
> >> Fabio
> >>
> >> Il giorno ven 29 set 2023 alle ore 07:52 Alberto Cammozzo via nexa
> >> nexa(a)server-nexa.polito.it ha scritto:
> >>
> >>>
> https://www.theguardian.com/australia-news/2023/sep/28/australian-books-tra…
> >>>
> >>> Thousands of books from some of Australia’s most celebrated authors
> have potentially been caught up in what Booker prize-winning novelist
> Richard Flanagan has called “the biggest act of copyright theft in history”.
> >>>
> >>> The works have allegedly been pirated by the US-based Books3 dataset
> and used to train generative AI for corporations such as Meta and Bloomberg.
> >>>
> >>> Flanagan, who found 10 of his works, including the multi-international
> award-winning 2013 novel The Narrow Road to the Deep North, on the Books3
> dataset, told Guardian Australia he was deeply shocked by the discovery
> made several days ago.
> >>>
> >>> “I felt as if my soul had been strip mined and I was powerless to stop
> it,” he said in a statement.
> >>>
> >>> “This is the biggest act of copyright theft in history.”
> >>>
> >>> AI could ‘turbo-charge fraud’ and be monopolised by tech companies,
> Andrew Leigh warns
> >>>
> >>> The Australian Publishers Association confirmed to Guardian Australia
> on Wednesday that as many as 18,000 fiction and nonfiction titles with
> Australian ISBNs (unique international standard book numbers) appeared to
> be affected by the copyright infringement, although it is not yet clear
> what proportion of these are Australian editions of internationally
> authored books.
> >>>
> >>> “We’re still working through [the data] to work out the impact in
> terms of Australian authors,” APA spokesperson Stuart Glover said.
> >>>
> >>> “This is a massive legal and ethical challenge for the publishing
> industry and for authors globally.”
> >>>
> >>> A search tool published on Monday by US media platform The Atlantic
> and uploaded by the US Authors Guild on Wednesday revealed the works of
> Peter Carey, Helen Garner, Kate Grenville, Anna Funder, Christos Tsiolkas
> and Thomas Keneally, as well as Flanagan and dozens of other high-profile
> Australian authors, were included in the pirated dataset containing more
> than 180,000 titles.
> >>>
> >>> On Thursday, the Australian Society of Authors issued a statement
> saying it was “horrified” to learn that the works of Australian writers
> were being used to train artificial intelligence without permission from
> the authors.
> >>>
> >>> ASA chief executive, Olivia Lanchester, described the Books3 dataset
> as piracy on an industrial scale.
> >>>
> >>> “Authors appropriately feel outraged,” Lanchester said. “The fact is
> this technology relies upon books, journals, essays written by authors, yet
> permission was not sought nor compensation granted.”
> >>>
> >>> Lanchester said the Australian literary industry, while not objecting
> per se to emerging technologies such as AI, was deeply concerned about the
> lack of transparency evident in the development and monetisation of AI by
> global tech companies.
> >>>
> >>> “Turning a blind eye to the legitimate rights of copyright owners
> threatens to diminish already precarious creative careers,” she said.
> >>>
> >>> “The enrichment of a few powerful companies is at the cost of
> thousands of individual creators. This is not how a fair market functions.”
> >>>
> >>> Josephine Johnston, chief executive of Australia’s Copyright Agency,
> described the Books3 development as “a free kick to big tech” at the
> expense of Australia’s creative and cultural life.
> >>>
> >>> “We’re going to need greater transparency – how these tools have been
> developed, trained, how they operate – before people can truly understand
> what their legal rights might be,” she said.
> >>>
> >>> “We seem to be in this terrible position now where content owners –
> remembering that the vast majority of them will be individual authors – may
> actually have to take out court cases to enforce their rights.”
> >>>
> >>> Australian copyright law protects creators of original content from
> data scraping.
> >>>
> >>> Litigation in the US against ChatGPT creator OpenAI over use of
> allegedly pirated book datasets, Books1 and Books2 (which do not appear to
> be affiliated with Books3) has already commenced.
> >>>
> >>> In July, North American horror/fantasy writers Mona Awad (author of
> Bunny) and Paul Tremblay (author of The Cabin at the End of the World)
> filed a lawsuit in a San Francisco federal court, alleging ChatGPT
> unlawfully digested their books as part of its AI training data.
> >>>
> >>> On 28 August, OpenAI filed a motion to dismiss the lawsuit, arguing
> that the authors “misconceive the scope of copyright, failing to take into
> account the limitations and exceptions (including fair use) that properly
> leave room for innovations like the large language models now at the
> forefront of artificial intelligence”.
> >>>
> >>> On 19 September the Writers Guild and 17 of its members, including
> bestselling novelists John Grisham, George RR Martin and Jodi Picoult,
> filed a complaint in a New York district court against OpenAI, seeking
> redress for “flagrant and harmful infringements” of guild members’
> registered copyrights.
> >>>
> >>> In a statement on its website, the guild says while it is aware that
> companies such as Meta and Bloomberg have used the Books3 dataset to train
> their LLMs, it is not yet clear whether OpenAI is using Books3 to train its
> ChatGPT models GPT 3.5 or GPT 4.
> >>>
> >>> Democracies face ‘truth decay’ as AI blurs fact and fiction, warns
> head of Australia’s military
> >>>
> >>> Guardian Australia has sought comment from OpenAI, which has yet to
> officially respond to the guild’s complaint, and Meta.
> >>>
> >>> On 4 September, US technology magazine Wired reported that a Danish
> anti-piracy group called Rights Alliance had been told by Bloomberg that
> the company did not plan to train future versions of its BloombergGPT using
> Books3.
> >>>
> >>> Bloomberg declined to respond to the Guardian’s queries.
> >>>
> >>> The APA said the global nature of the issue would present significant
> challenges in enforcement and prosecution, and has joined the authors’
> society in calling for AI technologies to be regulated.
> >>>
> >>> Consultation closed last month for a Department of Industry, Science
> and Resources discussion paper on supporting responsible AI.
> >>>
> >>> A parliamentary inquiry is under way examining the use of generative
> artificial intelligence in the Australian education system.
> >>>
> >>> Flanagan said it was up to the Australian government to act to protect
> Australia’s writers.
> >>>
> >>> “It has power and we do not,” he said.
> >>>
> >>> “If it cares for our culture it must now stand up and fight for it.”
> >>>
> >>> _______________________________________________
> >>> nexa mailing list
> >>> nexa(a)server-nexa.polito.it
> >>> https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
> >>
> >> _______________________________________________
> >> nexa mailing list
> >> nexa(a)server-nexa.polito.it
> >> https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
> > _______________________________________________
> > nexa mailing list
> > nexa(a)server-nexa.polito.it
> > https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
> _______________________________________________
> nexa mailing list
> nexa(a)server-nexa.polito.it
> https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
>
Sept. 29, 2023
Re: [nexa] R: R: R: ‘Biggest act of copyright theft in history’: thousands of Australian books allegedly used to train AI model | Australia news | The Guardian
by Stefano Zacchiroli
On Fri, Sep 29, 2023 at 05:33:04PM +0200, Lorenzo Albertini wrote:
> Non so in usa (fair use, che ha ambito applicativo alquanto vasto?)
La posizione (per ora non verificata in tribunale) di Microsoft/GitHub
per Copilot è esattamente che il training di modelli ML secondo il
copyright americano costituisca fair use. Nel caso dei libri questo non
evacua però (credo, IANAL, etc.) la domanda di Stefano Quintarelli su se
un contratto d'uso di un ebook impedisca comunque il training. (Nel caso
di Microsoft/GitHub la domanda invece non si poneva, perché avevano già
tutto il codice che "gentilmente" milioni di sviluppatori gli hanno
chiesto di ospitare...)
--
Stefano Zacchiroli . zack(a)upsilon.cc . https://upsilon.cc/zack _. ^ ._
Full professor of Computer Science o o o \/|V|\/
Télécom Paris, Polytechnic Institute of Paris o o o </> <\>
Co-founder & CTO Software Heritage o o o o /\|^|/\
https://twitter.com/zacchiro . https://mastodon.xyz/@zacchiro '" V "'
Sept. 29, 2023
R: R: R: ‘Biggest act of copyright theft in history’: thousands of Australian books allegedly used to train AI model | Australia news | The Guardian
by Lorenzo Albertini
Sul presupposto che cmq ricorra <riproduzione>, sarebbe difficile individuare (in UE) un'eccezione.
Non invocabile direi l'art. 4 dir. Copyright 790 :
≪Articolo 4 Eccezioni o limitazioni ai fini dell'estrazione di testo e di dati
1. Gli Stati membri dispongono un'eccezione o una limitazione ai diritti di cui all'articolo 5, lettera a), e all'articolo 7, paragrafo 1, della direttiva 96/9/CE, all'articolo 2 della direttiva 2001/29/CE, all'articolo 4, paragrafo 1, lettere a) e b), della direttiva 2009/24/CE e all'articolo 15, paragrafo 1, della presente direttiva per le riproduzioni e le estrazioni effettuate da opere o altri materiali cui si abbia legalmente accesso ai fini dell'estrazione di testo e di dati.
2. Le riproduzioni e le estrazioni effettuate a norma del paragrafo 1 possono essere conservate per il tempo necessario ai fini dell'estrazione di testo e di dati.
3. L'eccezione o la limitazione di cui al paragrafo 1 si applica a condizione che l'utilizzo delle opere e di altri materiali di cui a tale paragrafo non sia stato espressamente riservato dai titolari dei diritti in modo appropriato, ad esempio attraverso strumenti che consentano lettura automatizzata in caso di contenuti resi pubblicamente disponibili online.
4. Il presente articolo non pregiudica l'applicazione dell'articolo 3 della presente direttiva≫.
Non so in usa (fair use, che ha ambito applicativo alquanto vasto?)
>>-----Messaggio originale-----
>>Da: Stefano Quintarelli <stefano(a)quintarelli.it>
>>Inviato: venerdì 29 settembre 2023 17:08
>>A: Lorenzo Albertini <lorenzoalbertini.vr(a)gmail.com>; 'Nexa' <nexa@server-
>>nexa.polito.it>
>>Oggetto: Re: [nexa] R: R: ‘Biggest act of copyright theft in history’: thousands
>>of Australian books allegedly used to train AI model | Australia news | The
>>Guardian
>>
>>grazie
>>
>>ma il punto focale del mio quesito non e' il training ma, prima del training, la
>>genesi dei testi usati per il training
>>
>>ciao, s.
>>
>>On 29/09/23 16:36, Lorenzo Albertini wrote:
>>> §§ 54-64della citazione in giudizio (facilmente reperibile ,_ad es.
>>>
>>qui_<https://www.google.com/url?sa=t&rct=j&q=&esrc=s&source=web&cd= <https://www.google.com/url?sa=t&rct=j&q=&esrc=s&source=web&cd=&ved=2ahUKEwi…>
>>&ved=2ahUKEwiDh_vygtCBAxX4SPEDHZlhAxMQFnoECBYQAQ&url=https%3A
>>%2F%2Fwww.classaction.org%2Fmedia%2Fauthors-guild-et-al-v-openai-inc-
>>et-al.pdf&usg=AOvVaw1tUMb6Gk10kZCsvoAo0PH6&opi=89978449>):
>>>
>>> <<54. Recent generative AI systems designed to recognize input text
>>> and generate
>>>
>>> output text are built on “large language models” or “LLMs.”
>>>
>>> 55. LLMs use predictive algorithms that are designed to detect
>>> statistical patterns in
>>>
>>> the text datasets on which they are “trained” and, on the basis of
>>> these patterns, generate
>>>
>>> responses to user prompts. “Training” an LLM refers to the process by
>>> which the parameters that
>>>
>>> define an LLM’s behavior are adjusted through the LLM’s ingestion and
>>> analysis of large
>>>
>>> “training” datasets.
>>>
>>> 56. Once “trained,” the LLM analyzes the relationships among words in
>>> an input
>>>
>>> prompt and generates a response that is an approximation of similar
>>> relationships among words
>>>
>>> in the LLM’s “training” data. In this way, LLMs can be capable of
>>> generating sentences,
>>>
>>> paragraphs, and even complete texts, from cover letters to novels.
>>>
>>> 57. “Training” an LLM requires supplying the LLM with large amounts of
>>> text for
>>>
>>> the LLM to ingest—the more text, the better. That is, in part, the
>>> large in large language model.
>>>
>>> 58. As the U.S. Patent and Trademark Office has observed, LLM
>>> “training” “almost
>>>
>>> by definition involve[s] the reproduction of entire works or
>>> substantial portions thereof.”4
>>>
>>> 59. “Training” in this context is therefore a technical-sounding
>>> euphemism for
>>>
>>> “copying and ingesting.”
>>>
>>> 60. The quality of the LLM (that is, its capacity to generate
>>> human-seeming responses
>>>
>>> to prompts) is dependent on the quality of the datasets used to “train” the
>>LLM.
>>>
>>> 61. Professionally authored, edited, and published books—such as those
>>> authored by
>>>
>>> Plaintiffs here—are an especially important source of LLM “training” data.
>>>
>>> 62. As one group of AI researchers (not affiliated with Defendants)
>>> has observed,
>>>
>>> “[b]ooks are a rich source of both fine-grained information, how a
>>> character, an object or a scene
>>>
>>> looks like, as well as high-level semantics, what someone is thinking,
>>> feeling and how these
>>>
>>> states evolve through a story.”5
>>>
>>> 63. In other words, books are the high-quality materials Defendants
>>> want, need, and
>>>
>>> have therefore outright pilfered to develop generative AI products
>>> that produce high-quality
>>>
>>> results: text that appears to have been written by a human writer.
>>>
>>> 64. This use is highly commercial.>>.
>>>
>>>
>>> _______________
>>>
>>> Le informazioni contenute nella presente comunicazione e nei documenti
>>> ad essa allegati potrebbero essere tutelate dal segreto professionale
>>> e sono comunque confidenziali e ad uso esclusivo del destinatario
>>> sopra indicato. Qualora la presente comunicazione non fosse destinata
>>> a Voi, Vi preghiamo di tener presente che la divulgazione,
>>> distribuzione o riproduzione di qualunque informazione contenuta nella
>>> presente comunicazione o nei documenti ad essa allegati sono vietate.
>>> Se avete ricevuto la presente comunicazione per errore, Vi preghiamo di
>>volerci avvertire immediatamente e di distruggere quanto ricevuto senza
>>leggerlo. Grazie per la collaborazione.
>>>
>>>
>>> The information contained in this email and any documents attached to
>>> it may be legally privileged and confidential. The information is
>>> intended only for the use of the individual or entity named above. If
>>> you are not the intended recipient, you are hereby notified that any
>>> use, dissemination, distribution or reproduction of any information
>>> contained in or attached to this email is prohibited. If you have
>>> received this email in error, please immediately notify us by reply email or by
>>telephone, and destroy the original transmission and its attachments without
>>reading them. Thank you.
>>>
>>>>>-----Messaggio originale-----
>>>
>>>>>Da: nexa <nexa-bounces(a)server-nexa.polito.it <mailto:nexa-bounces@server-nexa.polito.it> > Per conto di Rossana
>>>
>>>>>Morriello
>>>
>>>>>Inviato: venerdì 29 settembre 2023 16:08
>>>
>>>>>A: Nexa <nexa(a)server-nexa.polito.it <mailto:nexa@server-nexa.polito.it> >
>>>
>>>>>Oggetto: [nexa] R: ‘Biggest act of copyright theft in history’:
>>>>>thousands of
>>>
>>>>>Australian books allegedly used to train AI model | Australia news |
>>>>>The
>>>
>>>>>Guardian
>>>
>>>>>
>>>
>>>>>Non sono una giurista ma credo che questa rassegna possa essere utile
>>>>>alla
>>>
>>>>>discussione
>>>
>>>>>
>>>
>>>>>https://www.thefashionlaw.com/from-chatgpt-to-deepfake-creating- <https://www.thefashionlaw.com/from-chatgpt-to-deepfake-creating-apps->
>>apps-
>>>>>a-<https://www.thefashionlaw.com/from-chatgpt-to-deepfake-creating-
>>ap
>>>>>ps-a-running-list-of-key-ai-lawsuits/>
>>>
>>>>>running-list-of-key-ai-lawsuits/
>>>
>>>>>
>>>
>>>>>
>>>
>>>>>Saluti
>>>
>>>>>Rossana Morriello
>>>
>>>>>
>>>
>>>>>
>>>
>>>>>
>>>
>>>>>
>>>
>>>>>-----Messaggio originale-----
>>>
>>>>>Da: nexa
>>>>><nexa-bounces(a)server-nexa.polito.it<mailto:nexa-bounces@server-
>>nexa.p
>>>>>olito.it>>
>>> Per conto di Stefano
>>>
>>>>>Quintarelli
>>>
>>>>>Inviato: venerdì 29 settembre 2023 15:21
>>>
>>>>>Cc: Nexa
>>>>><nexa(a)server-nexa.polito.it<mailto:nexa@server-nexa.polito.it <mailto:nexa@server-nexa.polito.it<mailto:nexa@server-nexa.polito.it> >>
>>>
>>>>>Oggetto: Re: [nexa] ‘Biggest act of copyright theft in history’:
>>>>>thousands of
>>>
>>>>>Australian books allegedly used to train AI model | Australia news |
>>>>>The
>>>
>>>>>Guardian
>>>
>>>>>
>>>
>>>>>Ho una domanda per i giuristi (anzi, piu' di una)
>>>
>>>>>
>>>
>>>>>per allenare un modello, ho bisogno di un file con la versione
>>>>>digitale di un
>>>
>>>>>testo.
>>>
>>>>>(cosnsidero ovviamente testi non PD, CC0, ecc.)
>>>
>>>>>
>>>
>>>>>la versione digitale di un testo la posso ottenere da un ebook (gia'
>>>>>digitale),
>>>
>>>>>togliendo il probabile DRM.
>>>
>>>>>ma un ebook non e' unbene ma e' un servizio soggetto a licenza d'uso,
>>>>>quindi
>>>
>>>>>se non e'
>>>
>>>>>prevista nella licenza d'uso la facolta' di estrarre il testo
>>>>>digitale per allenarci un
>>>
>>>>>modello, mi sembra che ci sia gia' una violazione della licenza, per
>>>>>cui, credo,
>>>
>>>>>non possa essere usato come base di un allenamento, tanto piu' se il
>>>>>fine di
>>>
>>>>>tale allenamento e'
>>>
>>>>>commerciale (se vendo un servizio basato su quel modello).
>>>
>>>>>
>>>
>>>>>se e' cosi', per allenare il mio modello devo allora prednere il
>>>>>testo digitale
>>>
>>>>>facendo scan/ocr di un testo cartaceo.
>>>
>>>>>ma cio' e' possibile, se non erro, solo per uso personale e non commerciale.
>>>
>>>>>
>>>
>>>>>se questo e' corretto, non mi pare ci sia un modo per prendere un
>>>>>testo digitale
>>>
>>>>>senza infrangere una licenza d'uso/copyright
>>>
>>>>>
>>>
>>>>>dove e' la fallacia del ragionamento ?
>>>
>>>>>
>>>
>>>>>grazie, s.
>>>
>>>>>
>>>
>>>>>On 29/09/23 15:00, Stefano Borroni Barale wrote:
>>>
>>>>> > Buongiorno lista,
>>>
>>>>> >
>>>
>>>>> >> L'idea che istruire un modello su dei testi coperti da copyright
>>>>> >> sia
>>>
>>>>> >> una violazione del suddetto copyright è altamente opinabile
>>>
>>>>> >
>>>
>>>>> > Fin qui, ho l'impressione che tutti i legali in lista concorderanno.
>>>
>>>>> >
>>>
>>>>> >> ragionamento è in realtà abbastanza semplice: se istruirsi su un
>>>
>>>>> >> testo ne violasse il copyright, saremmo tutti dei criminali.
>>>
>>>>> >
>>>
>>>>> > Ma siccome noi siamo umani e quello che produciamo non è - salvo i
>>>>> > discorsi
>>>
>>>>>dei politici(*) - ontologicamente identico alla produzione di esseri
>>>>>tecnici non
>>>
>>>>>viventi, logica vuole che quanto si applica a noi non possa
>>>>>applicarsi a un LLM,
>>>
>>>>>tanto quanto la legge sul copyright non si applica pedissequamente
>>>>>all'utilizzo
>>>
>>>>>di testi umani per creare modelli linguistici.
>>>
>>>>> >
>>>
>>>>> > Questo è il motivo per il quale tutti i tentativi di "proteggere
>>>>> > via copyright" il
>>>
>>>>>prodotto di software generativi sono falliti miseramente, e con
>>>>>motivazioni
>>>
>>>>>scritte in sentenze; che per il diritto credo abbiano un peso assai
>>>>>maggiore del
>>>
>>>>>sito di CC.
>>>
>>>>> >
>>>
>>>>> > La mia impressione è che la questione terrà impegnati legali,
>>>>> > informatici,
>>>
>>>>>filosofi e società ancora moooooolto a lungo.
>>>
>>>>> > SBB
>>>
>>>>> >
>>>
>>>>> > (*) Come sanno bene i bambini degli anni '80 che hanno giocato con
>>>
>>>>> > questo spassoso giocattolo:
>>>
>>>>> >https://www.enricodalbosco.it/giochi/tubolario/<https://www.enricod
>>>>> >albosco.it/giochi/tubolario/>
>>>
>>>>> >
>>>
>>>>> >
>>>
>>>>> > Di quei testi
>>>
>>>>> >> non c'è fisicamente traccia all'interno dei modelli, non viene
>>>
>>>>> >> copiato niente. I modelli sono un'opera trasformativa di quei
>>>>> >> testi,
>>>
>>>>> >> non derivativa.
>>>
>>>>> >>
>>>
>>>>> >> Lo argomenta molto bene Creative Commons:
>>>
>>>>> >>https://creativecommons.org/2023/02/17/fair-use-training-generativ
>>>>> >>e-a<https://creativecommons.org/2023/02/17/fair-use-training-gener
>>>>> >>ative-a>
>>>
>>>>> >> i/
>>>
>>>>> >>
>>>
>>>>> >> Detto questo, cito le parole di un altro autore, Jeff Jarvis:
>>>
>>>>> >>
>>>
>>>>>https://www.facebook.com/jeff.jarvis/posts/pfbid0LMFeqdTYoxnGHQAZ <https://www.facebook.com/jeff.jarvis/posts/pfbid0LMFeqdTYoxnGHQAZwp5<>
>>wp5<
>>>>>https://www.facebook.com/jeff.jarvis/posts/pfbid0LMFeqdTYoxnGHQAZ <https://www.facebook.com/jeff.jarvis/posts/pfbid0LMFeqdTYoxnGHQAZwp5H>
>>wp5H
>>>>>>
>>>
>>>>>H
>>>
>>>>> >> MmeeVqgMSjL2dkcwMcBojkb2cinBpgYTHyc7Fhq1B9NPl
>>>
>>>>> >>
>>>
>>>>> >> «I, for one, am not complaining about my books being in in large
>>>
>>>>> >> language model training sets. I write to enter ideas into public
>>>
>>>>> >> discourse. I prefer informed over ignorant AI. I believe it is
>>>>> >> fair
>>>
>>>>> >> use for anyone to read & use books for transformative work. In
>>>>> >> fact,
>>>
>>>>> >> I'd probably feel snubbed if my books were not there. I'm happy
>>>>> >> when
>>>
>>>>> >> they are in libraries. I'm fine that they're here.»
>>>
>>>>> >>
>>>
>>>>> >> Fabio
>>>
>>>>> >>
>>>
>>>>> >> Il giorno ven 29 set 2023 alle ore 07:52 Alberto Cammozzo via
>>>>> >> nexa
>>>
>>>>> >>nexa(a)server-nexa.polito.it<mailto:nexa@server-nexa.polito.it>ha <mailto:nexa@server-nexa.polito.it<mailto:nexa@server-nexa.polito.it>ha>
>>scritto:
>>>
>>>>> >>
>>>
>>>>> >>>https://www.theguardian.com/australia- <https://www.theguardian.com/australia-news/2023/sep/28/australian>
>>news/2023/sep/28/australian
>>>>> >>>-<https://www.theguardian.com/australia-
>>news/2023/sep/28/australi
>>>>> >>>an-bo>
>>>
>>>>>bo
>>>
>>>>> >>> oks-training-ai-books3-stolen-pirated
>>>
>>>>> >>>
>>>
>>>>> >>> Thousands of books from some of Australia’s most celebrated
>>>>> >>> authors
>>>
>>>>>have potentially been caught up in what Booker prize-winning novelist
>>>>>Richard
>>>
>>>>>Flanagan has called “the biggest act of copyright theft in history”.
>>>
>>>>> >>>
>>>
>>>>> >>> The works have allegedly been pirated by the US-based Books3
>>>>> >>> dataset
>>>
>>>>>and used to train generative AI for corporations such as Meta and
>>Bloomberg.
>>>
>>>>> >>>
>>>
>>>>> >>> Flanagan, who found 10 of his works, including the
>>>>> >>> multi-international
>>>
>>>>>award-winning 2013 novel The Narrow Road to the Deep North, on the
>>>
>>>>>Books3 dataset, told Guardian Australia he was deeply shocked by the
>>>
>>>>>discovery made several days ago.
>>>
>>>>> >>>
>>>
>>>>> >>> “I felt as if my soul had been strip mined and I was powerless to stop
>>it,”
>>>
>>>>>he said in a statement.
>>>
>>>>> >>>
>>>
>>>>> >>> “This is the biggest act of copyright theft in history.”
>>>
>>>>> >>>
>>>
>>>>> >>> AI could ‘turbo-charge fraud’ and be monopolised by tech
>>>>> >>> companies,
>>>
>>>>> >>> Andrew Leigh warns
>>>
>>>>> >>>
>>>
>>>>> >>> The Australian Publishers Association confirmed to Guardian
>>>>> >>> Australia on
>>>
>>>>>Wednesday that as many as 18,000 fiction and nonfiction titles with
>>>
>>>>>Australian ISBNs (unique international standard book numbers)
>>>>>appeared to
>>>
>>>>>be affected by the copyright infringement, although it is not yet
>>>>>clear what
>>>
>>>>>proportion of these are Australian editions of internationally authored
>>books.
>>>
>>>>> >>>
>>>
>>>>> >>> “We’re still working through [the data] to work out the impact
>>>>> >>> in terms of
>>>
>>>>>Australian authors,” APA spokesperson Stuart Glover said.
>>>
>>>>> >>>
>>>
>>>>> >>> “This is a massive legal and ethical challenge for the
>>>>> >>> publishing industry
>>>
>>>>>and for authors globally.”
>>>
>>>>> >>>
>>>
>>>>> >>> A search tool published on Monday by US media platform The
>>>>> >>> Atlantic and
>>>
>>>>>uploaded by the US Authors Guild on Wednesday revealed the works of
>>>>>Peter
>>>
>>>>>Carey, Helen Garner, Kate Grenville, Anna Funder, Christos Tsiolkas
>>>>>and
>>>
>>>>>Thomas Keneally, as well as Flanagan and dozens of other high-profile
>>>
>>>>>Australian authors, were included in the pirated dataset containing
>>>>>more than
>>>
>>>>>180,000 titles.
>>>
>>>>> >>>
>>>
>>>>> >>> On Thursday, the Australian Society of Authors issued a
>>>>> >>> statement saying
>>>
>>>>>it was “horrified” to learn that the works of Australian writers were
>>>>>being used
>>>
>>>>>to train artificial intelligence without permission from the authors.
>>>
>>>>> >>>
>>>
>>>>> >>> ASA chief executive, Olivia Lanchester, described the Books3
>>>>> >>> dataset as
>>>
>>>>>piracy on an industrial scale.
>>>
>>>>> >>>
>>>
>>>>> >>> “Authors appropriately feel outraged,” Lanchester said. “The
>>>>> >>> fact is this
>>>
>>>>>technology relies upon books, journals, essays written by authors,
>>>>>yet
>>>
>>>>>permission was not sought nor compensation granted.”
>>>
>>>>> >>>
>>>
>>>>> >>> Lanchester said the Australian literary industry, while not
>>>>> >>> objecting per se
>>>
>>>>>to emerging technologies such as AI, was deeply concerned about the
>>>>>lack of
>>>
>>>>>transparency evident in the development and monetisation of AI by
>>>>>global
>>>
>>>>>tech companies.
>>>
>>>>> >>>
>>>
>>>>> >>> “Turning a blind eye to the legitimate rights of copyright
>>>>> >>> owners threatens
>>>
>>>>>to diminish already precarious creative careers,” she said.
>>>
>>>>> >>>
>>>
>>>>> >>> “The enrichment of a few powerful companies is at the cost of
>>>>> >>> thousands
>>>
>>>>>of individual creators. This is not how a fair market functions.”
>>>
>>>>> >>>
>>>
>>>>> >>> Josephine Johnston, chief executive of Australia’s Copyright
>>>>> >>> Agency,
>>>
>>>>>described the Books3 development as “a free kick to big tech” at the
>>>>>expense
>>>
>>>>>of Australia’s creative and cultural life.
>>>
>>>>> >>>
>>>
>>>>> >>> “We’re going to need greater transparency – how these tools have
>>>>> >>> been
>>>
>>>>>developed, trained, how they operate – before people can truly
>>>>>understand
>>>
>>>>>what their legal rights might be,” she said.
>>>
>>>>> >>>
>>>
>>>>> >>> “We seem to be in this terrible position now where content
>>>>> >>> owners –
>>>
>>>>>remembering that the vast majority of them will be individual authors
>>>>>– may
>>>
>>>>>actually have to take out court cases to enforce their rights.”
>>>
>>>>> >>>
>>>
>>>>> >>> Australian copyright law protects creators of original content
>>>>> >>> from data
>>>
>>>>>scraping.
>>>
>>>>> >>>
>>>
>>>>> >>> Litigation in the US against ChatGPT creator OpenAI over use of
>>>>> >>> allegedly
>>>
>>>>>pirated book datasets, Books1 and Books2 (which do not appear to be
>>>
>>>>>affiliated with Books3) has already commenced.
>>>
>>>>> >>>
>>>
>>>>> >>> In July, North American horror/fantasy writers Mona Awad (author
>>>>> >>> of
>>>
>>>>>Bunny) and Paul Tremblay (author of The Cabin at the End of the
>>>>>World) filed a
>>>
>>>>>lawsuit in a San Francisco federal court, alleging ChatGPT unlawfully
>>>>>digested
>>>
>>>>>their books as part of its AI training data.
>>>
>>>>> >>>
>>>
>>>>> >>> On 28 August, OpenAI filed a motion to dismiss the lawsuit,
>>>>> >>> arguing that
>>>
>>>>>the authors “misconceive the scope of copyright, failing to take into
>>>>>account
>>>
>>>>>the limitations and exceptions (including fair use) that properly
>>>>>leave room for
>>>
>>>>>innovations like the large language models now at the forefront of
>>>>>artificial
>>>
>>>>>intelligence”.
>>>
>>>>> >>>
>>>
>>>>> >>> On 19 September the Writers Guild and 17 of its members,
>>>>> >>> including
>>>
>>>>>bestselling novelists John Grisham, George RR Martin and Jodi
>>>>>Picoult, filed a
>>>
>>>>>complaint in a New York district court against OpenAI, seeking
>>>>>redress for
>>>
>>>>>“flagrant and harmful infringements” of guild members’ registered
>>copyrights.
>>>
>>>>> >>>
>>>
>>>>> >>> In a statement on its website, the guild says while it is aware
>>>>> >>> that
>>>
>>>>>companies such as Meta and Bloomberg have used the Books3 dataset to
>>>
>>>>>train their LLMs, it is not yet clear whether OpenAI is using Books3
>>>>>to train its
>>>
>>>>>ChatGPT models GPT 3.5 or GPT 4.
>>>
>>>>> >>>
>>>
>>>>> >>> Democracies face ‘truth decay’ as AI blurs fact and fiction,
>>>>> >>> warns
>>>
>>>>> >>> head of Australia’s military
>>>
>>>>> >>>
>>>
>>>>> >>> Guardian Australia has sought comment from OpenAI, which has yet
>>>>> >>> to
>>>
>>>>>officially respond to the guild’s complaint, and Meta.
>>>
>>>>> >>>
>>>
>>>>> >>> On 4 September, US technology magazine Wired reported that a
>>>>> >>> Danish
>>>
>>>>>anti-piracy group called Rights Alliance had been told by Bloomberg
>>>>>that the
>>>
>>>>>company did not plan to train future versions of its BloombergGPT
>>>>>using
>>>
>>>>>Books3.
>>>
>>>>> >>>
>>>
>>>>> >>> Bloomberg declined to respond to the Guardian’s queries.
>>>
>>>>> >>>
>>>
>>>>> >>> The APA said the global nature of the issue would present
>>>>> >>> significant
>>>
>>>>>challenges in enforcement and prosecution, and has joined the authors’
>>>
>>>>>society in calling for AI technologies to be regulated.
>>>
>>>>> >>>
>>>
>>>>> >>> Consultation closed last month for a Department of Industry,
>>>>> >>> Science and
>>>
>>>>>Resources discussion paper on supporting responsible AI.
>>>
>>>>> >>>
>>>
>>>>> >>> A parliamentary inquiry is under way examining the use of
>>>>> >>> generative
>>>
>>>>>artificial intelligence in the Australian education system.
>>>
>>>>> >>>
>>>
>>>>> >>> Flanagan said it was up to the Australian government to act to
>>>>> >>> protect
>>>
>>>>>Australia’s writers.
>>>
>>>>> >>>
>>>
>>>>> >>> “It has power and we do not,” he said.
>>>
>>>>> >>>
>>>
>>>>> >>> “If it cares for our culture it must now stand up and fight for it.”
>>>
>>>>> >>>
>>>
>>>>> >>> _______________________________________________
>>>
>>>>> >>> nexa mailing list
>>>
>>>>> >>>nexa(a)server-nexa.polito.it<mailto:nexa@server-nexa.polito.it <mailto:nexa@server-nexa.polito.it<mailto:nexa@server-nexa.polito.it> >
>>>
>>>>> >>>https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa<https
>>>>> >>>://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa>
>>>
>>>>> >>
>>>
>>>>> >> _______________________________________________
>>>
>>>>> >> nexa mailing list
>>>
>>>>> >>nexa(a)server-nexa.polito.it<mailto:nexa@server-nexa.polito.it <mailto:nexa@server-nexa.polito.it<mailto:nexa@server-nexa.polito.it> >
>>>
>>>>> >>https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa<https:
>>>>> >>//server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa>
>>>
>>>>> > _______________________________________________
>>>
>>>>> > nexa mailing list
>>>
>>>>> >nexa(a)server-nexa.polito.it<mailto:nexa@server-nexa.polito.it <mailto:nexa@server-nexa.polito.it<mailto:nexa@server-nexa.polito.it> >
>>>
>>>>> >https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa<https:/
>>>>> >/server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa>
>>>
>>>>>_______________________________________________
>>>
>>>>>nexa mailing list
>>>
>>>>>nexa(a)server-nexa.polito.it<mailto:nexa@server-nexa.polito.it <mailto:nexa@server-nexa.polito.it<mailto:nexa@server-nexa.polito.it> >
>>>
>>>>>https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa<https://s
>>>>>erver-nexa.polito.it/cgi-bin/mailman/listinfo/nexa>
>>>
>>>>>_______________________________________________
>>>
>>>>>nexa mailing list
>>>
>>>>>nexa(a)server-nexa.polito.it<mailto:nexa@server-nexa.polito.it <mailto:nexa@server-nexa.polito.it<mailto:nexa@server-nexa.polito.it> >
>>>
>>>>>https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa<https://s
>>>>>erver-nexa.polito.it/cgi-bin/mailman/listinfo/nexa>
>>>
>>>
>>> _______________________________________________
>>> nexa mailing list
>>> nexa(a)server-nexa.polito.it <mailto:nexa@server-nexa.polito.it>
>>> https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
Sept. 29, 2023
Re: [nexa] R: R: ‘Biggest act of copyright theft in history’: thousands of Australian books allegedly used to train AI model | Australia news | The Guardian
by Stefano Quintarelli
grazie
ma il punto focale del mio quesito non e' il training ma, prima del training, la genesi
dei testi usati per il training
ciao, s.
On 29/09/23 16:36, Lorenzo Albertini wrote:
> §§ 54-64della citazione in giudizio (facilmente reperibile ,_ad es.
> qui_<https://www.google.com/url?sa=t&rct=j&q=&esrc=s&source=web&cd=&ved=2ahUKEwi…>):
>
> <<54. Recent generative AI systems designed to recognize input text and generate
>
> output text are built on “large language models” or “LLMs.”
>
> 55. LLMs use predictive algorithms that are designed to detect statistical patterns in
>
> the text datasets on which they are “trained” and, on the basis of these patterns, generate
>
> responses to user prompts. “Training” an LLM refers to the process by which the parameters
> that
>
> define an LLM’s behavior are adjusted through the LLM’s ingestion and analysis of large
>
> “training” datasets.
>
> 56. Once “trained,” the LLM analyzes the relationships among words in an input
>
> prompt and generates a response that is an approximation of similar relationships among words
>
> in the LLM’s “training” data. In this way, LLMs can be capable of generating sentences,
>
> paragraphs, and even complete texts, from cover letters to novels.
>
> 57. “Training” an LLM requires supplying the LLM with large amounts of text for
>
> the LLM to ingest—the more text, the better. That is, in part, the large in large language
> model.
>
> 58. As the U.S. Patent and Trademark Office has observed, LLM “training” “almost
>
> by definition involve[s] the reproduction of entire works or substantial portions thereof.”4
>
> 59. “Training” in this context is therefore a technical-sounding euphemism for
>
> “copying and ingesting.”
>
> 60. The quality of the LLM (that is, its capacity to generate human-seeming responses
>
> to prompts) is dependent on the quality of the datasets used to “train” the LLM.
>
> 61. Professionally authored, edited, and published books—such as those authored by
>
> Plaintiffs here—are an especially important source of LLM “training” data.
>
> 62. As one group of AI researchers (not affiliated with Defendants) has observed,
>
> “[b]ooks are a rich source of both fine-grained information, how a character, an object or
> a scene
>
> looks like, as well as high-level semantics, what someone is thinking, feeling and how these
>
> states evolve through a story.”5
>
> 63. In other words, books are the high-quality materials Defendants want, need, and
>
> have therefore outright pilfered to develop generative AI products that produce high-quality
>
> results: text that appears to have been written by a human writer.
>
> 64. This use is highly commercial.>>.
>
>
> _______________
>
> Le informazioni contenute nella presente comunicazione e nei documenti ad essa allegati
> potrebbero essere tutelate dal segreto professionale e sono comunque confidenziali e ad
> uso esclusivo del destinatario sopra indicato. Qualora la presente comunicazione non fosse
> destinata a Voi, Vi preghiamo di tener presente che la divulgazione, distribuzione o
> riproduzione di qualunque informazione contenuta nella presente comunicazione o nei
> documenti ad essa allegati sono vietate. Se avete ricevuto la presente comunicazione per
> errore, Vi preghiamo di volerci avvertire immediatamente e di distruggere quanto ricevuto
> senza leggerlo. Grazie per la collaborazione.
>
>
> The information contained in this email and any documents attached to it may be legally
> privileged and confidential. The information is intended only for the use of the
> individual or entity named above. If you are not the intended recipient, you are hereby
> notified that any use, dissemination, distribution or reproduction of any information
> contained in or attached to this email is prohibited. If you have received this email in
> error, please immediately notify us by reply email or by telephone, and destroy the
> original transmission and its attachments without reading them. Thank you.
>
>>>-----Messaggio originale-----
>
>>>Da: nexa <nexa-bounces(a)server-nexa.polito.it> Per conto di Rossana
>
>>>Morriello
>
>>>Inviato: venerdì 29 settembre 2023 16:08
>
>>>A: Nexa <nexa(a)server-nexa.polito.it>
>
>>>Oggetto: [nexa] R: ‘Biggest act of copyright theft in history’: thousands of
>
>>>Australian books allegedly used to train AI model | Australia news | The
>
>>>Guardian
>
>>>
>
>>>Non sono una giurista ma credo che questa rassegna possa essere utile alla
>
>>>discussione
>
>>>
>
>>>https://www.thefashionlaw.com/from-chatgpt-to-deepfake-creating-apps-a-<https://www.thefashionlaw.com/from-chatgpt-to-deepfake-creating-apps-a-runn…>
>
>>>running-list-of-key-ai-lawsuits/
>
>>>
>
>>>
>
>>>Saluti
>
>>>Rossana Morriello
>
>>>
>
>>>
>
>>>
>
>>>
>
>>>-----Messaggio originale-----
>
>>>Da: nexa <nexa-bounces(a)server-nexa.polito.it<mailto:nexa-bounces@server-nexa.polito.it>>
> Per conto di Stefano
>
>>>Quintarelli
>
>>>Inviato: venerdì 29 settembre 2023 15:21
>
>>>Cc: Nexa <nexa(a)server-nexa.polito.it<mailto:nexa@server-nexa.polito.it>>
>
>>>Oggetto: Re: [nexa] ‘Biggest act of copyright theft in history’: thousands of
>
>>>Australian books allegedly used to train AI model | Australia news | The
>
>>>Guardian
>
>>>
>
>>>Ho una domanda per i giuristi (anzi, piu' di una)
>
>>>
>
>>>per allenare un modello, ho bisogno di un file con la versione digitale di un
>
>>>testo.
>
>>>(cosnsidero ovviamente testi non PD, CC0, ecc.)
>
>>>
>
>>>la versione digitale di un testo la posso ottenere da un ebook (gia' digitale),
>
>>>togliendo il probabile DRM.
>
>>>ma un ebook non e' unbene ma e' un servizio soggetto a licenza d'uso, quindi
>
>>>se non e'
>
>>>prevista nella licenza d'uso la facolta' di estrarre il testo digitale per allenarci un
>
>>>modello, mi sembra che ci sia gia' una violazione della licenza, per cui, credo,
>
>>>non possa essere usato come base di un allenamento, tanto piu' se il fine di
>
>>>tale allenamento e'
>
>>>commerciale (se vendo un servizio basato su quel modello).
>
>>>
>
>>>se e' cosi', per allenare il mio modello devo allora prednere il testo digitale
>
>>>facendo scan/ocr di un testo cartaceo.
>
>>>ma cio' e' possibile, se non erro, solo per uso personale e non commerciale.
>
>>>
>
>>>se questo e' corretto, non mi pare ci sia un modo per prendere un testo digitale
>
>>>senza infrangere una licenza d'uso/copyright
>
>>>
>
>>>dove e' la fallacia del ragionamento ?
>
>>>
>
>>>grazie, s.
>
>>>
>
>>>On 29/09/23 15:00, Stefano Borroni Barale wrote:
>
>>> > Buongiorno lista,
>
>>> >
>
>>> >> L'idea che istruire un modello su dei testi coperti da copyright sia
>
>>> >> una violazione del suddetto copyright è altamente opinabile
>
>>> >
>
>>> > Fin qui, ho l'impressione che tutti i legali in lista concorderanno.
>
>>> >
>
>>> >> ragionamento è in realtà abbastanza semplice: se istruirsi su un
>
>>> >> testo ne violasse il copyright, saremmo tutti dei criminali.
>
>>> >
>
>>> > Ma siccome noi siamo umani e quello che produciamo non è - salvo i discorsi
>
>>>dei politici(*) - ontologicamente identico alla produzione di esseri tecnici non
>
>>>viventi, logica vuole che quanto si applica a noi non possa applicarsi a un LLM,
>
>>>tanto quanto la legge sul copyright non si applica pedissequamente all'utilizzo
>
>>>di testi umani per creare modelli linguistici.
>
>>> >
>
>>> > Questo è il motivo per il quale tutti i tentativi di "proteggere via copyright" il
>
>>>prodotto di software generativi sono falliti miseramente, e con motivazioni
>
>>>scritte in sentenze; che per il diritto credo abbiano un peso assai maggiore del
>
>>>sito di CC.
>
>>> >
>
>>> > La mia impressione è che la questione terrà impegnati legali, informatici,
>
>>>filosofi e società ancora moooooolto a lungo.
>
>>> > SBB
>
>>> >
>
>>> > (*) Come sanno bene i bambini degli anni '80 che hanno giocato con
>
>>> > questo spassoso giocattolo:
>
>>> >https://www.enricodalbosco.it/giochi/tubolario/<https://www.enricodalbosco.it/giochi/tubolario/>
>
>>> >
>
>>> >
>
>>> > Di quei testi
>
>>> >> non c'è fisicamente traccia all'interno dei modelli, non viene
>
>>> >> copiato niente. I modelli sono un'opera trasformativa di quei testi,
>
>>> >> non derivativa.
>
>>> >>
>
>>> >> Lo argomenta molto bene Creative Commons:
>
>>> >>https://creativecommons.org/2023/02/17/fair-use-training-generative-a<https://creativecommons.org/2023/02/17/fair-use-training-generative-a>
>
>>> >> i/
>
>>> >>
>
>>> >> Detto questo, cito le parole di un altro autore, Jeff Jarvis:
>
>>> >>
>
>>>https://www.facebook.com/jeff.jarvis/posts/pfbid0LMFeqdTYoxnGHQAZwp5<https://www.facebook.com/jeff.jarvis/posts/pfbid0LMFeqdTYoxnGHQAZwp5H>
>
>>>H
>
>>> >> MmeeVqgMSjL2dkcwMcBojkb2cinBpgYTHyc7Fhq1B9NPl
>
>>> >>
>
>>> >> «I, for one, am not complaining about my books being in in large
>
>>> >> language model training sets. I write to enter ideas into public
>
>>> >> discourse. I prefer informed over ignorant AI. I believe it is fair
>
>>> >> use for anyone to read & use books for transformative work. In fact,
>
>>> >> I'd probably feel snubbed if my books were not there. I'm happy when
>
>>> >> they are in libraries. I'm fine that they're here.»
>
>>> >>
>
>>> >> Fabio
>
>>> >>
>
>>> >> Il giorno ven 29 set 2023 alle ore 07:52 Alberto Cammozzo via nexa
>
>>> >>nexa(a)server-nexa.polito.it<mailto:nexa@server-nexa.polito.it>ha scritto:
>
>>> >>
>
>>> >>>https://www.theguardian.com/australia-news/2023/sep/28/australian-<https://www.theguardian.com/australia-news/2023/sep/28/australian-bo>
>
>>>bo
>
>>> >>> oks-training-ai-books3-stolen-pirated
>
>>> >>>
>
>>> >>> Thousands of books from some of Australia’s most celebrated authors
>
>>>have potentially been caught up in what Booker prize-winning novelist Richard
>
>>>Flanagan has called “the biggest act of copyright theft in history”.
>
>>> >>>
>
>>> >>> The works have allegedly been pirated by the US-based Books3 dataset
>
>>>and used to train generative AI for corporations such as Meta and Bloomberg.
>
>>> >>>
>
>>> >>> Flanagan, who found 10 of his works, including the multi-international
>
>>>award-winning 2013 novel The Narrow Road to the Deep North, on the
>
>>>Books3 dataset, told Guardian Australia he was deeply shocked by the
>
>>>discovery made several days ago.
>
>>> >>>
>
>>> >>> “I felt as if my soul had been strip mined and I was powerless to stop it,”
>
>>>he said in a statement.
>
>>> >>>
>
>>> >>> “This is the biggest act of copyright theft in history.”
>
>>> >>>
>
>>> >>> AI could ‘turbo-charge fraud’ and be monopolised by tech companies,
>
>>> >>> Andrew Leigh warns
>
>>> >>>
>
>>> >>> The Australian Publishers Association confirmed to Guardian Australia on
>
>>>Wednesday that as many as 18,000 fiction and nonfiction titles with
>
>>>Australian ISBNs (unique international standard book numbers) appeared to
>
>>>be affected by the copyright infringement, although it is not yet clear what
>
>>>proportion of these are Australian editions of internationally authored books.
>
>>> >>>
>
>>> >>> “We’re still working through [the data] to work out the impact in terms of
>
>>>Australian authors,” APA spokesperson Stuart Glover said.
>
>>> >>>
>
>>> >>> “This is a massive legal and ethical challenge for the publishing industry
>
>>>and for authors globally.”
>
>>> >>>
>
>>> >>> A search tool published on Monday by US media platform The Atlantic and
>
>>>uploaded by the US Authors Guild on Wednesday revealed the works of Peter
>
>>>Carey, Helen Garner, Kate Grenville, Anna Funder, Christos Tsiolkas and
>
>>>Thomas Keneally, as well as Flanagan and dozens of other high-profile
>
>>>Australian authors, were included in the pirated dataset containing more than
>
>>>180,000 titles.
>
>>> >>>
>
>>> >>> On Thursday, the Australian Society of Authors issued a statement saying
>
>>>it was “horrified” to learn that the works of Australian writers were being used
>
>>>to train artificial intelligence without permission from the authors.
>
>>> >>>
>
>>> >>> ASA chief executive, Olivia Lanchester, described the Books3 dataset as
>
>>>piracy on an industrial scale.
>
>>> >>>
>
>>> >>> “Authors appropriately feel outraged,” Lanchester said. “The fact is this
>
>>>technology relies upon books, journals, essays written by authors, yet
>
>>>permission was not sought nor compensation granted.”
>
>>> >>>
>
>>> >>> Lanchester said the Australian literary industry, while not objecting per se
>
>>>to emerging technologies such as AI, was deeply concerned about the lack of
>
>>>transparency evident in the development and monetisation of AI by global
>
>>>tech companies.
>
>>> >>>
>
>>> >>> “Turning a blind eye to the legitimate rights of copyright owners threatens
>
>>>to diminish already precarious creative careers,” she said.
>
>>> >>>
>
>>> >>> “The enrichment of a few powerful companies is at the cost of thousands
>
>>>of individual creators. This is not how a fair market functions.”
>
>>> >>>
>
>>> >>> Josephine Johnston, chief executive of Australia’s Copyright Agency,
>
>>>described the Books3 development as “a free kick to big tech” at the expense
>
>>>of Australia’s creative and cultural life.
>
>>> >>>
>
>>> >>> “We’re going to need greater transparency – how these tools have been
>
>>>developed, trained, how they operate – before people can truly understand
>
>>>what their legal rights might be,” she said.
>
>>> >>>
>
>>> >>> “We seem to be in this terrible position now where content owners –
>
>>>remembering that the vast majority of them will be individual authors – may
>
>>>actually have to take out court cases to enforce their rights.”
>
>>> >>>
>
>>> >>> Australian copyright law protects creators of original content from data
>
>>>scraping.
>
>>> >>>
>
>>> >>> Litigation in the US against ChatGPT creator OpenAI over use of allegedly
>
>>>pirated book datasets, Books1 and Books2 (which do not appear to be
>
>>>affiliated with Books3) has already commenced.
>
>>> >>>
>
>>> >>> In July, North American horror/fantasy writers Mona Awad (author of
>
>>>Bunny) and Paul Tremblay (author of The Cabin at the End of the World) filed a
>
>>>lawsuit in a San Francisco federal court, alleging ChatGPT unlawfully digested
>
>>>their books as part of its AI training data.
>
>>> >>>
>
>>> >>> On 28 August, OpenAI filed a motion to dismiss the lawsuit, arguing that
>
>>>the authors “misconceive the scope of copyright, failing to take into account
>
>>>the limitations and exceptions (including fair use) that properly leave room for
>
>>>innovations like the large language models now at the forefront of artificial
>
>>>intelligence”.
>
>>> >>>
>
>>> >>> On 19 September the Writers Guild and 17 of its members, including
>
>>>bestselling novelists John Grisham, George RR Martin and Jodi Picoult, filed a
>
>>>complaint in a New York district court against OpenAI, seeking redress for
>
>>>“flagrant and harmful infringements” of guild members’ registered copyrights.
>
>>> >>>
>
>>> >>> In a statement on its website, the guild says while it is aware that
>
>>>companies such as Meta and Bloomberg have used the Books3 dataset to
>
>>>train their LLMs, it is not yet clear whether OpenAI is using Books3 to train its
>
>>>ChatGPT models GPT 3.5 or GPT 4.
>
>>> >>>
>
>>> >>> Democracies face ‘truth decay’ as AI blurs fact and fiction, warns
>
>>> >>> head of Australia’s military
>
>>> >>>
>
>>> >>> Guardian Australia has sought comment from OpenAI, which has yet to
>
>>>officially respond to the guild’s complaint, and Meta.
>
>>> >>>
>
>>> >>> On 4 September, US technology magazine Wired reported that a Danish
>
>>>anti-piracy group called Rights Alliance had been told by Bloomberg that the
>
>>>company did not plan to train future versions of its BloombergGPT using
>
>>>Books3.
>
>>> >>>
>
>>> >>> Bloomberg declined to respond to the Guardian’s queries.
>
>>> >>>
>
>>> >>> The APA said the global nature of the issue would present significant
>
>>>challenges in enforcement and prosecution, and has joined the authors’
>
>>>society in calling for AI technologies to be regulated.
>
>>> >>>
>
>>> >>> Consultation closed last month for a Department of Industry, Science and
>
>>>Resources discussion paper on supporting responsible AI.
>
>>> >>>
>
>>> >>> A parliamentary inquiry is under way examining the use of generative
>
>>>artificial intelligence in the Australian education system.
>
>>> >>>
>
>>> >>> Flanagan said it was up to the Australian government to act to protect
>
>>>Australia’s writers.
>
>>> >>>
>
>>> >>> “It has power and we do not,” he said.
>
>>> >>>
>
>>> >>> “If it cares for our culture it must now stand up and fight for it.”
>
>>> >>>
>
>>> >>> _______________________________________________
>
>>> >>> nexa mailing list
>
>>> >>>nexa(a)server-nexa.polito.it<mailto:nexa@server-nexa.polito.it>
>
>>> >>>https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa<https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa>
>
>>> >>
>
>>> >> _______________________________________________
>
>>> >> nexa mailing list
>
>>> >>nexa(a)server-nexa.polito.it<mailto:nexa@server-nexa.polito.it>
>
>>> >>https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa<https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa>
>
>>> > _______________________________________________
>
>>> > nexa mailing list
>
>>> >nexa(a)server-nexa.polito.it<mailto:nexa@server-nexa.polito.it>
>
>>> >https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa<https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa>
>
>>>_______________________________________________
>
>>>nexa mailing list
>
>>>nexa(a)server-nexa.polito.it<mailto:nexa@server-nexa.polito.it>
>
>>>https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa<https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa>
>
>>>_______________________________________________
>
>>>nexa mailing list
>
>>>nexa(a)server-nexa.polito.it<mailto:nexa@server-nexa.polito.it>
>
>>>https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa<https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa>
>
>
> _______________________________________________
> nexa mailing list
> nexa(a)server-nexa.polito.it
> https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
Sept. 29, 2023
Foresight & Data Protection a Milano il 6 Ottobre
by Stefano Leucci
Cari tutti,
scrivo per invitarvi all’evento che si terrà il prossimo venerdì *6 ottobre
a Milano, alle ore 14:00*.
Questa occasione speciale sarà dedicata all’esplorazione dei punti di
contatto tra due discipline complesse e che so essere molto vicine agli
interessi di Nexa: *il foresight e la data protection*.
Il programma prevede un intervento iniziale da parte del manager del
laboratorio d'innovazione del CNIL (l’Autorità Privacy Francese).
Successivamente, una tavola rotonda arricchirà il dibattito con la
partecipazione di speaker provenienti dal mondo pubblico, privato ed
accademico, i quali si occuperanno di esaminare e discutere strategie
“proattive e anticipatorie” nel settore della privacy.
*Mi risulta essere la prima volta che si organizza un evento con questo
focus specifico in Italia*, un motivo in più per partecipare a questo
momento di apprendimento e confronto.
Per maggiori dettagli sull'evento, vi invito a visitare il seguente
link: https://www.privacyforfutures.it/2023/09/27/privacy-for-futures-i-possibili…
Inoltre, se siete interessati ad avere un assaggio preliminare sui temi che
verranno trattati, abbiamo iniziato una discussione qui:
https://www.agendadigitale.eu/sicurezza/privacy/i-futuri-della-privacy-perc…
Avremo anche una diretta online. Non appena sarà disponibile, vi farò avere
anche il link per connettervi in caso non riusciste ad essere presenti di
persona.
Un caro saluto a tutti,
Stefano
Sept. 29, 2023