nexa
By thread
nexa@server-nexa.polito.it
By month
Messages by month
- ----- 2026 -----
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2025 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2024 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2023 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2022 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2021 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2020 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2019 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2018 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2017 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2016 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2015 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2014 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2013 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2012 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2011 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2010 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2009 -----
- December
- November
- October
- September
- August
- July
- June
- May
October 2023
- 51 participants
- 263 messages
Re: [nexa] ‘Biggest act of copyright theft in history’: thousands of Australian books allegedly used to train AI model | Australia news | The Guardian
by 380°
Buongiorno Giacomo,
Giacomo Tesio <giacomo(a)tesio.it> writes:
[...]
>> stai dicendo che quelle parti di testo, che sono espresse in /forma/
>> difficilmente estraibile, sarebbero plagio (ampie parti dei testi
>> originali)?
>
> Se siano plagio o semplicemente opere derivata create e distribuite
> senza il permesso dell'autore è una valutazione giuridica che non so
> fare.
Io invece /credo/ di saperla fare (voglio l'Orso d'oro in faccia tosta)
sulla base di quello che osservo e mi pare non ci sia nessuna delle
fattispecie che indichi
...solo un processo potrà dirlo
> Sto semplicemente dicendo che quei testi sono in gran parte presenti nel LLM seppure
> codificati con perdita di informazione.
OK, su cosa succede tecnicamente, ovvero sul tipo di elaborazione e
immagazzinamento dei testi _elaborari_, credo sia tutto sufficientemente
chiaro.
> Una similitudine tecnicamente più attinente di uno zip sarebbe un jpeg
> o un mp3, ma non volevo confondere ulteriormente il mio interlocutore.
Ottimo, vedo che tecnicamente siamo allineati :-)
[...]
> Ripeto, a me può anche stare bene purché io possa nello stesso modo disassemblare
> Microsoft Windows o Microsoft Office e distribuirne il codice sotto
> GPL,
Ma Giacomo! Non solo /tu/ puoi farlo, è *già* stato fatto ed è
perfettamente lagale nonostante quello che "si dice in giro"; la tecnica
si chiama "binary reverse engineering" [1].
Chissà se un giorno qualcuno avrà il coraggio di programmare un sistema
di machine learning, magari basato su bLLM (binary large language model)
per aiutare i ricercatori ad applicare quella tecnica :-O
> magari attribuendolo a Mickey Mouse.
Ha beccato proprio il nome *perfetto* per attirare l'attenzione del
Censore Intergalattico, Topo Gigio darebbe meno nell'occhio :-D
> Non sono contrario alla abolizione delle varie forme di "proprietà
> intellettuale",
Non dovresti usare quella bestemmia! ...ma ti perdono :-)
> voglio solo sia esplicita e valga per tutti.
>
> Ma finché non posso usare come mi piace il codice di Microsoft, Microsoft non
> deve usare come le pare il mio.
Se tu potessi leggere il codice sorgente di qualsiasi software
proprietario senza essere costretto a firmare un NDA, lo potresti
_rielaborare_ *anche* usando lo stesso linguaggio di programmazione e
distribuire quel codice rielaborato: sono *certo* che tu e almeno altre
decine di migliaia di bravi programmatori sareste in grado di
modificarne la forma espressiva in modo tale che non risulti manco come
opera derivata.
...per tutto il resto c'è il binary reverse engineering (che costa
troppa fatica!)
>> > Violazioni particolarmente gravi perché il codice GPL viene poi
>> > incluso in prodotti proprietari.
>>
>> se permetti, sono stracavolacci di quelli che copia-incollano l'output
>> da CopyALot, non ho verificato ma scommetto un fiorino che è pure
>> scritto chiaramente nelle condizioni di utilizzo del servizio
>
> No 380: se io voglio riservare alla collettività un mio pezzo di codice utilizzando
> una licenza copyleft (metti la AGPL) e Microsoft lo distribuisce senza attribuzione
> e con licenza sbagliata, buttandolo in mezzo ad un software proprietario di un
> proprio cliente pagante, se permetti "sono stracavolacci" miei.
>
> Subisco un danno morale ed economico.
Sì ma solo se il cliente pagante spegne il cervello e usa
pedissequamente l'output del servizio di turno
>
> E quel che è peggio, non ho alcun modo di individuare precisamente quanto grave
> sia questo danno morale, ovvero in quanti software proprietari che aborro
> il mio lavoro sia stato inserito.
>
> La responsabilità del programmatore che riceve il mio codice da CopyALot viene dopo:
> prima Microsoft ha realizzato un opera derivata dal mio codice (il "modello" di Copilot)
> che distribuisce il mio codice senza
> permesso quando deve "rompere il ghiaccio".
>
>
>> quale sarebbe la violazione del diritto d'autore, se non c'è plagio *e*
>> chi usa quei testi per usarli in una elaborazione ha *pagato* "i libri"?
>
> Anche ne avesse comprate un milione di copie, non avrebbe alcun diritto di creare
> opere derivate.
>
> Quanti testi CC-BY ND sono stati usati da OpenAI per programmare statisticamente
> ChatGPT?
>
> Quanti CC-BY SA?
>
>
> IMHO ha violato e sta violando entrambe le licenze.
>
> E quel che è peggio è che sta violando i diritti morali degli autori.
>
>
> Giacomo
[1] https://en.wikipedia.org/wiki/Reverse_engineering#Binary_software
--
380° (Giovanni Biscuolo public alter ego)
«Noi, incompetenti come siamo,
non abbiamo alcun titolo per suggerire alcunché»
Disinformation flourishes because many people care deeply about injustice
but very few check the facts. Ask me about <https://stallmansupport.org>.
Oct. 1, 2023
Authors shocked to find AI ripoffs of their books being sold on Amazon | Artificial intelligence (AI) | The Guardian
by Alberto Cammozzo
<https://www.theguardian.com/technology/2023/sep/30/authors-shocked-to-find-…>
Altri esempi di digital information pollution [1] dovuti alla industrializzazione della produzione di testi con LLM.
Alberto
[1] <https://en.m.wikipedia.org/wiki/Information_pollution>
Oct. 1, 2023
Has Google’s monopoly on the search engine market finally timed out? | John Naughton | The Guardian
by Alberto Cammozzo
<https://www.theguardian.com/commentisfree/2023/sep/30/google-antitrust-us-d…>
Although you’d never guess it from mainstream media, the most significant antitrust case in more than 20 years is under way in Washington. In it, the US justice department, alongside the attorneys general of eight states, is suing Google for abusively monopolising digital advertising technologies, thereby subverting competition through “serial acquisitions” and anti-competitive auction manipulation. Or, to put it more prosaically, arguing that Google – which has between 90% and 95% of the search market – has maintained its monopoly not by making a better product, but by locking down almost every avenue through which consumers might find a different search engine and making sure they only see Google wherever they look.
Why is this significant? Basically, because the US government has been asleep at the wheel for almost a quarter of a century and has finally woken up to its democratic responsibilities. The last time it stirred itself to take on an aggressive monopolist was in 2001, when it sued Microsoft for illegally tying its Internet Explorer browser to Windows as part of a (successful) campaign to destroy Netscape, maker of the first distinctive commercial web browser, which Bill Gates and co perceived as a potentially lethal competitive threat. In an eerie echo of that earlier lawsuit, the justice department is now accusing Google of similar tactics – for example, illegally tying the company’s search engine to its Android smartphone operating system and its Chrome browser. And the government is seeking to break up the company, just as it once sought to break up Microsoft.
The parallels between the two cases are striking. In 2001, for example, Microsoft Windows had 93% of the global market for operating systems. In 2023, Google has 92% of the market for its search engine.
In the 1990s, Microsoft had been slow to appreciate the significance of the web and was late to the market with a mediocre browser – Internet Explorer – much inferior to the Netscape alternative. But if you were a manufacturer of PCs in those days, you couldn’t get a licence to install Windows on them without also bundling Explorer, making the Microsoft browser the default, which was about as anti-competitive as you could get.
Where Google has the power, it makes its search engine the default; where it doesn’t, it uses money
Now Google, according to the justice department, is also in the default-setting game. Where it has the necessary power – as with the Android operating system that it controls, or its now-dominant Chrome browser – it makes Google’s search engine the default. Where it lacks ownership, it uses money – for example, paying $10bn a year for privileges such as making Google the default search engine on Apple iOS. How Google squares this lavish expenditure with its insistence that the dominance of its (free) search engine confirms its excellence is one of the intriguing mysteries of the trial. Is Apple taking Google for a lucrative ride? Or is Google worried that if it were not the default search, iPhone users might, er, defect?
As well they might. In its early days, Google’s search engine was a breath of fresh air, so much so that some people divided internet eras into BG (Before Google) and AG. But over the years, it has morphed into a laughable betrayal of its co-founders’ original high-mindedness about the evils of advertising. In 2020, a randomised trial found that Google-associated results (ads for, or links to, the company’s other services) constituted more than 60% of the “first screen” – what is visible initially on a smartphone – of an average Google search result. And in one of five searches, the entire first screen was Google results.
I gave up using Google years ago, but even occasional visits have tended to confirm its decay, which is understandable given that 57% of its revenue now comes from search ads. Google is going the way of all monopolists, morphing from innovation into rent-seeking. A bit like Microsoft, in other words, before the latter discovered AI.
There is, however, one big difference between the 2001 Microsoft case and the one now in progress in Washington – the absence of media coverage. Back when Microsoft had its back to the wall, the trial was widely covered by mainstream media. But the Google case is getting relatively little airtime. In part, this may be because there is, alas, zero public interest in antitrust. But events so far suggest a more worrying explanation – namely, the apparent deference of Judge Amit Mehta to Google’s neurotic demands to keep as much as possible of the evidence presented in court out of the public eye.
Early in the proceedings, for example, he denied a third-party motion to broadcast a publicly accessible audio feed of the trial. As a consequence, the hearing is only available to people who can attend in person. And even if you can attend, as the writer and former policymaker Matt Stoller reports on his BIG newsletter, “it’s hard to see the trial because huge portions are fully sealed”. This is no way for a democracy to go about checking unaccountable corporate power. Justice needs to be seen to be believed, even if Google disagrees
Oct. 1, 2023