nexa
By thread
nexa@server-nexa.polito.it
By month
Messages by month
- ----- 2026 -----
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2025 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2024 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2023 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2022 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2021 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2020 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2019 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2018 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2017 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2016 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2015 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2014 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2013 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2012 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2011 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2010 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2009 -----
- December
- November
- October
- September
- August
- July
- June
- May
- 42 participants
- 30624 messages
Re: [nexa] Spreadsheets Considered Harmful.
by 380°
Buongiorno,
ATTENZIONE, messaggio nella bottiglia per tutti gli utenti di fogli di
calcolo sulla terra: se pensate che non stiate già /programmando/, vi
sbagliate.
So che questo thread è già diventato abbastanza incasinato ma la
superficialità con la quale a volte anche gli informatici trattano
l'argomento in oggetto merita un approfondimento.
Grazie Andrea per i tuoi commenti, che io intendo /approfondire/ un
pochino affinché non rimangano dannosi equivoci.
Andrea Trentini <andrea.trentini(a)unimi.it> writes:
> On 27/10/2023 10:58, Giuseppe Attardi wrote:
>>> On 26 Oct 2023, at 17:12, nexa-request(a)server-nexa.polito.it wrote:
>>> Aneddoto: lo sapete che la politica di "austerity" implementata
>>> dalla EU fu frutto di un errore in un foglio di calcolo [3], vero?
>>> :-O ...già, i fogli di calcolo fanno già abbastanza danni,
>>
>> Non sono i fogli di calcolo a fare danni, ma chi li usa malamente e a
>> sostegno di tesi preconcette:
No, no, no, assolutamente no: il problema è che **non esiste un modo per
usare bene i fogli di calcoro**, ovvero usare fogli di calcolo è
/dannoso/, perché i fogli di calcolo _sono_ una forma di programma E il
linguaggio (chiamatelo sistema, se volete) di programmazione usato per
generarli... fa pena. Punto.
Usare fogli di calcolo da solo l'illusione di aver sotto controllo il
processo di programmazione, oltretutto incasina oltre ogni decenza la
vita di chi (anche sè stessi) deve metter mano ai casini che i fogli di
calcolo /nascondono/.
...ma a sQuola si continua a usare: "word" per il trattamento testi,
"excel" per la gestione dei dati... e "power point" per le
presentazioni; continuiamo a /fare/ del male, avanti!
C'è un libro dal titolo «Cleaning Data for Effective Data Science» nel
quale un intero capitolo è dedicato ai fogli di calcolo: «Spreadsheets
Considered Harmful».
[1] https://gnosis.cx/cleaning/tabular.html#spreadsheets
Questo è un "executive summary" dei concetti:
--8<---------------cut here---------------start------------->8---
* Non-enforced field/column identity
* Computational opacity
* Semi-tabular data
* Non-contiguous data
* Invisible data and data type discrepancies
* User interface as attractive nuisance
[...] spreadsheets in general, and Excel in particular, are anathema to
effective data science. While perhaps not as much as in CSV files, a
great share of the world's data lives in Excel spreadsheets. There are
numerous kinds of data corruption that are the special realm of
spreadsheets. [...]
--8<---------------cut here---------------end--------------->8---
N.B.: qui si sta parlando _solo_ dei problemi legati all'uso del
"paradigma" del foglio di calcolo per la gestione dei dati, poi
ovviamente c'è tutto il resto dei /problemi/ che si applicano a tutti
gli altri "paradigmi" (CSV, database) di gestione dati. [2]
> il discorso "chi li usa malamente" non regge secondo me, è lo stesso
> di chi difende il C dicendo che "basta saperlo usare e non si fanno
> danni"
Superficialmente il paragone con un linguaggio di programmazione
potrebbe sembrare fuori luogo _ma_ grattando appena la superficie si
"scopre" presto che il processo chiamato "gestione dei dati"
è... programmazione:
--8<---------------cut here---------------start------------->8---
In procedural programming (including object-oriented programming),
actions flow sequentially through code, with clear locations for
branches or function calls; even in functional paradigms, compositions
are explicitly stated. In spreadsheets it is anyone's guess what
computation depends on what else, and what data ranges are actually
included. Errors can occasionally be found accidentally, but program
analysis and debugging are *nearly* impossible.
--8<---------------cut here---------------end--------------->8---
(rif. [1])
In altre parole, i programmi per lo sviluppo e l'esecuzione dei fogli di
calcolo sono delle IDE (Integrated Development Environment) «for
computation, organization, analysis and storage of data in tabular
form.» [3]
Dopotutto, il primo foglio di calcolo della storia, il LANPAR del 1969,
per esteso si chiamava: «LANguage for Programming Arrays at Random».
[...]
> come esiste una reciproca influenza fra lingua e pensiero[1] esiste
> anche una reciproca influenza fra linguaggio (e Excel lo è)
Sì Excel è una IDE per fogli di calcolo con funzioni (librerie) e
linguaggi incorporati (embedded), che combina i paradigmi di
programmazione visuale con quelli di programmazione "macro" in Visual
Basic e (in futuro?) Python.
Nel 2021, con l'introduzione di LAMBDA, Excel è diventato un ambiente
(contenente un linguaggio) di programmazione "turing complete":
--8<---------------cut here---------------start------------->8---
Ever since it was released in the 1980s, Microsoft Excel has changed how
people organize, analyze, and visualize their data, providing a basis
for decision-making for the millions of people who use it each day. It’s
also the world’s most widely used programming language.
Excel formulas are written by an order of magnitude more users than all
the C, C++, C#, Java, and Python programmers in the world
combined. Despite its success, considered as a programming language
Excel has fundamental weaknesses.
--8<---------------cut here---------------end--------------->8---
(https://www.microsoft.com/en-us/research/blog/lambda-the-ultimatae-excel-wo…)
Quindi: Excel è il linguaggio di programmazione di gran lunga più usato
al mondo.
...giusto per ribadire in che stato siamo messi.
> e sviluppo di strumenti (un foglio di calcolo è uno strumento)
Uno strumento *sbagliato*, che /facilita/ alcune cose e /complica/ TUTTO
il resto, **incluso** il _debuging_ (perché è programmazione, punto):
--8<---------------cut here---------------start------------->8---
Most of what spreadsheets do to make themselves convenient for their
users makes them bad for scientfic reproducibility, data science,
statistics, data analysis, and related areas. Spreadsheets have apparent
rows and columns in them, but nothing enforces consistent use of those,
even within a single sheet. Some particular feature often lives in
column F for some rows, but the equivalent thing is in column H for
other rows, for example. Contrast this with a CSV file or an SQL table;
for these latter formats, while all the data in a column is not
necessarily good data, it generally must pertain to the same feature.
Another danger of spreadsheets is not around data ingestion, per se, at
all. Computation within spreadsheets is spread among many cells in no
obvious or easily inspectable order, leading to numerous large-scale
disasterous consequences [...]
--8<---------------cut here---------------end--------------->8---
(rif. [1])
Se penso che MS ha creato un intero linguaggio di programmazione sul
paradigma dei fogli di calcolo
(https://en.wikipedia.org/wiki/Microsoft_Power_Fx) mi vengono i
brividi. :-D
Lo dico in un'altro modo: se avete necessità di imparare a programmare,
invece che sbattervi per anni /inutilmente/ ad imparare (a programmare
in) Excel - /anche/ con LAMBDA - o Microsoft Fx, vi consiglio
_vivamente_ di usare un linguaggio sano, da Python in su va bene
qualsiasi, R è il più indicato per il trattamento dei dati (si veda
sotto).
[...]
> un foglio elettronico, come il C, è un linguaggio in cui è facilissimo
> "spararsi in un piede"[2] e, a differenza del C, lo usano tutti, anche
> chi non ha la minima idea della sua pericolosità ("error-proneness")
Perché pensano che no, loro non stanno programmando (anche il più banale
dei fogli di calcolo è un programma).
I gravi difetti dei fogli di calcolo sono descritti anche su Wikipedia [3]
https://en.wikipedia.org/wiki/Spreadsheet#Shortcomings: ci sono così
tante cose _specifiche_ e ben documentate che non è possibile fare alcun
riassunto, quello che c'è scitto è già un riassunto del riassunto.
C'è una e una sola soluzione al complesso di problemi che l'uso dei
fogli di calcolo introducono nel processo di gestione (che comprende
l'elaborazione) dei dati: NON usarli.
Quella per i fogli di calcolo è una dipendenza dalla quale si può e si
deve guarire:
--8<---------------cut here---------------start------------->8---
The perception of the ease-of-use of spreadsheets is to some extent an
illusion. It is dead easy to get an answer from a spreadsheet, however,
it is not necessarily easy to get the right answer. Thus the distorted
view.
The difficulty of using alternatives to spreadsheets is overestimated by
many people. Safety features can give the appearance of difficulty when
in fact these are an aid.
The hard way looks easy, the easy way looks hard.
[...] Perhaps the best alternative to spreadsheets for their
computational function is the R language. R is a version of the S
language that was first created at Bell Labs. Though this is often
thought of as just being for statistics, that is not true. It was
designed for computing with data — precisely what spreadsheets are used
for.
[...] Many people will gasp in horror at the thought of using a
programming language instead of a spreadsheet. The fact is that a
spreadsheet is a programming language — it is just one that you are used
to.
--8<---------------cut here---------------end--------------->8---
(https://www.burns-stat.com/documents/tutorials/spreadsheet-addiction/)
--8<---------------cut here---------------start------------->8---
Main points:
* what is done in spreadsheets can be done in R
* the vast majority of people who could benefit from R are not using it
* spreadsheets are dangerous for complex analyses
* debugging spreadsheets is close to impossible
* spreadsheets are slow
* spreadsheets (for some reason) tend to have ugly graphics
--8<---------------cut here---------------end--------------->8---
(
https://www.burns-stat.com/documents/presentations/3-5-reasons-to-switch-fr…
)
--8<---------------cut here---------------start------------->8---
There are many reasons to make the switch from spreadsheets to R.
But for me there is only one elephant in the room. The elephant is
safety
--8<---------------cut here---------------end--------------->8---
(https://www.burns-stat.com/pages/Present/Excel_to_R_annotated.pdf)
[...]
Saluti, 380°
> [1] evito sterili bibliografie accademiche e vi segnalo invece un gustoso sketch del bravissimo
> Gioele Dix che "disamina" un banale avviso ferroviario anche in funzione delle culture+lingue:
> https://www.youtube.com/watch?v=8xuAw0_6qwE
>
> [2] ancora humor... https://www.eng.uwaterloo.ca/~comp03a/misc/humour/shootfoot.html
>
> [3] che hanno dato origine ai vari NUSAP, Reproducibility Project, F.A.I.R., ecc.
[2] è un peccato che l'autore, nel 2021, abbia volontariamente deciso di
non trattare i c.d. "graph databases" (e.g. RDF), perché sarebbe ora che
i "data scientists" cominciassero a usarli /massicciamente/
(rif. https://gnosis.cx/cleaning/hierarchical.html, capitolo NoSQL
Databases)
[3] https://en.wikipedia.org/wiki/Spreadsheet
--
380° (Giovanni Biscuolo public alter ego)
«Noi, incompetenti come siamo,
non abbiamo alcun titolo per suggerire alcunché»
Disinformation flourishes because many people care deeply about injustice
but very few check the facts. Ask me about <https://stallmansupport.org>.
Oct. 30, 2023
166° Mercoledì di Nexa | 8 novembre 2023, ore 17.00
by Valeria Bergantino
Gentilissime, gentilissimi,
Vi invitiamo a partecipare al *166° Mercoledì di Nexa*, che si terrà
*mercoledì 8 novembre*, alle ore 17.00,
con un incontro dal titolo /*"Etica del digitale ed euristiche mentali"*/.
Ospite dell'incontro: *Guglielmo Tamburrini* (Università di Napoli
Federico II)/*.*/
_L'incontro si terrà IN PRESENZA e ONLINE._
_SEDE FISICA_ dell'incontro: Centro Nexa su Internet e Società,
Politecnico di Torino, Via Boggio 65/a, Torino (1° piano).
Per accedere alla sala si raccomanda di suonare al citofono *Portineria*
e di seguire le indicazioni segnalate lungo il percorso.
QUI <https://nexa.polito.it/contatti> maggiori informazioni su come
raggiungerci.
_STANZA VIRTUALE_ dell'incontro:
https://didattica.polito.it/VClass/NexaEvent
Di seguito maggiori dettagli:
Se non visualizzi correttamente questo messaggio clicca qui
<https://nexa.polito.it/mercoledi-166>
NEXA
166° Mercoledì di Nexa
Mercoledì 8 novembre 2023, ore 17.00 - 19.00
Politecnico di Torino
https://nexa.polito.it/mercoledi-166
/Etica del digitale ed euristiche mentali/
*GUGLIELMO TAMBURRINI (Università di Napoli Federico II)*
*L'INCONTRO SI TERRÀ IN PRESENZA E ONLINE*
*SEDE FISICA*: Centro Nexa su Internet e Società, Politecnico di Torino,
Via Boggio 65/a, Torino (1° piano). Suonare al citofono *Portineria* -
Seguire le indicazioni segnalate dai cartelli lungo il percorso. (Per
maggiori informazioni su come raggiungerci clicca QUI
<https://nexa.polito.it/contatti>)
*STANZA VIRTUALE*: https://didattica.polito.it/VClass/NexaEvent
Per mettere in pratica le *raccomandazioni etiche sulla progettazione e
l’uso dei sistemi digitali* è fondamentale affidarsi a *processi
decisionali analitici e ponderati*. Ma i modelli psicologici del doppio
processo decisionale affermano che spesso affidiamo le nostre scelte e
azioni a valutazioni intuitive, più rapide ed emotivamente cariche. Da
qui emergono difficoltà significative per la messa in pratica di linee
di condotta analiticamente ponderate ed eticamente motivate per la
progettazione e l’uso dei sistemi digitali. Nel corso dell’incontro si
discuteranno, in particolare: la *difficoltà di perseguire*, in assenza
di esperienze vissute e reazioni viscerali al riscaldamento climatico,
il *contenimento dell’impronta di carbonio del settore digitale*; la
*difficoltà di esercitare la prudenza e la frugalità nella sfera
digitale*, differendo la fruizione immediata di stati emotivi positivi;
la *difficoltà di attuare un controllo umano veramente significativo sui
sistemi dell’IA nel settore militare* – soprattutto nel campo delle armi
autonome e della difesa nucleare – allorché si opacizzano i processi
deliberativi e si accelerano le decisioni, ostacolando così il
dispiegamento di ragionamenti analitici e ponderati da parte dei
decisori militari e politici. Si prenderanno anche in considerazione
alcune strategie per attenuare queste difficoltà, e per *conferire
all’etica del digitale un ruolo più incisivo nell’ambito della ragione
pratica*.
BIOGRAFIA e informazioni supplementari:
[mercoledì166]
*Guglielmo TAMBURRINI* è professore ordinario di filosofia della scienza
e della tecnologia all’Università di Napoli Federico II. Coordinatore
del primo progetto europeo sull’etica della robotica (Ethicbots,
2005-08), vincitore del Premio Internazionale Giulio Preti nel 2014 per
il suo lavoro didattico e di ricerca sull’etica della robotica e
dell’intelligenza artificiale, è membro dell’ICRAC (International
Committee for Robot Arms Control) e del Consiglio scientifico dell’USPID
(Unione degli Scienziati per il Disarmo). Tra le sue pubblicazioni
recenti in lingua italiana, il volume /Etica delle macchine. Dilemmi
morali per la robotica e l’intelligenza artificiale/ (Roma, 2020), vari
capitoli sia del volume /Automi e persone. Introduzione all’etica
dell’intelligenza artificiale e della robotica/ (Roma, 2021), che ha
curato insieme a F. Fossa e V. Schiaffonati, sia del volume /open access
Dai droni alle armi autonome/ (a cura di F. Farruggia, Roma 2023).
*Letture consigliate:*
* P. Slovic, H. S. Lin (2020), /The caveman and the bomb in the
digital age/, disponibile al LINK
<https://www.hoover.org/sites/default/files/research/docs/trinkunas_threetwe…>
* F. Farruggia (a cura di) (2023), /Dai droni alle armi autonome/, in
particolare i capp. 4, 5 e 9. Roma, Franco Angeli, disponibile in
open access al LINK
<https://series.francoangeli.it/index.php/oa/catalog/book/948>
* G. Tamburrini (2022), The AI carbon footprint and responsibilities
of AI scientists, Philosophies, 7(1), 4 (open access) al LINK
<https://doi.org/10.3390/philosophies7010004>
* D. Kahneman (2011), /Thinking, Fast and Slow/, trad. It. Pensieri
lenti e veloci, Milano, Mondadori, 2012
Che cosa sono il Centro Nexa e i cicli di incontri “Mercoledì di Nexa” e
“Nexa Lunch Seminar”
<https://www.facebook.com/nexa.center/> <https://twitter.com/nexacenter>
<https://www.youtube.com/user/NexaCenter>
<https://www.instagram.com/nexa_center/>
<https://www.linkedin.com/company/3054864/admin/>
#nexawednesday <https://twitter.com/search?q=%23nexawednesday&src=typd>
#UniversitàdiNapoliFedericoII <https://twitter.com/UninaIT>
Il Centro Nexa su Internet & Società del Politecnico di Torino
(Dipartimento di Automatica e Informatica), fondato nel 2006, è un
centro di ricerca interdisciplinare che, in collaborazione con
l’Università di Torino (in particolare il Dipartimento di
Giurisprudenza), studia le tecnologie digitali e il loro rapporto con la
società. Maggiori informazioni all'indirizzo: http://nexa.polito.it/about.
Durante i “Mercoledì di Nexa”, che si tengono *ogni 2° mercoledì del
mese alle ore 17 in punto*, il Centro Nexa su Internet e Società apre le
sue porte non solo agli esperti e a tutti coloro i quali lavorano con
Internet, ma anche a semplici appassionati e cittadini. Il ciclo di
incontri intende approfondire, con un linguaggio preciso ma accessibile,
i temi legati alla Rete: “intelligenza artificiale”, reti sociali,
software libero, capitalismo della sorveglianza, neutralità della rete,
libertà di espressione, privacy, condivisione di file, "big data" e
"open data", “smart cities”, e molto altro.
Al centro della maggior parte degli incontri un ospite pronto a
dialogare con i direttori del Centro Nexa, il Prof. Juan Carlos De
Martin del Politecnico di Torino, i Proff. Marco Ricolfi e Maurizio
Borghi dell'Università di Torino, lo staff, i Fellows del Centro Nexa e
tutti i presenti.
Maggiori informazioni sui Mercoledì di Nexa, incluso un elenco di tutti
i “Mercoledì” passati, sono disponibili all'indirizzo:
http://nexa.polito.it/mercoledi. Le registrazioni degli incontri sono
disponibili qui: https://nexa.polito.it/audio-video.
Si segnala inoltre che dal maggio 2012 *ogni 4° mercoledì* del mese
*dalle ore 13 alle ore 14* il Centro Nexa organizza anche i "*Nexa Lunch
Seminar*". Una lista di tutti i “Lunch Seminar” passati è disponibile
all'indirizzo: http://nexa.polito.it/lunch-seminars.
See our events calendar <http://nexa.polito.it/events> if you're curious
about future luncheons, discussions, lectures, and conferences not
listed in this email. Our events are free and open to the public, unless
otherwise noted.
Responsabile Comunicazione Centro Nexa su Internet & Società: *Valeria
Bergantino*, tel: +39 011 090 7219, Mob: +39 347 344 3585,
valeria.bergantino(a)polito.it.
Maggiori informazioni sui Mercoledì di Nexa e i Nexa Lunch Seminar, sono
disponibili all'indirizzo: http://nexa.polito.it/events. Weekly Events
Newsletter. Sign up <http://nexa.polito.it/mailing-lists> to receive
this newsletter if this email was forwarded to you. To manage your
subscription preferences, please click here
<https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa>.
Connect & get involved: Jobs, internships, and more
<http://nexa.polito.it/get-involved>.
Nexa Center for Internet and Society Newsletter
Cordiali saluti,
--
Valeria Bergantino
Communication Officer
Nexa Center for Internet & Society
Politecnico di Torino - DAUIN
Via Pier Carlo Boggio, 65/A - 10138 Torino
web: https://nexa.polito.it/ <https://nexa.polito.it/>
mail: valeria.bergantino(a)polito.it <mailto:valeria.bergantino@polito.it>
tel: 3473443585
Oct. 30, 2023
Re: [nexa] opinioni corrette Vs. scienza Vs. fogli di calcolo (was Re: IA, lavoro, immaginari)
by Giuseppe Attardi
> On 27 Oct 2023, at 12:39, Fabio Alemagna <falemagn(a)gmail.com> wrote:
>
>> Non esistono dati oggettivi, ma solo dati scelti, e la scelta influenza le conclusioni che se ne possono trarre.
>
> Se non esistono "dati oggettivi", allora non ha valore di oggettività
> la tua precedente affermazione per la quale "i fatti" "hanno
> dimostrato" che l'austerity non funziona.
>
Il discorso che facevo qui era sul confronto tra previsioni e fatti, entrambi basati sugli stessi dati.
Un confronts alla pari, al netto della loro oggettività o soggettività.
— Beppe
Oct. 30, 2023
This new data poisoning tool lets artists fight back against generative AI | MIT Technology Review
by Alberto Cammozzo
<https://www.technologyreview.com/2023/10/23/1082189/data-poisoning-artists-…>
A new tool lets artists add invisible changes to the pixels in their art before they upload it online so that if it’s scraped into an AI training set, it can cause the resulting model to break in chaotic and unpredictable ways.
The tool, called Nightshade, is intended as a way to fight back against AI companies that use artists’ work to train their models without the creator’s permission. Using it to “poison” this training data could damage future iterations of image-generating AI models, such as DALL-E, Midjourney, and Stable Diffusion, by rendering some of their outputs useless—dogs become cats, cars become cows, and so forth. MIT Technology Review got an exclusive preview of the research, which has been submitted for peer review at computer security conference Usenix.
AI companies such as OpenAI, Meta, Google, and Stability AI are facing a slew of lawsuits from artists who claim that their copyrighted material and personal information was scraped without consent or compensation. Ben Zhao, a professor at the University of Chicago, who led the team that created Nightshade, says the hope is that it will help tip the power balance back from AI companies towards artists, by creating a powerful deterrent against disrespecting artists’ copyright and intellectual property. Meta, Google, Stability AI, and OpenAI did not respond to MIT Technology Review’s request for comment on how they might respond.
Zhao’s team also developed Glaze, a tool that allows artists to “mask” their own personal style to prevent it from being scraped by AI companies. It works in a similar way to Nightshade: by changing the pixels of images in subtle ways that are invisible to the human eye but manipulate machine-learning models to interpret the image as something different from what it actually shows.
The team intends to integrate Nightshade into Glaze, and artists can choose whether they want to use the data-poisoning tool or not. The team is also making Nightshade open source, which would allow others to tinker with it and make their own versions. The more people use it and make their own versions of it, the more powerful the tool becomes, Zhao says. The data sets for large AI models can consist of billions of images, so the more poisoned images can be scraped into the model, the more damage the technique will cause.
A targeted attack
Nightshade exploits a security vulnerability in generative AI models, one arising from the fact that they are trained on vast amounts of data—in this case, images that have been hoovered from the internet. Nightshade messes with those images.
Related Story
This artist is dominating AI-generated art. And he’s not happy about it.
Greg Rutkowski is a more popular prompt than Picasso.
Artists who want to upload their work online but don’t want their images to be scraped by AI companies can upload them to Glaze and choose to mask it with an art style different from theirs. They can then also opt to use Nightshade. Once AI developers scrape the internet to get more data to tweak an existing AI model or build a new one, these poisoned samples make their way into the model’s data set and cause it to malfunction.
Poisoned data samples can manipulate models into learning, for example, that images of hats are cakes, and images of handbags are toasters. The poisoned data is very difficult to remove, as it requires tech companies to painstakingly find and delete each corrupted sample.
The researchers tested the attack on Stable Diffusion’s latest models and on an AI model they trained themselves from scratch. When they fed Stable Diffusion just 50 poisoned images of dogs and then prompted it to create images of dogs itself, the output started looking weird—creatures with too many limbs and cartoonish faces. With 300 poisoned samples, an attacker can manipulate Stable Diffusion to generate images of dogs to look like cats.
Generative AI models are excellent at making connections between words, which helps the poison spread. Nightshade infects not only the word “dog” but all similar concepts, such as “puppy,” “husky,” and “wolf.” The poison attack also works on tangentially related images. For example, if the model scraped a poisoned image for the prompt “fantasy art,” the prompts “dragon” and “a castle in The Lord of the Rings” would similarly be manipulated into something else.
Zhao admits there is a risk that people might abuse the data poisoning technique for malicious uses. However, he says attackers would need thousands of poisoned samples to inflict real damage on larger, more powerful models, as they are trained on billions of data samples.
“We don’t yet know of robust defenses against these attacks. We haven’t yet seen poisoning attacks on modern [machine learning] models in the wild, but it could be just a matter of time,” says Vitaly Shmatikov, a professor at Cornell University who studies AI model security and was not involved in the research. “The time to work on defenses is now,” Shmatikov adds.
Gautam Kamath, an assistant professor at the University of Waterloo who researches data privacy and robustness in AI models and wasn’t involved in the study, says the work is “fantastic.”
The research shows that vulnerabilities “don’t magically go away for these new models, and in fact only become more serious,” Kamath says. “This is especially true as these models become more powerful and people place more trust in them, since the stakes only rise over time.”
A powerful deterrent
Junfeng Yang, a computer science professor at Columbia University, who has studied the security of deep-learning systems and wasn’t involved in the work, says Nightshade could have a big impact if it makes AI companies respect artists’ rights more—for example, by being more willing to pay out royalties.
AI companies that have developed generative text-to-image models, such as Stability AI and OpenAI, have offered to let artists opt out of having their images used to train future versions of the models. But artists say this is not enough. Eva Toorenent, an illustrator and artist who has used Glaze, says opt-out policies require artists to jump through hoops and still leave tech companies with all the power.
Toorenent hopes Nightshade will change the status quo.
“It is going to make [AI companies] think twice, because they have the possibility of destroying their entire model by taking our work without our consent,” she says.
Autumn Beverly, another artist, says tools like Nightshade and Glaze have given her the confidence to post her work online again. She previously removed it from the internet after discovering it had been scraped without her consent into the popular LAION image database.
“I’m just really grateful that we have a tool that can help return the power back to the artists for their own work,” she says.
Oct. 29, 2023
Google paid a whopping $26.3 billion in 2021 to be the default search engine everywhere
by Daniela Tafani
Google paid a whopping $26.3 billion in 2021 to be the default search engine everywhere
/ We knew Google paid handsomely to be the default browser in Safari, Firefox, and elsewhere. Now we know, after years of guessing, exactly what it cost.
By David Pierce, editor-at-large and Vergecast co-host with over a decade of experience covering consumer tech. Previously, at Protocol, The Wall Street Journal, and Wired.
Oct 27, 2023, 6:56 PM GMT+2|
The US v. Google antitrust trial is about many things, but more than anything, it’s about the power of defaults. Even if it’s easy to switch browsers or platforms or search engines, the one that appears when you turn it on matters a lot. Google obviously agrees and has paid a staggering amount to make sure it is the default: testimony in the trial revealed that Google spent a total of $26.3 billion in 2021 to be the default search engine in multiple browsers, phones, and platforms.
That number, the sum total of all of Google’s search distribution deals, came out during the Justice Department’s cross-examination of Google’s search head, Prabhakar Raghavan. It was made public after a debate earlier in the week between the two sides and Judge Amit Mehta over whether the figure should be redacted. Mehta has begun to push for more openness in the trial in general, and this was one of the most significant new pieces of information to be shared openly.
Just to put that $26.3 billion in context: Alphabet, Google’s parent company, announced in its recent earnings report that Google Search ad business brought in about $44 billion over the last three months and about $165 billion in the last year. Its entire ad business — which also includes YouTube ads — made a bit under $90 billion in profit. This is all back-of-the-napkin math, but essentially, Google is giving up about 16 percent of its search revenue and about 29 percent of its profit to those distribution deals.
Google is giving up about 16 percent of its search revenue and about 29 percent of its profit to those distribution deals
Most of that money, of course, goes to Apple. The New York Times recently reported that Google’s deal to be the default search engine in Safari across Google products cost the company about $18 billion in 2021. (Apple’s outsize percentage of the total is why that particular deal has been such a focus of the first weeks of the trial.) In addition, Google pays Mozilla for default placement in Firefox; it pays Samsung for the same on its devices; and it has deals with many device makers, wireless carriers, and other platforms to be the default as well.
Until now, these numbers have been closely held secrets, leaving competitors and analysts to speculate about exactly what it’s worth to Google to be the near-universal default choice. The information also comes as Google is beginning its defense portion of the trial, which started with Raghavan testifying that Google is at perpetual risk of losing its cool — and its users — to platforms like TikTok and ChatGPT. Raghavan said that some users call his search engine “Grandpa Google.” (Raghavan has been saying stuff like this for a while now.) He also said that he sees Yelp and Amazon as competitors and that, in such a hot market, Google has to do everything it can to stay relevant and compete. The Justice Department, on the other hand, is making the case that spending $26.3 billion on securing default status everywhere is actually a way to make sure the market isn’t competitive. After a few more weeks of testimony, Mehta will have to decide who’s right.
https://www.theverge.com/2023/10/27/23934961/google-antitrust-trial-default…
Oct. 28, 2023
Controformazione digitale
by de petra giulio
Molte volte in lista è emersa la necessità, l’urgenza e la difficoltà di
diffondere competenza critica tra i non addetti ai lavori. E’ questa
l’intenzione del libro ‘Dati digitali. Guida per un uso consapevole’, un
tentativo di fare “controformazione digitale” a partire dai dati.
https://themiscrime.com/it/edizioni-themis/digitale-societa/item/588-i-dati…
Chi è interessato può leggere qui l’introduzione
https://centroriformastato.it/controformazione-digitale/
Oct. 28, 2023
Re: [nexa] https://dontspy.eu - severa ma giusta
by Andrea Trentini
ah ok, mi era sfuggito il motivo del set, chiedo scusa
--
Sent from my Android device with K-9 Mail. Please excuse my brevity.
Oct. 28, 2023
Re: [nexa] https://dontspy.eu - severa ma giusta
by Claudio Agosti
Grazie per il feedback Andrea! Al consiglio d'Europa siedono i ministri di
questo governo. Evidentemente non è chiaro, cosí come non lo è il blogpost
che spiega perché ci sono solo 4 ministri (+ la figura responsabile per la
trattativa dell'AI Act, per stato membro, nel nostro caso Butti).
Rendere questi requisiti comprensibili, ma tenere la campagna semplice e
accessibile, é un lavoro di compromesso e di economia dell'informazione
più complesso di molti altri task 😅
Tant'é che dietro le quinte esiste un db di figure politiche registrate
(che dovrebbero divenire 5 per stato membro, solo l'Italia li ha tutti e 5
per ora), e poi un db di foto referenziato solo a queste figure.
Ciao!
Claudio
On Sat, 28 Oct 2023, 09:50 Andrea Trentini, <andrea.trentini(a)unimi.it>
wrote:
> ho guardato solo Italia per ora, mi pare ci siano solo (pochissimi) membri
> di Fratelli d'Italia,
> forse prima di partire col battage avrei popolato un po' più
> "neutralmente" il db
>
> ho provato a caricare immagine (Elly Schlein) ma non ho capito come
> aggiungere un soggetto (mi fa
> scegliere solo tra quelli esistenti)
>
> per il resto ottima iniziativa
>
> --
> Andrea Trentini ⠠⠵
> http://atrent.it
> public key ID: 0xA7A91E3B
> Dip.to di Informatica
> Università degli Studi di Milano
>
>
>
Oct. 28, 2023
Re: [nexa] https://dontspy.eu - severa ma giusta
by Andrea Trentini
ho guardato solo Italia per ora, mi pare ci siano solo (pochissimi) membri di Fratelli d'Italia,
forse prima di partire col battage avrei popolato un po' più "neutralmente" il db
ho provato a caricare immagine (Elly Schlein) ma non ho capito come aggiungere un soggetto (mi fa
scegliere solo tra quelli esistenti)
per il resto ottima iniziativa
--
Andrea Trentini ⠠⠵
http://atrent.it
public key ID: 0xA7A91E3B
Dip.to di Informatica
Università degli Studi di Milano
Oct. 28, 2023