nexa
By thread
nexa@server-nexa.polito.it
By month
Messages by month
- ----- 2026 -----
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2025 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2024 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2023 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2022 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2021 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2020 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2019 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2018 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2017 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2016 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2015 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2014 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2013 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2012 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2011 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2010 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2009 -----
- December
- November
- October
- September
- August
- July
- June
- May
- 40 participants
- 30621 messages
White Paper - Rethinking Privacy in the AI Era: Policy Provocations for a Data-Centric World
by Daniela Tafani
White Paper
February 22, 2024
Rethinking Privacy in the AI Era: Policy Provocations for a Data-Centric World
Jennifer King, Caroline Meinhardt
Executive Summary
➜ In this paper, we present a series of arguments and predictions about how existing and future privacy and data protection regulation will impact the development and deployment of AI systems.
➜ Data is the foundation of all AI systems. Going forward, AI development will continue to increase developers’ hunger for training data, fueling an even greater race for data acquisition than we have already seen in past decades.
➜ Largely unrestrained data collection poses unique risks to privacy that extend beyond the individual level—they aggregate to pose societal-level harms that cannot be addressed through the exercise of individual data rights alone.
➜ While existing and proposed privacy legislation, grounded in the globally accepted Fair Information Practices (FIPs), implicitly regulate AI development, they are not sufficient to address the data acquisition race as well as the resulting individual and systemic privacy harms.
➜ Even legislation that contains explicit provisions on algorithmic decision-making and other forms of AI does not provide the data governance measures needed to meaningfully regulate the data used in AI systems.
➜ We present three suggestions for how to mitigate the risks to data privacy posed by the development and adoption of AI:
1. Denormalize data collection by default by shifting away from opt-out to opt-in data collection. Data collectors must facilitate true data minimization through “privacy by default” strategies and adopt technical standards and infrastructure for meaningful consent mechanisms.
2. Focus on the AI data supply chain to improve privacy and data protection. Ensuring dataset transparency and accountability across the entire life cycle must be a focus of any regulatory system that addresses data privacy.
3. Flip the script on the creation and management of personal data. Policymakers should support the development of new governance mechanisms and technical infrastructure (e.g., data intermediaries and data permissioning infrastructure) to support and automate the exercise of individual data rights and preferences.
https://hai.stanford.edu/white-paper-rethinking-privacy-ai-era-policy-provo…
March 3, 2024
Re: [nexa] Vi racconto com'è “vendere” la propria voce per addestrare un'intelligenza artificiale
by 380°
Buongiorno,
chissà che la giornalista praticante che ha scritto il pezzo per Wired è
stata pagata più di 25€ per il suo lavoro. :-O
don Luca Peyron <dluca.universitari(a)gmail.com> writes:
> Tra costume e futuro
Oh poveri noi: un futuro di bullshit jobs precari? (Stormshit jobs?!?)
--8<---------------cut here---------------start------------->8---
"bullshit jobs": "a form of paid employment that is so completely
pointless, unnecessary, or pernicious that even the employee cannot
justify its existence even though, as part of the conditions of
employment, the employee feels obliged to pretend that this is not the
case."[2] Many people who are working these bullshit or pointless jobs
know that they are working jobs that do not contribute to society in a
thoughtful way.
--8<---------------cut here---------------end--------------->8---
(https://en.wikipedia.org/wiki/Bullshit_Jobs)
...perché con competenze del genere ("creative workers"?), lavorare così
precariamente (freelance? lavoratore indipendente? non precario?) da
essere costretti a benedire «un'entrata in più (a cui, n.d.r.) non si
rinuncia mai» di 25€ lordi per poco meno di due ore di lavoro, con
cessione dei diritti d'autore, firma di un NDA e di una liberatoria per
fornire i dati personali raccolti (quali?) a terze parti, per di più per
"robe" che ben presto verranno buttate nel cestino (tipo gli "assistenti
articifialmente intelligenti")... come può essere definito altrimenti?!?
Molto più utile, molto meno alienante e altrettanto remunerativo andare
mezza giornata a fare le pulizie... anche senza ricorrere ai mirabolanti
servigi [1] di robe come https://prontopro.it/, basta il passaparola!
Credo che "Il Lavoro" /alienato/ - quello sul quale è fondata niente po'
po' di meno che la repubblica italiana - sia ormai giunto a uno stato
*vegetativo* irreversibile: meglio staccare la spina.
Oltretutto noi occidentali siamo abituati a pensare che la
schiavizzazione, ops alienzaione, dei lavoratori precarizzati "delle
piattaforme" sia un fenomeno marginale "da noi" e riguardi solo il
"terzo mondo", mentre comincia a emergere come questo fenomeno stia
diventando strutturale ovunque: altro materiale per Antonio Casilli :-)
> https://www.wired.it/article/intelligenza-artificiale-addestramento-voce-fr…
A futura memoria, ecco cosa ha scritto:
--8<---------------cut here---------------start------------->8---
Chiara Venuto 02.03.2024
1 Vi racconto com'è “vendere” la propria voce per addestrare un'intelligenza artificiale
════════════════════════════════════════════════════════════════════════════════════════
Il racconto in presa diretta di una giornalista che ha prestato la sua
voce per la raccolta di dati utili al machine learning. Tra pronunce
giuste “scorrette” e podcast proclamati senza un pubblico ad ascoltare
Lo scorso 19 gennaio ho passato due ore a *farneticare dei temi più
disparati davanti a un microfono*. L'obiettivo? *Addestrare un'*
[intelligenza artificiale]. Un'esperienza buffa, a tratti
straniante. Ma in pochi ascoltano con tanta attenzione quanto una
macchina.
[intelligenza artificiale]
<https://www.wired.it/article/creare-musica-con-intelligenza-artificiale-sit…>
1.1 Cercasi addestratori
────────────────────────
Andiamo con ordine. Io sono una giornalista praticante, ma spesso mi
capita di lavorare come copywriter e traduttrice freelance. Per questo
motivo sono iscritta ormai da anni a un portale che si chiama
*Upwork*. Qui, i *lavoratori indipendenti* appartenenti a vari settori
(dall'editoria all'informatica) trovano i loro “clienti”, ovvero le
aziende che hanno bisogno di un collaboratore occasionale per un
progetto. Le imprese pubblicano annunci in cui spiegano ciò di cui
hanno bisogno e i freelance inviano una presentazione e una proposta.
Negli ultimi mesi sono fioriti annunci di progetti legati al *settore
dell'intelligenza artificiale*. Ogni giorno vengono pubblicati quasi
duecento post che hanno a che fare con quest'ambito. Diversi
riguardano la figura di *“addestratore” di programmi di apprendimento
automatico*. A un'entrata in più non si rinuncia mai e, considerando
il mio background da laureata di lingue, il mio profilo è stato
apprezzato subito dai reclutatori delle aziende coinvolte nella
realizzazione di software per *[assistenti vocali] basati sull'AI*.
Chi ti istruisce su come svolgere questo lavoro può essere più o meno
esigente. La prima volta, lo scorso gennaio, il mio compito era quello
di *ripetere 570 parole o piccole frasi*. A volte c'era scritto di
scandire piano, altre volte dovevo parlare velocemente. “ /Ciao, nome
dell'assistente vocale!/”, “ /Devo andare al Colosseo/” e così via. Si
trattava, a quanto ho potuto apprendere, di registrazioni finalizzate
ad addestrare un *software per veicoli “intelligenti”*. Per inviare le
mie registrazioni dovevo usare *Voicelinku*, un'app nata con lo scopo
di acquisire questo materiale. La persona che mi ha trovata su Upwork
è di *Chengdu (Cina)* e secondo i dati pubblicati sul portale
lavorerebbe per *un'azienda che si chiama Bkvoice*. [Attraverso una
ricerca su Weibo] (un famoso [social network] cinese) ho scoperto che
potrebbe trattarsi di una realtà che fa capo a Bokai Jiayin Dubbing,
azienda che si occupa di *doppiaggio in diverse lingue*. Il
condizionale è d'obbligo, perché la catena di fornitura in questi
settori è sempre lunga e non è sempre chiaro chi sia l'utilizzatore
finale delle informazioni ricevute.
[assistenti vocali] <https://www.wired.it/topic/assistenti-vocali>
[Attraverso una ricerca su Weibo] <https://weibo.com/u/6395244498>
[social network] <https://www.wired.it/topic/social-network>
1.2 Spaghetti English
─────────────────────
Ciascuno dei miei audio veniva sottoposto a una verifica delle
condizioni di registrazione e della *correttezza della mia
pronuncia*. In caso di *problemi di comprensione (ce ne sono stati
71)* dopo pochi minuti l'applicazione stessa mi diceva da dove
ricominciare. In alcuni casi, registrazioni vocali assolutamente
normali venivano interpretati come rumorosi, in altri in un inglese
sgrammaticato mi veniva spiegato di *pronunciare meglio alcune
parole*. Solitamente si trattava di annotazioni un po' assurde. Il mio
modo di parlare veniva *etichettato come “poco italiano”* perché
leggevo le parole in inglese con una pronuncia – appunto –
inglese. Secondo gli sviluppatori il mio “Hello!” doveva suonare più
come un “ /Ello!/” (rinunciando totalmente all'aspirazione della h) e
il mio “ /Siri/” doveva essere “Si /R/i”, con una r vibrante, ossia
come verrebbe pronunciato da chi ha un *accento italiano molto
marcato*.
Questo perché *le macchine devono imparare a capire anche chi ha una
pronuncia diversa da quella standard*. Un processo un po' snervante,
roba di un'ora e mezza, pochi soldi (25 dollari lordi) ma tutto
sommato facili. Alla fine dei giochi il recruiter mi ha chiesto se
conoscessi qualcun altro disposto a collaborare, così ho indicato un
paio di amici. [Oggi leggo su Reddit che in passato delle persone si
sono chieste se si trattasse di una truffa]. Qualcuno non è stato
pagato, ma nel nostro caso tutto è filato liscio e i nostri dati
personali non sembrano essere stati rubati.
[Oggi leggo su Reddit che in passato delle persone si sono chieste se si
trattasse di una truffa]
<https://www.reddit.com/r/Upwork/comments/18m2hdi/are_voice_recording_jobs_l…>
1.3 A ruota libera
──────────────────
L'ultima volta, questo gennaio, è andata diversamente. *Ho dovuto
registrare otto [podcast] da 15 minuti* fingendo di avere davvero un
pubblico davanti. Dopo ho caricato tutto su un sito. I temi erano i
più disparati: lavoro, salute e viaggi. Su questi argomenti potevo
creare un discorso in autonomia: agli sviluppatori non interessava
molto ciò che dicevo, quanto il /come/. Nessuna regola: potevo
mantenere la mia cantilena messinese e parlare di qualsiasi cosa mi
venisse in mente. Importanti erano solamente gli accorgimenti legati
alla *tecnica di registrazione* e all'assenza di *rumori di
fondo*. Grande spazio, invece, per aneddoti e creatività.
Ho parlato del caso di Giovanna Pedretti, la ristoratrice di Lodi
trovata morta dopo che era stata accusata di avere pubblicato sulla
pagina [Google] della propria attività una recensione falsa dai
contenuti omofobi e abilisti e la relativa risposta di condanna, che
secondo alcuni sarebbe stata parte di un'operazione di marketing. Ho
poi discusso dei *migliori viaggi in Italia, di liste d'attesa
chilometriche per una visita specialistica* e persino dell'ultima
serie Rai che ho visto: /La Storia/, tratta dall'omonimo capolavoro di
Elsa Morante e proposta per il piccolo schermo con la regia di
Francesca Archibugi. Insomma, un esercizio di stile per me che forse
un giorno mi troverò a registrare podcast per lavoro.
[podcast] <https://www.wired.it/topic/podcast>
[Google] <https://www.wired.it/topic/google>
1.4 Il lavoro sporco
────────────────────
Il mio ruolo è l'ultimo di una *catena molto più complessa* fatta di
sviluppatori e ingegneri informatici, linguisti computazionali e così
via. Mentre porto avanti i progetti che mi vengono affidati mi rendo
però conto di *non sapere quasi nulla delle aziende a cui affido
qualcosa di così sensibile come la mia voce* e i miei pensieri. È
facile *delegare il “lavoro sporco” a freelancer* che sanno pochissimo
di cosa stanno facendo, chiedendo poi loro di firmare accordi ambigui
per la *cessione dei [diritti]*. La seconda azienda con la quale ho
collaborato mi ha fatto siglare un contratto di non divulgazione
insieme a quello sulla [*privacy*]. Così per un anno io non potrò dire
chi sono e che cosa fanno, mentre loro potranno *fornire i miei dati e
le mie registrazioni* anche ad aziende terze.
[diritti] <https://www.wired.it/topic/diritti>
[*privacy*] <https://www.wired.it/topic/privacy>
--8<---------------cut here---------------end--------------->8---
Saluti, 380°
[1] https://prontoproit.zendesk.com/hc/it/articles/10080806589468-Quali-sono-i-…
--8<---------------cut here---------------start------------->8---
è previsto un costo di contatto, cioè l'importo che il professionista
paga quando decide di inviare un preventivo in risposta ad una
richiesta.
--8<---------------cut here---------------end--------------->8---
OK, ma quanto costa "un contatto"?
https://prontoproit.zendesk.com/hc/it/articles/10080800395292-Qual-%C3%A8-i…
--8<---------------cut here---------------start------------->8---
Qual è il costo del contatto?
3 mesi fa Aggiornato
Il costo del contatto (o Lead Cost) è l'importo che si paga quando si
decide di inviare un'offerta ad un'opportunità. È possibile visualizzare
ed accettare il costo prima di inviare un preventivo. ProntoPro
garantisce l'invio di tutti i preventivi ai clienti. Tuttavia, è sempre
a discrezione del cliente scegliere o meno il preventivo.
--8<---------------cut here---------------end--------------->8---
OK, ma allora pigliate per il cu...?!?
Per approfindimenti: https://accordoutilizzo.prontopro.it/
In merito al GDPR, i server del sito sono su Amazon:
https://bgp.he.net/dns/prontopro.it :-)
--
380° (Giovanni Biscuolo public alter ego)
«Noi, incompetenti come siamo,
non abbiamo alcun titolo per suggerire alcunché»
Disinformation flourishes because many people care deeply about injustice
but very few check the facts. Ask me about <https://stallmansupport.org>.
March 2, 2024
Vi racconto com'è “vendere” la propria voce per addestrare un'intelligenza artificiale
by don Luca Peyron
Tra costume e futuro
https://www.wired.it/article/intelligenza-artificiale-addestramento-voce-fr…
Buona lettura
dl
_________________________
don Luca Peyron
Pastorale Universitaria - Apostolato Digitale
Arcidiocesi di Torino
www.universitari.to.it
via XX settembre 83, Torino
tel. 011 5156239
March 2, 2024
Court orders maker of Pegasus spyware to hand over code to WhatsApp
by Antonio
NSO Group, the maker of one the world's most sophisticated cyber weapons, has been ordered by a US court to hand its code for Pegasus and other spyware products to WhatsApp as part of the company's ongoing litigation. The decision by Judge Phyllis Hamilton is a major legal victory for WhatsApp, the Meta-owned communication app which has been embroiled in a lawsuit against NSO since 2019, when it alleged that the Israeli company's spyware had been used against 1,400 WhatsApp users over a two-week period.
NSO's Pegasus code, and code for other surveillance products it sells, is seen as a closely and highly sought state secret. NSO is closely regulated by the Israeli ministry of defense, which must review and approve the sale of all licences to foreign governments. In reaching her decision, Hamilton considered a plea by NSO to excuse it of all its discovery obligations in the case due to "various US and Israeli restrictions."
Ultimately, however, she sided with WhatsApp in ordering the company to produce"all relevant spyware" for a period of one year before and after the two weeks in which WhatsApp users were allegedly attacked: from 29 April 2018 to 10 May 2020. NSO must also give WhatsApp information "concerning the full functionality of the relevant spyware." Hamilton did, however, decide in NSO's favor on a different matter: the company will not be forced at this time to divulge the names of its clients or information regarding its server architecture.
https://www.theguardian.com/technology/2024/feb/29/pegasus-surveillance-cod…
Qualche vecchio link per rinfrescare la memoria su NSO:
https://www.amnesty.org/en/documents/doc10/4516/2021/en/
https://globalfreedomofexpression.columbia.edu/wp-content/uploads/2020/07/4…
https://citizenlab.ca/2018/09/hide-and-seek-tracking-nso-groups-pegasus-spy…
A.
March 1, 2024
La raccolta dati su Facebook e Instagram è "massiva" e "illegale"
by Antonio
Otto gruppi europei [1] di consumatori hanno depositato un reclamo per la violazione delle norme sulla privacy da parte di Meta.
Comunicato stampa: https://www.beuc.eu/press-releases/consumer-groups-launch-complaints-agains…
Riassunto del reclamo: https://www.beuc.eu/sites/default/files/publications/BEUC-X-2024-020_How_Me…
A.
[1] dTest (Czech Republic), Forbrugerrådet Tænk (Denmark), EKPIZO (Greece), UFC-Que Choisir (France), Forbrukerrådet (Norway), Spoločnosť ochrany spotrebiteľov (S.O.S.) Poprad (Slovakia), Zveza Potrošnikov Slovenije – ZPS (Slovenia) and CECU (Spain)
March 1, 2024
The “End of Programming” will look a lot like programming
by Lorenzo
Segnalo questa opinione:
https://ben11kehoe.medium.com/the-end-of-programming-will-look-a-lot-like-p…
"In general, a lot of the AI takes I see assert that AI will be able to assume the entire responsibility for a given task for a person, and implicitly assume that the person’s accountability for the task will just sort of…evaporate? Like, if the AI got it wrong, it’s not your fault? But if you have no real way to ensure the task is correctly performed, they’re probably going to find someone else to accomplish that task after it’s failed a few times."
March 1, 2024
Re: [nexa] Microsoft, Mistral AI e l'AI Act
by Giacomo Tesio
Salve 380,
credo di aver abbastanza chiara la differenza fra
- creazione di opere derivate (come sono i "modelli AI" di cui parliamo)
- distribuzione di opere derivate (come sono, transitivamente, gli output di tali software)
- ridistribuzione di opere originali o loro parti
Creative Commons in effetti confonde i due temi, confrontando Google Books
(che distribuisce verbatim, parti di testi coperti da copyright) con i "modelli AI"
che "imparano" dalle opere protette (ROTFL!!!).
Non mi sembra di aver fatto lo stesso errore, ma se qualcosa che ho scritto ti sembra
evidenziare ignoranza giuridica in merito, ti sarei grato se volessi chiarire quale passaggio
esattamente ti ha dato questa impressione e possibilmente qualche riferimento
per colmare la lacuna in questione.
A presto!
Giacomo
Il 29 Febbraio 2024 12:09:10 CET, "380°" <g380(a)biscuolo.net> ha scritto:
> Giacomo Tesio <giacomo(a)tesio.it> writes:
>
> [...]
>
> >> https://creativecommons.org/2023/02/17/fair-use-training-generative-ai/
> >
> > Sono avvocati, mica informatici. :-)
>
> [...]
>
> > L'ignoranza informatica diffusa è talmente profonda che nemmeno rischiano
> > di perdere la faccia!
>
> Anche l'ignoranza del diritto, in questo caso d'autore, è così diffusa e
> profonda, anche tra gli informatici, che c'è estrema confusione tra
> distribuzione di opere dell'ingengo ottenute tramite (ri)elaborazione
> (ovviamente locale) di testi, sui quali negli USA esiste la disciplina
> del "fair use" (in EU siamo più arzigogolati), e _redistribuzione_ dei
> testi orignali (che sai benissimo sono tutelati _esattamente_ NELLA loro
> forma originale).
>
> Se vuoi contestare agli avvocati di Creative Commons ignoranza
> informatica fai pure, ma non mi pare sia la strata migliore per
> contestare i loro giudizi; lasciatelo dire da un informatico che ha
> dovuto molto precocemente fare i conti con la propria ingnoranza nel
> diritto d'autore.
>
> [...]
>
> Saluti, 380°
>
Feb. 29, 2024
Re: [nexa] public-inbox? (was Re: funzione di ricerca nelle mail della lista?)
by Marco A. Calamari
On gio, 2024-02-29 at 11:53 +0100, 380° wrote:
> Buongiorno,
>
> Antonio <antonio(a)piumarossa.it> writes:
>
> > > sì mail-archive ne trova 5392, ma in tutte le mailing-list che ha in
> > > pancia...
> > > M
> >
> > devi mettere [nexa] nel Subject, non nel From.
> > In questo modo te ne trova 437, vedi qui:
> > https://www.mail-archive.com/search?a=1&l=all&haswords=scuola&x=0&y=0&from=…
>
> Questa è la query corretta, considerato che per fortuna il software che
> gestisce la lista aggiunge [nexa] a tutti i subject.
>
> Comunque un giorno riuscirò a convincere il mio alter ego a trovare il
> tempo di metter su una istanza di public-inbox [1] per la lista Nexa e
> per quella AISA, perché sono risorse _preziosissime_ e meritano di
> essere preservate per bene.
Qui posso postare un suggerimento, visto che di archiviazione a lungo termine
mi sono occupato di recente.
https://medium.com/@calamarim/list/archivismi-la-serie-689254e647ad
"Preservare per bene" secondo me dovrebbe far prendere in considerazione
anche un'archiviazione di questo tipo, ad esempio su Internet Archive.
Anche soltanto caricandoci un pdf, i processi automatici lo rendono ricercabile
e selezionabile, anche se è fatto di pagine scansionate.
L'OCR che usano fa letteralmente paura!
Poi, se traguardiamo i secoli, c'è sempre l'Arctic World Archive"!
JM2C. Buona giornata a tutti. Marco
>
> Cercando di rispettare le linee guida di marketing [2] non userò
> superlativi o altre super*, ma public-inbox batte di alcuni ordini di
> grandezza mail-archive, specialmente in potenza di ricerca (e /quindi/ è
> lo strumento perfetto per i ricercatori)
>
> Per capirci, con public-inbox si hanno a disposizione questi criteri di
> ricerca (configurabili per istanza):
>
> --8<---------------cut here---------------start------------->8---
>
> s: match within Subject e.g. s:"a quick brown fox"
> d: match date-time range, git "approxidate" formats supported
> Open-ended ranges such as `d:last.week..' and
> `d:..2.days.ago' are supported
> b: match within message body, including text attachments
> nq: match non-quoted text within message body
> q: match quoted text within message body
> n: match filename of attachment(s)
> t: match within the To header
> c: match within the Cc header
> f: match within the From header
> a: match within the To, Cc, and From headers
> tc: match within the To and Cc headers
> l: match contents of the List-Id header
> bs: match within the Subject and body
>
> [...]
>
> --8<---------------cut here---------------end--------------->8---
> (da https://yhetil.org/guix-devel/_/text/help/)
>
> É /quasi/ come avere una interfaccia web di Notmuch [3] dedicata a una
> mailing list; nulla può battere un database locale Notmuch [4], ma
> public-inbox è un ottimo strumento sussidiario per chi non ha voglia di
> installarselo localmente.
>
> Saluti, 380°
>
>
> [1] https://public-inbox.org/README.html
>
> [2] https://public-inbox.org/marketing.html
>
> [3] https://notmuchmail.org/
>
> [4] la ricerca "query:nexa scuola" su un database di più di 800K email ci
> mette circa 3 secondi (la query:nexa restringe la ricerca ai messaggi
> con header ListId:".*nexa.*"); il solo conteggio dei messaggi
> corrispondenti meno di mezzo secondo, io ne ho 636 nel mio archivio.
>
>
> P.S.: l'email è stata data per morta troppo presto... e un po' troppo
> superficialmente :-D
Feb. 29, 2024
Re: [nexa] Microsoft, Mistral AI e l'AI Act
by 380°
Buongiorno Giuseppe,
(non so esattamente da cosa dipenda - usi la modalità digest? - ma il
tuo client email continua a spezzare i thread e questo rende le
discussioni in lista estremamente più difficoltose)
Giuseppe Attardi <attardi(a)di.unipi.it> writes:
> Secondo Creative Commons, l’utilizzo di pagine web per l’addestramento
> di modelli, costituisce “fair use”:
> https://creativecommons.org/2023/02/17/fair-use-training-generative-ai/
Attenzione che Stefano si riferisce alla _redistribuzione_ del dataset
di training, non del solo LLM
>> From: Stefano Zacchiroli <zack(a)upsilon.cc>
>>> On Tue, Feb 27, 2024 at 09:17:10AM +0100, Giuseppe Attardi wrote:
>>> Facciamolo con fondi pubblici un modello davvero completamente Open,
>>> dai dati di apprendimento, al codice, ai pesi del modello, ai test di
>>> valutazione.
>>
>> Concordo con l'obiettivo e sul fatto che una AI che possa dirsi "open"
>> (o meglio: "libera") dovrebbe esserlo in tutto: dataset di training,
>> codice di training, codice di inferenza, pesi del modello.
>>
>> Ma attenzione al fatto che, a leggi vigenti, tale obiettivo non è
>> raggiungibile per modelli a-la ChatGPT. Il motivo è che includono nei
>> loro dataset di training grandi parti del Web (solitamente ottenute via
>> crawling fatto in casa), che nessuna parte terza può legittimamente
>> redistribuire, dato che solo una piccolissima parte del Web è
>> disponibile sotto licenze libere.
Quindi: siccome nei dataset di training c'è "roba" non libera, quella
"roba" deve essere esclusa da un ipotetico dataset da redistribuire con
una licenza libera.
>> Una AI "libera", secondo i criteri accennati sopra, ha quindi oggi uno
>> svantaggio competitivo enorme rispetto a quelle chiuse --- il che è
>> molto deprimente.
A meno che, invece che distribuire la "roba" proprietaria, non si
forniscano le "ricette" necessarie affinché il codice di training sia in
grado di andare a "leggerselo da solo" il materiale sul web: quello
sarebbe "fair use", che è la stessa identica cosa che fanno quelli che
sviluppano LLM proprietari
Se non c'è la "roba" proprietaria ma solo "la ricetta" non c'è
redistribuzione.
[...]
Saluti, 380°
--
380° (Giovanni Biscuolo public alter ego)
«Noi, incompetenti come siamo,
non abbiamo alcun titolo per suggerire alcunché»
Disinformation flourishes because many people care deeply about injustice
but very few check the facts. Ask me about <https://stallmansupport.org>.
Feb. 29, 2024
LMs and AI make software development harder
by Daniela Tafani
LLMs and AI make software development harder
LLMs and AI make software development harder. Wait, what? Isn’t the whole point of AI to make writing code easier? Well, yes. But writing code is the easy part of software development. The hard part is understanding the problem, designing business logic and debugging tough bugs. And that’s where AI code assistants like copilot or chatgpt make our job harder, as they strip a way the easy parts of our job and only leave us with the hard parts and make it harder for new developers to master the craft of software development.
Coding is the easy part?
Is coding really that easy? No, not exactly easy - mastering a programming language still takes years of practice. But when looking at software development as a whole, writing code is one of the easier part and it is no wonder that chatgpt and copilot can write decent code. First, they have trained on millions of lines of code and second, code is by its nature very easy to understand for a machine as programming languages are very structured languages with limited vocabularies. For a LLM it is probably much easier to learn than natural language.
Programming languages are just very powerful tools that we use to solve problems
In the end, programming languages are just very powerful tools that we use to solve problems. And the hard part is not the learning tool, but understanding the problem and designing a solution for it. This is instantly obvious as most software engineering problems could be solved by a lot of different programming languages, which one to pick is a matter of context or even personal preference.
Another indicator that programming is that easy part is, that the more senior a software developer gets, the less time they usually spend writing code. Instead seniors spend more time understanding the problem, designing the solution, jumping in to debug tough bugs or doing design decisions and of course mentoring junior team members. While this might not be true for every senior developer, when looking at my software development bubble this is a clear trend.
The hard parts of software development
Copilot and other AI assistants are a great help for developers, but they are not flawless. A part of it is natural, as they are trained on existing code without any context and there are also some bad habits from the training data that code assistants might have picked up. And while this might get optimized over time, at the moment it means that developers still have to review the code that is generated by the AI code assistants. And reviews are hard - especially if one cannot query the author of the code for their intent.
And even if the code is good enough, it might still introduce flaws into the control logic of a program, might be missing edge cases or introduce a regression bug when integrated into an existing code base. This means that developers have to debug the code that is generated by the AI code assistants in case of an error. And debugging is hard - especially for these kind of problems where it might be hard to recreate the circumstances that cause the bug in the first place.
As the generated code heavily depends on the context we give the AI code assistants, this means that we have to be very precise in our descriptions which means that we have to understand the problem very well, which requires domain knowledge and context awareness on the side of the developer. Even if we just focus on the technical part, being aware of the surrounding architecture and the existing code is crucial to get good results.
Granted we could ask LLMs like chatgpt for help with integration into the codebase or we could just pass it the whole codebase and let it redesign everything. But apart from requiring lot of input to give enough context debugging in an unfamiliar codebase is even tougher than debugging stuff that you wrote yourself.
And then there is the whole thing about figuring out what exactly our product should do, how it should behave and how it should look like. At the moment this still requires a lot of human smarts and while AI tools might allow us to iterate faster on figuring out what we want to build in the end it is still a human that has to make the decision.
AI generated software development is exhausting
It seems a given that AI assistants will change our job by automating away writing code and even helping us with some design decisions. It is very convenient that we can ask chatgpt questions regarding system design and get reasonable answers. What is still left to us is making the decision on which answer to pick and which prompt to give to the LLM to get the results that we need. And this is very exhausting - decision fatigue is a thing and it is very real. Already before AI code assistants the limiting factor in the speed of delivering software was not the often the decision making process of an organization or a team - not writing the code.
The limiting factor in delivery speed is decision making, not writing code
On top of that is that current company structures will most likely still hold software developers accountable for the code that is running in a product, not the AI code assistants that wrote them in the first place. This will add another layer of stress on it, not just do we need to make more decisions faster, we are also to blame if the AI code assistants make a mistake.
And if there is a mistake then the debugging needs to be done, which often needs a lot of context and background knowledge to be efficient. AI tools are of less help there, because they cannot figure out context changes by themselves. They might help us with the easy part of debugging like running tests with different variations, to narrow down the cause but finding the prompt for an LLM to generate the fix will still be on us.
Are AI tools replacing developers?
AI assistants might lower the initial hurdle to get into software development, but they will not make it easier to become a good, experienced software developer. Most of the senior developers I know gained the background knowledge and context needed to formulate complex solutions from years of slogging through (bad) code and learning from their mistakes. This might be an inefficient way of learning but it is very effective in building up the domain knowledge that is needed for software development. This knowledge is also something that is very hard to teach in a formal way, as books or online tutorials by nature are somewhat generic and and adaption to real life situations still needs hands-on experience.
As I see it, broad usage of AI tools will change the the skill distribution of software developers. We might end up with a lot more junior developers that are able to write code - or at least prompt the LLMs write the code - but lack the deep understanding of software development to be efficient in decision making. On the other hands senior developers that have acquired the context and domain knowledge will be fewer and fewer as the effort to acquire this knowledge will be higher as AI tools will hide away the parts that would enable us to learn unless the generated code is reviewed in-depth, which then raises the question if we gain that much efficiency through the tools at all.
So are AI tools replacing developers? Currently no, they will transform the job of a developer but they will not replace them. The question is how we as an industry will make sure that we retain the knowledge and experience that we have gained over the years. It will also raise the question how we handle the human side of software development, as the job will either become more boring because we just feed machines with prompts, yet more stressful because we have to make more hard decisions faster. Or maybe AI tools are really just a hype and a fad and nothing will change at all.
Written on November 3, 2023
https://dominikberner.ch/ai-tools-make-our-job-harder/
Feb. 29, 2024