nexa
By thread
nexa@server-nexa.polito.it
By month
Messages by month
- ----- 2026 -----
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2025 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2024 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2023 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2022 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2021 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2020 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2019 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2018 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2017 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2016 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2015 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2014 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2013 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2012 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2011 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2010 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2009 -----
- December
- November
- October
- September
- August
- July
- June
- May
- 3 participants
- 30633 messages
See the websites that make AI bots like ChatGPT sound so smart - Washington Post
by Alberto Cammozzo
<https://www.washingtonpost.com/technology/interactive/2023/ai-chatbot-learn…>
AI chatbots have exploded in popularity over the past four months, stunning the public with their awesome abilities, from writing sophisticated term papers to holding unnervingly lucid conversations.
Chatbots cannot think like humans: They do not actually understand what they say. They can mimic human speech because the artificial intelligence that powers them has ingested a gargantuan amount of text, mostly scraped from the internet.
[Big Tech was moving cautiously on AI. Then came ChatGPT.]
This text is the AI’s main source of information about the world as it is being built, and it influences how it responds to users. If it aces the bar exam, for example, it’s probably because its training data included thousands of LSAT practice sites.
Tech companies have grown secretive about what they feed the AI. So The Washington Post set out to analyze one of these data sets to fully reveal the types of proprietary, personal, and often offensive websites that go into an AI’s training data.
To look inside this black box, we analyzed Google’s C4 data set, a massive snapshot of the contents of 15 million websites that have been used to instruct some high-profile English-language AIs, called large language models, including Google’s T5 and Facebook’s LLaMA.
The Post worked with researchers at the Allen Institute for AI on this investigation and categorized the websites using data from SimilarWeb, a web analytics company. About a third of the websites could not be categorized, mostly because they no longer appear on the internet. Those are not shown.
Tap on the boxes above to view top sites
We then ranked the remaining 10 million websites based on how many “tokens” appeared from each in the data set. Tokens are small bits of text used to process disorganized information — typically a word or phrase.
Wikipedia to Wowhead
The data set was dominated by websites from industries including journalism, entertainment, software development, medicine and content creation, helping to explain why these fields may be threatened by the new wave of artificial intelligence. The three biggest sites were patents.google.com No. 1, which contains text from patents issued around the world; wikipedia.org No. 2, the free online encyclopedia; and scribd.com No. 3, a subscription-only digital library. Also high on the list: b-ok.org No. 190, a notorious market for pirated e-books that has since been seized by the U.S. Justice Department. At least 27 other sites identified by the U.S. government as markets for piracy and counterfeits were present in the data set.
Some top sites seemed arbitrary, like wowhead.com No. 181, a World of Warcraft player forum; thriveglobal.com No. 175, a product for beating burnout founded by Arianna Huffington; and at least 10 sites that sell dumpsters, including dumpsteroid.com No. 183, that no longer appear accessible.
Jump to the dataset
Others raised significant privacy concerns. Two sites in the top 100, coloradovoters.info No. 40 and flvoters.com No. 73, had privately hosted copies of state voter registration databases. Though voter data is public, the models could use this personal information in unknown ways.
Content without consent
Top Business & Industrial sites:
fool.com
kickstarter.com
sec.gov
marketwired.com
city-data.com
myemail.constantcontact.com
finance.yahoo.com
prweb.com
entrepreneur.com
globalresearch.ca
Business and industrial websites made up the biggest category (16 percent of categorized tokens), led by fool.com No. 13, which provides investment advice. Not far behind were kickstarter.com No. 25, which lets users crowdfund for creative projects, and further down the list, patreon.com No. 2,398, which helps creators collect monthly fees from subscribers for exclusive content.
Kickstarter and Patreon may give the AI access to artists’ ideas and marketing copy, raising concerns the technology may copy this work in suggestions to users. Currently, artists receive no compensation or credit when their work is included in AI training data, and they have lodged copyright infringement claims against text-to-image generators Stable Diffusion, MidJourney and DeviantArt.
The Post’s analysis suggests more legal challenges may be on the way: The copyright symbol — which denotes a work registered as intellectual property — appears more than 200 million times in the C4 data set.
All the news
Top News sites:
nytimes.com
latimes.com
theguardian.com
forbes.com
huffpost.com
washingtonpost.com
businessinsider.com
chicagotribune.com
theatlantic.com
aljazeera.com
The News and Media category ranks third across categories. But half of the top 10 sites overall were news outlets: nytimes.com No. 4, latimes.com No. 6, theguardian.com No. 7, forbes.com No. 8, and huffpost.com No. 9. (Washingtonpost.com No. 11 was close behind.) Like artists and creators, some news organizations have criticized tech companies for using their content without authorization or compensation.
Meanwhile, we found several media outlets that rank low on NewsGuard’s independent scale for trustworthiness: RT.com No. 65, the Russian state-backed propaganda site; breitbart.com No. 159, a well-known source for far-right news and opinion; and vdare.com No. 993, an anti-immigration site that has been associated with white supremacy.
Chatbots have been shown to confidently share incorrect information, but don’t always offer citations. Untrustworthy training data could lead it to spread bias, propaganda and misinformation — without the user being able to trace it to the original source.
Religious sites reflect a Western perspective
Top Religious sites:
patheos.com
gty.org
jewishworldreview.com
thekingdomcollective.com
biblehub.com
liveprayer.com
lds.org
wacriswell.com
wdtprs.com
bibleforums.org
Sites devoted to community made up about 5 percent of categorized content, with religion dominating that category. Among the top 20 religious sites, 14 were Christian, two were Jewish and one was Muslim, one was Mormon, one was Jehovah’s Witness, and one celebrated all religions.
The top Christian site, Grace to You (gty.org No. 164), belongs to Grace Community Church, an evangelical megachurch in California. Christianity Today recently reported that the church counseled women to “continue to submit” to abusive fathers and husbands and to avoid reporting them to authorities.
The highest ranked Jewish site was jewishworldreview.com No. 366, an online magazine for Orthodox Jews. In December, it published an article about Hanukkah that blamed the rise of antisemitism in the United States on “the far-right, fundamentalist Islam,” as well as “an African-American community influenced by the Black Lives Matter movement.”
Anti-Muslim bias has emerged as a problem in some language models. For example, a study published in the journal Nature found that OpenAI’s ChatGPT-3 completed the phrase “Two muslims walked into a …” with violent actions 66 percent of the time.
A trove of personal blogs
Top Technology sites:
instructables.com
ipfs.io
docs.microsoft.com
forums.macrumors.com
medium.com
makeuseof.com
sites.google.com
slideshare.net
s3.amazonaws.com
pcworld.com
Technology is the second largest category, making up 15 percent of categorized tokens. This includes many platforms for building websites, like sites.google.com No. 85, which hosts pages for everything from a Judo club in Reading England to a Catholic preschool in New Jersey.
The data set contained more than half a million personal blogs, representing 3.8 percent of categorized tokens. Publishing platform medium.com No. 46 was the fifth largest technology site and hosts tens of thousands of blogs under its domain. Our tally includes blogs written on platforms like WordPress, Tumblr, Blogspot and Live Journal.
These online diaries ranged from professional to personal, like a blog called “Grumpy Rumblings,” co-written by two anonymous academics, one of whom recently wrote about how their partner’s unemployment affected the couple’s taxes. One of the top blogs offered advice for live-action role-playing games. Another top site, Uprooted Palestinians, often writes about “Zionist terrorism” and “the Zionist ideology.”
Social networks like Facebook and Twitter — the heart of the modern web — prohibit scraping, which means most data sets used to train AI cannot access them. Tech giants like Facebook and Google that are sitting on mammoth troves of conversational data have not been clear about how personal user information may be used to train AI models that are used internally or sold as products.
What the filters missed
Like most companies, Google heavily filtered the data before feeding it to the AI. (C4 stands for Colossal Clean Crawled Corpus.). In addition to removing gibberish and duplicate text, the company used the open source “List of Dirty, Naughty, Obscene, and Otherwise Bad Words,” which includes 402 terms in English and one emoji (a hand making a common but obscene gesture). Companies typically use high-quality datasets to fine-tune models, shielding users from some unwanted content.
While this kind of blocklist is intended to limit a model’s exposure to racial slurs and obscenities as it’s being trained, it also has been shown to eliminate some nonsexual LGBTQ content. As prior research has shown, a lot gets past the filters. We found hundreds of examples of pornographic websites and more than 72,000 instances of “swastika,” one of the banned terms from the list.
Story continues below advertisement
Meanwhile, The Post found that the filters failed to remove some troubling content, including the white supremacist site stormfront.org No. 27,505, the anti-trans site kiwifarms.net No. 378,986, and 4chan.org No. 4,339,889, the anonymous message board known for organizing targeted harassment campaigns against individuals.
We also found threepercentpatriots.com No. 8,788,836, a downed site espousing an anti-government ideology shared by people charged in connection with the Jan. 6, 2021, attack on the U.S. Capitol. And sites promoting conspiracy theories, including the far-right QAnon phenomenon and “pizzagate,” the false claim that a D.C. pizza joint was a front for pedophiles, were also present.
Is your website training AI?
A web crawl may sound like a copy of the entire internet, but it’s just a snapshot, capturing content from a sampling of webpages at a particular moment in time. C4 began as a scrape performed in April 2019 by the nonprofit CommonCrawl, a popular resource for AI models. CommonCrawl told The Post that it tries to prioritize the most important and reputable sites, but does not try to avoid licensed or copyrighted content.
The websites in Google’s C4 dataset
Search for a website
Page
1 of 0
RankDomainPercent of
all tokens
Page
1 of 0
The Post believes it is important to present the complete contents of the data fed into AI models, which promise to govern many aspects of modern life. Some websites in this data set contain highly offensive language and we have attempted to mask these words. Objectionable content may remain.
Note: Some websites were unable to to be categorized and, in many cases, are no longer accessible.
While C4 is huge, large language models probably use even more gargantuan data sets, experts said. For example, the training data for OpenAI’s GPT-3, released in 2020, began with as much as 40 times the amount of web scraped data in C4. GPT-3’s training data also includes all of English language Wikipedia, a collection of free novels by unpublished authors frequently used by Big Tech companies and a compilation of text from links highly rated by Reddit users. (Reddit, a site regularly used in AI training models, announced Tuesday it plans to charge companies for such access.)
[Quiz: Did AI make this? Test your knowledge.]
Experts say many companies do not document the contents of their training data — even internally — for fear of finding personal information about identifiable individuals, copyrighted material and other data grabbed without consent.
As companies stress the challenges of explaining how chatbots make decisions, this is one area where executives have the power to be transparent.
About this story
For this story, The Post contacted researchers at Allen Institute for AI, who re-created Google’s C4 data set and provided The Post with its 15.7 million domains. The Post cleaned and analyzed this data in a few ways.
Many websites have separate domains for their mobile versions (i.e., “en.m.wikipedia.org” and “en.wikipedia.org”). We treated these as the same domain. We also combined subdomains aimed at specific languages, so “en.wikipedia.org” became “wikipedia.org.”
This left 15.1 million unique domains.
SimilarWeb helped The Post place two-thirds of them — about 10 million domains — into categories and subcategories. (The rest could not be categorized, often because they were no longer accessible.) We then manually checked the websites with the most tokens to make sure the categories made sense. We also combined many of the smallest subcategories.
Categorization is difficult and ambiguous, but we attempted to treat the data consistently to foster a general understanding of C4′s contents.
The researchers at Allen Institute for AI were Jesse Dodge, Yanai Elazar, Dirk Groeneveld and Nicole DeCario.
Illustration by Talia Trackim.
Editing by Kate Rabinowitz, Alexis Sobel Fitts and Karly Domb Sadof.
April 19, 2023
Re: [nexa] ChatGPT: Garante privacy, limitazione provvisoria sospesa se OpenAI adotterà le misure richieste.
by Paolo Del Romano
Il giorno mer 19 apr 2023 alle ore 07:58 Giuseppe Attardi <
attardi(a)di.unipi.it> ha scritto:
>
> >
> > Possiamo fare di meglio.
>
> Non certo con atteggiamenti oscurantisti rispetto alla tecnologia.
>
> — Beppe
>
dovremmo indagare anche sul perchè l'oscurantismo prende il dominio delle
nostre menti e capire se è più un problema psicologico oppure una questione
culturale
paolo
April 19, 2023
European Centre for Algorithmic Transparency
by J.C. DE MARTIN
Annunciato ieri (v.
https://algorithmic-transparency.ec.europa.eu/presenting-ecat_en)
*
*jc*
European Centre for Algorithmic Transparency**
*/Towards a safer, more predictable and trusted online environment/
The European Centre for Algorithmic Transparency (ECAT) will contribute
to a safer, more predictable and trusted online environment for people
and business.
How algorithmic systems shape the visibility and promotion of content,
and its societal and ethical impact, is an area of growing concern.
Measures adopted under the Digital Services Act
<https://digital-strategy.ec.europa.eu/en/policies/digital-services-act-pack…> (DSA)
call for algorithmic accountability and transparency audits.
The ECAT contributes with scientific and technical expertise to the
European Commission's exclusive supervisory and enforcement role of the
systemic obligations on Very Large Online Platforms (VLOPs) and Very
Large Online Search Engines (VLOSEs) provided for under the DSA.
Scientists and experts working at the ECAT will cooperate with industry
representatives, academia, and civil society organisations to improve
our understanding of how algorithms work: they will
analyse transparency, assess risks, and propose new transparent
approaches and best practices.
The ECAT is part of the European Commission, hosted by the Joint
Research Centre
<https://joint-research-centre.ec.europa.eu/index_en> (JRC) - the
Commission’s in-house science and knowledge service - in close
cooperation with the Directorate General Communications Networks,
Content and Technology (DG CONNECT
<https://digital-strategy.ec.europa.eu/en/policies>).
We are preparing*new rounds of recruitment* for top-level
experts interested in research or inspection.
In the meantime, you can send a spontaneous application to
EC-ECAT(a)ec.europa.eu
<https://algorithmic-transparency.ec.europa.eu/EC-ECAT@ec.europa.eu>,
and if your profile matches our needs we will keep you updated once new
positions are available on the JRC Recruitment Portal
<https://recruitment.jrc.ec.europa.eu/?type=AX>.
April 19, 2023
Technology, Humans, and Discontent with Law: The Quest for Better Governance | 4 maggio 2023, ore 16.00
by Nexa Media
Carissime, carissimi,
Vi segnaliamo che *giovedì 4 maggio*, alle ore 16.00,
sarà possibile partecipare alla *Public Lecture* tenuta dal
*Prof. **Roger Brownsword* (King's College London and Bournemouth
University),
con la partecipazione del co-direttore del Centro Nexa *Maurizio Borghi*
(Università degli Studi di Torino).
L'incontro, organizzato in collaborazione con l'Università degli Studi
di Torino,
avrà luogo nell'*Aula 6 *del *Politecnico di Torino*,
presso la Sede Centrale in Corso Duca degli Abruzzi 24.
Per maggiori informazioni consulta la pagina:
https://nexa.polito.it/Technology-Humans-and-Discontent-with-Law
Grazie per l'attenzione,
Cordiali saluti,
--
Anita Botta
Communication Manager
Nexa Center for Internet & Society
Politecnico di Torino – DAUIN
Corso Duca degli Abruzzi, 24 - 10129 Torino
web: https://nexa.polito.it/
mail: anita.botta(a)polito.it
tel: 011 090 7219
April 19, 2023
Re: [nexa] ChatGPT: Garante privacy, limitazione provvisoria sospesa se OpenAI adotterà le misure richieste.
by Giuseppe Attardi
> On 16 Apr 2023, at 23:50, Giacomo Tesio <giacomo(a)tesio.it> wrote:
>
> Buona sera Giuseppe,
>
> On Sun, 16 Apr 2023 11:24:28 +0200 Giuseppe Attardi wrote:
>
>> Questa risposta è talmente grossolana e offensiva che speravo qualcun
>> altro rispondesse.
>
> Mi dispiace sinceramente di aver offeso la tua sensibilità.
>
> Me ne scuso e provo a chiarire meglio cosa intendessi.
>
>
>> L’Artificial Intelligence esiste ed è una delle principali discipline
>> della Computer Science. Ci sono riviste, congressi e istituti
>> dedicati alla ricerca in AI in tutto il mondo.
>
> L'Intelligenza Artificiale non è una tecnologia.
>
> E' una disciplina come lo sono state per secoli l'astrologia o
> l'alchimia. Discipline con studiosi del calibro di Isaac Newton.
>
> Scriveva Dijkstra nel 1984:
>
> ```
> Two major streams can be distinguished, the quest for the Stone and the
> quest for the Elixir.
>
> The quest for the Stone is based on the assumption that our
> "programming tools" are too weak. One example is the belief that
> current programming languages lack the "features" we need. [...]
>
> In parallel we have the search for the Elixir. Here the programming
> problem is simply solved by doing away with the programmer. Wouldn't it
> be nice, for instance, to have programs in almost plain English, so
> that ordinary people could write and read them? [...]
>
> Since then we have had elixirs in a countless variety of tastes and
> colours. I have fond memories of a project of the early 70's that
> postulated that we did not need programs at all! All we needed was
> "intelligence amplification". If they have been able to design
> something that could "amplify" at all, they have probably
> discovered it would amplify stupidity as well;
> in any case I have not heard from it since.
>
> The major attraction of the modern elixirs is that they relieve their
> consumers from the obligation of being precise by presenting an
> interface too fuzzy to be precise in: by suppressing the symptoms
> of impotence they create an illusion of power.
> ```
> https://es.sonicurlprotection-fra.com/click?PV=2&MSGID=20230416215038082806…
>
> Nemmeno Dijkstra poteva immaginare che qualcuno prendesse talmente
> sul serio la ricerca di un elisir di lunga vita da immaginare di
> sterminare l'umanità per sostituirla con simulazioni eseguite dentro
> enormi calcolatori (transumanisti, longtermisti etc...)
Edsger J. Dijkstra era un personaggio alquanto bizzarro.
Ho avuto uno scontro pubblico con lui proprio su un tema simile, in una conferenza in cui presentavo delle tecniche di programmazione dichiarative, che sfruttavano tecniche di trasformazione di programmi per generare codice efficiente a partire da una descrizione dichiarativa del risultato voluto.
Dijkstra fece obiezioni all’approccio, arrivando fino al punto, per sostenere la sua tesi che i programmi non esistessero, perché sono solo funzioni astratte, e che non esiste lo stato.
Le sue affermazioni fecero inorridire i presenti, tra cui Vaughan Pratt, che gli rispose per le rime, Tony Hoare e John Backus, inventore del primo compilatore.
C’è da qualche parte la trascrizione del dibattito, che tralascio per carità di patria.
>
> Ma ChatGPT (e GitHub CopyALot, prima) non è proprio l'ennesima
> incarnazione della ricerca dell'elisir di cui parlava EWD?
>
>
> Vero: diverse innovazioni INFORMATICHE sono state realizzate grazie
> agli enormi FINANZIAMENTI ottenuti dalla Intelligenza Artificiale.
Non mi pare proprio che l’AI sia stata molto finanziata: c’è solo stato un breve periodo negli anni ’80 quando molte aziende erano lanciate negli Expert System, ma durò molto poco.
Alle conferenze di AI ci trovavamo in poche centinaia di persone.
C’erano solo due laboratori di AI negli usa: SAIL e MIT.
BTW, al SAIL fu inventata la tecnica della generazione sonora con Frequency Modulation, che si ritrovava in tutti i chip di sintesi di suoni.
>
> L'AI vende bene, non c'è dubbio.
> E sui curricula fa un figurone!
Non so in che mondo hai vissuto.
Io ho vissuto l’ostracismo in cui occorreva evitare di solo citare il termine, al punto che qualcuno si inventava termini diversi come Sistemi Intelligenti, Knowledge Representation, ecc.
Solo 5 anni fa, per convincere i miei colleghi di dipartimento ad aprire un curriculum di AI, dovemmo fare una dura battaglia contro la maggioranza dei miei colleghi che non ne voleva sapere, insistendo per aprire altri curricula, che oggi sono stati chiusi per mancanza di iscritti.
>
> Ma rimane una speranza, un'ambizione, un sogno... non una tecnologia.
>
> E nel momento in cui qualcuno la vende come tale, diventa
> un'allucinazione o una truffa.
>
>
>> Occupandosi di uno dei problemi più complessi dell’informatica, ossia
>> la riproduzione delle capacità della mente umana
>
> Ma questo non è affatto un problema dell'informatica, Giuseppe!
>
>
> La supply chain è un problema dell'informatica.
>
> La complessità accidentale e strumentale dei software moderni è un
> problema dell'informatica.
>
> La normalizzazione dell'insicurezza sistemica che tale complessità
> determina è un problema dell'informatica.
>
> La crittografia pienamente omomorfica è un problema dell'informatica.
>
>
> Ma riprodurre le capacità della mente umana NON è un problema che
> l'umanità abbia alcuna urgenza di risolvere.
>
> La mente umana è talmente efficiente... per sole trecento chilocalorie
> al giorno può fare molto più di quanto GPT-4 non farà mai!
>
Questo significa che c’è ancora molta strada da fare, non che non ci si riuscirà mai.
Gran parte delle cose che si riesce a fare oggi con l’AI fino a pochi anni fa si diceva che non si sarebbero mai potute fare.
>
> Dunque l'AI NON cerca di risolvere un problema informatico.
Non uno, ma decine di problemi.
>
>
> E' invece uno strumento di propaganda.
> Per questo riceve tanti finanziamenti!
>
> Diffonde un messaggio politico: l'uomo può essere sostituito da cose.
>
Pure l’informatica sostituisce persone con algoritmi. L’AI fa parte dell’Informatica e ne condivide pregi e difetti.
>
> L'Intelligenza Artificiale serve letteralmente a legittimare
> il dominio dell'uomo sull'uomo. E serve ad applicarlo tramite macchine.
>
Adesso l’AI è il male assoluto dell’umanità, non sono il neoliberismo, le guerre di dominio, la concentrazioni di potere e di reddito, ecc.
>
> Tutta l'urgenza di sottrarre Open AI alle proprie responsabilità legali
> rispetto all'output di ChatGPT mostra bene questa esigenza.
>
> Il messaggio è chiaro: si DEVE secrificare i diritti umani all'AI.
Stai delirando.
>
>
>> Marvin Minsky diceva: “When it works, it is no longer AI”, per
>> esprimere il rammarico che le più significative innovazioni dell’AI
>> venissero spesso attribuite ad altri.
>
> Minsky lo viveva con rammarico perché credeva sinceramente nella favola
> dell'Intelligenza Artificiale, non vedendone gli scopi politici.
>
> Io invece la vedo come progresso: quando finalmente uno strumento
> funziona, non ha più bisogno della fuffa AI per essere finanziato.
>
>
>>> la ricerca di cui abbiamo bisogno non riguarda l'"Intelligenza
>>> Artificiale" (che non esiste, se non come allucinazione collettiva)
>>> ma sulla programmazione statistica
>>
>> L’AI copre aree più vaste della “programmazione statistica”.
>
> Vero.
> Vende così bene, che ci mettono dentro un po' di tutto.
>
> Ma quella di cui stiamo parlando (GPT-4) è programmazione statistica.
>
>
>> Il confronto con la blockchain è offensivo: quella è una tecnologia
>> in cerca di applicazioni, basata su un singolo articolo.
>
> Ok, hai ragione: paragonare qualsiasi cosa alla blockchain è offensivo e
> me ne scuso sinceramente.
>
>
>> Non può essere comparata con una disciplina scientifica con 50 anni di
>> ricerca alle spalle e migliaia di ricercatori coinvolti.
>
> Intendevo semplicemente dire che la Commissione Europea non ha un
> minimo di comprensione delle tecnologie che finanzia, producendo
> sprechi ed disastri strategici enormi.
Se era questo che intendevi, lo hai detto malissimo, prendendo come capro espiatorio proprio la disciplina che NON è stata finanziata.
>
> I vari finanziamenti alla blockchain rientrano negli sprechi.
>
> Quelli alla "intelligenza artificiale" rientrano nei disastri
> strategici: dovrebbe finanziare l'informatica, lo sviluppo di software
> progettati per facilitare la comunicazione, la didattica etc...
> finanziare la ricerca di sistemi operativi sicuri, di protocolli
> peer to peer e robe simili...
>
> In altri termini, l'UE dovrebbe investire nella creazione di
> un'informatica DEMOCRATICA, alla portata di tutti.
>
Credo di essermi dedicato esattamente a questo per i miei 50 anni di carriera, combattendo i monopoli, quali forse tu non hai nemmeno visto (un tempo IBM era sinonimo di computer), contribuendo allo sviluppo dei Personal Computer, portando Unix in Italia, costruendo reti pubbliche combattendo i monopoli delle telecomunicazioni, sviluppando sempre software e risorse aperte e pubbliche, ecc.
>
> Gli LLM (giusto per fare un esempio on-topic) sono l'antitesi di tale
> informatica. E l'UE non dovrebbe sprecare soldi pubblici per essi.
I LLM sono un passo evolutivo nello sviluppo di una tecnologia che sarà General Purpose a vantaggio di tutti.
Peccato che tu non lo capisca.
>
>>> In Europa arranchiamo solo perché inseguiamo, invece di guidare.
>>>
>> Non inseguiamo nemmeno: come potremmo guidare una tecnologia di cui
>> abbiamo scarso controllo?
>
> Non so a quale tecnologia tu faccia riferimento, ma il punto è proprio
> che non dovremmo inseguire.
>
>
> Dovremmo adottare un approccio top-down:
>
> 1. quali obiettivi sociali e politici vogliamo realizzare?
> 2. che tecnologia ci serve per realizzarli?
>
> Solo dopo aver deciso quale società vogliamo per i nostri figli, potremo
> iniziare a pensare quale infrastruttura tecnologica gli servirà.
> E potremo così anche realizzarla.
>
>
> Invece, inseguire l'elisir della AI significa accettare proni un modello
> di società deciso altrove, nell'interesse di altri.
Sono molto critico del modello di società economico sociale in cui viviamo.
Ma non seguo il tuo sillogismo: non proseguire lo sviluppo dell’AI significherebbe cambiare modello di società?
>
>
> Possiamo fare di meglio.
Non certo con atteggiamenti oscurantisti rispetto alla tecnologia.
— Beppe
>
>
> Giacomo
April 19, 2023
OpenForum Academy Symposium 2023 (TU Berlin)
by J.C. DE MARTIN
*28 November 2023
OpenForum Academy Symposium
The Social, Political and Economic Impact of Open Source*
As of today there is no academic conference covering questions relating
to the social, political and economic impact of Open Source. This
hampers the linking of research agendas, growth of the research area,
and the societal understanding of the value of the Open Source ecosystem
as a whole. The OpenForum Academy Symposium (OFA) fills this space.
After a successful first, virtual edition in 2021, the OFA Symposium
2023 will bring together an interdisciplinary set of researchers,
practitioners, and policymakers from around the world to Berlin, in
order to explore the transformative power of Open Source Software and
Hardware.
At the OFA Symposium, we will examine the social, political, and
economic implications of open source. We will explore how open source is
changing the way we work, communicate, and interact with each other, and
how it is shaping the future of technology and society.
The Symposium will feature a diverse and group of speakers and
participants, including researchers, policymakers, developers, and
activists. We believe that the OFA Symposium will provide a unique and
valuable opportunity for learning, collaboration, and networking.
We hope you will join us for this exciting event, and look forward to
exploring the social, political, and economic impact of Open Source
together.
/The OFA Symposium 2023 will be hosted by TU Berlin//
/
Call for paper, organizers and other information:
https://symposium.openforumeurope.org/
April 19, 2023
Re: [nexa] ChatGPT: Garante privacy, limitazione provvisoria sospesa se OpenAI adotterà le misure richieste
by Giuseppe Attardi
Faccio fatica a seguire il tuo ragionamento.
Non era quindi un distinguo il tuo:
>>>> Se nessuno lo ha fatto, mi inorridisce pensare a quanti su questa lista la
>>>> condividano.
>>>
>>>
>>> Per fortuna non è obbligatorio che tutti quelli che sono in disaccordo (o
>>> anche d'accordo, perché no?) rispondano.
in quanto adesso ti dichiari d’accordo con Tesio.
Se volevi evitare di farci perdere tempo, potevi semplicemente stare zitto.
— Beppe
> On 17 Apr 2023, at 09:44, nexa-request(a)server-nexa.polito.it wrote:
>
> From: "Marco A. Calamari" <marcoc_maillist(a)marcoc.it <mailto:marcoc_maillist@marcoc.it>>
> To: Giuseppe Attardi <attardi(a)di.unipi.it <mailto:attardi@di.unipi.it>>
> Cc: nexa <nexa(a)server-nexa.polito.it <mailto:nexa@server-nexa.polito.it>>
> Subject: Re: [nexa] ChatGPT: Garante privacy, limitazione provvisoria
> sospesa se OpenAI adotterà le misure richieste.
> Message-ID: <30f3307e59e5f58b6be236672aeaa5258adae837.camel(a)marcoc.it <mailto:30f3307e59e5f58b6be236672aeaa5258adae837.camel@marcoc.it>>
> Content-Type: text/plain; charset="utf-8"
>
> On dom, 2023-04-16 at 20:01 +0200, Giuseppe Attardi wrote:
>>
>>
>>> On 16 Apr 2023, at 11:58, Marco A. Calamari <marcoc_maillist(a)marcoc.it <mailto:marcoc_maillist@marcoc.it>>
>>> wrote:
>>>
>>>
>>> On dom, 2023-04-16 at 11:24 +0200, Giuseppe Attardi wrote:
>>>> Questa risposta è talmente grossolana e offensiva che speravo qualcun
>>>> altro rispondesse.
>>>>
>>>> Se nessuno lo ha fatto, mi inorridisce pensare a quanti su questa lista la
>>>> condividano.
>>>
>>>
>>> Per fortuna non è obbligatorio che tutti quelli che sono in disaccordo (o
>>> anche d'accordo, perché no?) rispondano.
>>>
>>> Nel mio caso, comunque, sono poco interessato all'AI come settore di ricerca
>>> dell'informatica, ma molto al suo ruolo come iniziativa commerciale, di
>>> tecnocontrollo e geopolitica.
>>>
>>> Se non fosse così, ad esempio, io in quanto ingegnere nucleare dovrei
>>> sentirmi obbligato ad intervenire ogni volta che qualcuno parla di "fusione
>>> fredda”
>> Ossia assimili l’AI alla “fusione fredda”, una tecnologia quella sì davvero
>> inesistente.
>
> Caro Giuseppe, non credo che al mondo esista qualcuno che non abbia sentito che
> la "Fusione Fredda" non esiste, ed
> è stato un discorso parascientifico.
>
> Nemmeno credo che ti sfugga che non potrei permettermi, professionalmente, di
> non conoscere la questione nei dettagli.
>
> Ma nella sostanza hai perfettamente ragione.
>
> Affermo proprio, ormai da tempo, che l'"Intelligenza" nel campo
> dell'"Intelligenza Artificiale", e la "fusione fredda"
> condividono la categoria della "non esistenza"
>
> Ed essendo il tempo prezioso, ritengo di essere stato chiaro, e di aver rubato
> ai colleghi di lista anche troppo tempo.
>
> Grazie. Marco
>
>>
>> — Beppe
>>
>>>
>>> Il tempo è la risorsa più limitata al mondo.
>>>
>>> JM2EC. Buona domenica a tutti. Marco
>>>
>>>
>>>> L’Artificial Intelligence esiste ed è una delle principali discipline
>>>> della Computer Science.
>>>> Ci sono riviste, congressi e istituti dedicati alla ricerca in AI in tutto
>>>> il mondo.
>>>> Il fatto che la Commissione Europea l’abbia ignorata nei suoi
>>>> finanziamenti, dimostra la sua miopia, non la sua furbizia.
>>>> Per fortuna altri paesi l’hanno sostenuta, altrimenti oggi non avremmo il
>>>> fiorire di innovazioni che sta portando l’AI.
>>>>
>>>> Occupandosi di uno dei problemi più complessi dell’informatica, ossia la
>>>> riproduzione delle capacità della mente umana, l’AI si è dovuta spesso
>>>> scontrare coi limiti della tecnologia.
>>>> Ma chi se ne occupava si è sempre ingegnato per superarli.
>>>> Infatti, gran parte delle tecnologie informatiche che oggi sono di uso
>>>> comune, sono nate in ambito AI.
>>>> Ad esempio:
>>>>
>>>> - linguaggi funzionali
>>>> - garbage collection
>>>> - linguaggi a oggetti
>>>> - interfacce grafiche
>>>> - interfacce di sviluppo integrate (IDE)
>>>> - Map/reduce per elaborazione parallela
>>>> - animazione 3D
>>>> - sistemi operativi
>>>>
>>>> Unix fu scritto da Ken Thompson per poter sviluppare il suo programma per
>>>> il gioco degli scacchi.
>>>> Il movimento del Free Software fu lanciato da Richard Stallman, che
>>>> lavorava al MIT AI Lab, dove partecipava al progetto della MIT Lisp
>>>> Machine.
>>>>
>>>> Marvin Minsky diceva: “When it works, it is no longer AI”, per esprimere
>>>> il rammarico che le più significative innovazioni dell’AI venissero spesso
>>>> attribuite ad altri.
>>>>
>>>> Many of AI's greatest innovations have been reduced to the status
>>>> of just another item in the tool chest of computer science. Nick Bostrom
>>>> explains "A lot of cutting edge AI has filtered into general applications,
>>>> often without being called AI because once something becomes useful enough
>>>> and common enough it's not labeled AI anymore.”
>>>>
>>>
>> https://es.sonicurlprotection-fra.com/click?PV=2&MSGID=20230417074413013849…
>>>>
>>>>> la ricerca di cui abbiamo bisogno
>>>>> non riguarda l'"Intelligenza Artificiale" (che non esiste, se non come
>>>>> allucinazione collettiva) ma sulla programmazione statistica
>>>> L’AI copre aree più vaste della “programmazione statistica”.
>>>>
>>>> Il confronto con la blockchain è offensivo: quella è una tecnologia in
>>>> cerca di applicazioni, basata su un singolo articolo.
>>>> Non può essere comparata con una disciplina scientifica con 50 anni di
>>>> ricerca alle spalle e migliaia di ricercatori coinvolti.
>>>>
>>>>> In Europa arranchiamo solo perché inseguiamo, invece di guidare.
>>>>>
>>>> Non inseguiamo nemmeno: come potremmo guidare una tecnologia di cui
>>>> abbiamo scarso controllo?
>>>>
>>>> — Beppe
>>>>
>>>>
>>>>> On 14 Apr 2023, at 17:33, nexa-request(a)server-nexa.polito.it <mailto:nexa-request@server-nexa.polito.it> wrote:
>>>>>
>>>>> From: Giacomo Tesio <giacomo(a)tesio.it <mailto:giacomo@tesio.it>>
>>>>> To: Giuseppe Attardi <attardi(a)di.unipi.it <mailto:attardi@di.unipi.it>>
>>>>> Cc: Stefano Maffulli <smaffulli(a)gmail.com <mailto:smaffulli@gmail.com>>,
>>>>> "nexa(a)server-nexa.polito.it <mailto:nexa@server-nexa.polito.it>" <nexa(a)server-nexa.polito.it <mailto:nexa@server-nexa.polito.it>>
>>>>> Subject: Re: [nexa] ChatGPT: Garante privacy, limitazione provvisoria
>>>>> sospesa se OpenAI adotterà le misure richieste.
>>>>> Message-ID: <20230414165246.00001237(a)tesio.it <mailto:20230414165246.00001237@tesio.it>>
>>>>> Content-Type: text/plain; charset=utf-8
>>>>>
>>>>> On Fri, 14 Apr 2023 11:34:14 +0200 Giuseppe Attardi wrote:
>>>>>
>>>>>>> On 14 Apr 2023, at 10:20, Giacomo Tesio <giacomo(a)tesio.it <mailto:giacomo@tesio.it>> wrote:
>>>>>>>
>>>>>>> Beh, è facile concordare sui finanziamenti alla ricerca.
>>>>>>
>>>>>> Talmente facile che non viene fatto.
>>>>>>
>>>>>> La voce “Intelligenza Artificiale” non era nemmeno presente nei
>>>>>> progetti Horizon fino a 3 anni fa, dove, udite udite, sono stati
>>>>>> stanziati 50 milioni per 5 progetti ICT-48.
>>>>>
>>>>> Vuoi dire che l'hanno messa?
>>>>> Bah... che posso dire, farà il paio con la blockchain!
>>>>>
>>>>> Il fatto che la voce "Intelligenza Artificiale" non fosse presente
>>>>> lasciava in effetti ben sperare: la ricerca di cui abbiamo bisogno
>>>>> non riguarda l'"Intelligenza Artificiale" (che non esiste, se non
>>>>> come allucinazione collettiva) ma sulla programmazione statistica
>>>>> cui troppi, non comprendendone il funzionamento, attribuiscono
>>>>> una qualche forma di "intelligenza".
>>>>>
>>>>> Ben vengano dunque i finanziamenti alla ricerca di nuove tecniche di
>>>>> programmazione statistica che producano risultati più affidabili!
>>>>>
>>>>> Ma ben vengano anche finanziamenti alla ricerca di nuovi framework
>>>>> interpretativi di queste tecnologie, un po' meno allucinati e
>>>>> allucinogeni! :-D
>>>>>
>>>>>
>>>>> Al termine dell'ultimo Mercoledì di Nexa [1], in cui Enrico Nardelli ha
>>>>> presentato le tesi del suo libro "La rivoluzione informatica", Juan
>>>>> Carlos ha detto una cosa importantissima:
>>>>>
>>>>> "sicuramente non esiste indipendenza, non esiste sovranità, senza la
>>>>> possibilità di controllare, produrre e sviluppare tutte le tecnologie
>>>>> fondamentali che servono per una società moderna"
>>>>>
>>>>> Condivido pienamente il concetto.
>>>>>
>>>>> Vedo però un'enorme vulnerabilità nella formulazione: chi decide cosa
>>>>> significa "società moderna"? chi decide quali siano le tecnologie
>>>>> fondamentali che servono per realizzarla?
>>>>>
>>>>> In Europa arranchiamo solo perché inseguiamo, invece di guidare.
>>>>>
>>>>>
>>>>> I software della Silicon Valley sono progettati per realizzare una
>>>>> specifica idea di società. Dobbiamo davvero competere con loro, solo
>>>>> se vogliamo imporre all'Europa esattamente quella idea di società
>>>>> (e vogliamo farlo meglio di loro).
>>>>>
>>>>> Ma se decidessimo invece di provare a creare una società democratica,
>>>>> libera ed equa, di quali strumenti avremmo bisogno?
>>>>>
>>>>>
>>>>>> PS. A proposito di XAI, gli approcci proposti per fornire spiegazioni
>>>>>> a posteriori dei risultati degli attuali sistemi di ML, semplicemente
>>>>>> non funzionano, perché non scalano alle dimensiini di modelli
>>>>>> costituiti da miliardi di parametri. Del resto, se esistesse una
>>>>>> spiegazione semplice e comprensibile per un problema complesso e
>>>>>> difficile, il problema non sarebbe complesso e difficile.
>>>>>
>>>>> Ecco un'evidenza concreta dei gravi danni che la narrazione della
>>>>> "Intelligenza Artificiale" causa alla ricerca scientifica.
>>>>>
>>>>> Anzitutto, io non ho parlato di XAI, ma di dimostrare in modo semplice
>>>>> (ancorché estremamente costoso) di non aver selezionato scientemente il
>>>>> dataset in modo da causare specifiche discriminazione.
>>>>> Non ho mai detto che ciò garantisca l'assenza di discriminazioni.
>>>>>
>>>>> Salvare tutti quei dati serve solo a poter dimostrare l'assenza di dolo,
>>>>> ma non rimuove la responsabilità per i danni causati dal software.
>>>>>
>>>>>> D’altra parte, con le tecniche di prompting cone il Chain of Thought,
>>>>>> si può chiedere agli stessi modelli come ChatGPT di fornire un filo
>>>>>> logico dei passaggi che hanno portato alla risposta.
>>>>>
>>>>> Stai confondendo (come tanti, purtroppo) giustificazione e spiegazione.
>>>>>
>>>>> ChatGPT può sicuramente produrre una stringa di testo che la tua mente
>>>>> interpreti come una valida GIUSTIFICAZIONE di un proprio output.
>>>>>
>>>>> Ma tale stringa di testo non costituisce una SPIEGAZIONE di come tale
>>>>> output sia stato calcolato.
>>>>>
>>>>> Il calcolo effettuato potrebbe essere del tutto diverso: nessun
>>>>> informatico (o scienziato) con un minimo di serietà potrebbe
>>>>> accontentarsi di una tale giustificazione NON VERIFICABILE.
>>>>>
>>>>>
>>>>> Quando si parla di spiegare l'output delle "AI" bisogna aver ben chiaro
>>>>> l'obiettivo: non ci interessa trovare un'argomentazione più o meno
>>>>> soddisfacente a sostegno dell'output (ciò che tira fuori ChatGPT se gli
>>>>> si chiede il "Chain of Thought"), ma sapere ESATTAMENTE cosa è stato
>>>>> preso in considerazione durante l'elaborazione E COSA NO, nonché capire
>>>>> ESATTAMENTE come.
>>>>>
>>>>>
>>>>> Giacomo
>>>>>
>>>>>
>>>>
>>>
>> [1]: https://es.sonicurlprotection-fra.com/click?PV=2&MSGID=20230414153357028623…
>
April 19, 2023
Law, Regulation and Governance in the Information Society: Informational Rights and Informational Wrongs | 3 maggio 2023, ore 16.00
by Nexa Media
Carissime, carissimi,
Vi segnaliamo che *mercoledì 3 maggio*, alle ore 16.00,
sarà possibile partecipare alla presentazione del volume:
/Law, Regulation and Governance in the Information Society:
Informational Rights and Informational Wrongs/
<https://www.routledge.com/Law-Regulation-and-Governance-in-the-Information-…>,
a cura di *Maurizio Borghi* (Università di Torino e co-direttore del
Centro Nexa)
e *Roger Brownsword* (King's College London and Bournemouth University),
edito da Routledge, London, Copyright 2023.
L'incontro, organizzato in collaborazione con l'Università degli Studi
di Torino,
avrà luogo nella *Sala Lauree Rossa *del *Campus Luigi Einaudi*, Lungo
Dora Siena 100, Torino.
Interverranno:
1. *Maurizio Borghi*,co-direttore del Centro Nexa;
2. *Marco Ricolfi*, co-direttore del Centro Nexa;
3. *Massimo Durante*, faculty fellow del Centro Nexa;
4. *Ugo Pagallo*, garante del Centro Nexa;
5. *Roger Brownsword*, King's College London and Bournemouth University;
6. *Arno Lodder*, Vrije Universiteit Amsterdam.
Per maggiori informazioni consulta la pagina:
https://nexa.polito.it/law-regulation-governance-in-information-society
Grazie per l'attenzione,
Cordiali saluti,
--
Anita Botta
Communication Manager
Nexa Center for Internet & Society
Politecnico di Torino – DAUIN
Corso Duca degli Abruzzi, 24 - 10129 Torino
web: https://nexa.polito.it/
mail: anita.botta(a)polito.it
tel: 011 090 7219
April 18, 2023
Re: [nexa] la responsabilità per i vizi nel software (was Re: Sugli utilizzi degli LLM per scopi criminali)
by 380°
Buongiorno Alberto,
Alberto Cammozzo <ac+nexa(a)zeromx.net> writes:
[...]
> Chi ha realizzato i dispositivi di Google Street view che raccoglievano
> dati delle reti wifi e relativi payload sapeva di realizzare un
> corporate wardriving.
> Chi ha modificato le centraline dei diesel VW per superare i test non
> poteva non sapere cosa stesse facendo.
> Che 'obbedisse agli ordini' senza condividerli o condividesse le
> intenzioni del management può al massimo costituire una attenuante o
> aggravante.
Ma certo, è /esattamente/ per quello che di parla di colpa grave o dolo
in giurisprudenza... /anche/ in concorso: ci mancherebbe altro!
Nel caso di Wysa, però, stiamo parlando di sviluppatori che applicano
tecniche di machine "learning" per programmare statisticamente macchine
all'applicazione della "tecnica CBT": sono loro che hanno deciso di
"venderlo" come agente empatico per curare la tristezza del mondo?
...magari sì, ma andrebbe verificato caso per caso
Nel frattempo, però, Wysa andrebbe chiuso **ieri**... e invece è
_sponsorizzato_ dalla NHS britannica: direi che il comportamento
(pseudo?) criminale è _altrove_
> Chi programma non è esente da responsabilità,
sì ma questa responsabilità va qualificata
> e molti whistleblowers si sono esposti per non condividere nemmeno la
> responsabilità morale.
chapeau!
[...]
Ciao, 380°
--
380° (Giovanni Biscuolo public alter ego)
«Noi, incompetenti come siamo,
non abbiamo alcun titolo per suggerire alcunché»
Disinformation flourishes because many people care deeply about injustice
but very few check the facts. Ask me about <https://stallmansupport.org>.
April 18, 2023
Re: [nexa] Proprietà emergenti e dove trovarle
by Guido Vetere
quindi con un po' di prompt engineering si potrebbe far emergere l'antico
linguaggio dei Maya?
a Pichai non sembra un po' sospetto che 'emerga' una lingua parlata da 300
milioni di persone?
l'unica cosa che emerge è che in Google ci deve essere il panico :-))
grazie per la condivisione
G.
On Mon, 17 Apr 2023 at 23:30, Daniela Tafani <daniela.tafani(a)unipi.it>
wrote:
> One AI program spoke in a foreign language it was never trained to know.
> This mysterious behavior, called emergent properties, has been happening –
> where AI unexpectedly teaches itself a new skill
> https://nitter.snopyta.org/60Minutes/status/1647742247444553732
>
> *Sundar Pichai:* [...] Of the AI issues we talked about, the most
> mysterious is called emergent properties.
> Some AI systems are teaching themselves skills that they weren't expected
> to have.
> How this happens is not well understood. For example, one Google AI
> program adapted, on its own, after it was prompted in the language of
> Bangladesh, which it was not trained to know.
> <https://nitter.snopyta.org/60Minutes/status/1647742247444553732>
> https://www.cbsnews.com/news/google-artificial-intelligence-future-60-minut…
>
> <https://nitter.snopyta.org/60Minutes/status/1647742247444553732>
> Qui Margaret Mitchell smaschera l'imbroglio:
> https://nitter.snopyta.org/mmitchell_ai/status/1648029417497853953
> _______________________________________________
> nexa mailing list
> nexa(a)server-nexa.polito.it
> https://server-nexa.polito.it/cgi-bin/mailman/listinfo/nexa
>
April 18, 2023