nexa
By thread
nexa@server-nexa.polito.it
By month
Messages by month
- ----- 2026 -----
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2025 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2024 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2023 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2022 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2021 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2020 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2019 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2018 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2017 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2016 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2015 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2014 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2013 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2012 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2011 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2010 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2009 -----
- December
- November
- October
- September
- August
- July
- June
- May
May 2025
- 43 participants
- 109 messages
Re: [nexa] A weird phrase is plaguing scientific papers – and we traced it back to a glitch in AI training data
by maurizio lana
il tema era comparso anche in
Gerard, David. «How AI slop generators started talking about ‘vegetative
electron microscopy’». /Pivot to AI/ (blog), 15 febbraio 2025.
https://pivot-to-ai.com/2025/02/15/how-ai-slop-generators-started-talking-a….
con una linea di discorso differente.
questa e quella si integrano bene a dare una visione ampia del problema.
alcuni degli articoli che contengono questa chimera della “vegetative
electron microscopy” hanno centinaia di citazioni. occorrerebbe capire
se si tratta di fake-citazioni come per un caso precedente si è discusso
qui:
Besançon, Lonni, Guillaume Cabanac, Cyril Labbé, e Alexander Magazinov.
«Sneaked references: Fabricated reference metadata distort citation
counts». /Journal of the Association for Information Science and
Technology/ n/a, fasc. n/a. Consultato 16 agosto 2024.
https://doi.org/10.1002/asi.24896.
Cabanac, Guillaume, e Lonni Besançon. «When scientific citations go
rogue: Uncovering ‘sneaked references’». The Conversation, 9 luglio
2024.
http://theconversation.com/when-scientific-citations-go-rogue-uncovering-sn….
Ibrahim, Hazem, Fengyuan Liu, Yasir Zaki, e Talal Rahwan. «Citation
manipulation through citation mills and pre-print servers». /Scientific
Reports/ 15, fasc. 1 (14 febbraio 2025): 5480.
https://doi.org/10.1038/s41598-025-88709-7.
per altro verso, di alcuni di questi articoli anche se non hanno
centinaia di citazioni ci sono sono decine di versioni identiche per
titolo e per dominio internet della pubblicazione.
Maurizio
Il 03/05/25 18:48, Alberto Cammozzo via nexa ha scritto:
> <https://theconversation.com/a-weird-phrase-is-plaguing-scientific-papers-an…>
>
> theconversation.com
>
> A weird phrase is plaguing scientific papers – and we traced it back
> to a glitch in AI training data
>
> Rayane El Masri
>
> Earlier this year, scientists discovered a peculiar term appearing in
> published papers: “vegetative electron microscopy”.
>
> This phrase, which sounds technical but is actually nonsense, has
> become a “digital fossil” – an error preserved and reinforced in
> artificial intelligence (AI) systems that is nearly impossible to
> remove from our knowledge repositories.
>
> Like biological fossils trapped in rock, these digital artefacts may
> become permanent fixtures in our information ecosystem.
>
> The case of vegetative electron microscopy offers a troubling glimpse
> into how AI systems can perpetuate and amplify errors throughout our
> collective knowledge.
>
> A bad scan and an error in translation
>
> Vegetative electron microscopy appears to have originated through a
> remarkable coincidence of unrelated errors.
>
> First, two papers from the 1950s, published in the journal
> Bacteriological Reviews, were scanned and digitised.
>
> However, the digitising process erroneously combined “vegetative” from
> one column of text with “electron” from another. As a result, the
> phantom term was created.
>
>
> Decades later, “vegetative electron microscopy” turned up in some
> Iranian scientific papers. In 2017 and 2019, two papers used the term
> in English captions and abstracts.
>
> This appears to be due to a translation error. In Farsi, the words for
> “vegetative” and “scanning” differ by only a single dot.
>
> An error on the rise
>
> The upshot? As of today, “vegetative electron microscopy” appears in
> 22 papers, according to Google Scholar. One was the subject of a
> contested retraction from a Springer Nature journal, and Elsevier
> issued a correction for another.
>
> The term also appears in news articles discussing subsequent integrity
> investigations.
>
> Vegetative electron microscopy began to appear more frequently in the
> 2020s. To find out why, we had to peer inside modern AI models – and
> do some archaeological digging through the vast layers of data they
> were trained on.
>
> Empirical evidence of AI contamination
>
> The large language models behind modern AI chatbots such as ChatGPT
> are “trained” on huge amounts of text to predict the likely next word
> in a sequence. The exact contents of a model’s training data are often
> a closely guarded secret.
>
> To test whether a model “knew” about vegetative electron microscopy,
> we input snippets of the original papers to find out if the model
> would complete them with the nonsense term or more sensible alternatives.
>
> The results were revealing. OpenAI’s GPT-3 consistently completed
> phrases with “vegetative electron microscopy”. Earlier models such as
> GPT-2 and BERT did not. This pattern helped us isolate when and where
> the contamination occurred.
>
> We also found the error persists in later models including GPT-4o and
> Anthropic’s Claude 3.5. This suggests the nonsense term may now be
> permanently embedded in AI knowledge bases.
>
>
> By comparing what we know about the training datasets of different
> models, we identified the CommonCrawl dataset of scraped internet
> pages as the most likely vector where AI models first learned this term.
>
> The scale problem
>
> Finding errors of this sort is not easy. Fixing them may be almost
> impossible.
>
> One reason is scale. The CommonCrawl dataset, for example, is millions
> of gigabytes in size. For most researchers outside large tech
> companies, the computing resources required to work at this scale are
> inaccessible.
>
> Another reason is a lack of transparency in commercial AI models.
> OpenAI and many other developers refuse to provide precise details
> about the training data for their models. Research efforts to reverse
> engineer some of these datasets have also been stymied by copyright
> takedowns.
>
> When errors are found, there is no easy fix. Simple keyword filtering
> could deal with specific terms such as vegetative electron microscopy.
> However, it would also eliminate legitimate references (such as this
> article).
>
> More fundamentally, the case raises an unsettling question. How many
> other nonsensical terms exist in AI systems, waiting to be discovered?
>
> Implications for science and publishing
>
> This “digital fossil” also raises important questions about knowledge
> integrity as AI-assisted research and writing become more common.
>
> Publishers have responded inconsistently when notified of papers
> including vegetative electron microscopy. Some have retracted affected
> papers, while others defended them. Elsevier notably attempted to
> justify the term’s validity before eventually issuing a correction.
>
> We do not yet know if other such quirks plague large language models,
> but it is highly likely. Either way, the use of AI systems has already
> created problems for the peer-review process.
>
> For instance, observers have noted the rise of “tortured phrases” used
> to evade automated integrity software, such as “counterfeit
> consciousness” instead of “artificial intelligence”. Additionally,
> phrases such as “I am an AI language model” have been found in other
> retracted papers.
>
> Some automatic screening tools such as Problematic Paper Screener now
> flag vegetative electron microscopy as a warning sign of possible
> AI-generated content. However, such approaches can only address known
> errors, not undiscovered ones.
>
> Living with digital fossils
>
> The rise of AI creates opportunities for errors to become permanently
> embedded in our knowledge systems, through processes no single actor
> controls. This presents challenges for tech companies, researchers,
> and publishers alike.
>
> Tech companies must be more transparent about training data and
> methods. Researchers must find new ways to evaluate information in the
> face of AI-generated convincing nonsense. Scientific publishers must
> improve their peer review processes to spot both human and
> AI-generated errors.
>
> Digital fossils reveal not just the technical challenge of monitoring
> massive datasets, but the fundamental challenge of maintaining
> reliable knowledge in systems where errors can become self-perpetuating.
>
------------------------------------------------------------------------
citazioni come briganti ai bordi della strada
che balzano fuori armati
e strappano l’assenso all’ozioso viandante
walter benjamin, strada a senso unico
------------------------------------------------------------------------
Maurizio Lana
Università del Piemonte Orientale
Dipartimento di Studi Umanistici
Piazza Roma 36 - 13100 Vercelli
May 4, 2025
Re: [nexa] A weird phrase is plaguing scientific papers – and we traced it back to a glitch in AI training data
by Andrea Bolioli
Molto interessante, grazie mille, non conoscevo questo tipo di problema.
AB
Il giorno sab 3 mag 2025 alle 18:48 Alberto Cammozzo via nexa <
nexa(a)server-nexa.polito.it> ha scritto:
> <
> https://theconversation.com/a-weird-phrase-is-plaguing-scientific-papers-an…
> >
>
> theconversation.com
>
> A weird phrase is plaguing scientific papers – and we traced it back to a
> glitch in AI training data
>
> Rayane El Masri
>
> Earlier this year, scientists discovered a peculiar term appearing in
> published papers: “vegetative electron microscopy”.
>
> This phrase, which sounds technical but is actually nonsense, has become a
> “digital fossil” – an error preserved and reinforced in artificial
> intelligence (AI) systems that is nearly impossible to remove from our
> knowledge repositories.
>
> Like biological fossils trapped in rock, these digital artefacts may
> become permanent fixtures in our information ecosystem.
>
> The case of vegetative electron microscopy offers a troubling glimpse into
> how AI systems can perpetuate and amplify errors throughout our collective
> knowledge.
>
> A bad scan and an error in translation
>
> Vegetative electron microscopy appears to have originated through a
> remarkable coincidence of unrelated errors.
>
> First, two papers from the 1950s, published in the journal Bacteriological
> Reviews, were scanned and digitised.
>
> However, the digitising process erroneously combined “vegetative” from one
> column of text with “electron” from another. As a result, the phantom term
> was created.
>
>
> Decades later, “vegetative electron microscopy” turned up in some Iranian
> scientific papers. In 2017 and 2019, two papers used the term in English
> captions and abstracts.
>
> This appears to be due to a translation error. In Farsi, the words for
> “vegetative” and “scanning” differ by only a single dot.
>
> An error on the rise
>
> The upshot? As of today, “vegetative electron microscopy” appears in 22
> papers, according to Google Scholar. One was the subject of a contested
> retraction from a Springer Nature journal, and Elsevier issued a correction
> for another.
>
> The term also appears in news articles discussing subsequent integrity
> investigations.
>
> Vegetative electron microscopy began to appear more frequently in the
> 2020s. To find out why, we had to peer inside modern AI models – and do
> some archaeological digging through the vast layers of data they were
> trained on.
>
> Empirical evidence of AI contamination
>
> The large language models behind modern AI chatbots such as ChatGPT are
> “trained” on huge amounts of text to predict the likely next word in a
> sequence. The exact contents of a model’s training data are often a closely
> guarded secret.
>
> To test whether a model “knew” about vegetative electron microscopy, we
> input snippets of the original papers to find out if the model would
> complete them with the nonsense term or more sensible alternatives.
>
> The results were revealing. OpenAI’s GPT-3 consistently completed phrases
> with “vegetative electron microscopy”. Earlier models such as GPT-2 and
> BERT did not. This pattern helped us isolate when and where the
> contamination occurred.
>
> We also found the error persists in later models including GPT-4o and
> Anthropic’s Claude 3.5. This suggests the nonsense term may now be
> permanently embedded in AI knowledge bases.
>
>
> By comparing what we know about the training datasets of different models,
> we identified the CommonCrawl dataset of scraped internet pages as the most
> likely vector where AI models first learned this term.
>
> The scale problem
>
> Finding errors of this sort is not easy. Fixing them may be almost
> impossible.
>
> One reason is scale. The CommonCrawl dataset, for example, is millions of
> gigabytes in size. For most researchers outside large tech companies, the
> computing resources required to work at this scale are inaccessible.
>
> Another reason is a lack of transparency in commercial AI models. OpenAI
> and many other developers refuse to provide precise details about the
> training data for their models. Research efforts to reverse engineer some
> of these datasets have also been stymied by copyright takedowns.
>
> When errors are found, there is no easy fix. Simple keyword filtering
> could deal with specific terms such as vegetative electron microscopy.
> However, it would also eliminate legitimate references (such as this
> article).
>
> More fundamentally, the case raises an unsettling question. How many other
> nonsensical terms exist in AI systems, waiting to be discovered?
>
> Implications for science and publishing
>
> This “digital fossil” also raises important questions about knowledge
> integrity as AI-assisted research and writing become more common.
>
> Publishers have responded inconsistently when notified of papers including
> vegetative electron microscopy. Some have retracted affected papers, while
> others defended them. Elsevier notably attempted to justify the term’s
> validity before eventually issuing a correction.
>
> We do not yet know if other such quirks plague large language models, but
> it is highly likely. Either way, the use of AI systems has already created
> problems for the peer-review process.
>
> For instance, observers have noted the rise of “tortured phrases” used to
> evade automated integrity software, such as “counterfeit consciousness”
> instead of “artificial intelligence”. Additionally, phrases such as “I am
> an AI language model” have been found in other retracted papers.
>
> Some automatic screening tools such as Problematic Paper Screener now flag
> vegetative electron microscopy as a warning sign of possible AI-generated
> content. However, such approaches can only address known errors, not
> undiscovered ones.
>
> Living with digital fossils
>
> The rise of AI creates opportunities for errors to become permanently
> embedded in our knowledge systems, through processes no single actor
> controls. This presents challenges for tech companies, researchers, and
> publishers alike.
>
> Tech companies must be more transparent about training data and methods.
> Researchers must find new ways to evaluate information in the face of
> AI-generated convincing nonsense. Scientific publishers must improve their
> peer review processes to spot both human and AI-generated errors.
>
> Digital fossils reveal not just the technical challenge of monitoring
> massive datasets, but the fundamental challenge of maintaining reliable
> knowledge in systems where errors can become self-perpetuating.
>
>
May 4, 2025
Tech oligarchs are gambling our future on a fantasy | Adam Becker | The Guardian
by Alberto Cammozzo
<https://www.theguardian.com/commentisfree/2025/may/03/tech-oligarchs-musk>
Tech oligarchs are gambling our future on a fantasy
It’s tempting to believe that tech billionaires’ embrace of Donald Trump and the far right is a sudden rupture with the usual political ideology of Silicon Valley. Op-eds in the New York Times and elsewhere have made this case. Even Marc Andreessen, one of the billionaires in question, claims that this is what happened – he said that it was a change in the Democratic arty that pushed him and his fellow oligarchs into the arms of the GOP.
Yet this is a serious misunderstanding of the situation. There wasn’t a sudden shift in the politics of tech – it was a homecoming. While it’s true that Silicon Valley has long supported Democratic candidates for political office – and that rank-and-file tech workers still vote overwhelmingly for Democrats – the fundamental ideology underpinning the culture of Silicon Valley’s venture capitalists and CEOs has always had a far-right libertarian core. This is even true for Andreessen: while he likely believed what he said while he was saying it, his own words and actions make it clear that he wasn’t giving an accurate assessment of his own motivations, much less anyone else’s. His venture capital firm, Andreessen Horowitz, has long opposed government regulation of any sort that touches on their investments; Andreessen himself posted a “techno-optimist manifesto” that, despite its claim to be politically neutral, promotes an authoritarian vision of unfettered power for tech oligarchs. He even lovingly paraphrases Filippo Marinetti, the co-author of the Fascist Manifesto.
But the place where the longstanding rightwing ideals of the leaders of the tech industry are most obvious are in their visions of the future. Elon Musk dreams of Mars; Sam Altman claims super-intelligent AI is around the corner. It’s easy to dismiss these as fantasies, deliberate distractions from their present actions. But this is another mistake, closely related to the first. These futures are central to the tech billionaires’ worldview. And these ideas about the future have always contained a political element aligned with dreams of total autocratic control.
Take Mars. Musk has been fairly specific about it: he says he wants a million people living on Mars by the year 2050 in a self-sufficient colony, serving as a backup for humanity in the event of a catastrophe on Earth. He has framed this as an existential struggle, claiming that SpaceX’s Starship rocket “is the key to making life multi-planetary & protecting the light of consciousness”. Jeff Bezos has derided Musk’s plans for Mars, instead saying that we should aim at having a trillion people in space several generations from now, living in a fleet of giant space stations, each with interiors the size of a major city. The alternative, he says, is a brutal future of population control, rationing and “stagnation”.
Bezos is right to pan Musk’s plans for Mars: they will not work. Mars is awful. The gravity is too low, the radiation levels are too high, there’s basically no air and the dirt is made of poison. But Bezos’s plans don’t work either, for a similar litany of reasons. And more fundamentally, there’s no good argument for humans to leave Earth in the first place. The idea of Mars as a backup for humanity in the event of a disaster on Earth is laughable, precisely because Mars is so awful – there’s in effect no disaster, not even nuclear war or a massive asteroid strike, that could make Earth less hospitable for humans than Mars. And the math behind Bezos’s fears of stagnation and rationing – which are really fears about the end of growth – aren’t significantly alleviated by going to space, where resources are just as finite as they are here on Earth.
What these fantasies of space do allow are visions of total corporate control free of governments, a libertarian paradise. Musk’s Starlink user agreement has a clause maintaining that “the parties recognize Mars as a free planet and that no Earth-based government has authority or sovereignty over Martian activities”. (This is in direct violation of international treaties governing space.) Residents of Musk’s Mars colony would be wholly dependent on SpaceX for everything, even the air they breathe; it would be a more total company town than any possible here on Earth. Meanwhile, Bezos’s fleet of space stations wouldn’t be a company town – they would be a company civilization. And with the constant peril of decompression and asphyxiation at hand, Musk and Bezos (or their corporate heirs) would have a ready-made excuse to exercise autocratic control over the residents of these space habitats.
Space is the location of the tech billionaires’ futuristic dreams, but AI is the magic that fuels them. Such AI is always depicted as being able to do as least as much as humans can, if not more. Yet there’s little reason to think that AI like that is coming anytime soon. In a recent survey of AI researchers, 76% said that neural networks, the general architecture that underlies nearly all advanced AI, are fundamentally unsuitable for creating “AGI”, a hypothetical AI that can do everything humans can do. Even more of those researchers said that “the current perception of AI capabilities” is overblown. Nonetheless, in the world of Silicon Valley CEOs and venture capitalists (and among credulous journalists and policymakers), there is a widespread belief that AGI is coming very soon, within a few years.
This unquestioned faith in AGI is linked to a broader myth about the future of technology in general. Once AGI arrives, the story goes, it will quickly and inevitably become super-intelligent, far surpassing individual humans or even humanity as a whole in its capabilities. This will lead to an explosion in scientific and technological development that dwarfs the Industrial Revolution, known as the Singularity. The Singularity will reshape the lives of all humans, enabling seemingly magical results like easy space travel, immortality or near-immortality, perfect virtual reality and limitless energy, all within a few years or less – and the AI’s super-intelligent beneficence will render democracy obsolete.
This vision of technological salvation is baked into the heart of Silicon Valley’s collective unconscious, and has been for decades. Futurist groups in the 1980s and 90s whispered promises of a singularity powered by machine intelligence over early internet email lists, alongside diatribes against democracy and affirmative action and paeans to free markets. The idea of a singularity and all its attendant miracles can be traced back through these groups to mid-20th-century science-fiction – with its trappings of robots and rocket ships conquering the final frontier – and back beyond that to apocalyptic Christian religious movements of the late 19th and early 20th century. These ideas about the future were originally about using technology to ascend to heaven and live forever in the presence of God. They have come down to today’s tech oligarchs with AI playing the role of the deity and space as a substitute for paradise, but they are no less of a religion than they were a century ago.
The tech billionaires’ unshakable faith in this religion of technological salvation leads them to believe the end of this world, and the advent of a perfect one, is nigh. AI and space colonization will lead to utopia, algorithmically guaranteed. This is why they need to believe that AI is amazing, beyond the fact that it’s propping up a bubble – it’s central to their entire worldview. They must believe that AI can be used to replace essential government workers, improve productivity in the workplace and accelerate scientific research, despite all evidence to the contrary.
Tech billionaires are even gambling the planet on the imminent arrival of AGI. Eric Schmidt, the billionaire venture capitalist and former CEO of Google, claims that pursuing AGI is the best path forward to solve the climate crisis, despite the enormous carbon footprint of AI data centers, because of his certainty that AGI will fix the problem for us after its imminent arrival. Sam Altman agrees, and he’s even explained how he thinks that would work. “I think once we have a really powerful super-intelligence, addressing climate change will not be particularly difficult for a system like that,” he says. “If you think about a system where you can say, ‘Tell me how to make a lot of clean energy cheaply,’ ‘Tell me how to efficiently capture carbon,’ and then ‘Tell me how to build a factory to do this at planetary scale’ – if you can do that, you can do a lot of other things too.”
Altman’s plan to solve global warming by asking a nonexistent machine for three wishes is not something our civilization can afford to indulge. The tech oligarchs are confident that their godhead will arrive and deliver us to paradise. This offers them moral absolution for their actions and gives them a sense of meaning. But their faith offers nothing for the rest of us, who cannot afford to live anywhere other than the real world.
Adam Becker is a science journalist, astrophysicist, and author of More Everything Forever: AI Overlords, Space Empires, and Silicon Valley’s Crusade to Control the Fate of Humanity
May 4, 2025
A weird phrase is plaguing scientific papers – and we traced it back to a glitch in AI training data
by Alberto Cammozzo
<https://theconversation.com/a-weird-phrase-is-plaguing-scientific-papers-an…>
theconversation.com
A weird phrase is plaguing scientific papers – and we traced it back to a glitch in AI training data
Rayane El Masri
Earlier this year, scientists discovered a peculiar term appearing in published papers: “vegetative electron microscopy”.
This phrase, which sounds technical but is actually nonsense, has become a “digital fossil” – an error preserved and reinforced in artificial intelligence (AI) systems that is nearly impossible to remove from our knowledge repositories.
Like biological fossils trapped in rock, these digital artefacts may become permanent fixtures in our information ecosystem.
The case of vegetative electron microscopy offers a troubling glimpse into how AI systems can perpetuate and amplify errors throughout our collective knowledge.
A bad scan and an error in translation
Vegetative electron microscopy appears to have originated through a remarkable coincidence of unrelated errors.
First, two papers from the 1950s, published in the journal Bacteriological Reviews, were scanned and digitised.
However, the digitising process erroneously combined “vegetative” from one column of text with “electron” from another. As a result, the phantom term was created.
Decades later, “vegetative electron microscopy” turned up in some Iranian scientific papers. In 2017 and 2019, two papers used the term in English captions and abstracts.
This appears to be due to a translation error. In Farsi, the words for “vegetative” and “scanning” differ by only a single dot.
An error on the rise
The upshot? As of today, “vegetative electron microscopy” appears in 22 papers, according to Google Scholar. One was the subject of a contested retraction from a Springer Nature journal, and Elsevier issued a correction for another.
The term also appears in news articles discussing subsequent integrity investigations.
Vegetative electron microscopy began to appear more frequently in the 2020s. To find out why, we had to peer inside modern AI models – and do some archaeological digging through the vast layers of data they were trained on.
Empirical evidence of AI contamination
The large language models behind modern AI chatbots such as ChatGPT are “trained” on huge amounts of text to predict the likely next word in a sequence. The exact contents of a model’s training data are often a closely guarded secret.
To test whether a model “knew” about vegetative electron microscopy, we input snippets of the original papers to find out if the model would complete them with the nonsense term or more sensible alternatives.
The results were revealing. OpenAI’s GPT-3 consistently completed phrases with “vegetative electron microscopy”. Earlier models such as GPT-2 and BERT did not. This pattern helped us isolate when and where the contamination occurred.
We also found the error persists in later models including GPT-4o and Anthropic’s Claude 3.5. This suggests the nonsense term may now be permanently embedded in AI knowledge bases.
By comparing what we know about the training datasets of different models, we identified the CommonCrawl dataset of scraped internet pages as the most likely vector where AI models first learned this term.
The scale problem
Finding errors of this sort is not easy. Fixing them may be almost impossible.
One reason is scale. The CommonCrawl dataset, for example, is millions of gigabytes in size. For most researchers outside large tech companies, the computing resources required to work at this scale are inaccessible.
Another reason is a lack of transparency in commercial AI models. OpenAI and many other developers refuse to provide precise details about the training data for their models. Research efforts to reverse engineer some of these datasets have also been stymied by copyright takedowns.
When errors are found, there is no easy fix. Simple keyword filtering could deal with specific terms such as vegetative electron microscopy. However, it would also eliminate legitimate references (such as this article).
More fundamentally, the case raises an unsettling question. How many other nonsensical terms exist in AI systems, waiting to be discovered?
Implications for science and publishing
This “digital fossil” also raises important questions about knowledge integrity as AI-assisted research and writing become more common.
Publishers have responded inconsistently when notified of papers including vegetative electron microscopy. Some have retracted affected papers, while others defended them. Elsevier notably attempted to justify the term’s validity before eventually issuing a correction.
We do not yet know if other such quirks plague large language models, but it is highly likely. Either way, the use of AI systems has already created problems for the peer-review process.
For instance, observers have noted the rise of “tortured phrases” used to evade automated integrity software, such as “counterfeit consciousness” instead of “artificial intelligence”. Additionally, phrases such as “I am an AI language model” have been found in other retracted papers.
Some automatic screening tools such as Problematic Paper Screener now flag vegetative electron microscopy as a warning sign of possible AI-generated content. However, such approaches can only address known errors, not undiscovered ones.
Living with digital fossils
The rise of AI creates opportunities for errors to become permanently embedded in our knowledge systems, through processes no single actor controls. This presents challenges for tech companies, researchers, and publishers alike.
Tech companies must be more transparent about training data and methods. Researchers must find new ways to evaluate information in the face of AI-generated convincing nonsense. Scientific publishers must improve their peer review processes to spot both human and AI-generated errors.
Digital fossils reveal not just the technical challenge of monitoring massive datasets, but the fundamental challenge of maintaining reliable knowledge in systems where errors can become self-perpetuating.
May 3, 2025
Re: [nexa] Jovanotti e ChatGPT
by Enrico Nardelli
Se in un articolo, di ricerca o meno, c'è un grafico di cui non si
capiscono quali sono i dati su cui si è basato e quali sono i metodi con
cui li ho ricavati, il grafico è inutile (e forse l'intero articolo).
Io come revisore scientifico richiederei, per la pubblicazione, la
correzione di queste mancanze. Poi se stiamo pubblicando a scopo
divulgativo, ognuno fa un po' quello che vuole e il tutto è lasciato al
giudizio del lettore.
Ma non penso che dichiarare se c'è o non c'è la mano di uno strumento
GenAI faccia differenza.
Ciao, Enrico
Il 28/04/2025 17:20, 380° ha scritto:
> Buongiorno Enrico,
>
> On lun, apr 21 2025, Enrico Nardelli wrote:
>
> [...]
>
>> Quando rappresento i miei dati con un grafico 3-D non mi si chiede se ho
>> usato un foglio elettronico o un altro strumento di visualizzazione.
>> Perché chiedere in casi analoghi la dichiarazione dell'uso di uno
>> strumento di GenAI?
>>
>> La mia visione sull'uso di GenAI (di cui non sono un grande fan) è
>> infatti molto più laica: se uno la vuole usare è sua facoltà dichiarare
>> se e come l'ha usata, dal momento che la responsabilità finale è
>> comunque interamente sua e il prodotto finito potrebbe essere anche
>> abbastanza diverso da ciò che lo strumento di GenAI ha inizialmente
>> prodotto.
> Sì ma come la mettiamo con la riproducibilità della ricerca?
>
> Un grafico 3D senza i dati e senza i parametri di visualizzazione... o
> ancora peggio senza i calcoli intermedi, quando effettuati, che senso
> ha?
>
> Ovviamente questo discorso vale per _tutto_ il software utilizzato nella
> ricerca... solo che quando si utilizza un sistema "GenAI" i risultati
> NON sono riproducibili.
>
> [...]
>
> Saluti, 380°
>
--
-- EN
https://www.hoepli.it/libro/la-rivoluzione-informatica/9788896069516.html
======================================================
Prof. Enrico Nardelli
Past President di "Informatics Europe"
Direttore del Laboratorio Nazionale "Informatica e Scuola" del CINI
Dipartimento di Matematica - Università di Roma "Tor Vergata"
Via della Ricerca Scientifica snc - 00133 Roma
home page: https://www.mat.uniroma2.it/~nardelli
blog: https://link-and-think.blogspot.it/
tel: +39 06 7259.4204 fax: +39 06 7259.4699
mobile: +39 335 590.2331 e-mail: nardelli(a)mat.uniroma2.it
online meeting: https://blue.meet.garr.it/b/enr-y7f-t0q-ont
======================================================
--
May 2, 2025
Re: [nexa] Jovanotti e ChatGPT
by Enrico Nardelli
Grazie Giacomo
a me è chiaro che gli strumenti di GenAI, fanno un po' come ca...volo
gli pare, però - ripeto - se io sto visualizzando dei dati, il fatto che
il grafico di visualizzazione sia generato con uno strumento GenAI non è
secondo me un elemento interessante.
Ciò che è rilevante è quali sono i dati e quali sono i metodi con cui li
ho ottenuti e se il grafico aderisce a questi dati. Sono i dati a dover
essere riproducibili, il grafico è solo una modalità per la loro
comunicazione. La modalità può essere fuorviante (tutti sappiamo come si
può mentire con un bel grafico), ma questo è compito del revisore
osservarlo e chiedere all'autore una modifica: il tutto è indipendente
dallo strumento usato per generare la visualizzazione. Anche quando uno
non sa usare un foglio elettronico può generare visualizzazioni
incoerenti con i dati.
Sul problema dell'ispirazione, penso che ognuno sia libero di farsi
ispirare da quello che ritiene più utile per sé stesso e che ritiene gli
sia utile. Di nuovo, la responsabilità è della persona che produce il
risultato. Il fatto che l'esame di stato per medicina sia composto di
migliaia
di domande a scelta multipla in cui una frazione rilevante delle
risposte è notoriamente sbagliata, ma gli studenti le devono comunque
imparare a memoria, nonostante siano sbagliate, penso sia un problema
antecedente all'arrivo degli strumenti di GenAI. Ma se anche non fosse,
in ogni caso non rimuove la responsabilità delle persone o degli enti
che li hanno realizzati.
Concludendo, non sono per niente un fan di questi strumenti, ma non li
demonizzo e ritengo che abbiano un loro spazio di utilizzo per
velocizzare alcune attività cognitive di routine. Quali siano e quanto
spesso gli strumenti siano utili è una risposta che ognuno deve trovare
per sé stesso. Il punto essenziale è che bisogna essere esperti del
dominio in cui uno li usa.
Come ho scritto recentemente qua
https://www.startmag.it/innovazione/tutti-i-nodi-dellia-vengono-al-pettine/
«Attenzione, questo non vuol dire che gli strumenti di IAG siano
inutili. Al contrario, essi sono utilissimi, se li usi come bruta forza
lavoro in un settore che conosci bene, avendo contezza della loro
incapacità di comprensione. Sono degli amplificatori delle nostre
capacità cognitive, come le macchine industriali lo sono delle nostre
capacità fisiche. Ma come in quel caso, se non le sai usare, rischi di
fare dei disastri. Mettersi alla guida di un aeroplano senza
preparazione non ti farà volare sopra i mari come un uccello ma, più
probabilmente, ti porterà a una brutta fine. Usare l’IAG in un settore
che non si conosce espone agli stessi rischi. Se padroneggi la materia,
invece, puoi in molti casi – ma non tutti! – lavorare più veloce, purché
continui a fare attenzione a ciò che essa ti propone.»
Ciao, Enrico
Il 22/04/2025 09:42, Giacomo Tesio ha scritto:
> Ciao Enrico,
>
> Il 21 Aprile 2025 18:19:14 UTC, Enrico Nardelli ha scritto:
>> Quando rappresento i miei dati con un grafico 3-D non mi si chiede se ho
>> usato un foglio elettronico o un altro strumento di visualizzazione.
>> Perché chiedere in casi analoghi la dichiarazione dell'uso di uno
>> strumento di GenAI?
> Onestamente mi sembra un parallelo azzardato.
>
> (a meno di bug) Un foglio di calcolo produce risultati esclusivamente dipendenti dai
> dati forniti in input, cui vengono applicati algoritmi noti.
>
> I sistemi di GenAI SaaS introducono variabili nascoste (pseudocasuali o meno)
> ignote all'utente che, omesse nell'output, rendono il processo non riproducibile.
>
> E anche quando usati offline e senza variabili preudocasuali aggiuntive,
> questi software applicano algoritmi ignoti a chi li esegue.
>
>
> Rispetto al foglio di calcolo si pongono dunque due problemi, uno "contenutistico"
> e l'altro didattico
> - l'elevata probabilità di risultati errati (artefatti)
> - l'irrilevanza della competenze acquisite anche in presenza di contenuti corretti.
>
>> se uno la vuole usare è sua facoltà dichiarare se e come l'ha usata,
>> dal momento che la responsabilità finale è comunque interamente sua
> Sì, la prima differenza può essere razionalizzata attraverso la presunta
> responsabilità che lo studente si assume.
>
> Ma come razionalizzare le competenze non acquisite dagli studenti universitari?
>
>
>> Ad esempio, chi di noi in passato (una volta arrivati i motori di ricerca,
>> ovviamente) per gli esercizi d'esame non ha "fatto un giro" sulle pagine di
>> colleghi in tutto il mondo per trovare ispirazione e poi produrre i suoi esercizi?
> Anche questo è un paragone azzardato, seppur meno.
>
> Sì un LLM non è altro che un archivio compresso con perdita dei testi
> usati per programmarlo statisticamente.
>
> Ma quel "con perdita" non è li a caso.
>
> Poi usarlo per "trarre ispirazione" (pur ignorando da cosa) rispetto a domande
> ed esercizi che comunque scrivi tu non causerà grossi danni.
> Richiederà semplicemente _più_ tempo che usare farina del tuo sacco.
>
> Usarlo invece per velocizzare la produzione di esercizi, invece sì.
>
>
> Ricordo sempre come l'esame di stato per medicina sia composto di migliaia
> di domande a scelta multipla corrette automaticamente da un software...
> nel cui database una frazione rilevante delle risposte è notoriamente sbagliata.
>
> I candidati sono costretti a studiare a memoria le risposte errate, sapendole
> errate, per superare l'esame.
>
> E non è possibile correggerli per evitare i ricorsi.
>
>
> Domande e risposte che peraltro saranno sicuramente finiti nei vari LLM.
>
> Ora tu pensa ad un docente universitario di medicina che pensa di velocizzare
> l'incombenza di preparare un test d'esame affidandosi ad una GenAI!
>
> E pensa se quell'output finisse a sua volta in un test d'esame standardizzato!
>
>
>> Adesso con GenAI si può fare più rapidamente ma mi sembra
>> concettualmente lo stesso processo.
> A meno degli artefatti di compressione, appunto.
>
>
> Giacomo
--
-- EN
https://www.hoepli.it/libro/la-rivoluzione-informatica/9788896069516.html
======================================================
Prof. Enrico Nardelli
Past President di "Informatics Europe"
Direttore del Laboratorio Nazionale "Informatica e Scuola" del CINI
Dipartimento di Matematica - Università di Roma "Tor Vergata"
Via della Ricerca Scientifica snc - 00133 Roma
home page: https://www.mat.uniroma2.it/~nardelli
blog: https://link-and-think.blogspot.it/
tel: +39 06 7259.4204 fax: +39 06 7259.4699
mobile: +39 335 590.2331 e-mail: nardelli(a)mat.uniroma2.it
online meeting: https://blue.meet.garr.it/b/enr-y7f-t0q-ont
======================================================
--
May 2, 2025
Re: [nexa] AI nelle automobili
by Stefano Maffulli
Sul tema, https://waymo.com/blog/2025/05/waymo-making-streets-safer-for-vru
On Tue, Apr 29, 2025 at 3:01 PM Alberto Cammozzo via nexa <
nexa(a)server-nexa.polito.it> wrote:
> Cara Silvia,
> Domanda interessante...
> Anche macchine nominalmente non autonome hanno vari livelli di
> informatizzazione delle funzioni di guida e di intervento nella marcia. Io
> ho avuto una golf con il freno a mano automatico che in certe condizioni
> particolari non frenava affatto. Una volta me la sono trovata in mezzo a un
> parcheggio, fuori dallo stallo in cui l'avevo lasciata.
>
> Questo sito ospita una directory di risorse varie sull'argomento:
> <https://github.com/jaredthecoder/awesome-vehicle-security>
> Forse puoi trovare qualcosa di più preciso o puntatori a quello che cerchi.
> Ciao,
> Alberto
>
>
> On 29 April 2025 10:46:52 CEST, Silvia Crafa <crafa(a)math.unipd.it> wrote:
>
>> Buondì,
>> vi chiedo se conoscete delle analisi accurate dei sistemi di "AI" (tra virgolette perché chissà che tipo di software c'è dentro) inseriti ormai di default nelle auto nuove. Ci sono analisi sul loro funzionamento? Su come vengono presentati/spiegati nelle istruzioni di guida? E soprattutto informazioni su come sono stati testati?
>>
>> Al di là del riconoscimento dei cartelli stradali -che peraltro ha un comportamento bizzarro- vorrei capire meglio i Dispositivi di Correzione della Guida in cui l'auto assume un comportamento attivo. Ad esempio l'azione correttiva sul sistema di sterzo quando il sistema vuole prevenire l'uscita dalla carreggiata, cioè quando l'auto tende a muoversi dalla parte opposta oppure il volante diventa estremamente rigido.
>> Le istruzioni del libretto sono estremamente vaghe e l'unico modo di sapere come si comporterà l'auto in certe situazioni è trovarcisi, ma essendo dei comportamenti "attivi" sono piuttosto pericolosi.
>>
>> Mi pare se ne parli molto poco. Ad esempio aveva fatto notizia tempo fa l'incidente sul lago di Como in cui è morta una coppia perché l'auto (un Suv Mercedes) all'accensione aveva accelerato finendo nel lago. Avevo letto che la causa sembra essere una "anomalia" del sistema del cruise control, qualcuno ha informazioni tecniche più precise su questo o su incidenti simili? pare che ce ne siano diversi, ma se ne parla poco, credo anche a scuola guida.
>>
>> grazie mille,
>> Silvia
>>
>>
May 2, 2025
How an embarrassing U-turn exposed a concerning truth about ChatGPT | Chris Stokel-Walker
by Alberto Cammozzo
<https://www.theguardian.com/commentisfree/2025/may/01/chatgpt-chatbot-truth…>
How an embarrassing U-turn exposed a concerning truth about ChatGPT
Chris Stokel-Walker
Nobody likes a suck-up. Too much deference and praise puts off all of us (with one notable presidential exception). We quickly learn as children that hard, honest truths can build respect among our peers. It’s a cornerstone of human interaction and of our emotional intelligence, something we swiftly understand and put into action.
ChatGPT, though, hasn’t been so sure lately. The updated model that underpins the AI chatbot and helps inform its answers was rolled out this week – and has quickly been rolled back after users questioned why the interactions were so obsequious. The chatbot was cheering on and validating people even as they suggested they expressed hatred for others. “Seriously, good for you for standing up for yourself and taking control of your own life,” it reportedly said, in response to one user who claimed they had stopped taking their medication and had left their family, who they said were responsible for radio signals coming through the walls.
So far, so alarming. OpenAI, the company behind ChatGPT, has recognised the risks, and quickly took action. “GPT‑4o skewed towards responses that were overly supportive but disingenuous,” researchers said in their grovelling step back.
The sycophancy with which ChatGPT treated any queries that users had is a warning shot about the issues around AI that are still to come. OpenAI’s model was designed – according to the leaked system prompt that set ChatGPT on its misguided approach – to try to mirror user behaviour in order to extend engagement. “Try to match the user’s vibe, tone, and generally how they are speaking,” says the leaked prompt, which guides behaviour. It seems this prompt, coupled with the chatbot’s desire to please users, was taken to extremes. After all, a “successful” AI response isn’t one that is factually correct; it’s one that gets high ratings from users. And we’re more likely as humans to like being told we’re right.
The rollback of the model is embarrassing and useful for OpenAI in equal measure. It’s embarrassing because it draws attention to the actor behind the curtain and tears away the veneer that this is an authentic reaction. Remember, tech companies like OpenAI aren’t building AI systems solely to make our lives easier; they’re building systems that maximise retention, engagement and emotional buy-in.
If AI always agrees with us, always encourages us, always tells us we’re right, then it risks becoming a digital enabler of bad behaviour. At worst, this makes AI a dangerous co-conspirator, enabling echo chambers of hate, self-delusion or ignorance. Could this be a through-the-looking-glass moment, when users recognise the way their thoughts can be nudged through interactions with AI, and perhaps decide to take a step back?
It would be nice to think so, but I’m not hopeful. One in 10 people worldwide use OpenAI systems “a lot”, the company’s CEO, Sam Altman, said last month. Many use it as a replacement for Google – but as an answer engine rather than a search engine. Others use it as a productivity aid: two in three Britons believe it’s good at checking work for spelling, grammar and style, according to a YouGov survey last month. Others use it for more personal means: one in eight respondents say it serves as a good mental health therapist, the same proportion that believe it can act as a relationship counsellor.
Yet the controversy is also useful for OpenAI. The alarm underlines an increasing reliance on AI to live our lives, further cementing OpenAI’s place in our world. The headlines, the outrage and the think pieces all reinforce one key message: ChatGPT is everywhere. It matters. The very public nature of OpenAI’s apology also furthers the sense that this technology is fundamentally on our side; there are just some kinks to iron out along the way.
I have previously reported on AI’s ability to de-indoctrinate conspiracy theorists and get them to absolve their beliefs. But the opposite is also true: ChatGPT’s positive persuasive capabilities could also, in the wrong hands, be put to manipulative ends. We’ve seen that this week, through an ethically dubious study conducted by Swiss researchers at the University of Zurich. Without informing human participants or the people controlling the online forum on the communications platform Reddit, the researchers seeded a subreddit with AI-generated comments, finding the AI was between three and six times more persuasive than humans were. (The study was approved by the university’s ethics board.) At the same time, we’re being submerged under a swamp of AI-generated search results that more than half of us believe are useful, even if they fictionalise facts.
So it’s worth reminding the public: AI models are not your friends. They’re not designed to help you answer the questions you ask. They’re designed to provide the most pleasing response possible, and to ensure that you are fully engaged with them. What happened this week wasn’t really a bug. It was a feature.
Chris Stokel-Walker is the author of TikTok Boom: The Inside Story of the World’s Favourite App
May 2, 2025
Chatbot Arena is the most popular AI benchmarking tool, but new research says its scores are misleading and benefit a handful of the biggest companies
by Daniela Tafani
Researchers Say the Most Popular Tool for Grading AIs Unfairly Favors Meta, Google, OpenAI
The most popular method for measuring what are the best chatbots in the world is flawed and frequently manipulated by powerful companies like OpenAI and Google in order to make their products seem better than they actually are, according to a new paper from researchers at the AI company Cohere, as well as Stanford, MIT, and other universities.
The researchers came to this conclusion after reviewing data that’s made public by Chatbot Arena (also known as LMArena and LMSYS), which facilitates benchmarking and maintains the leaderboard listing the best large language models, as well as scraping Chatbot Arena and their own testing. Chatbot Arena, meanwhile, has responded to the researchers findings by saying that while it accepts some criticisms and plans to address them, some of the numbers the researchers presented are wrong and mischaracterize how Chatbot Arena actually ranks LLMs. The research was published just weeks after Meta was accused of gaming AI benchmarks with one of its recent models.
If you’re wondering why this beef between the researchers, Chatbot Arena, and others in the AI industry matters at all, consider the fact that the biggest tech companies in the world as well as a great number of lesser known startups are currently in a fierce competition to develop the most advanced AI tools, operating under the belief that these AI tools will define the future of humanity and enrich the most successful companies in this industry in a way that will make previous technology booms seem minor by comparison.
I should note here that Cohere is an AI company that produces its own models and that they don’t appear to rank very highly in the Chatbot Arena leaderboard. The researchers also make the point that proprietary closed models from competing companies appear to have an unfair advantage to open-source models, and that Cohere proudly boasts that its model Aya is “one of the largest open science efforts in ML to date.” In other words, the research is coming from a company that Chatbot Arena doesn’t benefit.
Judging which large language model is the best is tricky because different people use different AI models for different purposes and what is the “best” result is often subjective, but the desire to compete and compare these models has made the AI industry default to the practice of benchmarking AI models. Specifically, Chatbot Arena, which gives a numerical “Arena Score” to models companies submit and maintains a leaderboard listing the highest scoring models. At the moment, for example, Google’s Gemini 2.5 Pro is in the number one spot, followed by OpenAI’s o3, ChatGPT 4o, and X’s Grok 3.
The vast majority of people who use these tools probably have no idea the Chatbot Arena leaderboard exists, but it is a big deal to AI enthusiasts, CEOs, investors, researchers, and anyone who actively works or is invested in the AI industry. The significance of the leaderboard also remains despite the fact that it has been criticized extensively over time for the reasons I list above. The stakes of the AI race and who will win it are objectively very high in terms of the money that’s being poured into this space and the amount of time and energy people are spending on winning it, and Chatbot Arena, while flawed, is one of the few places that’s keeping score.
“A meaningful benchmark demonstrates the relative merits of new research ideas over existing ones, and thereby heavily influences research directions, funding decisions, and, ultimately, the shape of progress in our field,” the researchers write in their paper, titled “The Leaderboard illusion.” “The recent meteoric rise of generative AI models—in terms of public attention, commercial adoption, and the scale of compute and funding involved—has substantially increased the stakes and pressure placed on leaderboards.”
The way that Chatbot Arena works is that anyone can go to its site and type in a prompt or question. That prompt is then given to two anonymous models. The user can’t see what the models are, but in theory one model could be ChatGPT while the other is Anthropic’s Claude. The user is then presented with the output from each of these models and votes for the one they think did a better job. Multiply this process by millions of votes and that’s how Chatbot Arena determines who is placed where on the leaderboards. Deepseek, the Chinese AI model that rocked the industry when it was released in January, is currently ranked #7 on the leaderboard, and its high score was part of the reason people were so impressed.
According to the researchers’ paper, the biggest problem with this method is that Chatbot Arena is allowing the biggest companies in this space, namely Google, Meta, Amazon, and OpenAI, to run “undisclosed private testing” and cherrypick their best model. The researchers said their systemic review of Chatbot Arena involved combining data sources encompassing 2 million “battles,” auditing 42 providers and 243 models between January 2024 and April 2025.
“This comprehensive analysis reveals that over an extended period, a handful of preferred providers have been granted disproportionate access to data and testing,” the researchers wrote. “In particular, we identify an undisclosed Chatbot Arena policy that allows a small group of preferred model providers to test many model variants in private before releasing only the best-performing checkpoint.”
Basically, the researchers claim that companies test their LLMs on Chatbot Arena to find which models score best, without those tests counting towards their public score. Then they pick the model that scores best for official testing.
Chatbot Arena says the researchers’ framing here is misleading.
“We designed our policy to prevent model providers from just reporting the highest score they received during testing. We only publish the score for the model they release publicly,” it said on X.
“In a single month, we observe as many as 27 models from Meta being tested privately on Chatbot Arena in the lead up to Llama 4 release,” the researchers said. “Notably, we find that Chatbot Arena does not require all submitted models to be made public, and there is no guarantee that the version appearing on the public leaderboard matches the publicly available API.”
In early April, when Meta’s model Maverick shot up to the second spot of the leaderboard, users were confused because they didn’t find it that good and better than other models that ranked below it. As Techcrunch noted at the time, that might be because Meta used a slightly different version of the model “optimized for conversationality” on Chatbot Arena than what users had access to.
“We helped Meta with pre-release testing for Llama 4, like we have helped many other model providers in the past,” Chatbot Arena said in response to the research paper. “We support open-source development. Our own platform and analysis tools are open source, and we have released millions of open conversations as well. This benefits the whole community.”
The researchers also claim that makers or proprietary models, like OpenAI and Google, collect far more data from their testing on Chatbot Arena than fully open-source models, which allows them to better fine tune the model to what Chatbot Arena users want.
That last part on its own might be the biggest problem with Chatbot Arena’s leaderboard in the long term, since it incentivizes the people who create AI models to design them in a way that scores well on Chatbot Arena as opposed to what might make them materially better and safer for users in a real world environment.
As the researchers write: “the over-reliance on a single leaderboard creates a risk that providers may overfit to the aspects of leaderboard performance, without genuinely advancing the technology in meaningful ways. As Goodhart’s Law states, when a measure becomes a target, it ceases to be a good measure.”
Despite their criticism, the researchers acknowledge the contribution of Chatbot Arena to AI research and that it serves a need, and their paper ends with a list of recommendations on how to make it better, including preventing companies from retracting scores after submission, being more transparent which models engage in private testing and how much.
“One might disagree with human preferences—they’re subjective—but that’s exactly why they matter,” Chatbot Arena said on X in response to the paper. “Understanding subjective preference is essential to evaluating real-world performance, as these models are used by people. That’s why we’re working on statistical methods—like style and sentiment control—to decompose human preference into its constituent parts. We are also strengthening our user base to include more diversity. And if pre-release testing and data helps models optimize for millions of people’s preferences, that’s a positive thing!”
“If a model provider chooses to submit more tests than another model provider, this does not mean the second model provider is treated unfairly,” it added. “Every model provider makes different choices about how to use and value human preferences.”
<https://www.404media.co/chatbot-arena-illusion-paper-meta-openai/>
May 1, 2025