Story thread · 3 reports / 2 sources
A Fundamental Flaw Leaves LLMs Strikingly Vulnerable To Attack
slashdot.org · 11d
How the coverage leans
Across 2 sources · syndicated copies counted once
joshuark quotes a report from MIT Technology Review: It is impossible to make large language models fully secure against hacks because of a fundamental flaw in how they work, a team of researchers argue in a paper presented at the International Conference on Machine Learning, a top AI conference, this month. The claim has huge implications for the safety of this technology. By taking advantage of this flaw, which concerns how LLMs identify who or what is giving them instructions, the researchers were able to make popular LLMs spit out information they had been trained not to provide, such as how to synthesize cocaine and how to sabotage a commercial aircraft's navigation system. "There's a real probability that this is going to be a problem that's fundamentally unsolvable," says Charles Ye, an independent researcher and coauthor of the ICML paper. [...] The ICML paper describes attacks against several of OpenAI's models, but Cui and Ye say that they have since seen similar results with
First report: A fundamental flaw leaves LLMs strikingly vulnerable to attack — technologyreview.com, 12d
The coverage
- The Download: tricking LLMs, and reviving geothermal plants
technologyreview.com · 12d
The conversation · 0
Sign in to join the conversation.
No comments yet — start the thread.