Story thread · 3 reports / 2 sources

A Fundamental Flaw Leaves LLMs Strikingly Vulnerable To Attack

slashdot.org · 11d

How the coverage leans

Across 2 sources · syndicated copies counted once

joshuark quotes a report from MIT Technology Review: It is impossible to make large language models fully secure against hacks because of a fundamental flaw in how they work, a team of researchers argue in a paper presented at the International Conference on Machine Learning, a top AI conference, this month. The claim has huge implications for the safety of this technology. By taking advantage of this flaw, which concerns how LLMs identify who or what is giving them instructions, the researchers were able to make popular LLMs spit out information they had been trained not to provide, such as how to synthesize cocaine and how to sabotage a commercial aircraft's navigation system. "There's a real probability that this is going to be a problem that's fundamentally unsolvable," says Charles Ye, an independent researcher and coauthor of the ICML paper. [...] The ICML paper describes attacks against several of OpenAI's models, but Cui and Ye say that they have since seen similar results with

First report: A fundamental flaw leaves LLMs strikingly vulnerable to attack technologyreview.com, 12d

The coverage

  1. The Download: tricking LLMs, and reviving geothermal plants

    technologyreview.com · 12d

The conversation · 0

Sign in to join the conversation.

No comments yet — start the thread.