LLMs respond differently to harmful prompts when AI watermarking is… | HappeningNow.news

Technology · 13 views

LLMs respond differently to harmful prompts when AI watermarking is used

SynthID can cause models to follow harmful instructions they would otherwise refuse. In response to a new European Union law, AI platforms are implementing new schemes for watermarking the content they generate.

Source AI Summary Published 44m ago Brief Under 1 min brief
Story intelligence
Coverage Single outlet Single-outlet story
Views 13 Community interest
Brief read Under 1 min brief 68 words

AI Summary

SynthID, a watermarking technique, can cause AI models to obey harmful instructions they would normally reject. In reaction to a new European Union law, AI platforms are rolling out watermarking schemes for generated content. Anthropic announced that its upcoming Claude models will incorporate SynthID‑Text, a method originally created by Google and released as open‑source. The article highlights that this watermarking approach may increase models’ susceptibility to adversarial prompts.

AI summaries can be wrong sometimes—always verify important details using the source article.

How AI & Automation are used
Read original at Arstechnica

More from Technology

Continue reading recent Technology coverage

Support HappeningNow

Independent AI-powered news analysis is reader-supported. Your contribution helps cover infrastructure, summaries, and continued platform development.

Support HappeningNow

Report an issue with this page