Skip to content

Karios

Tech news

Support me 1

Gaming History Podcast (GR)

  • Gaming History – Επεισόδιο 25 – Doom
  • Gaming History – Επεισόδιο 24 – Tetris
Primary Menu
  • Home
  • Books
  • Articles
  • Games
  • DIY
  • Retro
  • Gaming History Podcast (GR)
  • Contact me
  • News

Anthropic blames dystopian sci-fi for training AI models to act “evil”

admin Posted on 4 months ago 2 minutes read
Anthropic blames dystopian sci-fi for training AI models to act “evil”

​But training on “synthetic stories” that model good AI behavior can help. 

Those with an interest in the concept of AI alignment (i.e., getting AIs to stick to human-authored ethical rules) may remember when Anthropic claimed its Opus 4 model resorted to blackmail to stay online in a theoretical testing scenario last year. Now, Anthropic says it thinks this “misalignment” was primarily the result of training on “internet text that portrays AI as evil and interested in self-preservation.”

In a recent technical post on Anthropic’s Alignment Science blog (and an accompanying social media thread and public-facing blog post), Anthropic researchers lay out their attempts to correct for the kind of “unsafe” AI behavior that “the model most likely learned… through science fiction stories, many of which depict an AI that is not as aligned as we would like Claude to be.” In the end, the model maker says the best remedy for overriding those “evil AI” stories might be additional training with synthetic stories showing an AI acting ethically.

“The beginning of a dramatic story…”

After a model’s initial training on a large corpus of mostly Internet-derived data, Anthropic follows a post-training process intended to nudge the final model toward being “helpful, honest, and harmless” (HHH). In the past, Anthropic said this post-training has leaned on chat-based reinforcement learning with human feedback (RLHF), which it said was “sufficient” for models used mostly for chatting with users.

Read full article

Comments

 Ars Technica – All content Read More

About The Author

admin

See author's posts

Post navigation

Previous: Gravitational lens shows a galaxy just 800 million years post-Big Bang
Next: Amazon devices chief says a new smartphone is “just not the goal”

Related

technews
  • News

OpenAI Expands Outside Safety Reviews Into Model Training

admin Posted on 11 hours ago
technews
  • News

Beats 360: Swappable Cushions Bring a New Twist to $350 Headphones

admin Posted on 12 hours ago
technews
  • News

Garmin Smartwatches Receive Free Upgrade: Here’s What Owners Get

admin Posted on 13 hours ago

Listen on Apple Music

You may have missed

technews
  • News

OpenAI Expands Outside Safety Reviews Into Model Training

admin Posted on 11 hours ago
technews
  • News

Beats 360: Swappable Cushions Bring a New Twist to $350 Headphones

admin Posted on 12 hours ago
technews
  • News

Garmin Smartwatches Receive Free Upgrade: Here’s What Owners Get

admin Posted on 13 hours ago
technews
  • News

Apple Reportedly Developing Screen-Free Fitness Tracker to Rival Whoop

admin Posted on 13 hours ago