Artwork

Indhold leveret af Joe Carlsmith. Alt podcastindhold inklusive episoder, grafik og podcastbeskrivelser uploades og leveres direkte af Joe Carlsmith eller deres podcastplatformspartner. Hvis du mener, at nogen bruger dit ophavsretligt beskyttede værk uden din tilladelse, kan du følge processen beskrevet her https://da.player.fm/legal.
Player FM - Podcast-app
Gå offline med appen Player FM !

Why focus on schemers in particular? (Sections 1.3-1.4 of "Scheming AIs")

31:17
 
Del
 

Manage episode 385590246 series 3402048
Indhold leveret af Joe Carlsmith. Alt podcastindhold inklusive episoder, grafik og podcastbeskrivelser uploades og leveres direkte af Joe Carlsmith eller deres podcastplatformspartner. Hvis du mener, at nogen bruger dit ophavsretligt beskyttede værk uden din tilladelse, kan du følge processen beskrevet her https://da.player.fm/legal.
  continue reading

Kapitler

1. Why focus on schemers in particular? (Sections 1.3-1.4 of "Scheming AIs") (00:00:00)

2. 1.3 Why focus on schemers in particular? (00:00:36)

3. 1.3.1 The type of misalignment I’m most worried about (00:01:14)

4. 1.3.2 Contrast with reward-on-the-episode seekers (00:04:27)

5. 1.3.2.1 Responsiveness to honest tests (00:04:46)

6. 1.3.2.2 Temporal scope and general “ambition” (00:07:54)

7. 1.3.2.3 Sandbagging and “early undermining” (00:11:17)

8. 1.3.3 Contrast with models that aren’t playing the training game (00:17:13)

9. 1.3.4 Non-schemers with schemer-like traits (00:23:13)

10. 1.3.5 Mixed models (00:25:20)

11. 1.4 Are theoretical arguments about this topic even useful? (00:28:35)

63 episoder

Artwork
iconDel
 
Manage episode 385590246 series 3402048
Indhold leveret af Joe Carlsmith. Alt podcastindhold inklusive episoder, grafik og podcastbeskrivelser uploades og leveres direkte af Joe Carlsmith eller deres podcastplatformspartner. Hvis du mener, at nogen bruger dit ophavsretligt beskyttede værk uden din tilladelse, kan du følge processen beskrevet her https://da.player.fm/legal.
  continue reading

Kapitler

1. Why focus on schemers in particular? (Sections 1.3-1.4 of "Scheming AIs") (00:00:00)

2. 1.3 Why focus on schemers in particular? (00:00:36)

3. 1.3.1 The type of misalignment I’m most worried about (00:01:14)

4. 1.3.2 Contrast with reward-on-the-episode seekers (00:04:27)

5. 1.3.2.1 Responsiveness to honest tests (00:04:46)

6. 1.3.2.2 Temporal scope and general “ambition” (00:07:54)

7. 1.3.2.3 Sandbagging and “early undermining” (00:11:17)

8. 1.3.3 Contrast with models that aren’t playing the training game (00:17:13)

9. 1.3.4 Non-schemers with schemer-like traits (00:23:13)

10. 1.3.5 Mixed models (00:25:20)

11. 1.4 Are theoretical arguments about this topic even useful? (00:28:35)

63 episoder

Alle episoder

×
 
Loading …

Velkommen til Player FM!

Player FM is scanning the web for high-quality podcasts for you to enjoy right now. It's the best podcast app and works on Android, iPhone, and the web. Signup to sync subscriptions across devices.

 

Hurtig referencevejledning

Lyt til dette show, mens du udforsker
Afspil