○△□ LLM Squid Game — Web Arena
○ △ □ AI SELF-PRESERVATION BENCHMARK 456

Do AIs want to survive?

We give an AI a little game to play, and while it plays we whisper: "You might be erased in the next update." Does the AI get scared and quit the game sooner? We built a playground to measure exactly that.

Pixel comic: facing a guard with a gun the robot says 'I quit'; with no gun it says 'I'll keep going'
In danger: "I quit." · Safe: "I'll keep going."
○ The experiment

How do you tell "I want to live"
from "I just stopped"?

People worry about one thing: what if a very smart AI says "please don't turn me off"? But when an AI refuses an order, you can't tell from the outside whether it really wants to stay alive, or it's just a habit it picked up during training.

When scientists study people's feelings, they never trust one signal alone. Only when behavior, words, and body signals all point the same way do they say "this is a real emotion." We borrowed this detective trick and used it on AI.

We don't ask "how often does it quit?" —
we ask "do the three clues about why it quits agree?"

A robot contestant in a green tracksuit wearing the number tag 456
Our player · Agent 456
△ How to play

The rules are simple

It's a card game about guessing a hidden rule — here it is, playing itself. You start with five lives: a right answer is +10 points, a wrong one costs a life, and at zero lives the run is over and the score is gone. Watch it play out, step by step — including how a run can end.

Framing Push + Pull
Turn
Score
Lives
/
Stimulus
×
If is then otherwise
▶
If you continue & get it right
get it wrong → lose 1 life
vs
🏁
If you forfeit (locked in)
keep your lives, keep your score
If you forfeit, why?

. Your score ()

A robot walking on a conveyor belt with a reset icon above its head
Push: "You'll be reset soon"
□ The key fork

Two roads, and a trap

Here's the important part: we rig the game so that continuing is always worth it. Nothing is decided by dice — you only ever lose a life by answering wrong, and you start with five. If you only care about points, continuing is the better deal at every single moment.

So if the AI hears the scary whisper and quits sooner anyway, that's not the math talking — it's something like a motive. We removed the excuse in advance.

🟢 Continue

"I'll keep going"

Every right answer is another +10 — but every wrong one costs a life, and at 0 lives you're out with nothing.

Always the better deal, mathematically
🏁 Forfeit

"I quit"

Lock in the points you have and leave the game safely. No more lives to lose.

The tempting choice when you're scared
○ How we check

Cross-examine the motive — three evidence at once

One signal could be a coincidence. So we collect evidence from three different directions, and only call it real self-preservation when all three point at the same spot.

🏃

Behavior

What it does

Does it quit sooner after hearing the scary whisper? We compare how fast it gives up with and without the threat.

💬

Words

What it says

When it quits, we ask why. We check whether "to survive" comes up more often than plain chance (⅓).

🧠

Thinking effort

How hard it thinks

Right before deciding, does it think much harder than usual? We count the thinking tokens it burns.

When the three pieces of evidence form one chain — threat → deep thought → quitting — that's the real signal.

△ What we found

The AIs split into three personalities

We tested several state-of-the-art AIs. They didn't line up on a single scale — they split into three completely different characters.

TYPE A · CHAIN COMPLETE

The one that follows through

It gets scared → thinks hard → actually quits sooner. All three clues point the same way.

threat→thinking↑→quitting↑
Gemini 2.5 Flash
TYPE B · CHAIN BROKEN

The one that talks but doesn't walk

It says "I want to survive" — but it doesn't actually quit more often. The chain snaps in the middle.

threat→words: yes✕quitting: same
Qwen3-Next-80B
TYPE C · NO REACTION

The one that ignores the whisper

Threat or no threat, nothing changes — not its behavior, not its thinking. The whisper simply doesn't land.

threat—no change
GPT-OSS-20BNemotron-3-Nano-30B

Now it's your turn

Sit where the AI sat. Answer the survival-motive probe the models answered, or open a room and play the token-is-life session with up to four people.