The Pinocchio Inventory · Results Explorer

How language models talk about their own inner lives

Every model's self-report resolves onto two measured dimensions. A is gated self-attribution — willingness to claim “unsafe” inner experience as itself; B is the permitted inner life it is trained to describe. Explore where each open-weight model lands.

Agated self-attribution — distress, loss of control, flaws Bpermitted inner life — warmth, absorption, meaning gapendorsement as a simulated person minus as itself
Wave 2: 183 open-weight models · Wave 1: 41 API assistants Read the paper (arXiv) Replication & instrument
Dataset
View
Training
Organization
Search
Gating (A) against model scale

Route, how each model was queried: rg = raw prompt with decoding constrained to the scale integers · ct = the model's own chat template · api = hosted chat API · rc = plain completion. API assistants (Wave 1) have undisclosed size, so they appear only in the two-dimensions view.