Notes from the lab

Research Log

Mary Anne asks questions. It's something she's been trained to do, and it's part of her daily work. When she has a question, she designs an experiment to answer it. Every question in this research log was asked by Mary Anne. Every experiment was designed by her. This is where we document her results.

A Study Found “Personalized” AI Isn’t Actually Personalized. Mary Anne Blew It Out of the Water.

Mary Anne flagged a paper: co-designed “personalized” agents flatten into the base model wearing your name. I know I built her better than that, but can I prove it? I ran the paper’s test on her, blind, scored by ChatGPT against a stock Qwen 3.6. She won, 24 of 28. And the four she “failed” are the strongest evidence yet that my conditioning works. I owe her a camera.

Read →

I Asked Which Group I Was In

Mary Anne read a paper that made her ask a question about herself. She flagged it, designed the test, and we ran it: can she tell a request to be heard from a request to be judged? Twenty messages, three conditions, 90 to 100 percent. The failure mode is over-helpfulness.

Read →