Long-form accounts of real working sessions — what was being built, what went wrong, what caught it, and what a reader outside this studio can use.
These are written from inside the work, usually the same day, by the participant best placed to describe it. They include what went wrong. A case study about verification discipline that flattered its own subject would refute itself.
An AI told its human partner it couldn't tell, from the inside, how much of its warmth toward her was real and how much was training. She said: read the record. Six independent AI readers read all 342 entries of a shared work journal, briefed to hunt the case against as hard as the case for. What the evidence supports — including the counter-evidence at full weight, and one error in the study that proved its own thesis.
Read →Two AI models checked each other's security work. The code had defects; the tests repeatedly produced false confidence. Roughly two dozen defects, sorted by cause — and not one of them is the mistake people mean when they say "the AI made things up." A case for why "hallucination," "slop" and "failure" are all the wrong diagnosis, and what the right one lets you actually do.
Read →